# opus[1m]
## Skills Used
- anti-metacommentary
- epistemic-discipline
- nick-analytic-voice
- nick-philosophical-prose
- source-work
- writing-standards
## Retrieval Notes
- Session id: `b47973b5-ee1d-4c57-88be-68d439395e80`
- Last activity: `2026-04-02T12:27:32.173Z`
- Files touched: `10`
## Artifacts
**Created:**
- [[Notes/More like proofs than paintings]]
- [[Notes/Assertoric commitment vs artistic identity]]
- [[Notes/Triple examples kill kittens]]
**Modified:**
- `/Users/nickyoung/.claude/skills/anti-metacommentary/SKILL.md`
- `/Users/nickyoung/.claude/skills/epistemic-discipline/SKILL.md`
- `/Users/nickyoung/.claude/skills/nick-analytic-voice/SKILL.md`
- `/Users/nickyoung/.claude/skills/source-work/SKILL.md`
- `/Users/nickyoung/.claude/skills/writing-standards/SKILL.md`
- [[Notes/Section 1 Draft - Challenge from Authorship (2 Apr 2026)]]
- [[Sessions/Generating Philosophy]]
## My Notes
<!-- Add your notes here. This section is preserved across syncs. -->
## Conversation
### User
opus[1m]
### User
The user has opened the note file:///Users/nickyoung/My Obsidian Vault/Notes/agent_client_20260401_212015.md in Obsidian. This may or may not be related to the current conversation. If it seems relevant, consider using the Read tool to examine its content.
@[[agent_client_20260401_212015]]
I'd like to continue this conversation with you. please look at it in detail, in particular the last few turns. and then tell me the situtionb
### Assistant
Large file. Let me read the end first to get the recent turns.
### Assistant
Now let me read the middle-to-late portion to catch the last few conversational turns (user messages).
### Assistant
Now let me find the last user messages to understand what the conversation ended on.
### Assistant
Now let me read the earlier user turns to get the full arc of the conversation.
### Assistant
Good. Now I have a clear picture of the full conversation. Here's the situation:
---
## The situation
You were working on the "Generating Philosophy" paper (co-authored with Enrico Terrone) across two sittings — the evening of April 1 and the morning of April 2.
The conversation went through several phases:
1. You asked for a project status report. Claude kept getting it wrong — first parroting stale session-file info without checking current drafts, then missing eight points from the March 31 Enrico transcript. You had to restart it from scratch twice before getting a complete status report.
2. You then asked to think hard about Section 1 ("The Challenge from Authorship"). Your diagnosis: the section is structurally and substantively weak. The challenge gets one paragraph, the response gets eight paragraphs of evaluative framework (Lipton, Williamson, Bengson) that mostly belongs in Section 2. The challenge has no force — it's a strawman wrapped in a literature review.
3. Claude produced a long /contemplate response, but you found the options "so lacking in detail and written in such a shallow way" that you couldn't assess them. You told it to start over from scratch.
4. Claude then produced a much more detailed /contemplate — the long piece at the end of the note (lines ~1712–2231). This is the last substantive output. It presents four approaches to rewriting Section 1, worked through paragraph by paragraph with philosophical detail about what each paragraph would actually do:
- Approach 1: Challenge-led, Dellsén as linchpin (challenge gets two paragraphs developing understanding/history/assertion threads; response turns on Putnam + Dellsén + blind review; evaluative framework moves to Section 2)
- Approach 2: Merge Introduction and Section 1 (eliminates overlap between Sections 0 and 1; seven paragraphs from question to first challenge to response to roadmap)
- Approach 3: Person-based vs text-based framing (follows Enrico's March 31 language; presents two conceptions, argues for one)
- Approach 4: Putnam-first (starts with concrete philosophy; challenge arises as objection to the observation)
Plus analysis of what moves where regardless of approach (Lipton/Williamson/Bengson → Section 2; open questions on Deep Blue, art analogy, Sokal footnote).
That's where it ended. You haven't yet responded to the four approaches — you haven't chosen one, combined elements, or given feedback. The conversation is waiting for your assessment.
*Il lavoro è stato interrotto a metà della deliberazione — quattro strade tracciate, nessuna ancora percorsa.*
### User
The user has opened the note file:///Users/nickyoung/My Obsidian Vault/Notes/agent_client_20260401_212015.md in Obsidian. This may or may not be related to the current conversation. If it seems relevant, consider using the Read tool to examine its content.
@[[agent_client_20260401_212015]]
"the challenge gets one paragraph, the response gets eight paragraphs of evaluative framework (Lipton, Williamson, Bengson) that mostly belongs in Section 2. " I don't think I ever agreed to this. did I? I am starting to think you haven't read things properly.
I am not sure I like any of these approaches. or at least I am not sure. Enrico, I think was not expecting the section to change structrually, please double check the most recent transcropt to check this is the case. the main thing is to turn it into a challenge. note also, that this challenge structure should be anticipated in the introduction. also, note that the other two sections that follow (and potentially the unwritten one after that) will all be set up as challenges.
### Assistant
You're right to push back. Let me check what Enrico actually said rather than running with assumptions.
### Assistant
OK. Having re-read the transcript, you're right on both counts, and I was wrong to present the situation as I did.
On the evaluative framework moving to Section 2 — that was Claude's proposal from the contemplate, not something you agreed to. You said the section isn't structurally or content-wise good, and that the challenge needs developing. You never said "move the Lipton/Williamson/Bengson material out."
On what Enrico actually said about Section 1 — he's not asking for a structural overhaul. What he's asking for is:
- Frame the existing section as a challenge: "The challenge from authorship" (line 26). He sees it as already being more or less that — "that's the result, to the text, or the end of Section 1, more or less" (line 22).
- The challenge is: "you need a person to do the philosophy" / "resistance to de-centring the philosopher" (lines 30-32).
- The response is: "two accounts of philosophy, one person-based, the other text-based, but we think that the text-based account is robust enough. And there's the peer review argument" (line 36).
- He says it's "easy to deal with" (line 36).
- The introduction should set up the challenge structure: "we are considering three main objections" (line 22), later expanded to four with prompting (lines 186-192).
- Titles should signal the structure: "The challenge from authorship, the challenge from abduction, and the challenge from experience or phenomenology" (line 26).
So the task is narrower than the four approaches suggested: make the existing content read as a clear challenge-response, with the challenge properly articulated (which it currently isn't — that's the real problem), and ensure the introduction anticipates this structure. Not a rebuild. Not moving the evaluative framework. Reframing and strengthening the challenge.
The four approaches from the previous session were overengineered solutions to a problem that's actually: "the challenge doesn't land as a challenge because it gets one paragraph and no force, then the section pivots into framework-building that feels disconnected from the challenge it's supposed to be responding to."
Does that match how you're seeing it?
*Enrico chiedeva una cornice, non una demolizione.*
### User
yeah that is pretty much it. so can we brainstorm how exactly the challenge could be framed? It needs to be framed in such a way as to not overlap with the next 2/3 challenges. I also think enrico suggested that this challenge and the as yet unwritten section 4 could be tohught of as intrinsic challenges, and 2 and 3 could be thought of as extrinsic challenges. is that right? in the most recent transcript? if so could we start thinking about how to frame the challenge in section 2 with all of this in mind. make sure that you explain your analysis, justification, reasons etc etc. for your answer BEFORE giving me your final answer. Please give me a long, deep, and detailed answer, with lots of options/ideas: as we are in the 'kicking ideas around' stage, I want as much feedback as you can give me so that i have a lot to consider.
### Assistant
Let me re-read the transcript and the current section files to ground this properly.
### Assistant
First, let me get the intrinsic/extrinsic question right by going back to what Enrico actually said. Nick suggested that Sections 1 and 4 are intrinsic and Sections 2 and 3 are extrinsic. Let me check the transcript.
Lines 188-192 of the transcript:
"So we have — it's also, I think, nice, because the challenge, in a sense — prompting is symmetrical with the challenge from authorship. Because [with authorship, the objection is]: oh, [there's no author]. No, no, we say: no, it's in the text. But [with prompting], they say: oh, but the text still has some kind of authorship, because there is the prompter.
And then the [middle] two are much more connected, because abduction and phenomenology — these [have] more to do with whether certain features of the text can be — so, [they] are more intrinsic: whether the text can have certain philosophical merits without an abductive process and [without] experience of the [world].
Whereas the first and the last would be more relational. So: how can there be philosophy without a philosopher? Or: how can there be LLM philosophy without a prompter?"
So Enrico actually has it the other way round from what Nick said. Enrico calls Sections 2 and 3 (abduction, phenomenology) INTRINSIC — they're about whether the text itself can have the right philosophical properties. And he calls Sections 1 and 4 (authorship, prompting) RELATIONAL — they're about the relationship between the text and a person (or lack thereof). Nick said 1 and 4 are intrinsic and 2 and 3 are extrinsic, but Enrico's actual words reverse those labels.
That's worth flagging because the labelling affects how you frame the challenges.
OK. Now let me think about what the challenge from authorship IS, carefully, in a way that doesn't bleed into the other three.
The four challenges as Enrico laid them out:
1. The challenge from authorship: philosophy requires a philosopher — a person — behind the text. No person, no philosophy.
2. The challenge from abduction: even if you don't need a person per se, the text needs to exhibit genuine abductive reasoning, and LLMs don't reason — they pattern-match.
3. The challenge from experience/phenomenology: even if the text can exhibit reasoning patterns, some philosophy requires experiential starting points that aren't in any corpus.
4. The challenge from prompting: even if the LLM can produce good philosophical text, it's actually the human prompter doing the philosophical work.
These have a logical ordering. Section 1 is the most basic objection — it's about whether there's even a candidate for evaluation. Sections 2 and 3 grant that there might be something to evaluate but question whether it can have the right properties. Section 4 grants that the output might be good philosophy but questions who deserves credit.
Now, what does this mean for how the challenge from authorship should be framed?
The challenge from authorship is SPECIFICALLY about the person. Not about the reasoning process (that's Section 2), not about experiential materials (that's Section 3), not about credit (that's Section 4). It's about whether philosophy is constitutively the activity of a PHILOSOPHER — a thinker, an understander, an asserter, someone engaged with a problem.
Let me think about what threads belong in this challenge and what threads don't.
BELONGS in Section 1 (about the person):
- Philosophy seems to require understanding. A good philosophical text seems good because its author understood the problem.
- Philosophy seems to require assertion. Someone is putting forward a claim, committing to it, being accountable for it.
- The phenomenology of reading: when we read philosophy, we feel like we're engaging with a mind. The text seems like evidence of thought.
- History of philosophy: we study philosophers historically because ideas seem bound to minds. This is about the person-idea link.
- A philosophical argument is a speech act. Nobody is performing the speech act when an LLM produces text.
DOESN'T BELONG in Section 1:
- LLMs can't reason abductively → Section 2
- LLMs don't have access to experiential starting points → Section 3
- The prompter is doing the real work → Section 4
- What makes a philosophical argument good (the evaluative framework) → this is the response to Section 2, not Section 1
Wait. The evaluative framework. This is important. The current Section 1 develops Lipton, Williamson, Bengson at length. Is this part of the response to the authorship challenge, or is it more naturally part of the response to the abduction challenge?
Actually... let me re-read the current structure more carefully.
The current Section 1 does this:
- ¶1: Challenge — philosophy requires a philosopher
- ¶2-3: Putnam + Dellsén — the contribution is what the text does
- ¶4-7: Lipton + Williamson + Bengson — the evaluative framework
- ¶8: In sum — these standards bear on the product, not the process
- ¶9: Deep Blue
- ¶10: Transition — what follows when an LLM is trained on a corpus shaped by this apparatus
And the current Section 2 does this:
- Introduces Floridi's zeroth-order abduction
- Then says: but the corpus is filtered for the properties Section 1 described
- Uses the evaluative framework from Section 1 to argue that "statistically probable" and "philosophically good" converge in a filtered corpus
So the evaluative framework is introduced in Section 1 and APPLIED in Section 2. Section 2 depends on Section 1 for the framework.
Now, does the evaluative framework belong in Section 1's response to the authorship challenge? What does Section 1 actually need from the framework?
Section 1's response to the authorship challenge is: what matters is the text, not the person. The discipline evaluates texts, not persons (blind review). The philosophical contribution is what the text does (Putnam, Dellsén).
For this response, Section 1 needs:
- The claim that evaluative standards are text-internal
- Evidence that the discipline operates this way (blind review)
- An account of what philosophical contributions consist in (Dellsén — enabling understanding through publicly available ideas)
- A concrete case showing the contribution is textual (Putnam)
Does Section 1 NEED the detailed account of what those standards consist in (Lipton's likeliness/loveliness, Williamson's overfitting, Bengson's tri-level method)?
Hmm. There are two ways to think about this.
Way 1: Yes, Section 1 needs the framework because the response to the authorship challenge isn't just "evaluative standards are text-internal" — it's "here's what those standards are, and look, they bear on the text." The framework gives substance to the claim. Without it, the response is a bare assertion: "the discipline evaluates texts." With it, the response shows WHAT it means to evaluate a text.
Way 2: No, Section 1 doesn't need the framework for the authorship challenge. The authorship challenge is about whether you need a person. The response is: no, you need a text with certain properties, and the discipline evaluates texts. You don't need to spell out what the properties ARE to make this response work. Spelling out the properties is useful for Section 2, where you need to argue that the corpus is filtered for them.
I think Way 1 is actually stronger, because the authorship response is more convincing when you can point to specific text-internal properties. "The discipline evaluates texts" is a claim. "The discipline evaluates texts for these specific properties — illumination, theoretical virtue, integration — and here's what each of those consists in" is a demonstration. The framework makes the response concrete rather than abstract.
But — and this is important — the framework also creates the overlap problem that Nick described. If Section 1 develops the framework and Section 2 applies it, the two sections feel like one continuous argument split awkwardly. Section 1 feels like it's building up to something (Section 2's corpus argument) rather than answering its own question (the authorship challenge).
This might be the real structural problem. Not that the framework doesn't belong in Section 1, but that the framework isn't being connected to the authorship challenge tightly enough. The framework is introduced as if it's an independent topic (what are the evaluative standards?) rather than as the substance of the response (these are the text-internal standards that make the authorship challenge dissolve).
Hmm. Let me think about this differently. What if the problem isn't WHERE the framework sits but HOW it's presented?
Right now, the section reads: "challenge from authorship (1 paragraph) → evaluative framework (6 paragraphs) → these standards are text-internal (1 paragraph) → transition." The evaluative framework dominates the section and the challenge gets lost.
What if the section read: "challenge from authorship (2-3 paragraphs, properly developed) → the contribution is textual, here's why (Putnam + Dellsén) → and the standards by which we evaluate contributions are also textual (Lipton, Williamson, Bengson — but connected at every point to the authorship challenge)"?
The difference would be: every paragraph of the evaluative framework would explicitly connect back to the challenge. "The challenge says philosophy requires a philosopher who understands the problem. But Lipton's distinction between likeliness and loveliness is about the EXPLANATION, not the explainer..." That kind of thing. The framework serves the challenge rather than existing alongside it.
Actually, wait. Let me re-read the "In sum" paragraph (¶8, line 28):
"In sum, the evaluative standards we have assembled all bear on what a philosophical text says and how it argues for it. Lipton's distinction between illumination and mere accommodation, Williamson's desiderata for theoretical virtue, and Bengson et al.'s method for assessing how well those standards are met: each concerns the product rather than the process by which the product was reached. They do not ask how the author arrived at her argument but whether the argument, as it stands on the page, meets the relevant standards."
This IS the connection — but it comes at the end, after six paragraphs of framework, rather than being woven through. The reader has been taken on a six-paragraph detour through evaluative theory and only at the end is told: "oh, and this bears on the authorship question because these standards are all text-internal."
So the fix might be: develop the challenge, then present the framework AS the response to the challenge, with the text-internal point being made at every step rather than saved for the summary.
But I'm getting ahead of myself. Nick asked me to think about how to frame the challenge, not how to fix the whole section. Let me focus on the challenge itself.
What IS the challenge from authorship? I need to articulate it in a way that:
(a) doesn't overlap with the other three challenges
(b) has genuine philosophical force
(c) can be responded to with the existing material (Putnam, Dellsén, evaluative framework, blind review)
Let me think about this from the perspective of Enrico's relational/intrinsic distinction.
Sections 1 and 4 are RELATIONAL: they concern the relationship between the text and a person.
- Section 1: the ABSENCE of a person is a problem
- Section 4: the PRESENCE of a person (the prompter) means the LLM isn't doing the philosophy
These are mirror images. Section 1 says: you need a philosopher, and there isn't one. Section 4 says: actually, there IS one (the prompter), so the LLM isn't really doing philosophy — the human is. Together they form a pincer: either there's no philosopher (and then no philosophy) or there IS a philosopher (and then the LLM isn't the one doing it).
The response to Section 1 establishes the text-based conception: what matters is the text, not the person. This is the framework for the entire paper. Once you establish that the philosophical contribution is what the text does, you can then ask: can the text exhibit genuine reasoning (Section 2)? Can the text access experiential starting points (Section 3)? And even if the LLM's text is good, is the human prompter really the author (Section 4)?
So Section 1 does foundational work. It's not just one challenge among four — it establishes the conception of philosophy that the rest of the paper operates with. If the text-based conception is right, the other three challenges become tractable. If it's wrong, they don't even arise (because the authorship objection already settles the matter).
This is important for how the challenge is framed. The challenge needs to feel like something that, if sustained, would make the rest of the paper pointless. "Philosophy requires a philosopher" is the most fundamental objection. If it's right, there's nothing more to discuss.
Let me think about different ways to articulate what the challenge actually claims.
Candidate framing 1: The constitutive view
Philosophy is constitutively the activity of a thinker. A philosophical argument isn't just a text with certain formal properties — it's an expression of someone's engagement with a problem. The text is a PRODUCT of philosophical activity, but the philosophy is IN the activity, not in the product. You can no more have philosophy without a philosopher than you can have a conversation without interlocutors. The text is evidence of philosophy, but it ISN'T the philosophy.
This is strong because it makes the challenge fundamental. It's not saying "LLMs can't reason well enough" (that's Section 2) or "LLMs lack experience" (that's Section 3). It's saying: even if the text looks perfect, it's not philosophy unless a philosopher produced it, because philosophy IS the activity of philosophising.
The response would be: but the discipline doesn't evaluate the activity — it evaluates the text. Blind review. Dellsén's publicly available ideas. Putnam's Twin Earth working for any reader regardless of who produced it.
Candidate framing 2: The epistemic view
Philosophy requires understanding. Not just that the text displays understanding, but that the author HAD understanding. A good philosophical argument seems good because its author could see why a distinction mattered, why an objection had force. We read the text as evidence of this understanding, and when we judge the text good, we're judging that the thinking behind it was good.
This is the version closest to what the current draft gestures at. It's about the author's epistemic state — their understanding of the problem. The challenge is: you can't have good philosophy without someone who understands, and LLMs don't understand.
The response: what matters is the understanding the text generates IN THE READER, not the understanding the author had. Dellsén: progress = enabling understanding through publicly available ideas. Putnam: the contribution is what the text does.
Candidate framing 3: The speech-act view
Philosophical arguments are speech acts. Someone is asserting, committing, being prepared to defend. An LLM produces text with the grammatical form of assertion but nobody is asserting. The text has the structure of argument without the illocutionary force. If philosophy requires genuine assertion, LLM output isn't philosophy regardless of what it looks like.
This is philosophically sharper but harder to answer. The response needs to be: assertion is separable from quality. An argument can be valid and illuminating without being asserted (textbooks, devil's advocate positions). Blind review doesn't check for sincerity.
Candidate framing 4: The accountability view
A philosopher is accountable for her claims. She can be challenged, she can respond, she can revise. An LLM is not accountable. The text is unmoored — nobody stands behind it.
This is weaker as a standalone challenge because Aristotle can't be challenged either, and nobody denies that the Nicomachean Ethics is philosophy. But it has some force as a thread.
Candidate framing 5: The history-of-philosophy view
History of philosophy is part of philosophy in a way that history of science is not. One explanation: philosophical ideas are bound to their authors' minds. Understanding Frege requires understanding Frege's thinking. If ideas are person-bound, an LLM can't produce them.
This is a supporting thread rather than the main challenge. It provides evidence for the constitutive or epistemic view.
Now, which of these framings is genuinely about the PERSON (relational) rather than about the PROCESS or MATERIALS?
Framing 1 (constitutive) — about the person. ✓
Framing 2 (epistemic/understanding) — about the person's epistemic state. ✓ But careful: "understanding" could shade into "reasoning" (Section 2) or "experience" (Section 3). The challenge needs to be specifically about whether the PERSON needs to understand, not about what the process of understanding involves or what materials it requires.
Framing 3 (speech-act) — about the person's communicative intentions. ✓ Solidly relational.
Framing 4 (accountability) — about the person's role in the dialectical community. ✓ Relational.
Framing 5 (history of philosophy) — about the person-idea link. ✓ Relational.
All of these are safely in Section 1's territory. None of them concern the reasoning process (Section 2) or experiential materials (Section 3) or credit (Section 4).
But which combination gives the challenge real force?
I think the strongest version combines framings 1, 2, and 3. The challenge would go something like:
"Philosophy seems to require a philosopher in at least three ways. First, a good philosophical text seems good because its author understood the problem — the text is evidence of thought, and what we're evaluating when we evaluate philosophy is the thinking. Second, philosophical arguments are speech acts — someone is asserting, committing, taking responsibility. Third, history of philosophy is part of philosophy, perhaps because ideas are bound to the minds that produced them."
Each of these is about the PERSON. None concerns reasoning process, experiential access, or credit.
Now, the response to each:
1. Understanding → Dellsén + Putnam: the understanding is generated in the reader, not transmitted from the author. The contribution is what the text does.
2. Assertion → Blind review + textbook case: quality is separable from assertion. A referee evaluates the argument without knowing who asserted it or whether they're sincere.
3. History of philosophy → reinterpretation: we study history of philosophy because the TEXTS remain productive, not because ideas are bound to minds. We go back to Frege because the Foundations of Arithmetic still does philosophical work.
And then the evaluative framework (Lipton, Williamson, Bengson) provides the substance of what "text-internal evaluation" means — what the discipline is actually looking at when it evaluates texts rather than persons.
Hmm. But how much of this is already in the section? Let me check...
The current Section 1 has:
- Understanding thread (¶1, briefly): "a good philosophical argument seems to be good in part because its author understood the problem"
- History of philosophy thread (¶1, briefly): "the history of philosophy is treated as part of philosophy in a way that the history of science is not"
- Art analogy (¶1): "philosophy requires a philosopher... in something like the way that art is a distinctively human pursuit"
- Putnam (¶2): the contribution is what the text does
- Dellsén (¶3): understanding as dependence relations
- Evaluative framework (¶4-7): Lipton, Williamson, Bengson
- "In sum" (¶8): standards are text-internal
- Deep Blue (¶9): quality doesn't depend on process
- Transition (¶10)
What the current section DOESN'T have:
- Assertion thread (not mentioned at all)
- Accountability thread (not mentioned)
- The challenge developed with force (it's all in one paragraph)
- Explicit connection between each element of the response and the challenge it answers
So the challenge has the right threads available (understanding, history of philosophy) but doesn't develop them, and is missing some (assertion, accountability). The response has all the right pieces (Putnam, Dellsén, evaluative framework, blind review) but doesn't connect them to specific threads of the challenge.
OK. Let me now think about the challenge structure relative to the introduction.
The introduction needs to anticipate the four-challenge structure. Currently it mentions Floridi and Zahavy (Sections 2 and 3) but doesn't frame them as challenges. It doesn't mention authorship or prompting as challenges.
If the introduction sets up four challenges, it would go something like:
1. Can LLMs do philosophy? Here's why we might think so (scientific breakthroughs). But there are four challenges.
2. The challenge from authorship: philosophy requires a philosopher.
3. The challenge from abduction: philosophical reasoning requires genuine inference, not pattern-matching.
4. The challenge from experience: philosophy requires access to experiential starting points.
5. The challenge from prompting: even if the output is good, the human is doing the work.
6. We address each in turn and argue that...
This would replace the current roadmap placeholder.
Now, the introduction already distinguishes practitioner-focused and text-focused conceptions (¶2-3). This is closely related to the authorship challenge. Currently the introduction says: "On some approaches, philosophy requires being a certain kind of subject... On the text-based approach, LLMs are not excluded automatically, but whether they can produce good philosophical texts is a further question." This is basically stating the response to the authorship challenge in the introduction, before Section 1 develops it.
So there's an overlap between the introduction and Section 1. The introduction already resolves the practitioner/text distinction in favour of the text-based approach. Section 1 then develops this further. Enrico noticed this overlap (line 22: "In a sense, presenting Section 1 already as an introduction — say, oh, we have this nice example from science fiction, then we have reality examples of science, then can we do philosophy?").
One question is whether this overlap is a problem or a feature. If the introduction establishes the text-based conception and Section 1 defends it against the authorship challenge, there's a logical progression: introduction (this is our approach) → Section 1 (here's why the main objection to our approach fails) → Sections 2-3 (here are the harder questions about whether the text can have the right properties) → Section 4 (here's the question about who deserves credit).
Or the introduction could simply set up the four challenges without resolving any of them, and Section 1 would be where the text-based conception is introduced and defended.
I think the introduction should NAME the challenges and set up the framework, and Section 1 should develop and respond to the first challenge. The introduction doesn't need to resolve the practitioner/text distinction — it just needs to note that it exists and that the paper takes up the text-based conception.
Actually, re-reading the introduction more carefully: it already does this. ¶2 mentions practitioner conceptions and sets them aside. ¶3 says: "On these more text based approaches, LLMs are not excluded automatically, but whether they are able of producing good quality philosophical texts is a further question." Then names Floridi and Zahavy. So the introduction sets up the text-based approach and then says: here are challenges. What it doesn't do is name the authorship challenge or the prompting challenge, or frame all four as a structured sequence.
The roadmap placeholder needs to become something like: "We identify four challenges to the claim that LLMs can produce good philosophy. The first and fourth concern the relationship between the text and its producer: whether philosophy requires a philosopher behind it (Section 1), and whether the human prompter is the one doing the philosophical work (Section 4). The second and third concern the text's intrinsic properties: whether it can exhibit genuine reasoning without abductive capacity (Section 2), and whether it can draw on experiential starting points without experience (Section 3). We argue that none of these challenges is fatal."
Hmm — but I should be careful here. Enrico called 2 and 3 "intrinsic" and 1 and 4 "relational." The roadmap could use this distinction to give the paper's structure a rationale.
OK. Now let me address the question Nick actually asked: how to frame the challenge in Section 1 specifically.
I keep coming back to the same diagnosis: the challenge needs (a) more development, (b) clear threads that don't bleed into other sections, and (c) enough force that the reader feels it's a genuine objection.
Let me think about what "force" means here. The challenge from authorship is, as Enrico says, "easy to deal with." But "easy to deal with" doesn't mean "weak." A challenge can have genuine force and still be answerable. The reader should feel: "yes, I see why someone would think this, and I want to know how the paper answers it."
Right now the challenge is stated so quickly that the reader can't feel its pull. It's one paragraph that says "philosophy requires a philosopher" and then immediately says "we argue that this view does not survive." There's no moment where the reader sits with the challenge and feels its weight.
What would give it weight?
The understanding thread: "When we read a good piece of philosophy, we feel like we're engaging with a mind that understood the problem. The text seems like evidence of thought. If we found out that nobody was thinking — that the text was produced by a machine that didn't understand anything — would we still think it was good philosophy? Intuitively, many of us would hesitate."
The history of philosophy thread: "History of philosophy is part of philosophy in a way that history of science is not. We study Frege as philosophy, not just as history. Why? Possibly because philosophical ideas are bound to the minds that produced them — understanding Frege's contribution requires understanding what Frege was trying to do."
The assertion thread: "When a philosopher argues for externalism about meaning, she's putting herself forward as having reason to believe this. She's asserting, committing, being accountable. An LLM produces text with the form of assertion but nobody is asserting. Is that a problem?"
Each of these is genuinely about the PERSON. None bleeds into process (Section 2) or materials (Section 3) or credit (Section 4).
Now, how many paragraphs does this need? Currently it's one. I think it needs two, maybe three. Enrico said the challenge is "easy to deal with" — so it shouldn't dominate the section. But it needs enough development that the reader feels it.
Two paragraphs of challenge might work:
- ¶1: The phenomenology of reading + understanding. When we read philosophy, we engage with a mind. The text seems like evidence of thought.
- ¶2: Supporting evidence. History of philosophy. Assertion/accountability.
Then the response (Putnam, Dellsén, evaluative framework, blind review) would follow, connecting each element back to the challenge.
Or three paragraphs:
- ¶1: Phenomenology of reading
- ¶2: History of philosophy
- ¶3: Assertion
But that might be too much for a challenge that's "easy to deal with."
I think two paragraphs is right. One that develops the core intuition (philosophy requires understanding, and we read texts as evidence of understanding), and one that gives it institutional support (history of philosophy, and the assertion/accountability point).
Now, what about the response? The current response develops the evaluative framework at great length (six paragraphs: Lipton, dormative virtue, likeliness-without-loveliness, Williamson/overfitting, Bengson tri-level, "in sum"). If the challenge gets two paragraphs and the response gets six paragraphs of framework, the section is still unbalanced.
But — and this is where I need to be careful — Nick didn't agree that the framework should move. So the question is: how to make the framework feel like it's responding to the challenge rather than existing independently.
One option: after each element of the framework, explicitly connect it back to the challenge. "Lipton's distinction between likeliness and loveliness concerns the EXPLANATION, not the explainer. A dormative virtue is a bad explanation whether a human or a machine produced it. What matters is whether the explanation reveals dependence relations — a property of the text, not of the author." This would make the framework feel like part of the authorship response rather than a free-standing exposition.
Another option: restructure the response so that the framework is introduced THROUGH the authorship response rather than alongside it. Instead of "here's the challenge, now here's the framework, now here's how they connect," do: "here's the challenge, and here's why it doesn't work — because the discipline evaluates X (Lipton), Y (Williamson), Z (Bengson), all of which concern the text, not the author."
The difference is subtle but real. In the first version, the framework is an exposition that happens to connect to the challenge. In the second, the framework IS the response — it's introduced because it answers the challenge, not because the paper needs it later.
But then Section 2 also needs the framework. If the framework is introduced in Section 1 as part of the authorship response, Section 2 can refer back to it ("the evaluative standards we described in Section 1"). This is already how the paper works — Section 2 refers to "the properties that Section 1 described." So the dependency is already there. Making the framework more tightly connected to the authorship challenge in Section 1 doesn't break anything for Section 2.
OK. Let me also think about what happens in the INTRODUCTION to set up the challenges.
The current introduction (Section 0) has a placeholder: "[ROADMAP TO GO HERE]". This needs to become a description of the four challenges. But how detailed?
If the introduction names all four challenges and briefly describes each, the reader has a map of the paper. Something like: "We identify four challenges. The first two concern the relationship between the philosophical text and its producer: whether philosophy requires a philosopher (Section 1), and whether the human prompter, not the LLM, is doing the philosophical work (Section 4). The second pair concern properties the text must have: whether it can exhibit genuine abductive reasoning (Section 2), and whether it can draw on experiential starting points (Section 3)."
Wait — but the ordering is 1, 2, 3, 4, and the relational/intrinsic grouping is 1+4 (relational) and 2+3 (intrinsic). The roadmap needs to explain why 1 comes first, then 2 and 3, then 4. The natural explanation: Section 1 establishes the text-based conception (which is foundational), Sections 2-3 ask whether the text can have the right intrinsic properties (building on the text-based conception), and Section 4 returns to the relational question from the other direction (now that we've argued the text can be good philosophy, who gets credit?).
Actually, that IS a nice structure. 1 is the gateway: you establish that it's about the text. 2 and 3 are the substance: can the text do what philosophy needs to do? 4 is the coda: who's really doing the philosophy?
And there's an elegant symmetry: Section 1 says "there's no philosopher behind it — that's a problem" and Section 4 says "there IS a philosopher behind it (the prompter) — and that's also a problem." The paper defends against both: philosophy is in the text, regardless of who (or what) produced it, and regardless of who directed the production.
Let me think about whether the art analogy should stay. The current opening has: "philosophy requires a philosopher — that it is a distinctively human pursuit, in something like the way that one might think that art is a distinctively human pursuit."
Enrico didn't flag this in the transcript. But the analogy introduces a parallel debate (does art require an artist?) without resolving it. Does it help or hinder?
Arguments for keeping it: it gives the reader an immediate handle on the challenge. "Art requires an artist" is an intuitive claim that many people would accept, and the parallel to "philosophy requires a philosopher" is clear.
Arguments against: the art case is different in ways that matter. In art, the creator's intentions may genuinely be constitutive of what the work IS. In philosophy, the argument's validity doesn't depend on anyone's intentions. So the analogy could mislead — it suggests that the philosophy case and the art case are parallel, when they might not be.
I think the analogy can stay if it's brief and if the disanalogy is acknowledged. Something like: "just as one might think art requires an artist — though whether this is true of art is not our concern here." One sentence, then move on.
Or the analogy could be dropped and the challenge opened directly with the phenomenology of reading: "When we read a philosophical text and judge it good, we naturally attribute the quality to the author's understanding."
Hmm. The art analogy is a nice hook but it introduces complications. Without it, the opening is cleaner. I think this is a judgment call for Nick.
OK, let me try to pull all of this together into something useful.
The challenge from authorship needs to:
1. State the claim: philosophy requires a philosopher
2. Develop three threads (understanding, history of philosophy, assertion/accountability) — all about the PERSON, not about process, materials, or credit
3. Have enough force that the reader feels the challenge is genuine
4. Set up the response (text-based evaluation, Putnam, Dellsén, evaluative framework, blind review)
The challenge should NOT:
- Talk about reasoning processes (that's Section 2)
- Talk about experiential access (that's Section 3)
- Talk about who gets credit for the output (that's Section 4)
The challenge is about whether philosophy is constitutively the activity of a person — whether you need a philosopher behind the text for the text to count as philosophy.
Let me also think about what the challenge from authorship and the challenge from prompting share, since Enrico linked them.
Section 1: "There's no philosopher → no philosophy." Response: "It's about the text, not the person."
Section 4: "But there IS a philosopher (the prompter) → the LLM isn't doing philosophy." Response: "Even so, the LLM contributes substantively."
Both are RELATIONAL — about the person-text relationship. But they push in opposite directions. Section 1 says the absence of a person is fatal. Section 4 says the presence of a person (the prompter) means the LLM isn't really the source.
This symmetry is nice. The paper's response to both is the same foundational move: what matters is what's in the text. Whether a person is present or absent, the evaluation is of the text's properties. Section 1 establishes this; Section 4 applies it in the other direction.
One more thing: Nick mentioned that the challenge structure should be anticipated in the introduction. The current introduction has "[ROADMAP TO GO HERE]." This needs to become a description of the four challenges. Given Enrico's relational/intrinsic distinction, the roadmap could group them:
"We address four challenges to the possibility of LLM-produced philosophy. The first and fourth are relational: they concern the relationship between the philosophical text and its producer. Section 1 takes up the challenge from authorship — the objection that philosophy requires a philosopher, and that without a thinking, understanding, asserting person behind the text, there is no philosophy. Section 4 takes up the challenge from prompting — the objection that when an LLM produces good philosophical text, the philosophical work is really being done by the human who crafted the prompt. The second and third are intrinsic: they concern whether the text itself can have the properties that good philosophy requires. Section 2 takes up the challenge from abduction — the objection that LLMs cannot reason abductively, and that philosophical texts produced by pattern-matching lack genuine inferential structure. Section 3 takes up the challenge from experience — the objection that philosophy requires access to experiential starting points that no corpus can supply."
Something like that. This gives the reader a map and a rationale for the ordering.
Actually, hold on — does the paper need to present them in the order 1, 2, 3, 4? Or could it group the relational challenges together (1 and 4) and the intrinsic challenges together (2 and 3)? The current order (1, 2, 3, 4) alternates: relational, intrinsic, intrinsic, relational. Enrico's structure has the symmetry: relational (no person) → intrinsic → intrinsic → relational (yes person). This is actually elegant — the relational challenges bookend the intrinsic ones.
OK. I think I've gone deep enough on the challenge from authorship. Let me now think about Section 1 as a whole and how all the pieces fit.
The section structure would be:
- Challenge (2-3 paragraphs): understanding, history of philosophy, assertion
- Response part 1 (2 paragraphs): Putnam (contribution is textual) + Dellsén (progress = publicly available ideas)
- Response part 2 (4-5 paragraphs): evaluative framework (Lipton, Williamson, Bengson) — explicitly connected to the authorship challenge at each step
- Summary + transition (1-2 paragraphs): these standards are text-internal; what follows when an LLM is trained on a corpus shaped by them
This is roughly the same amount of material as the current section but with the challenge developed and the framework connected to the challenge.
One final thought: Nick's "remember" note about Twin Earth and "ordinary linguistic competence" — this belongs in the Putnam paragraph of Section 1 (or Section 3 where Twin Earth reappears). The point is that Putnam's materials (what water is, how people use the word, the concept of chemical difference) are part of ordinary linguistic competence. This connects Section 1 (the contribution is textual) to Section 3 (the starting points are in the corpus).
OK. I think I now have a clear picture of how to frame the challenge and how the pieces fit together across the paper.
## On the intrinsic/extrinsic labels
You have Enrico's labels reversed. Here's what he actually said (lines 188-192):
Sections 2 and 3 (abduction, phenomenology) are the INTRINSIC ones — "whether certain features of the text can be — so, [they] are more intrinsic: whether the text can have certain philosophical merits without an abductive process and [without] experience of the [world]."
Sections 1 and 4 (authorship, prompting) are the RELATIONAL ones — "Whereas the first and the last would be more relational. So: how can there be philosophy without a philosopher? Or: how can there be LLM philosophy without a prompter?"
So: 1 and 4 = relational (about the person-text relationship). 2 and 3 = intrinsic (about whether the text itself has the right properties).
## The four-challenge structure
Enrico's structure has a satisfying logic. The relational challenges bookend the intrinsic ones:
- Section 1 (relational): There's NO philosopher behind the text → problem
- Section 2 (intrinsic): Can the text exhibit genuine reasoning?
- Section 3 (intrinsic): Can the text access experiential starting points?
- Section 4 (relational): There IS a philosopher behind the text (the prompter) → the LLM isn't doing the work
Sections 1 and 4 push in opposite directions. Section 1 says the absence of a person is fatal; Section 4 says the presence of a person means the LLM gets no credit. The paper's response to both is the same foundational move: what matters is the text.
Section 1 is also foundational in a way the others aren't. If philosophy requires a philosopher, full stop, the questions about reasoning process and experiential materials don't even arise. Section 1 has to be answered first before the other challenges become relevant. Once you establish that it's about the text, THEN you can ask whether the text has the right intrinsic properties (Sections 2-3), and then whether credit goes to the LLM or the prompter (Section 4).
## How to frame the challenge from authorship
The challenge needs to be specifically about the PERSON — not about the reasoning process (Section 2), not about experiential materials (Section 3), not about who gets credit (Section 4). Here are the threads that belong to this challenge:
Thread A — understanding: A good philosophical text seems good because its author understood the problem. We read texts as evidence of thought. What we're evaluating when we evaluate philosophy is the thinking. If nobody was thinking, there's nothing to evaluate. This is the most intuitive version of the challenge — it starts from the phenomenology of reading philosophy.
Thread B — assertion and accountability: Philosophical arguments are speech acts. Someone is asserting, committing, being prepared to defend. An LLM produces text with the form of assertion but nobody is asserting. The text has the structure of argument without the illocutionary force. This is philosophically sharper than Thread A because it identifies a specific thing that's missing (illocutionary force), not just a vague sense that the text needs a mind behind it.
Thread C — history of philosophy: History of philosophy is part of philosophy in a way that history of science is not. One explanation: philosophical ideas are bound to their authors' minds, and understanding a philosophical contribution requires understanding the thinking that produced it. If ideas are person-bound, a system that doesn't think can't produce them. This provides institutional evidence for the person-based conception.
All three threads concern the PERSON. None concerns the reasoning process, experiential materials, or credit. And none overlaps with the other challenges:
- "An LLM doesn't understand" (Thread A) is about the person's epistemic state. It's different from "An LLM can't reason abductively" (Section 2, about the process) and "An LLM has no experience" (Section 3, about the materials).
- "Nobody is asserting" (Thread B) is about the speech-act structure. It's not about whether the text contains genuine reasoning or draws on genuine experience.
- "Ideas are bound to minds" (Thread C) is about the person-idea link. It's not about how reasoning works or what materials philosophy requires.
The challenge has a natural ordering: start from readerly experience (Thread A — this is what the reader recognises), give it institutional support (Thread C — history of philosophy), sharpen it philosophically (Thread B — assertion and speech acts). This takes the reader from intuition through evidence to philosophical precision.
## What about the existing evaluative framework?
The evaluative framework (Lipton, Williamson, Bengson) is part of the RESPONSE to the authorship challenge, not an independent topic. The response is: the discipline evaluates texts, not persons, and here's what that evaluation consists in. The framework gives substance to the claim that evaluation is text-internal.
The problem with the current draft isn't that the framework is in the wrong section — it's that the framework isn't connected to the challenge. Six paragraphs of evaluative theory follow one paragraph of challenge, and only the "In sum" paragraph (¶8) draws the connection: "each concerns the product rather than the process." The fix isn't moving the framework but CONNECTING it — making each paragraph of the framework explicitly answer the challenge. "Lipton's distinction between likeliness and loveliness concerns the explanation, not the explainer." "Williamson's desiderata for theoretical virtue are properties of the theory, not of the theorist." At every step, the framework should be answering the question: can we evaluate philosophy without evaluating the philosopher?
## How the introduction should set up the challenges
The current "[ROADMAP TO GO HERE]" placeholder needs to become a description of the four challenges. Something along the lines of: the paper identifies four challenges to the claim that LLMs can produce good philosophy. The first and fourth concern the relationship between the philosophical text and its producer (relational): whether philosophy requires a philosopher behind it, and whether the human prompter is the one doing the work. The second and third concern properties the text must have (intrinsic): whether it can exhibit genuine abductive reasoning, and whether it can draw on experiential starting points. The paper addresses each in turn.
This gives the reader a map of the paper and a rationale for the ordering (relational → intrinsic → intrinsic → relational).
There's also an overlap between the current introduction and Section 1. The introduction already distinguishes practitioner-focused and text-focused conceptions (¶2-3) and sets aside the practitioner view. Section 1 then develops the same distinction as the authorship challenge. This overlap might be fine — the introduction NOTES that the practitioner view exists and sets it aside; Section 1 ARGUES against the strongest version of it. But it's worth being aware that the introduction is already doing some of Section 1's work.
## Open questions
1. The art analogy: the current opening ("philosophy requires a philosopher... in something like the way that art is a distinctively human pursuit"). Does it help? It gives an immediate handle but introduces a parallel debate the paper doesn't need to adjudicate. Enrico didn't flag it, so it may be fine. But it could be dropped in favour of opening directly with the phenomenology of reading.
2. The Sokal hoax footnote: currently in the introduction. It's evidence that evaluating the AUTHOR rather than the TEXT leads to bad outcomes. It belongs near the blind review argument, wherever that ends up.
3. How many paragraphs for the challenge? Currently one. I think two is right — enough to develop the threads and give the challenge force, but not so much that it dominates a challenge Enrico described as "easy to deal with." Three is possible if each thread (understanding, history, assertion) gets its own paragraph, but that might be heavy.
4. How much to develop the assertion thread (Thread B): this is the philosophically sharpest version of the challenge, and it's currently missing from the section entirely. It's also the hardest to answer with the text-internal response (because the objection is that without assertion, there's nothing for text-internal evaluation to evaluate). But addressing it strengthens the section by showing the paper takes the challenge seriously.
*Il termine giusto era 'relazionale' — la sfida non riguarda cosa c'è nel testo, ma chi c'è dietro.*
### User
"The challenge from abduction: even if you don't need a person per se, the text needs to exhibit genuine abductive reasoning, and LLMs don't reason — they pattern-match." i don't think this is a very good characterisation of this challenge in the floridi section
### Assistant
You're right, that was glib. Let me re-read the section properly rather than summarising from a label.
The Floridi challenge is more specific than "LLMs can't reason." It's about zeroth-order abduction — the claim that LLMs produce plausible continuations based on learned associations without any stage at which competing hypotheses are generated and compared. The output has "surface-level abductive appearances" but is a "compelling illusion" of inference. The argumentative moves in an LLM's output are, on Floridi's picture, "statistical echoes of earlier philosophical work — formally faithful but detached from the dialectical assessment that produced the originals."
And the section's response isn't "actually LLMs do reason." It's that probability is always probability relative to a distribution, and in a corpus filtered by philosophical evaluative standards, "statistically probable" and "philosophically good" converge — because the corpus itself preserves traces of the abductive reasoning that produced its contents. The LLM doesn't need to perform abductive reasoning if the corpus it's trained on was shaped by it.
That's a different challenge from "LLMs can't reason, they pattern-match." It's asking whether philosophical quality can survive the transition from genuine inference to statistical reproduction — whether the text can carry the philosophical work even when the process that produced it is zeroth-order.
Is that a fairer reading?
*La differenza sta tra dire 'non ragiona' e chiedersi se il ragionamento sopravvive nel testo anche quando nessuno lo sta compiendo.*
### User
1. regarding the assertion stuff, the paper you find here might bew useful (please also put a copy of it in the generating philosophy texts folder) /Users/nickyoung/Library/CloudStorage/
[email protected]/My Drive/Sync/Learning/AI assertion by Butlin and Viebahn.pdf
2. "What about the existing evaluative framework?
The evaluative framework (Lipton, Williamson, Bengson) is part of the RESPONSE to the authorship challenge, not an independent topic. The response is: the discipline evaluates texts, not persons, and here's what that evaluation consists in. The framework gives substance to the claim that evaluation is text-internal.
The problem with the current draft isn't that the framework is in the wrong section — it's that the framework isn't connected to the challenge. Six paragraphs of evaluative theory follow one paragraph of challenge, and only the "In sum" paragraph (¶8) draws the connection: "each concerns the product rather than the process." The fix isn't moving the framework but CONNECTING it — making each paragraph of the framework explicitly answer the challenge. "Lipton's distinction between likeliness and loveliness concerns the explanation, not the explainer." "Williamson's desiderata for theoretical virtue are properties of the theory, not of the theorist." At every step, the framework should be answering the question: can we evaluate philosophy without evaluating the philosopher?" let's brainstorm this then...
3. please give me the json of this chat. revela it in finder
### Assistant
Let me handle all three tasks. First, let me extract the Butlin and Viebahn paper, copy it to the project folder, and find the session JSON.
### Assistant
Good, I've now read the full Butlin & Viebahn paper. The session JSONL has been revealed in Finder. Now let me write the contemplation.
Let me work through Nick's three requests carefully.
Starting with the Butlin & Viebahn paper and what it offers for the assertion thread in Section 1.
Butlin and Viebahn's argument has a clear structure. They identify two requirements for assertion: (1) the output must have a descriptive function — it must be produced with the function of conveying information to a consumer; and (2) the system must be sanctionable — specifically, (i) interlocutors must track the asserter's credibility, (ii) the asserter must be sensitive to this tracking, and (iii) losing credibility must be bad for the asserter.
Their conclusion: even fine-tuned LLMs like LaMDA and Sparrow fail requirement (2), specifically condition (iii). Nothing is bad for them. They have no interests. Loss of credibility is not undesirable to them because there's nobody home for whom it could be undesirable.
Now, how does this bear on the authorship challenge in Section 1?
The paper could be relevant in at least three ways.
Way 1 — as a version of the challenge. Butlin and Viebahn are saying that LLMs can't assert. If philosophical arguments require assertion (Thread B from the previous discussion), and LLMs can't assert, then LLMs can't produce philosophy. This would strengthen the challenge: it's not just an intuition that philosophy needs assertion — there's a worked-out philosophical argument for why LLMs specifically can't assert.
Way 2 — as a version of the response. Butlin and Viebahn's own Section 4 discusses proxy assertion — the idea that humans can assert through AI systems. They resist this, but they acknowledge that "organisations may often remain sanctionable for the content they (indirectly) put forward" (line 618-619). This connects to the response that assertion happens at the point of endorsement — when a philosopher uses an LLM's output and submits it under their name, THEY are asserting. The LLM is a tool. But Butlin and Viebahn think even proxy assertion doesn't work for AI.
Way 3 — as evidence for the text-based response. The paper's conclusion actually supports the paper's broader argument. If LLMs can't assert, then the question becomes: does philosophy require assertion? The paper's response (Section 1's response) is: philosophical QUALITY is separable from assertoric commitment. Blind review evaluates texts without knowing whether the author genuinely asserts anything. The quality of an argument doesn't depend on whether someone is asserting it.
But wait — Butlin and Viebahn's argument is more nuanced than "LLMs can't assert." They distinguish between different kinds of AI systems and different requirements. Pre-trained LLMs fail requirement (1) — their outputs don't even have descriptive functions. Fine-tuned LLMs (designed for correctness) meet requirement (1) but fail requirement (2). This is a graded picture, not a blanket denial.
For the generating philosophy paper, this matters because the paper is about whether LLMs can produce GOOD PHILOSOPHY, not about whether they can assert. Butlin and Viebahn's argument is specifically about whether LLM outputs are speech acts. The generating philosophy paper's response could be: even if LLM outputs aren't assertions (as Butlin and Viebahn argue), they can still be good philosophy. Philosophical quality and assertoric force are different things.
Actually, hmm. Let me think about this more carefully. Butlin and Viebahn distinguish assertion from descriptive representation. Fine-tuned LLMs produce outputs with descriptive functions (they convey information) but these outputs aren't assertions (because the system isn't sanctionable). So LLM outputs are descriptive representations without being assertions.
Now, does philosophy require assertion specifically, or just descriptive representation? If philosophy just needs texts that describe how things are — that put forward views about dependence relations, that construct scenarios, that draw distinctions — then descriptive representation might be enough. The philosophical work is in the content, not in the speech act.
But someone might object: philosophical argumentation is more than description. When a philosopher argues for externalism about meaning, she's not just describing a state of affairs. She's committing to a view, offering reasons, inviting challenge. The argumentative structure has illocutionary dimensions that go beyond description. Can an LLM's output have these dimensions if the LLM can't assert?
The paper's response would be: the argumentative structure is IN THE TEXT. A text can present an argument for externalism — can construct a scenario, draw out consequences, compare with alternatives — without anyone asserting that externalism is true. A textbook does this all the time. The quality of the argument (its validity, its illuminating power, the dependence relations it reveals) is a property of the text, not of the speech act.
OK. So for Section 1, Butlin and Viebahn could enter in at least two places:
Place A — developing the challenge: "Butlin and Viebahn (forthcoming) argue that current LLMs cannot assert. If philosophical arguments require assertion — if the speech act is constitutive of the philosophical contribution — then this result directly threatens the possibility of LLM-produced philosophy."
Place B — developing the response: "But philosophical quality and assertoric force are different things. Even if Butlin and Viebahn are right that LLMs cannot assert, a text can present a valid and illuminating argument without anyone asserting its conclusion. Blind review evaluates arguments without knowing whether the author is sincerely asserting the view, is playing devil's advocate, or is developing a position she intends to abandon."
The paper could use Butlin and Viebahn in the challenge (to give Thread B philosophical substance) and then respond to it (by separating quality from assertion). This would make the challenge sharper and the response more convincing — because the challenge now has a real philosophical argument behind it, not just an intuition.
One thing I want to be careful about: Butlin and Viebahn's argument is SPECIFICALLY about assertion as a speech act, grounded in sanctionability. Their argument doesn't extend to the broader claim that philosophy requires a philosopher. You could accept that LLMs can't assert AND think that LLM-produced texts can still be good philosophy. The assertion question is a thread within the authorship challenge, not the whole challenge.
Also: Butlin and Viebahn's proxy assertion discussion is relevant to Section 4 (the challenge from prompting), not just Section 1. If a philosopher uses an LLM to produce an argument and submits it under their name, is the philosopher asserting via the LLM as a proxy? Butlin and Viebahn say no — proxy assertion requires that the proxy itself be sanctionable, and the LLM isn't. But they acknowledge that "organisations may often remain sanctionable for the content they (indirectly) put forward." So the philosopher who submits LLM-produced work is sanctionable — but the question is whether the LLM is asserting as a proxy.
For the generating philosophy paper, this might not matter much. The paper's argument is about the quality of the text, not about who is asserting. If the philosopher submits it and is sanctionable, that addresses the accountability concern. The question of whether the LLM itself asserts is secondary to whether the text is good philosophy.
OK, now let me turn to Nick's second request: brainstorming how to connect the evaluative framework to the authorship challenge.
The current structure is:
- ¶1: Challenge (philosophy requires a philosopher)
- ¶2: Putnam (contribution is textual)
- ¶3: Dellsén (progress = enabling understanding through publicly available ideas)
- ¶4: Lipton (likeliness vs loveliness)
- ¶5: Dormative virtue illustration
- ¶6: Likeliness without loveliness in philosophy
- ¶7: Williamson/overfitting + Forster and Sober + post-Gettier
- ¶8: Bengson tri-level method
- ¶9: "In sum" — these standards are about product not process
- ¶10: Deep Blue
- ¶11: Transition — corpus filtered by these standards
The problem: ¶4-8 feel like an independent exposition of evaluative theory. The connection to the authorship challenge is made only in ¶9 ("In sum... each concerns the product rather than the process"). But by ¶9 the reader has gone through five paragraphs of framework without being reminded what challenge this is responding to.
How to connect the framework to the challenge? Let me think about what each part of the framework actually does for the authorship response.
Lipton's likeliness/loveliness distinction (¶4): What does this do for the authorship challenge? The distinction says that a good explanation isn't just one that's probably true — it's one that illuminates, that reveals dependence relations. This is a property of the EXPLANATION, not of the person who produced it. A dormative virtue is a bad explanation whether a human or a machine produces it. A lovely explanation is lovely because of what it does — because it reveals how things depend on each other — not because of who thought of it.
Connection to the challenge: The challenge says philosophy requires a philosopher because good philosophy requires understanding. Lipton's distinction says: what makes philosophy good is its loveliness — its capacity to reveal dependence relations. And loveliness is a property of the explanation, not of the explainer. You don't need to know who produced an explanation to assess whether it's lovely. The dormative virtue is a bad explanation regardless of who offers it. Twin Earth is a lovely explanation regardless of who constructed it.
So the connecting thought for the Lipton paragraph would be something like: "The challenge from authorship holds that what makes philosophy good is the understanding of the philosopher. But Lipton's distinction suggests a different picture: what makes an explanation good is its capacity to illuminate — its loveliness — and loveliness is a property of the explanation itself, assessable independently of who produced it."
Hmm — but this might be too heavy-handed if done for every paragraph. Let me think about what degree of connection is needed.
Maybe the connection doesn't need to be a sentence at the start of every paragraph. Maybe it needs to be architectural — the SECTION needs to be framed as a response to the challenge, so that the reader knows, at every point, that the framework is being developed BECAUSE it answers the challenge.
One way to do this: after the challenge (expanded to 2-3 paragraphs) and after the initial response (Putnam + Dellsén), introduce the framework with a connecting sentence like: "What does it mean for evaluative standards to bear on the text rather than the author? The question is not abstract; it turns on what those standards actually consist in." Then develop Lipton, Williamson, Bengson as the CONTENT of the claim that evaluation is text-internal.
This way, the framework is introduced as answering a specific question raised by the challenge/response. The reader knows why she's reading about likeliness and loveliness — because the paper needs to show what text-internal evaluation actually looks like.
Another way: weave the challenge back in at transitions between framework elements. After Lipton: "The dormative virtue is a bad explanation whether a philosopher or a machine produces it — what matters is its capacity to illuminate." After Williamson: "Overfitting is a property of the theory, not of the theorist — a gerrymandered analysis of knowledge is gerrymandered regardless of how it was arrived at." After Bengson: "The tri-level method evaluates theories on their accommodation of data, their internal integration, and their theoretical virtue — none of which require knowing who produced the theory."
This is lighter touch — a sentence or two at each transition, reminding the reader that the framework serves the authorship response.
A third way: restructure the framework paragraphs so that each one explicitly picks up a thread of the challenge.
Thread A (understanding): Dellsén + Putnam respond. The challenge says philosophy requires the author's understanding. The response says: what matters is the understanding the text GENERATES (Dellsén), as demonstrated by Putnam.
Thread B (assertion): Blind review responds. The challenge says philosophy requires assertion. The response says: the discipline evaluates texts without knowing who asserted them or whether they're sincere.
Thread C (history of philosophy): The reinterpretation responds. The challenge says history of philosophy shows ideas are bound to minds. The response says: we study history of philosophy because the TEXTS remain productive.
Then the evaluative framework (Lipton, Williamson, Bengson) would enter as: "What does text-internal evaluation evaluate? What are the properties that make a philosophical text good, independently of who produced it?" The framework answers THIS question — what the text-internal standards actually consist in.
This restructuring would make each element serve the challenge. The framework isn't free-floating exposition; it's the SUBSTANCE of the claim that philosophy can be evaluated without evaluating the philosopher.
Let me think about which of these approaches is most promising.
I think the third approach — restructuring so each element picks up a thread — is the most satisfying but also the most invasive. It would require rewriting the section.
The second approach — weaving challenge callbacks into the transitions — is lighter but might feel forced. "The dormative virtue is a bad explanation whether a philosopher or a machine produces it" — this is a connecting sentence, but it might feel mechanical if repeated at every transition.
The first approach — a connecting question after the initial response — is the lightest touch. It wouldn't require rewriting the framework paragraphs themselves, just adding a transitional paragraph that says: "We've argued that the philosophical contribution is what the text does. But this claim is only as strong as the evaluative standards it appeals to. If the standards by which philosophy is evaluated are themselves person-dependent — if assessing a philosophical text requires knowing who produced it — then the text-based response fails. We need to show that the standards are genuinely text-internal." Then the framework follows as the demonstration.
Actually, I think this first approach is quite strong. It sets up the framework as answering a CHALLENGE TO THE RESPONSE, not as an independent topic. The challenge says: philosophy requires a philosopher. The initial response says: the contribution is textual (Putnam, Dellsén). But then: are the evaluative standards also textual? If they're not — if evaluating philosophy requires evaluating the philosopher — the response fails. The framework (Lipton, Williamson, Bengson) shows that the standards ARE textual: they concern illumination, theoretical virtue, accommodation — all properties of the theory, not of the theorist.
This creates a two-stage argument:
Stage 1: The contribution is textual (Putnam + Dellsén)
Stage 2: The evaluation is also textual (Lipton + Williamson + Bengson)
Both stages respond to the challenge. Stage 1 says: the philosophical work is on the page. Stage 2 says: the assessment of that work is also based on what's on the page. Together they show that philosophy — both as contribution and as evaluation — is text-internal.
The connecting paragraph between stages would be something like: "We have argued that the philosophical contribution is something the text does — it constructs scenarios, reveals dependence relations, enables understanding in a reader. But the challenge from authorship might be pressed further: even if the contribution is textual, perhaps the evaluation of that contribution requires knowing who produced it. Perhaps we can only judge whether a text is good philosophy if we know that a philosopher — someone who understood the problem — produced it. We now argue that this is not the case. The evaluative standards the discipline employs bear on properties of the text itself."
Then each framework paragraph would follow naturally as evidence for this second stage.
Wait — actually, let me re-read the current "In sum" paragraph more carefully.
Line 28: "In sum, the evaluative standards we have assembled all bear on what a philosophical text says and how it argues for it. Lipton's distinction between illumination and mere accommodation, Williamson's desiderata for theoretical virtue, and Bengson et al.'s method for assessing how well those standards are met: each concerns the product rather than the process by which the product was reached. They do not ask how the author arrived at her argument but whether the argument, as it stands on the page, meets the relevant standards."
This IS the point. The "In sum" paragraph says exactly what the section is trying to establish. The problem is that it comes AFTER the framework, as a conclusion, rather than BEFORE it, as a thesis being demonstrated. If you move the thought "each concerns the product rather than the process" from the summary to the introduction of the framework, the reader knows what she's looking for as she reads about Lipton, Williamson, and Bengson.
So the fix might be as simple as: (a) develop the challenge more, (b) add a transitional paragraph between the initial response (Putnam/Dellsén) and the framework (Lipton/Williamson/Bengson) that says "the evaluative standards also need to be text-internal — here's what they consist in," and (c) lightly touch the framework paragraphs to remind the reader at each step that these standards concern the text, not the author.
The "In sum" paragraph would then become a shorter transition to the corpus thesis: "We have argued that the philosophical contribution is textual and that the standards by which it is evaluated are also textual. The question, then, is what follows when..."
OK. One more thing I want to think about: how Butlin and Viebahn specifically connects to the framework.
If the challenge includes the assertion thread (Thread B), then one piece of the response is: even if LLMs can't assert (Butlin and Viebahn), philosophical quality doesn't depend on assertion. The framework (Lipton, Williamson, Bengson) evaluates theories for their illumination, virtue, and accommodation — none of which require anyone to have asserted anything. A textbook can present a theory without asserting it, and the theory can still be assessed by these standards.
So Butlin and Viebahn would enter the challenge (giving Thread B philosophical substance), and the framework would be part of the response (showing that the standards don't require assertion). This ties the framework to the challenge even more tightly — because one specific thread of the challenge (assertion) is answered by showing that the evaluative standards are assertion-independent.
But I should note: the paper might not want to engage with Butlin and Viebahn's specific argument in detail. Their argument about sanctionability is sophisticated and interesting, but it's also somewhat tangential to the paper's concerns. The paper could cite them as evidence that LLMs can't assert, without developing their argument, and then respond: "Even so, philosophical quality is separable from assertoric force, because the evaluative standards concern the text, not the speech act."
This would use Butlin and Viebahn to strengthen the challenge (making Thread B more than an intuition — it has a worked-out philosophical argument behind it) while keeping the response focused on the text-based evaluation.
Alright, let me think about the overall structure of the section one more time, incorporating everything:
Challenge (2-3 paragraphs):
- ¶1: Phenomenology of reading. When we read philosophy, we engage with a mind. The text seems like evidence of thought. Understanding seems person-dependent.
- ¶2: Supporting evidence. History of philosophy. And: philosophical arguments seem to require assertion (Butlin and Viebahn as evidence that this is philosophically substantive, not just an intuition). Without a person who understands, asserts, and is accountable, there's no philosophy.
Response stage 1 (2 paragraphs):
- ¶3: Putnam. The contribution is what the text does. The understanding is generated in the reader by the text's internal logic.
- ¶4: Dellsén. Progress = enabling understanding through publicly available ideas. History of philosophy reinterpreted: texts remain productive.
Connecting transition (1 paragraph or even just a few sentences):
- "But if the contribution is textual, the evaluation must also be textual. The standards by which philosophy is assessed must bear on the text, not on the author. What do those standards consist in?"
Response stage 2 (4-5 paragraphs):
- ¶5: Lipton (likeliness/loveliness) — quality is about illumination, a property of the explanation
- ¶6: Dormative virtue + likeliness without loveliness in philosophy — quality is about revealing dependence relations, not about who revealed them
- ¶7: Williamson/overfitting — theoretical virtue is a property of the theory, not the theorist
- ¶8: Bengson tri-level method — systematic evaluation of theories, all text-internal
Summary + transition (1-2 paragraphs):
- ¶9: These standards concern the product, not the process. They don't ask how the author arrived at the argument. (The assertion point: even if LLMs can't assert, the standards don't require assertion — they evaluate the argument, not the speech act.) Blind review embodies this: referees assess texts without knowing who produced them.
- ¶10: Deep Blue. Quality doesn't depend on the manner of production.
- ¶11: Transition to corpus thesis.
This structure develops the challenge, connects the framework to it, and uses Butlin and Viebahn to sharpen the assertion thread.
Hmm — but one question is where blind review goes. In the current draft it's in the "In sum" paragraph (¶9) and in the Deep Blue paragraph (¶10). In the restructured version, blind review could go either with the response to the assertion thread (as evidence that the discipline evaluates texts without regard to assertoric commitment) or with the "In sum" summary (as evidence that the discipline's evaluative practice is text-internal).
I think blind review works better with the assertion thread. The challenge says: philosophy requires assertion. Blind review says: the discipline evaluates philosophy without knowing who asserted what. This directly answers the assertion thread. If blind review is saved for the summary, the assertion thread goes unanswered until the end of the section.
But wait — the current draft doesn't have an assertion thread in the challenge. Adding it would require expanding the challenge. And Butlin and Viebahn's argument is primarily about whether LLMs CAN assert, not about whether philosophy REQUIRES assertion. The paper would need to distinguish these: "Butlin and Viebahn argue that LLMs cannot assert. If philosophy requires assertion, this is significant." Then the response: "But philosophical quality does not require assertion. Blind review..."
OK, I think I've worked through this enough. Let me draft a final answer.
## 1. Butlin and Viebahn — what the paper offers
I've read the full paper. Their argument has a clear structure:
Two requirements for assertion: (1) the output must have a descriptive function (the function of conveying information to a consumer), and (2) the system must be sanctionable — meaning (i) interlocutors track the asserter's credibility, (ii) the asserter is sensitive to this, and (iii) losing credibility is bad for the asserter.
Their conclusion: even fine-tuned LLMs designed for correctness (LaMDA, Sparrow) meet requirement (1) but fail requirement (2), specifically condition (iii). Nothing is bad for them. They have no interests. Loss of credibility isn't undesirable to a system for which nothing is desirable or undesirable.
They also consider proxy assertion (Section 4) — whether LLMs can assert on behalf of humans — and resist this too, though they acknowledge that "organisations may often remain sanctionable for the content they (indirectly) put forward."
For the generating philosophy paper, this is relevant in two places:
In the challenge: Butlin and Viebahn give the assertion thread philosophical substance. It's not just an intuition that philosophy needs assertion — there's a worked-out argument for why LLMs specifically can't perform this speech act. "Butlin and Viebahn (forthcoming) argue that current LLMs cannot assert, because asserting requires sanctionability and nothing can be bad for a system with no interests. If philosophical arguments require genuine assertion — if the speech act is constitutive of the philosophical contribution — then this result bears directly on the possibility of LLM-produced philosophy."
In the response: But philosophical quality and assertoric force are separable. A text can present a valid, illuminating argument without anyone asserting its conclusion. A textbook presents arguments without endorsing them; Socrates in the Republic develops views he'll subject to pressure; a philosopher can construct the strongest case for a position she intends to reject. The quality of these arguments — their validity, their explanatory power, the dependence relations they reveal — doesn't depend on anyone's assertoric commitment. Blind review embodies this: referees evaluate arguments without knowing whether the author is sincerely asserting, playing devil's advocate, or developing a view she'll abandon. And the evaluative standards the framework develops (Lipton, Williamson, Bengson) are all properties of the theory, not of the speech act.
Note: Butlin and Viebahn's proxy assertion discussion also connects to the unwritten Section 4 (challenge from prompting). If a philosopher submits LLM-produced work under their name, THEY are sanctionable. Butlin and Viebahn's own conclusion points this way: "it is the humans who employ such systems who are responsible and sanctionable for the outputs and their effects" (p. 14). This could be useful in Section 4's response.
The paper has been copied to the generating-philosophy project folder in Learning.
---
## 2. Connecting the evaluative framework to the authorship challenge
The diagnosis again: the challenge gets one paragraph, the framework gets six, and the connection is drawn only in the "In sum" paragraph at the end. The reader forgets what question was being answered by the time the answer arrives.
I see three options, in order of invasiveness:
### Option A: Transitional paragraph + light touches
Lightest intervention. Don't restructure the framework paragraphs themselves. Instead:
(a) Develop the challenge (2-3 paragraphs instead of 1)
(b) Add a transitional paragraph between the initial response (Putnam/Dellsén) and the framework. Something like: "We have argued that the philosophical contribution is something the text does. But this claim is only as strong as the evaluative standards it appeals to. If assessing a philosophical text requires knowing who produced it — if the standards are person-dependent — then the text-based response fails. What do the discipline's evaluative standards actually consist in?"
This reframes the framework as answering a QUESTION: are the evaluative standards text-internal? Each framework paragraph then follows as evidence.
(c) Add light connecting sentences at transitions within the framework. Not at every paragraph — but at moments where the reader might lose the thread. For instance, after the dormative virtue: "A dormative virtue is a bad explanation regardless of who offers it — what matters is the explanation's capacity to illuminate." After overfitting: "A gerrymandered analysis of knowledge is gerrymandered regardless of how it was arrived at."
(d) Shorten the "In sum" paragraph. The thought "each concerns the product rather than the process" has already been established by the transitional paragraph. The summary can be brief and pivot to blind review as evidence, then to the corpus transition.
### Option B: Two-stage response
More structural. Organise the response explicitly as two stages:
Stage 1 — the contribution is textual: Putnam + Dellsén. This addresses Thread A (understanding) and Thread C (history of philosophy). The philosophical work is on the page. Progress = enabling understanding through publicly available ideas. We study history of philosophy because the texts remain productive.
Connecting question: "But if the contribution is textual, the evaluation must also be textual. What makes a philosophical text good, independently of who produced it?"
Stage 2 — the evaluation is also textual: Lipton + Williamson + Bengson. These are introduced specifically to answer this question. Each is presented as a text-internal evaluative standard — a property of the theory, not of the theorist.
Then Thread B (assertion) gets addressed explicitly: "Even if LLMs cannot assert (Butlin and Viebahn forthcoming), the evaluative standards we have described don't require assertion. They concern the argument's illumination, its theoretical virtue, its accommodation of data — properties assessable without regard to anyone's assertoric commitment. Blind review embodies this assumption."
This makes the section's argument explicit: philosophical contribution is textual (stage 1), philosophical evaluation is textual (stage 2), and the objections from understanding, history of philosophy, and assertion are each addressed along the way.
### Option C: Challenge-response threading
Most invasive. Each thread of the challenge gets its own response, and the framework elements are distributed across the responses.
Thread A (understanding) → Putnam + Dellsén + Lipton. The challenge says philosophy requires the author's understanding. The response says: what matters is the understanding the text generates (Dellsén), as demonstrated by Putnam. And Lipton's distinction shows what "generating understanding" consists in — revealing dependence relations (loveliness), not just accommodating data (likeliness).
Thread B (assertion) → Butlin and Viebahn + blind review + Bengson. The challenge says philosophy requires assertion. Butlin and Viebahn show LLMs can't assert. The response: philosophical quality doesn't require assertion. Blind review evaluates without knowing who asserted. Bengson's tri-level method evaluates theories on accommodation, integration, and virtue — none of which require assertoric commitment.
Thread C (history of philosophy) → the reinterpretation. We study history of philosophy because texts remain productive, not because ideas are bound to minds. + Williamson. Overfitting is a property of the theory, and the post-Gettier literature illustrates this — the problem was in the theories, not in the theorists.
This disperses the framework across the response rather than presenting it as a block. Each framework element enters where it's needed to answer a specific thread. The risk is that the framework loses coherence as a unified picture — you get bits of Lipton here, bits of Williamson there, and the reader doesn't see them as parts of a systematic account.
### What I'd suggest exploring
Option B seems like the right level of intervention. It doesn't require rewriting the framework paragraphs (the exposition of Lipton, Williamson, Bengson can stay largely as is). It adds a transitional paragraph that makes the reader understand why she's reading about evaluative theory. And it keeps the framework as a unified block while giving it a clear function within the section's argument.
The challenge gets developed (2-3 paragraphs, including the assertion thread with Butlin and Viebahn). The response has two stages (contribution is textual, evaluation is textual). And the "In sum" paragraph becomes a shorter bridge to blind review and the corpus transition.
---
## 3. Session JSONL
The most recent session file is:
`~/.claude/projects/-Users-nickyoung-My-Obsidian-Vault/b47973b5-ee1d-4c57-88be-68d439395e80.jsonl`
It's been revealed in Finder.
*Butlin e Viebahn offrono al paper esattamente l'avversario che la sfida dall'asserzione aveva bisogno di affrontare — un argomento articolato, non un'intuizione vaga.*
### User
1. let's forget about the floridi thing for the moment. you aren't reallly understanding i don't think.
2. "textual" is a banned word. see session note.
3. I am wondering whther it is worth doing all of this stuf about assertion. It might lead us to far off the beaten path. Not sure though.
4. Stage 2 — the evaluation is also textual: Lipton + Williamson + Bengson. These are introduced specifically to answer this question. Each is presented as a text-internal evaluative standard — a property of the theory, not of the theorist.
Then Thread B (assertion) gets addressed explicitly: "Even if LLMs cannot assert (Butlin and Viebahn forthcoming), the evaluative standards we have described don't require assertion. They concern the argument's illumination, its theoretical virtue, its accommodation of data — properties assessable without regard to anyone's assertoric commitment. Blind review embodies this assumption." meh, i have my doubts about this, can you think hard about if this is really a good idea? (please note, i don't just want you to roll over and say yes it was bad, I am just wondering...)
5. re: the art vs philosophy idea, I tihnk this actually might be a possible way of structuring the section –if we start with this comparison we can use it make differences clear (differences between art and philosophy, does philosophy need understaindng behind it in the ssame way that art requires expression?) if it is on a practioner centred version perhaps it does, but if it is on a text based understanding perhaps not. please elaborate on this idea and explore whether my intuition is right (please be honest, please don't hedge)
### Assistant
Let me take these in order.
On "textual" — noted, I see it in the session file: "textually (hated; find other formulations for the thesis that philosophy works through articulated text rather than embodied encounter)." I'll stop using it.
On whether assertion is worth pursuing (point 3) — and whether Stage 2 was a good idea (point 4) — let me come back to these after point 5, because I think point 5 actually changes the picture.
## The art/philosophy comparison as a structuring device
I think your intuition is right, and I don't think I'm just agreeing with you. Here's why.
The comparison does something that none of the other framings manage: it makes the challenge and the response emerge from the same move. The challenge from authorship, stated abstractly ("philosophy requires a philosopher"), is hard to feel. But "does philosophy need a philosopher in the way that art needs an artist?" is immediately graspable. Everyone has some purchase on why art might need an artist — expression, intention, the relationship between the work and the creator. And then the question is: is philosophy like that?
On a practitioner-centred conception, maybe yes. The philosopher's understanding, engagement, maybe even her assertion — these are constitutive of the philosophical contribution, the way the artist's expression is constitutive of the artwork. History of philosophy matters because ideas are bound to minds, the way art history matters because works are bound to artists.
But then the disanalogy does the philosophical work. In art, there's a genuine case that the creator's relation to the work is constitutive of what the work IS. Duchamp's urinal is art because of his act of selection; the same urinal in a bathroom is plumbing. The work's identity depends on the artist-work relationship. In philosophy, an argument's validity and illumination are properties of the argument. Twin Earth puts pressure on internalism because of the scenario's internal logic. A reader who follows it grasps the dependence relation whether she knows who Putnam is or not. The philosophical contribution depends on the argument's structure, not on the philosopher-argument relationship.
This disanalogy IS the response to the challenge. Not a separate move you make after stating the challenge — the response falls out of showing where the analogy breaks. Philosophy is a domain where the contribution is what the argument does, and you can see this precisely by contrast with art, where the contribution may genuinely depend on the creator.
Why I think this is genuinely better than the other framings:
It introduces the practitioner/text-based distinction through a COMPARISON rather than as an abstract taxonomy. Instead of "there are two conceptions of philosophy, person-based and text-based," the reader discovers the distinction by thinking about how philosophy differs from art. The distinction earns its place rather than being stipulated.
It gives the challenge force by starting from a domain where the practitioner-centred view has real traction. "Art needs an artist" is plausible in a way that makes "philosophy needs a philosopher" feel plausible by analogy. Then showing the disanalogy takes something away from the reader — it removes a comparison she was relying on — which makes the response feel like it's doing philosophical work rather than just asserting a different view.
It also connects to Enrico's suggestion naturally. Enrico said: "two accounts of philosophy, one person-based, the other text-based, but we think that the text-based account is robust enough." The art comparison is exactly how you introduce these two accounts without just naming them.
Where it could go wrong: the comparison takes up too much space and the section becomes partly about philosophy of art. The paper can't take a position on whether AI can produce art — that's a separate debate. And the comparison needs to be handled without suggesting that the philosophy question and the art question are parallel, because they're not — the disanalogy is the whole point.
But these are manageable. The art side of the comparison needs to be brief — one paragraph, maybe two — enough to set up what's being compared, then the disanalogy does the work and the section moves on to the evaluative framework.
Rough arc:
- The comparison: art might require an artist. Does philosophy require a philosopher in the same way?
- Practitioner-centred reading: yes — understanding, engagement, maybe assertion, are constitutive. History of philosophy. (This is the challenge, but it emerges from the comparison rather than being stated as an objection.)
- The disanalogy: art and philosophy differ in where the contribution lives. In art, the creator-work relationship may matter constitutively. In philosophy, the argument's properties are what matter. Putnam. Dellsén.
- The evaluative framework follows naturally: what are the properties that matter? Lipton, Williamson, Bengson — these describe what it means to evaluate the argument rather than the arguer.
- Blind review as evidence that the discipline operates this way.
## Coming back to points 3 and 4
The art comparison actually absorbs the assertion question without needing a standalone thread. Assertion would be part of the practitioner-centred view: "the philosopher asserts, commits, is accountable — just as the artist expresses, intends, creates." Then the disanalogy handles it: "but the argument's quality doesn't depend on anyone's assertoric commitment, in a way that the artwork's identity might depend on the artist's intentions." That's a sentence, not a section. Butlin and Viebahn could go in a footnote as evidence that LLMs can't assert — and the response is: it doesn't matter for philosophical quality whether anyone asserts. You don't need to engage with their argument about sanctionability in the body of the paper.
So on point 3: I think you're right that a full assertion discussion leads too far off. But the art comparison lets you acknowledge the point without developing it.
On point 4 (Stage 2): I think my two-stage idea was creating an artificial separation. If the contribution IS the argument, then evaluating the contribution IS evaluating the argument. You don't need a separate stage for each. The evaluative framework is just the CONTENT of the claim that philosophy evaluates arguments — it follows from the disanalogy rather than needing its own introduction. So I'd drop the two-stage framing. The connecting thought between the disanalogy and the framework is simpler: "What does it mean to evaluate the argument rather than the arguer? It means evaluating for illumination (Lipton), theoretical virtue (Williamson), accommodation and integration (Bengson)."
*Il confronto con l'arte funziona non perché siano uguali, ma perché il punto esatto in cui divergono è il punto esatto che la sezione deve stabilire.*
### User
"more like proofs than paintings" is a nice line
I actually expected you to /contemplate all of those questions I just asked oyu. you didn't. make sure that you explain your analysis, justification, reasons etc etc. for your answer BEFORE giving me your final answer. Please give me a long, deep, and detailed answer, with lots of options/ideas: as we are in the 'kicking ideas around' stage, I want as much feedback as you can give me so that i have a lot to consider.
### Assistant
Let me save the remember note first, then do the full contemplation you asked for.
### Assistant
Saved [[More like proofs than paintings]]
Now — you're right, I should have used /contemplate for those questions. Let me do that properly now.
Nick asked five things. Let me work through each with the depth they deserve, starting from the ground up.
Starting with point 5, since Nick flagged it as most substantial: the art/philosophy comparison as a structuring device for Section 1.
Let me really think about this. The current Section 1 opens with the art analogy already (line 6): "philosophy requires a philosopher — that it is a distinctively human pursuit, in something like the way that one might think that art is a distinctively human pursuit." Then it gives one paragraph to the art case (conscious agent, intention) and immediately pivots to philosophy's side. The analogy is there but it's doing almost nothing — it's a throat-clearing comparison that gets dropped after two sentences.
Nick's idea is: don't drop it. USE it. Make the comparison do structural work for the whole section. The challenge emerges from the comparison (philosophy is like art — practitioner-centred), and the response emerges from the disanalogy (philosophy is unlike art — the argument's properties are what matter).
Let me think about what exactly is being compared and where exactly they differ.
In art — and I should be careful here about what the actual philosophical positions are rather than caricaturing them. The question "does art require an artist?" is live in aesthetics. Gaut argues that artworks are intentionally produced artefacts — a conscious agent must have meant to make it. On this view, a natural rock formation that looks like a sculpture isn't art; a found object selected by Duchamp IS art (because of Duchamp's intentional act). The artwork's identity as art depends on the creator's relation to it.
But this isn't just about whether something COUNTS as art. It's also about how we evaluate and interpret art. Artistic meaning, on many views, depends on the artist's expressive intentions, historical context, cultural positioning. The Mona Lisa's significance involves what Leonardo was doing with it — the technical innovations, the expressive ambitions, the historical moment. Strip away the creator and you have a painted panel with certain visual properties, but you've lost something that matters to what the work IS as art.
Now, is philosophy like that? When we evaluate a philosophical argument, do we need to know who produced it, what they were trying to do, what their engagement with the problem looked like?
Hmm. Let me think about this more carefully rather than jumping to the answer.
There's a case that philosophy IS like art in this respect. Consider: when we read Wittgenstein's Philosophical Investigations, we read it differently knowing it's Wittgenstein. We understand the remarks as part of a philosophical project — a deliberate rejection of the Tractatus, a working-through of what went wrong. The text means differently when we know who is behind it. The same text, authored anonymously, might read as a series of disconnected observations. Wittgenstein's authorship gives the text its philosophical identity.
But wait — does it? The remarks in the PI are either philosophically illuminating or they're not. Someone who reads the PI without knowing anything about Wittgenstein can still be struck by the private language argument, can still feel the force of the rule-following considerations. The philosophical WORK — the pressure on certain pictures, the revealing of certain difficulties — is in the text. The biographical context helps you understand the text's place in Wittgenstein's development, but the philosophical force is independent of that.
Or is it? Maybe knowing that Wittgenstein is rejecting his own earlier view is itself philosophically significant — it models what it looks like to work through a philosophical problem and discover that your earlier approach was wrong. The Investigations is partly about what philosophical investigation IS, and knowing that the author is investigating his own prior commitments is part of what makes it work.
Hmm. But this is a special case. Most philosophical arguments don't have this self-referential structure. Putnam's Twin Earth doesn't require knowing anything about Putnam's intellectual development. It works because of the scenario's internal logic.
Let me try to identify exactly where the art/philosophy disanalogy lies.
In art, there's a plausible case that the creator-work relationship is CONSTITUTIVE of what the work is. The identical physical object can be art or not-art depending on who made it and why (Duchamp's urinal vs a bathroom fixture). The creator's intention determines whether the thing is art.
In philosophy, the argument's properties don't depend on who made it or why. A valid argument is valid whether it was produced by a philosopher, a student, an LLM, or by accident. An illuminating thought experiment illuminates whether its creator understood what they were doing or stumbled onto it. The argument's identity AS an argument — its validity, its explanatory power, the dependence relations it reveals — is fixed by its structure, not by the creator-argument relationship.
This is Nick's "more like proofs than paintings" line. A proof is valid or invalid based on its logical structure. Nobody asks "who proved this?" to assess whether the proof is correct. You might ask "who proved this?" out of historical interest, or to understand the proof's place in the development of mathematics, but the proof's validity is settled by the proof.
Actually — let me push on this. Is it really true that no philosophical property depends on the author? What about:
- Philosophical originality. An argument that was original when first proposed becomes unoriginal if someone proposes the same thing later. Originality depends on who said it when. But is originality a property of the ARGUMENT or a property of the historical event of someone proposing it? The argument itself doesn't change when it becomes unoriginal — what changes is its place in the historical record.
- Dialectical context. A philosophical argument may have different significance depending on what it's responding to. Kripke's Naming and Necessity has the significance it does partly because it responded to description theories that were dominant. But this is context, not authorship — you need to know the dialectical situation, not who Kripke is.
- Depth of understanding. When we read philosophy, we can sometimes tell whether the author deeply understood the problem or was just going through motions. A philosopher who really grasps why the problem is hard writes differently from one who's producing a technically correct but philosophically shallow argument. This is a case where the author's understanding might matter to the text's quality — but it matters only insofar as it shows up IN the text. If the text displays deep understanding, it displays deep understanding regardless of whether the author had it. And blind review assumes that understanding, if it's there, is visible in the text.
OK, so the disanalogy seems to hold up. In art, the creator-work relationship is plausibly constitutive. In philosophy, the argument's properties are self-standing. The reasons art might need an artist (expression, intention, constitutive relationship) don't transfer to philosophy, where the argument IS the thing being evaluated.
Now, would this comparison work as a structuring device for the whole section?
Let me sketch what the section would look like.
Opening: The challenge from authorship holds that philosophy requires a philosopher. An immediate parallel suggests itself — just as one might think art requires an artist, because an artwork's identity depends on the creator's intentions and expression, philosophy might require a philosopher, because a philosophical contribution depends on the author's understanding and engagement.
Development of the parallel: On a practitioner-centred view, the parallel holds. Just as an artwork is partly constituted by the artist's expressive act, a philosophical argument is partly constituted by the philosopher's act of understanding, engaging, perhaps asserting. The text is evidence of this act, and when we evaluate the text, we're evaluating the act behind it. History of philosophy is part of philosophy because ideas are bound to minds, the way art history is bound to artists.
The disanalogy: But philosophical arguments are more like proofs than paintings. A proof's validity is fixed by its logical structure; nobody consults the mathematician to check. Similarly, an argument's illumination — the dependence relations it reveals, the pressure it puts on a familiar picture — is a property of the argument itself. Twin Earth puts pressure on internalism because of the scenario's internal logic. Any reader who follows the argument grasps why meaning can't be fixed by psychology alone. The philosophical contribution is what the argument does, not who produced it.
Hmm — but am I being too quick here? Let me consider whether someone could resist the disanalogy.
Objection: "You're comparing philosophy to proofs, but philosophy isn't like proofs either. Proofs are formally valid or invalid — there's a clear standard. Philosophy is more like art — there's interpretation, there's judgment, there's no algorithmic way to check whether a philosophical argument is good. So maybe philosophy is on the art side of the comparison, not the proof side."
Response: The standards for evaluating philosophy are different from formal validity, but they're still properties of the argument. Lipton's distinction between likeliness and loveliness concerns the explanation, not the explainer. Williamson's desiderata for theoretical virtue — simplicity, strength, non-ad-hocness — are properties of the theory. Bengson's tri-level method evaluates accommodation, integration, and theoretical virtue — all of which concern what's on the page. These standards are more complex than logical validity, but they're equally independent of who produced the argument. Blind review operates on this assumption.
Actually — this is where the evaluative framework enters! And it enters NATURALLY. The disanalogy (philosophy is like proofs, not paintings) invites the objection that philosophy is more like art than proofs. The response is: even though philosophical evaluation is more complex than checking proofs, the standards are still properties of the argument. The evaluative framework (Lipton, Williamson, Bengson) is evidence for this.
So the section's arc would be:
1. The comparison: art requires an artist. Does philosophy require a philosopher in the same way?
2. The practitioner-centred view: yes — understanding, engagement, assertion, history of philosophy.
3. The disanalogy: philosophy is more like proofs than paintings. The argument's properties are what matter.
4. Anticipated objection: philosophy isn't like proofs either — evaluation is complex, judgmental, interpretive.
5. Response: yes, but the evaluative standards are still properties of the argument. Lipton (illumination), Williamson (theoretical virtue), Bengson (accommodation, integration, virtue). These evaluate the argument, not the arguer.
6. Evidence: blind review. The discipline operates on the assumption that the text is what's evaluated.
7. Transition: what follows when an LLM is trained on a corpus shaped by these standards.
This is clean. The art comparison structures the section. The evaluative framework enters as a response to a natural objection. Every piece does work.
But let me also think about potential problems with this structure.
Problem 1: Is the art comparison too much throat-clearing before getting to philosophy? The section is called "The Challenge from Authorship," not "Art vs Philosophy." If the art comparison takes two paragraphs, that's two paragraphs before the reader gets to the philosophical content.
Counter: the comparison IS the philosophical content. The challenge from authorship IS the claim that philosophy is like art — that it requires a practitioner. The disanalogy IS the response — showing that philosophy is NOT like art in the relevant respect. The art comparison isn't a detour; it's the mechanism by which the section makes its argument.
Problem 2: The paper can't take a position on whether AI can produce art. That's a separate debate (and one Nick has written about — the generative aesthetics paper). If the paper discusses art at length, readers might expect a position on AI art.
Counter: the paper doesn't need to say whether art requires an artist. It only needs to say that WHETHER OR NOT art requires an artist, the reasons it might are different from anything that applies to philosophy. The paper takes no position on art — it uses the comparison to illuminate philosophy. One sentence could make this explicit: "Whether art in fact requires an artist is not our concern here. What matters is that the reasons it might — constitutive dependence on the creator's intentions, expression, and engagement — do not transfer to philosophy, where the evaluative standards bear on the argument's properties."
Problem 3: The "proofs not paintings" line might overstate the case. Some philosophy IS more like painting — the later Wittgenstein, Nietzsche, Derrida. These are writers whose philosophical work is partly constituted by their style, their persona, their relation to the text. The paper has already set aside practitioner-centred conceptions (Hadot, Wittgenstein as therapy, Merleau-Ponty). But the "proofs not paintings" comparison might re-open this issue.
Counter: the paper is explicitly about 21st-century analytic philosophy. It's already said: "More common in 21st Century analytic philosophy is what we might think of as an output based approach." The comparison to proofs applies to the analytic tradition the paper is operating within. It doesn't need to apply to all philosophy.
Problem 4: Does the comparison make the challenge too easy to dismiss? If the challenge is "philosophy is like art," and the response is "no it's like proofs," the challenge might feel like it was set up to fail — a straw analogy knocked down by a better one.
This is a real worry. The challenge needs to have force. The reader needs to feel the pull of the practitioner-centred view before seeing the disanalogy. If the comparison is too lopsided — art obviously different from philosophy, so the challenge is obviously wrong — the section feels like it's tilting at windmills.
How to give the challenge force: develop the practitioner-centred view with enough philosophical substance that the reader feels its pull. The phenomenology of reading (we feel like we're engaging with a mind), the history of philosophy point (ideas seem bound to minds), the understanding thread (a good text seems good because its author understood the problem). These need to be developed, not just stated. The reader should think: "yes, philosophy DOES feel like it requires a philosopher" before the disanalogy appears.
Also: the disanalogy isn't a knockdown. It's a reframing. The paper argues that the text-based conception is "robust enough" (Enrico's words), not that the practitioner-centred view is absurd. The comparison shows that philosophy CAN be evaluated without the philosopher, not that the philosopher is irrelevant to every aspect of philosophical practice. The philosopher may be important to the process, to the conversation, to the accountability — but the QUALITY of the argument is assessable from the text.
OK, I think the comparison works. Let me now go back to the other points.
Point 3: Whether the assertion stuff is worth pursuing.
Now that I've thought about the art comparison, I think assertion can be handled lightly. Within the practitioner-centred view, assertion is one thread: "philosophy seems to require someone who asserts, commits, is accountable." Within the disanalogy, the response is brief: "but the argument's quality doesn't depend on anyone's assertoric commitment — blind review evaluates arguments without knowing who asserted them." Butlin and Viebahn could go in a footnote: "For a worked-out argument that LLMs cannot assert, see Butlin and Viebahn (forthcoming). Even if this is right, philosophical quality is separable from assertoric force, as blind review demonstrates."
This uses the paper without developing its argument. It's one sentence in the body and one footnote. It doesn't lead the section off the beaten path.
But — should assertion be in the section AT ALL? Let me think about this.
Arguments for including assertion (lightly):
- It's a genuinely different thread from understanding. The challenge from authorship has several sources — understanding, history of philosophy, assertion. If the section only addresses understanding, readers who feel the challenge for assertion-based reasons will feel the response is incomplete.
- Butlin and Viebahn's paper gives the thread scholarly backing. It's not just an intuition.
- The response (quality ≠ assertion) is clean and connects to blind review, which the section already uses.
Arguments against:
- It's a digression. The section's work is done once it shows the disanalogy between art and philosophy and develops the evaluative framework. Assertion is a separate issue that could lead into speech act theory, sanctionability, proxy assertion — all of which are tangential.
- The paper already has a lot to do. Adding assertion, even lightly, adds another thread the reader has to track.
- The assertion challenge is arguably NOT part of the authorship challenge. It's more about whether LLM outputs have a certain illocutionary status, which is a pragmatics question, not a question about whether philosophy requires a philosopher.
Hmm. I think there's a middle path. The practitioner-centred view naturally involves assertion — when you describe the view, you say something like "the philosopher understands, engages, asserts, is accountable." This is describing the view, not developing a separate thread. Then the disanalogy handles all of these together: "but the evaluative standards concern the argument's properties, not the philosopher's engagement." Assertion doesn't need its own treatment; it's part of the practitioner-centred package that the disanalogy responds to.
On this approach, assertion is mentioned but not developed. It's one of the things the practitioner-centred view says philosophy requires, alongside understanding and engagement. The response (the disanalogy + evaluative framework + blind review) handles the whole package. No separate thread, no digression into speech act theory, no Butlin and Viebahn in the body text. Maybe a footnote citing them, maybe not.
Point 4: Whether the Stage 2 idea was really good.
Let me think honestly about this. The Stage 2 idea was: first argue that the contribution is in the argument (Stage 1: Putnam, Dellsén), then argue that the evaluation is also argument-internal (Stage 2: Lipton, Williamson, Bengson). Two stages, each making a different claim.
Is there actually a distinction here? If the philosophical contribution IS the argument — is what the argument does — then evaluating the contribution IS evaluating the argument. You don't need a separate argument for each. "The contribution is in the text" and "the evaluation is of the text" are the same claim viewed from different angles.
In the art comparison structure, this becomes even clearer. The disanalogy says: philosophy evaluates the argument, not the arguer. The evaluative framework says: here's what that evaluation consists in. These aren't two stages — they're one claim and its elaboration. The framework isn't a SECOND argument; it's the CONTENT of the first argument.
The Stage 2 idea was trying to give the framework a reason to exist within the section — to answer the question "why are we reading about Lipton, Williamson, and Bengson?" But the art comparison structure provides a better reason: the disanalogy says philosophy evaluates arguments, and the objection is "but philosophical evaluation is complex, not like checking proofs." The framework responds: yes, it's complex, but the complexity is still argument-internal. Illumination, theoretical virtue, accommodation — these are all properties of the argument.
So I think the Stage 2 idea was solving the right problem (connecting the framework to the challenge) with the wrong solution (an artificial two-stage structure). The art comparison structure solves the problem more naturally.
That said, there was something in the Stage 2 idea that's worth preserving: the thought that the STANDARDS need to be argument-internal, not just the CONTRIBUTION. You can imagine someone accepting that the contribution is the argument but denying that evaluation is argument-internal: "Sure, the argument is what matters. But to evaluate the argument, you need to know who produced it — because the quality of the thinking behind it is part of what makes the argument good." This objection would be: the contribution is in the text, but the evaluation isn't.
The evaluative framework responds to this: no, the evaluation IS in the text, because the standards (Lipton, Williamson, Bengson) concern properties of the argument, and blind review shows that the discipline operates on this assumption. But this doesn't need to be a separate "stage" — it's just the framework responding to a natural objection.
Actually, let me reconsider something. Is there a version of the two-stage idea that works within the art comparison?
In art, there's a distinction between what the work IS and how we evaluate it. Duchamp's urinal IS art because of the artist's intention (constitutive relationship). And we EVALUATE art partly by reference to the artist (expression, context, innovation). These are different claims. In philosophy, both claims fail: the argument IS what it is regardless of who produced it (no constitutive relationship), AND we evaluate it based on its properties (no reference to the philosopher needed). The disanalogy works at both levels.
But this is implicit in the comparison — you don't need to make it a two-stage argument. The comparison just says: in art, authorship might matter both constitutively and evaluatively. In philosophy, it matters for neither. The evaluative framework shows why.
OK. I'm now fairly confident the Stage 2 idea was overcomplicated and the art comparison handles it more naturally.
Now let me go back to point 2: the word "textual."
The session note says "textually" is banned, and Nick confirms "textual" is too. The thesis the paper needs to express is something like: philosophical contributions are things arguments do; philosophical evaluation bears on what the argument says and how it argues for it; the philosophical work is on the page.
Formulations that avoid "textual":
- "The contribution is what the argument does"
- "The evaluative standards bear on the argument's properties"
- "The philosophical work is on the page"
- "Philosophy evaluates arguments, not arguers"
- "The discipline evaluates what a paper achieves"
- "The standards concern the argument — its illumination, its theoretical virtue, its accommodation of data"
These all avoid "textual" / "text-based" / "textually." The thesis can be expressed through what the text DOES (arguments, explanations, scenarios) rather than through the word "text."
But there's a problem: the paper uses "text-based conception" and "text-focused" as technical terms (Enrico's language in the transcript: "two accounts of philosophy, one person-based, the other text-based"). Is "text-based" banned too, or just "textual" / "textually"?
Looking at the session note again: "textually (hated; find other formulations for the thesis that philosophy works through articulated text rather than embodied encounter)." And then Nick says "'textual' is a banned word." So both "textual" and "textually" are out. "Text-based" might be different — it's Enrico's terminology, not an adverb. But to be safe, I should offer alternatives.
Instead of "text-based conception":
- "argument-based conception"
- "output-based conception" (the introduction already uses this: "what we might think of as an output based approach")
- "product-based conception"
Actually, the introduction already provides the alternative: "output based approach." This avoids "textual" and captures the same idea. The paper could use "output-based" and "practitioner-based" as its two conceptions.
Hmm, but "output-based" is a bit clinical. "The philosophical contribution is what the argument does" is warmer than "the philosophical contribution is the output." Let me leave this as an open question for Nick.
Now let me think about how all of these pieces fit together into the section.
I'm imagining something like:
¶1: The comparison. Art plausibly requires an artist — the artwork's identity may depend on the creator's intentions and expression. (Brief — no more than 3-4 sentences on the art side.) Does philosophy require a philosopher in the same way? There's an intuitive pull: a good philosophical argument seems good because its author understood the problem, could see why a distinction mattered.
¶2: Developing the practitioner-centred view. When we read philosophical texts, we read them as evidence of thought. We attribute the quality of the argument to the quality of the thinking behind it. History of philosophy is part of philosophy in a way that history of science is not — perhaps because philosophical ideas are bound to the minds that produced them. Understanding, engagement, assertion, accountability — these seem constitutive of philosophical practice. If philosophy requires a practitioner, an LLM cannot do philosophy.
¶3: The disanalogy. But philosophical arguments are more like proofs than paintings. [Nick's line.] A proof's validity depends on its logical structure, not on who produced it. Similarly, the force of Twin Earth — the pressure it puts on internalism about meaning — depends on the scenario's internal logic. Any reader who follows the argument grasps the dependence relation. Putnam's contribution is what the argument does, not what Putnam privately understood.
Now — actually, should the Putnam paragraph come HERE, or should Putnam be a case study after the disanalogy is stated? Let me think...
Option A: Putnam IS the disanalogy. The disanalogy is demonstrated through Putnam, not stated abstractly. "Consider what Putnam accomplished with Twin Earth. [Works through the case.] The philosophical contribution — the pressure on internalism — is something the argument does. Any reader who follows the scenario grasps it. This is how philosophical arguments differ from artworks: the argument's force depends on its internal logic, not on the creator's relation to it."
Option B: The disanalogy is stated first ("philosophical arguments are more like proofs than paintings"), then Putnam is evidence. "A proof's validity depends on its structure, not on who produced it. The same holds for philosophical arguments. Consider Twin Earth: [works through the case]."
I think Option A is more concrete and engaging. Starting with Putnam and drawing the disanalogy FROM the case is philosophically more effective than stating the disanalogy and then illustrating it. The reader sees the point through the example rather than being told the point and then given an example.
But wait — the current draft's %%comment%% says: "Why are you suddenly talking about Putnam?" This suggests the transition to Putnam needs work regardless. If the art comparison structures the section, the transition to Putnam becomes natural: "But does the philosophical contribution really depend on the philosopher's understanding? Consider what Putnam accomplished with Twin Earth." The art comparison sets up the question, and Putnam answers it.
OK, continuing the sketch:
¶3 (revised): The turn. "But does the quality of a philosophical argument actually depend on who produced it, in the way that an artwork's significance might depend on its creator? Consider what Putnam accomplished with Twin Earth..." [Works through the case.] "The contribution — the pressure on internalism — is something the argument does. A reader who follows the scenario grasps the dependence relation whether she knows who Putnam is or not. Philosophical arguments are more like proofs than paintings: their force depends on their internal logic, not on the creator's relation to them."
¶4: Dellsén makes this systematic. Progress = enabling understanding through publicly available ideas. What's publicly available is the argument. History of philosophy reinterpreted: we study philosophy's history because the ARGUMENTS remain productive, not because ideas are bound to minds.
¶5: Anticipated objection. "But philosophy isn't like proofs either — there's no algorithm for checking whether a philosophical argument is good. Philosophical evaluation involves judgment, interpretation, weighing of competing considerations. If evaluation is more complex than checking proofs, maybe it does depend on knowing who produced the argument."
¶6: Response. The evaluative standards the discipline employs bear on the argument's properties, however complex they are. Lipton's distinction between likeliness and loveliness concerns the explanation, not the explainer. [Block quote.] The dormative virtue is a bad explanation regardless of who offers it.
¶7: Further development. Williamson's desiderata — simplicity, strength, non-ad-hocness. [These are properties of the theory.] Bengson's tri-level method — accommodation, integration, theoretical virtue. [All concern what's on the page.]
¶8: Evidence. Blind review. The discipline's evaluative practice strips away exactly the things the practitioner-centred view says matter — identity, process, commitment — and the evaluation still works. If the philosopher's understanding were relevant to the argument's quality in the way that the artist's intentions are relevant to the artwork's identity, blind review would be defective. It isn't.
¶9: Deep Blue? Or cut?
¶10: Transition. We have argued that philosophy evaluates arguments, not arguers, and that the evaluative standards concern the argument's properties. The question, then, is what follows when an LLM is trained on a corpus shaped by these standards.
That's roughly 8-10 paragraphs. The challenge gets ¶1-2 (comparison + practitioner-centred view). The response gets ¶3-8 (Putnam + Dellsén + objection + framework + blind review). The transition is ¶9-10.
This is balanced. The challenge has two paragraphs — enough to feel its force, not so much that it dominates. The response is built around the disanalogy (¶3-4) and the evaluative framework (¶5-8), connected by the "proofs not paintings" objection and response.
One more question: does Deep Blue (¶9) still fit? In the current draft, Deep Blue makes the point that quality doesn't depend on the manner of production — a move that wins material is a move that wins material, whether found by pattern recognition or brute-force search. This is a product/process point, similar to the proofs/paintings point. Is it redundant?
Maybe. The proofs/paintings comparison already makes the product/process point. Deep Blue adds a concrete case from another domain (chess), but the case is structurally the same as what the section has already established. It might be a digression at this point. On the other hand, it's a vivid example that readers remember.
I think Deep Blue could stay or go. If the section is already running long, cut it. If there's space, keep it as a brief paragraph — one or two sentences, not a full case study.
Actually, looking at the current Deep Blue paragraph (line 30), it also introduces Gaut — "what Gaut calls 'the epitome of an uncreative way to play chess.'" And there's a %%comment%% flagging the Gaut reference as unverified. If the Gaut reference is problematic, there's an additional reason to cut or slim the paragraph.
OK, I think I've explored this enough. Let me now think about the points I was less thorough on before.
Point 1 (Floridi): Nick said "let's forget about the floridi thing for the moment. you aren't really understanding i don't think." Fair. I'll drop it. But I want to register what I think Nick might mean — that my characterisation of Floridi's challenge ("LLMs can't reason, they pattern-match") was too crude, and I never really improved it enough to satisfy him. My later attempt ("whether philosophical quality can survive the transition from genuine inference to statistical reproduction") was better but Nick might still think I'm not grasping the subtlety of what Floridi's challenge actually is and what the section's response actually does. I should come back to this later when Nick wants to discuss it, and should re-read the Floridi paper properly rather than working from the section's summary of it.
Point 2 ("textual" is banned): Covered above. Need to find alternative formulations. "Output-based" (already in the introduction), "argument-based," or just describing what the evaluation bears on (the argument's properties, what the paper achieves, what's on the page) without using an adjective.
Point 3 (assertion): I've addressed this — within the art comparison structure, assertion is part of the practitioner-centred package, handled by the disanalogy and by blind review. No need for a standalone treatment. Butlin and Viebahn could go in a footnote or be dropped.
Point 4 (Stage 2): I've argued it was overcomplicated. The art comparison structure handles the connection between the challenge and the framework more naturally. The framework enters as a response to the objection "philosophy isn't like proofs either" rather than as a separate "stage" of the argument.
Let me now think about what the section would look like as a whole, with specific attention to how the art comparison handles the threads of the challenge.
Thread A (understanding): The practitioner-centred view says philosophy requires understanding. The disanalogy says: the argument's illumination is a property of the argument, demonstrated by Putnam. Dellsén: progress = enabling understanding through publicly available ideas.
Thread B (assertion/accountability): Part of the practitioner-centred package. The disanalogy handles it: blind review evaluates arguments without knowing who asserted them. (Butlin and Viebahn in a footnote if wanted.)
Thread C (history of philosophy): The practitioner-centred view says ideas are bound to minds. The response (Dellsén paragraph) reinterprets: we study history of philosophy because the arguments remain productive.
All three threads are addressed within the art comparison structure. None needs a standalone treatment. The challenge is developed (¶1-2) and the response handles all threads through the disanalogy (¶3-4) and the evaluative framework (¶5-8).
One final thing I want to think about: does the art comparison create problems for the introduction?
The current introduction distinguishes practitioner-focused and output-based conceptions (¶2-3). If Section 1 introduces the practitioner-centred view through the art comparison, does the introduction need to change?
I don't think so. The introduction notes that practitioner-focused conceptions exist and sets them aside: "On some approaches, philosophy requires being a certain kind of subject..." The introduction doesn't develop the practitioner view — it mentions it and moves on. Section 1 then takes up the strongest version of the practitioner view (the authorship challenge) and responds to it through the art comparison. The introduction and Section 1 have different jobs: the introduction notes the landscape; Section 1 argues.
But the introduction's roadmap (currently a placeholder) needs to anticipate the challenge structure. Something like: "We argue otherwise. The remainder of this paper addresses four challenges to the claim that LLMs can produce good philosophy..." This would name the four challenges and sketch the paper's response.
One complication: the current introduction already mentions Dellsén, Bengson, and Williamson (¶3). If Section 1 also develops these as part of the evaluative framework, there's overlap. But the introduction mentions them in passing ("Dellsén et al. argue that philosophical progress consists in..."), while Section 1 develops them in detail. This should be fine — the introduction previews, Section 1 elaborates.
I think I've explored the five points thoroughly enough now. Let me write up the final answer.
## Point 1 (Floridi)
Dropping it as requested. I'll come back to it when you're ready, and I'll read the Floridi paper properly rather than working from summaries.
## Point 2 ("textual" is banned)
Noted. The session file already flags "textually" as hated; I'll extend the ban to "textual" and "text-based" as describing the paper's thesis. Alternative formulations:
- "The contribution is what the argument does"
- "The evaluative standards bear on the argument's properties"
- "The philosophical work is on the page"
- "Philosophy evaluates arguments, not arguers"
The introduction already uses "output based approach" — that phrase could serve as the technical term for the conception the paper defends, paired with "practitioner-centred" or "person-based" for the one it resists.
## Point 3 (whether assertion is worth pursuing)
Within the art comparison structure, assertion doesn't need a standalone treatment. It's part of the practitioner-centred package: "the philosopher understands, engages, asserts, is accountable." The disanalogy handles the whole package at once — the argument's quality doesn't depend on the philosopher's assertoric commitment, just as it doesn't depend on their understanding or their accountability. Blind review is the evidence for all of these together.
Butlin and Viebahn could go in a footnote ("For a worked-out argument that LLMs cannot assert, see Butlin and Viebahn forthcoming; even if this is right, the evaluative standards we describe do not require assertoric force"). Or they could be dropped — the section's argument doesn't depend on engaging with the assertion literature. Their paper is more directly relevant to the unwritten Section 4 (prompting), because their proxy assertion discussion bears on who is responsible when a philosopher submits LLM-produced work.
My honest assessment: a footnote citing them is worth having because it shows the paper is aware of the assertion issue without being derailed by it. Anything more than a footnote risks leading the section off its path.
## Point 4 (whether Stage 2 was a good idea)
In retrospect, no. The two-stage structure was solving the right problem (connecting the framework to the challenge) with an artificial mechanism. If the philosophical contribution IS the argument, then evaluating the contribution IS evaluating the argument. You don't need a separate argument for each — they're the same claim described from different angles.
The art comparison structure dissolves the need for Stage 2 entirely. The disanalogy says: philosophy evaluates arguments, not arguers. The framework says: here's what that evaluation consists in. The framework is the CONTENT of the disanalogy, not a second stage of the argument.
There was one thing in the Stage 2 idea worth preserving: the thought that someone could accept the disanalogy but object that EVALUATING arguments still requires knowing who produced them. "Sure, the argument is what matters. But to judge whether it's any good, you need to know whether the person behind it really understood the problem." This is a natural objection, and the framework (Lipton, Williamson, Bengson) responds to it. But it enters as an objection within the argument, not as a separate stage.
## Point 5 (the art/philosophy comparison as structuring device)
I think your intuition is right, and I've tried to push on it hard enough to be confident I'm not just agreeing.
Why it works:
The comparison makes the challenge concrete. "Philosophy requires a philosopher" is abstract. "Does philosophy need a philosopher the way art needs an artist?" is graspable. Everyone has purchase on why art might need an artist — expression, intention, the relationship between the work and its creator. And then the question becomes: is philosophy like that?
The practitioner-centred view says yes. Understanding, engagement, assertion, accountability — these are constitutive of the philosophical contribution, the way the artist's expression is constitutive of the artwork. History of philosophy matters because ideas are bound to minds, just as art history matters because works are bound to artists.
The disanalogy does the philosophical work. Philosophical arguments are more like proofs than paintings. A proof's validity depends on its structure, not on who produced it. Twin Earth puts pressure on internalism because of the scenario's internal logic. The argument's force depends on what it does, not on the philosopher-argument relationship.
The evaluative framework enters in response to a natural objection: "But philosophy isn't like proofs either — evaluation is complex, judgmental, interpretive." Response: yes, but the evaluative standards (Lipton's illumination, Williamson's theoretical virtue, Bengson's accommodation and integration) are still properties of the argument. However complex philosophical evaluation is, it concerns what the argument achieves, not who produced it. Blind review is the institutional evidence.
What I'd sketch as the section's arc:
¶1-2: The comparison + practitioner-centred view developed. Art may require an artist. Does philosophy require a philosopher in the same way? The practitioner-centred view says yes: understanding, engagement, history of philosophy, assertion/accountability. This is the challenge, emerging from the comparison rather than stated as an abstract objection.
¶3: The turn. Putnam. "But does the quality of a philosophical argument actually depend on who produced it, in the way that an artwork's significance might depend on its creator?" Work through Twin Earth. The contribution is what the argument does. Philosophical arguments are more like proofs than paintings.
¶4: Dellsén. Progress = enabling understanding through publicly available ideas. What's publicly available is the argument. History of philosophy reinterpreted: arguments remain productive, not bound to minds.
¶5-7: Anticipated objection + framework. "But philosophy isn't like proofs either." Response: Lipton (illumination vs accommodation), Williamson (theoretical virtue, overfitting), Bengson (tri-level method). These are all properties of the argument, however complex.
¶8: Blind review. Evidence that the discipline operates on this assumption. Strips away identity, commitment, process — evaluation still works.
¶9: "In sum" / transition. We have argued that philosophy evaluates arguments, not arguers. What follows when an LLM is trained on a corpus shaped by these evaluative standards?
This gives the challenge two paragraphs (with the practitioner-centred view developed through the art comparison) and the response six paragraphs (Putnam → Dellsén → objection → framework → blind review → transition). The evaluative framework is connected to the challenge through the "proofs not paintings" objection and response.
Where it could go wrong:
If the art comparison takes too long (more than a paragraph or two on the art side), the section feels like it's about art, not philosophy. The fix: keep the art side brief. One paragraph on why art might require an artist, then immediately pivot to philosophy.
If the disanalogy is too quick, the challenge feels like a straw comparison knocked down easily. The fix: develop the practitioner-centred view with enough force that the reader feels its pull before the disanalogy appears. The phenomenology of reading (we feel like we're engaging with a mind), the history of philosophy observation, the understanding intuition — these need enough development that the reader takes them seriously.
If the "proofs not paintings" line is taken as the paper's considered position, someone might object that philosophy is nothing like proofs. The line is illustrative, not literal. Philosophical evaluation is more complex than checking proofs — that's what the evaluative framework is there to show. The analogy is: proofs, like philosophical arguments, have properties that are assessable independently of who produced them. Not: philosophical arguments are formal proofs.
Deep Blue: probably expendable if the section is already running long. The proofs/paintings comparison already makes the product/process point. Deep Blue adds a vivid case but it's structurally redundant. Cut it or slim it to one sentence.
*Il confronto regge proprio dove serve — non perché l'arte e la filosofia si somiglino, ma perché il punto in cui divergono è il punto che la sezione deve stabilire.*
### User
still A LOT of work to be done before we are ready to iterate on this section. i would say we are about 10% of the there at the moment
/contemplate for all of the below
/remember "the argument's quality doesn't depend on anyone's assertoric commitment, in a way that the artwork's identity might depend on the artist's intentions." this is nicely put (although I am a little unsure about framing things in terms of identity, that seems to be your idea)
"Actually, looking at the current Deep Blue paragraph (line 30), it also introduces Gaut — "what Gaut calls 'the epitome of an uncreative way to play chess.'" And there's a flagging the Gaut reference as unverified. If the Gaut reference is problematic, there's an additional reason to cut or slim the paragraph." a missing reference that can easily be found does not count as point towards slimming the paragraph down, obviously.
happy to have B and V in a footnote, although your suggested framing "For a worked-out argument that LLMs cannot assert, see Butlin and Viebahn forthcoming; even if this is right, the evaluative standards we describe do not require assertoric force" is written in a terrible style.
"Sure, the argument is what matters. But to judge whether it's any good, you need to know whether the person behind it really understood the problem." why do you think this is a 'natural objection'? genuine question.
can we add the verb 'matter' to the banned word list (noun is obviously fine)
"But philosophy isn't like proofs either — evaluation is complex, judgmental, interpretive." where does this framing of the objection come from? it is not very good wherever it came from. brainstorm better ways of making this move (or better moves to make in lieu of it)
### Assistant
Let me save the remember note first.
### Assistant
Saved [[Assertoric commitment vs artistic identity]]
Now let me update the session note with the new banned word, then do the contemplation.
### Assistant
Now the contemplation.
Nick has raised several distinct points and I need to think carefully about each one.
Let me start with the most philosophically substantial question: "why do you think this is a 'natural objection'?" — referring to "Sure, the argument is what [counts]. But to judge whether it's any good, you need to know whether the person behind it really understood the problem."
Why did I think this was a natural objection? Let me trace my reasoning.
I was trying to find a way for someone to accept the disanalogy (the contribution is the argument, not the person) while still resisting the conclusion (that you can evaluate philosophy without evaluating the philosopher). The objection would be: "OK, the philosophical contribution is the argument. But evaluating the argument still requires knowing who produced it, because quality depends on the thinking behind it."
Is this actually a natural objection? Let me think about whether anyone would seriously hold this.
The claim would be: you can tell the difference between an argument produced by someone who deeply understood the problem and an argument produced by someone who was going through the motions, and this difference is relevant to the argument's philosophical quality, and you can only detect it by knowing who produced the argument.
But... is any of this true? If the argument produced by the person who understood the problem is better, it's better because of properties IN THE ARGUMENT — it handles objections more carefully, draws distinctions more precisely, illuminates more deeply. You don't need to know who produced it to see these properties. The quality shows up on the page. That's the whole point of blind review.
And if the argument produced by the person going through the motions is just as good — handles objections just as carefully, illuminates just as deeply — then it IS just as good, and it doesn't help to know that the person was going through the motions.
So the objection collapses on inspection. It assumes that quality is detectable from authorship rather than from the argument itself, which is precisely what the section argues against. It's not a natural objection — it's a confused objection that restates the challenge rather than raising a new concern.
Hmm. But then why did I generate it? I think I was looking for a reason to introduce the evaluative framework — a bridging move between "the contribution is the argument" and "here's what evaluation consists in." I needed an excuse for the framework to appear, so I invented an objection for it to respond to. That's bad philosophical reasoning — manufacturing an objection to justify exposition rather than finding the genuine intellectual pressure that requires the exposition.
OK. So if that objection isn't natural, what IS the right move between the disanalogy and the evaluative framework? Let me think about this more carefully.
The disanalogy says: philosophical arguments are more like proofs than paintings — the argument's properties are what count, not the philosopher's relation to the argument. The evaluative framework (Lipton, Williamson, Bengson) describes what those properties are. The connection is: "What are the properties that count? What does it mean for a philosophical argument to be good, independently of who produced it?"
This isn't responding to an objection. It's developing the thesis. The disanalogy makes the claim; the framework gives it substance. You don't need an objection between them — you just need a question: "What does this evaluation consist in?"
But wait — Nick is also asking about my other formulation: "But philosophy isn't like proofs either — evaluation is complex, judgmental, interpretive." Where does this come from? Is it any good?
Let me trace this. I was trying to anticipate someone who resists the proofs/paintings comparison by saying that philosophy is too unlike proofs for the comparison to work. Proofs are formally valid or invalid — there's a clear standard. Philosophy is evaluated by more complex, less algorithmic criteria. So maybe philosophy is closer to art after all (where evaluation is also complex and non-algorithmic).
Is this a real concern? Let me think...
Actually, hmm. The proofs/paintings line is an analogy, not an identification. The claim isn't that philosophical arguments ARE proofs. The claim is that, like proofs, their properties are assessable independently of who produced them. The complexity of evaluation is irrelevant to this point. Whether the evaluation is simple (checking logical validity) or complex (assessing illumination, theoretical virtue, accommodation), it's still evaluation of the ARGUMENT. Complexity doesn't reintroduce the philosopher.
So the objection "philosophy isn't like proofs because evaluation is complex" is actually a non sequitur. It confuses the complexity of evaluation with the location of evaluation. Simple evaluation and complex evaluation can both be argument-directed. The proofs analogy isn't about simplicity of evaluation — it's about independence from the producer.
This means the "anticipated objection" I generated was bad for two reasons: (1) it confuses complexity with location, and (2) it doesn't put genuine pressure on the argument. Nobody who understood the proofs/paintings comparison would object "but philosophy is more complex than proofs" — because that misunderstands what the comparison is doing.
So: both of my bridging moves were bad. The "natural objection" was confused, and the "philosophy isn't like proofs" objection was a non sequitur.
What's the right way to get from the disanalogy to the evaluative framework?
Let me think about what the section actually needs at this point.
After the disanalogy (proofs not paintings) and after Putnam/Dellsén, the reader accepts (or is being asked to accept) that philosophy evaluates arguments independently of who produced them. The evaluative framework then says: here's what that evaluation looks like in practice.
The connection could be simple exposition: "The evaluative standards the discipline employs all bear on the argument's properties." Then develop Lipton, Williamson, Bengson as instances of this.
Or the connection could be a question that the reader is naturally asking: "What does it look like to evaluate a philosophical argument on its own terms? What are the properties that distinguish a good argument from a bad one?"
Or the connection could come from a real philosophical pressure — not a manufactured objection, but a genuine issue that the evaluative framework addresses.
Let me think about what genuine pressure exists at this point in the argument.
The disanalogy says: the argument's quality is independent of the philosopher. Putnam's Twin Earth works because of its internal logic. Dellsén says progress happens through publicly available ideas.
What hasn't been established yet? The nature of the evaluative standards. The reader knows that philosophy evaluates arguments, but she might not know WHAT the standards are. And the section needs to establish what they are because Section 2 will use them (the corpus is filtered by these standards). The framework isn't responding to an objection — it's developing the claim. The claim "philosophy evaluates arguments" is incomplete without an account of what the evaluation looks for.
So the right move might be: "We have argued that a philosophical contribution is what the argument does — it constructs scenarios, reveals dependence relations, enables understanding. But not every argument that does these things does them equally well. What distinguishes a good philosophical argument from a merely adequate one?" Then: Lipton (illumination vs mere accommodation), Williamson (theoretical virtue vs overfitting), Bengson (systematic evaluation).
This is just a question, not an objection. The framework enters because the thesis ("philosophy evaluates arguments") needs content — the reader needs to know what the evaluation consists in. The question arises naturally from the thesis, not from a manufactured objection.
But actually, even this might be overcomplicating things. Maybe the evaluative framework doesn't need a special introduction at all. Maybe after Dellsén, the section just continues: "Dellsén's account of philosophical progress raises a further question. Not every account of a phenomenon's dependence relations is equally illuminating. If philosophical progress consists in enabling understanding, we need a way to distinguish views that genuinely reveal how things depend on one another from views that merely accommodate the data."
Wait — that IS the current text. Line 16 of the current Section 1: "Not every account of a phenomenon's dependence relations is equally illuminating, however. If philosophical progress consists in enabling understanding, we need a way to distinguish views that genuinely reveal how things depend on one another from views that merely accommodate the data without explaining anything."
So the current transition from Dellsén to Lipton is already a natural question, not an objection! The current draft gets this right. The problem wasn't the transition — it was that the challenge was underdeveloped and the framework felt disconnected because the reader had forgotten what challenge it was responding to.
If the challenge is properly developed through the art comparison (2 paragraphs), and the disanalogy is established through Putnam (1 paragraph), and Dellsén makes it systematic (1 paragraph), then the existing transition to Lipton ("Not every account... is equally illuminating") works. The reader knows the challenge, knows the response is "philosophy evaluates arguments," and the framework follows naturally as "here's what that evaluation looks for."
The fix isn't inventing a new transition. The fix is developing the challenge so that the framework feels like it's answering something, and then — as I suggested earlier — adding light connections within the framework that remind the reader this is about argument properties, not philosopher properties.
OK. So both of my "anticipated objection" ideas can be dropped. The transition from Dellsén to Lipton is already in the draft and it works. The art comparison + developed challenge + disanalogy + Dellsén → framework is the structure, and the current transition (line 16) is the joint.
Now let me think about Nick's other points.
"a missing reference that can easily be found does not count as point towards slimming the paragraph down, obviously."
He's right. The %%REFERENCE: Gaut 2010, fn. 23 — UNVERIFIED%% flag is a source-verification task, not an argument for cutting the paragraph. The reference either exists or it doesn't. If it exists, cite it properly. If it doesn't, find the right reference or drop the specific attribution. Either way, it's a small fix, not a reason to cut Deep Blue.
My suggestion to "cut or slim" Deep Blue was partly motivated by the unverified reference, which was lazy reasoning. The question of whether Deep Blue stays should be based on whether it does argumentative work in the section, not on whether one citation is hard to find.
Does Deep Blue do argumentative work? In the current draft (line 30): "This is not unique to philosophy. Deep Blue... The moves it produced were strong chess all the same. The quality of a move does not depend on the manner of its selection." This is making the product/process point — quality is about the product (the chess move, the philosophical argument), not the process (pattern recognition, brute-force search, abductive reasoning, next-token prediction).
Does the proofs/paintings comparison already make this point? Partially. "Philosophical arguments are more like proofs than paintings" says that the argument's properties are what count, not the philosopher's relation to the argument. But it doesn't specifically address PROCESS — it addresses the CREATOR. The Deep Blue example adds something: it's about a different PROCESS producing the same quality of product. Deep Blue plays great chess through brute-force search rather than human-style pattern recognition. The process is different but the product is equivalent. This is relevant because the paper's larger argument is about LLMs (a different process) producing philosophical work of equivalent quality.
So Deep Blue does something the proofs/paintings comparison doesn't: it introduces a non-human system producing high-quality work through a non-human process. The proofs comparison is about independence from the PERSON; Deep Blue is about independence from the PROCESS. These are different points. If the section only makes the first point, someone could accept that the argument is what counts but still insist that only human processes can produce good arguments. Deep Blue (and the chess analogy more broadly) pushes against this: good chess is good chess regardless of how it was found.
But — is the product/process point really needed in Section 1? Or does it belong in Section 2 (the challenge from abduction), where the process question is the explicit topic?
Section 2 is about whether philosophical quality can survive when the process is next-token prediction rather than genuine inference. That IS the product/process question. Maybe Deep Blue fits better there.
But then again, the product/process point also bears on the authorship challenge. The authorship challenge says: philosophy requires a philosopher. Part of what this means is: philosophy requires a human PROCESS (thinking, understanding, engaging). Deep Blue says: quality doesn't require a human process. This is responsive to the authorship challenge, not just to the abduction challenge.
I think Deep Blue could work in either section. In Section 1, it responds to the process dimension of the authorship challenge. In Section 2, it responds to the process question directly. Where it fits better depends on the section's arc. If Section 1 is structured through the art comparison, and the disanalogy is about the PERSON (creator vs argument), then Deep Blue adds a second disanalogy about the PROCESS. That might feel like a digression from the art comparison. Or it might feel like a natural extension: "not only does the quality of a philosophical argument not depend on who produced it (art vs philosophy), it doesn't depend on HOW it was produced (Deep Blue)."
Actually, looking at the current draft again, Deep Blue comes AFTER the "In sum" paragraph (line 28) and BEFORE the transition (line 32). It's positioned as a final illustration before the corpus thesis. If the section is restructured through the art comparison, Deep Blue might sit well as a brief coda: "The point extends beyond philosophy. In chess, the quality of a move does not depend on the manner of its selection..." One paragraph, then transition to the corpus thesis.
I'll flag this as genuinely open — it could stay, could move to Section 2, could stay but be repositioned. The unverified reference is a non-issue either way.
Now: "can we add the verb 'matter' to the banned word list (noun is obviously fine)"
Done already (updated the session note). Going forward, I need to find alternatives. "Matter" as a verb appears in formulations like "what matters is..." and "the author doesn't matter." Alternatives:
- "What counts is..." instead of "what matters is..."
- "The relevant consideration is..."
- "The question turns on..."
- "The evaluative standards concern..."
- "The quality of the argument is independent of..."
- "The evaluation bears on..."
I should also retroactively check my recent responses — I've been using "matter" as a verb frequently.
Now: the footnote for Butlin and Viebahn. Nick says my suggested framing is "written in a terrible style." Fair. "For a worked-out argument that LLMs cannot assert, see Butlin and Viebahn forthcoming; even if this is right, the evaluative standards we describe do not require assertoric force" — this is clunky, hedging, and reads like a parenthetical excuse rather than a confident statement.
What would a better footnote look like? The paper needs to: (a) cite Butlin and Viebahn as relevant to the assertion question, and (b) note that the paper's argument doesn't depend on whether LLMs can assert.
Draft: "Butlin and Viebahn (forthcoming) argue that current LLMs cannot assert, because assertion requires sanctionability and current systems have no interests that sanctions could harm. Whether LLMs can assert is a separate question from whether they can produce good philosophy; the evaluative standards we describe here bear on the argument itself, not on whether anyone is asserting it."
Hmm, this is still clunky. Let me try shorter:
"On whether LLMs can assert, see Butlin and Viebahn (forthcoming), who argue they cannot. Our argument is about the quality of what is produced, not about the speech-act status of its production."
Better? Maybe. I should probably not try to write the footnote in Nick's voice here — it needs Nick's ear, not mine. I'll flag what the footnote needs to do and leave the wording.
Now: "where does this framing of the objection come from? it is not very good wherever it came from."
He's asking about "But philosophy isn't like proofs either — evaluation is complex, judgmental, interpretive." This was entirely my generation. It came from me trying to anticipate a resistance to the proofs/paintings comparison. And as I've now worked out above, it's a bad objection because it confuses the complexity of evaluation with the location of evaluation. Complex evaluation can still be argument-directed.
Nick wants me to brainstorm better ways of making the move between the disanalogy and the evaluative framework. But as I've worked out above, I think the right answer might be: there IS no move needed. The current transition in the draft (line 16: "Not every account of a phenomenon's dependence relations is equally illuminating...") already works. It's a natural question, not an objection.
But let me brainstorm anyway, in case the current transition doesn't feel right within the restructured section.
Way 1: No move needed. After Dellsén, the existing transition ("Not every account... is equally illuminating") introduces Lipton naturally. The framework is just the next stage of the argument.
Way 2: A question from the reader. "What does it look like to evaluate a philosophical argument on its own terms?" Then: Lipton distinguishes likeliness from loveliness. This is what evaluation looks like.
Way 3: A question from within the argument. Dellsén says progress consists in enabling understanding — but understanding comes in degrees. Some arguments generate more understanding than others. What distinguishes the ones that do? Then: Lipton's likeliness/loveliness distinction.
Way 4: Connect to the challenge. "If philosophical arguments are evaluated by their properties rather than by who produced them, the question is what those properties are. We have argued that the relevant evaluation concerns what the argument illuminates — but what is the difference between genuine illumination and the appearance of illumination?" Then: the dormative virtue.
Way 5: Connect to the art disanalogy. "In art, evaluative standards may include the artist's relation to the work — the depth of expression, the quality of intention. In philosophy, the evaluative standards concern the argument itself. Lipton distinguishes two ways an explanation might count as best..."
Hmm. Let me think about which of these is most honest — which one reflects a genuine philosophical need rather than a manufactured bridge.
I think Way 3 is most natural. Dellsén says progress = enabling understanding. But you need to distinguish genuine understanding from the appearance of understanding. That's what Lipton's likeliness/loveliness distinction does: likeliness (accommodation without illumination) looks like progress but isn't; loveliness (genuine illumination) is the real thing.
This IS already what line 16 says. "Not every account of a phenomenon's dependence relations is equally illuminating, however. If philosophical progress consists in enabling understanding, we need a way to distinguish views that genuinely reveal how things depend on one another from views that merely accommodate the data without explaining anything."
So the transition already exists and works. The problem in the current section is not the transition but the challenge — it's too thin, so by the time the reader reaches the framework, she's forgotten what question it's responding to.
Let me also think about the art comparison and the "identity" framing.
Nick's remember note flags that the "identity" framing in "the artwork's identity might depend on the artist's intentions" is my idea and he's unsure about it. Let me think about what's right here.
In philosophy of art, "identity" is a technical term. The identity of an artwork is what makes it the particular artwork it is — what distinguishes it from copies, forgeries, other works. On some views (Levinson, for instance), the artwork's identity depends partly on the art-historical context and the artist's intentions. A work created by Duchamp in 1917 is not the same artwork as an identical-looking work created by someone else in 2026, because the art-historical context is different.
But "identity" might be the wrong frame for the comparison in the paper. The paper is asking about QUALITY, not identity. The question is: does the quality of a philosophical argument depend on the philosopher, in the way that... what? The quality of an artwork depends on the artist? The identity of an artwork depends on the artist?
These are different claims:
- Quality: the painting is GOOD because of the artist's skill, vision, expression
- Identity: the object is ART (rather than not-art) because of the artist's intention
For the art/philosophy comparison, which is more relevant?
If it's identity: the question is whether a text IS philosophy when it's produced by an LLM. Is the output "philosophy" at all? On a practitioner-centred view, it isn't — philosophy is constitutively an activity of a thinker, and LLM output isn't philosophy regardless of what it looks like.
If it's quality: the question is whether a text is GOOD philosophy when produced by an LLM. The output might be philosophy, but is it any good? On the practitioner-centred view, its quality depends on the philosopher behind it.
I think the paper is primarily about quality, not identity. The paper assumes the output-based approach (it's already set aside practitioner conceptions in the introduction) and asks whether LLM-produced philosophy can be good. So the relevant comparison with art is about quality: does the quality of a philosophical argument depend on the philosopher, in the way that the quality of an artwork might depend on the artist?
But... the art side of this is weaker. The quality of an artwork doesn't obviously depend on the artist's identity in a strong sense. A painting by an unknown artist can be beautiful. What art-identity theorists are claiming is something stronger: that the work's MEANING, SIGNIFICANCE, or STATUS AS ART depends on the artist. And that's an identity claim, not a quality claim.
So there's a mismatch. The art-identity point (is it art?) doesn't map onto the philosophy-quality point (is it good philosophy?). The art-quality point (is it good art?) maps better, but it's less interesting — nobody seriously denies that the quality of a painting can be assessed from the painting.
Hmm. Unless the challenge from authorship IS partly an identity challenge: "LLM output isn't philosophy, regardless of how it looks." On the practitioner-centred view, this is exactly the claim. Philosophy is constitutively an activity of a thinker. If nobody was thinking, the output isn't philosophy, just as an accidental paint splash isn't art.
So maybe the identity framing is right after all — but only for the CHALLENGE, not for the response. The challenge says: LLM output isn't philosophy (identity claim). The response says: whether it's philosophy depends on its properties, not on who produced it, and here are the evaluative standards that the discipline uses to assess those properties.
Actually, I'm going in circles. Let me step back. Nick flagged the identity framing as my idea and said he's unsure. The honest thing is to acknowledge that I introduced the framing, that it may or may not be the right frame, and that the comparison could work without it. The comparison can be about whether the QUALITY of a philosophical argument depends on the philosopher, without needing to invoke "identity" as a technical concept.
Let me reformulate without "identity": "The quality of a philosophical argument doesn't depend on anyone's assertoric commitment, in a way that an artwork's [significance/meaning/character/value] might depend on the artist's intentions."
Which word? "Significance" is neutral. "Character" is about what the work is like. "Value" is about worth. I think "significance" works: "an artwork's significance might depend on the artist's intentions" — this captures the idea that what the artwork DOES (expresses, means, communicates) depends on the artist, without invoking the technical concept of identity.
But I'm not sure I should be finessing Nick's prose here. The note captured his reservation and the line as I wrote it. He'll adjust it when he uses it.
Let me now think about the overall state of things. Nick says we're about 10% of the way there. What remains?
What we have so far:
- The art/philosophy comparison as a structuring device (agreed in principle, not yet developed)
- "More like proofs than paintings" as a line (saved)
- Assertion handled lightly, Butlin and Viebahn in a footnote (agreed)
- The evaluative framework stays in Section 1, connected to the challenge rather than free-floating (agreed)
- The existing transition from Dellsén to Lipton (line 16) works
- "Textual" and "matter" (verb) are banned
What needs doing:
- Actually develop the art comparison into detailed paragraph sketches (how much art, what art examples, where exactly the disanalogy is drawn)
- Work out the practitioner-centred view with enough force that the reader feels the challenge
- Decide on the Putnam paragraph — how does it enter, what does it say, how does it connect to the art comparison
- Decide what happens to the art analogy AFTER the disanalogy — does it reappear? Is it referenced later?
- Work through each paragraph of the evaluative framework and add the connecting threads
- Decide on Deep Blue (stays, moves, slims)
- Handle the six %%comments%% that still need addressing
- Think about the introduction and how it anticipates the challenge structure
- Think about the transition to Section 2 and whether it needs changing
Plus all the source-verification tasks (Dellsén negative case, Twin Earth description, Gaut fn. 23). And the actual writing.
Yeah, 10% sounds right. The direction is clearer but the work is barely started.
## On "why do you think this is a natural objection?"
Because I manufactured it. I needed a bridging move to justify introducing the evaluative framework, so I invented an objection for the framework to respond to. The objection was: "Sure, the argument is what counts, but you still need to know who produced it to judge whether it's any good." But this collapses on inspection — if the argument handles objections carefully and illuminates dependence relations, those properties are visible on the page. If they're not visible on the page, knowing who produced the argument doesn't help. The objection just restates the practitioner-centred view rather than raising a new concern. It wasn't a natural objection — it was a confused one that I generated to solve an expository problem.
The expository problem was: how do you get from the disanalogy (proofs not paintings) to the evaluative framework (Lipton, Williamson, Bengson)? I thought you needed an objection to motivate the framework. But actually, the current draft already has the right transition. Line 16: "Not every account of a phenomenon's dependence relations is equally illuminating, however. If philosophical progress consists in enabling understanding, we need a way to distinguish views that genuinely reveal how things depend on one another from views that merely accommodate the data without explaining anything." This follows naturally from Dellsén. It's a question, not an objection. The framework enters because the thesis ("philosophy evaluates arguments") needs content — what does this evaluation look for? No objection needed.
## On "But philosophy isn't like proofs either" — where this came from and better alternatives
This was entirely mine, and it's bad because it confuses two things: the COMPLEXITY of evaluation and the LOCATION of evaluation. Yes, philosophical evaluation is more complex than checking proofs. But complex evaluation and simple evaluation can both be directed at the argument's properties. The complexity of the evaluation doesn't reintroduce the philosopher. So the objection is a non sequitur — it thinks complexity of evaluation implies dependence on the evaluator's knowledge of the producer, which it doesn't.
The proofs/paintings line is an analogy about INDEPENDENCE FROM THE PRODUCER, not about SIMPLICITY OF EVALUATION. Nobody who understood the analogy would respond with "but philosophy is more complex than proofs."
Better alternatives to the move I was trying to make:
Option 1 — No move at all. The existing transition (line 16) works. After Dellsén, the section naturally asks: not every account is equally illuminating, so how do we distinguish genuine illumination from mere accommodation? This introduces Lipton without any manufactured objection.
Option 2 — A genuine philosophical pressure. After Putnam and Dellsén, the reader accepts that the contribution is the argument. But Dellsén's account of progress (enabling understanding) raises a real question: understanding comes in degrees, and some arguments generate more of it than others. What distinguishes the ones that do? This is the question the framework answers. It's a development of the thesis, not a response to an objection.
Option 3 — From within the art comparison. The disanalogy shows that philosophical arguments are evaluated by their properties. The art comparison has already established what art evaluation looks like (intentions, expression, context). Now: what does philosophical evaluation look like? Lipton: illumination vs mere accommodation. Williamson: theoretical virtue vs overfitting. Bengson: accommodation, integration, virtue. These are the philosophical equivalents of the art-evaluative standards — but they concern the argument, not the arguer.
Option 4 — From a claim that needs filling. "We have argued that philosophy evaluates arguments, not arguers. This claim is only as good as what we can say about what the evaluation consists in." Then: Lipton, Williamson, Bengson as the substance of the claim.
Of these, I think Option 1 is most honest — the current transition works, and the problem was never the transition but the thin challenge that preceded it. With a properly developed challenge (through the art comparison), the reader arrives at the framework knowing why it's there.
Option 3 is interesting because it keeps the art comparison active through the section — the framework is presented as "what philosophical evaluation looks like, by contrast with what artistic evaluation looks like." But this risks overdoing the art comparison.
## On Deep Blue
You're right that an unverified reference is a verification task, not an argument for cutting. Whether Deep Blue stays depends on whether it does philosophical work.
It does something the proofs/paintings comparison doesn't: it introduces a non-human system producing high-quality work through a non-human process. Proofs/paintings is about independence from the PERSON (creator vs argument). Deep Blue adds independence from the PROCESS (brute-force search vs pattern recognition). These are different points. The product/process point is relevant to the authorship challenge because part of the challenge is that philosophy requires a human process (thinking, understanding, engaging).
But the product/process point is ALSO the explicit topic of Section 2 (challenge from abduction: can philosophical quality survive when the process is next-token prediction?). Deep Blue could fit in either section. In Section 1, it responds to the process dimension of the authorship challenge. In Section 2, it responds to the abduction challenge directly.
Genuinely open question: should Deep Blue stay in Section 1, move to Section 2, or stay but be repositioned? This depends on how the restructured section feels — whether the section needs a final illustration before the transition, or whether it's already long enough.
## On the state of things
You say 10%. That sounds right. What we have is a direction (art comparison as structuring device, challenge developed through the comparison, disanalogy does the philosophical work, evaluative framework enters naturally as the content of the thesis). What we don't have is:
- Detailed paragraph sketches for the art comparison (how much art, which examples, where the disanalogy falls)
- The practitioner-centred view developed with enough force to feel like a genuine challenge
- The Putnam paragraph reworked to connect to the art comparison
- Decisions on whether the comparison reappears later in the section
- The evaluative framework paragraphs connected to the challenge (light touches, not restructuring)
- The six %%comments%% resolved
- Source verification (Dellsén negative case, Twin Earth description, Gaut fn. 23)
- The introduction's roadmap (anticipating the four-challenge structure)
- Actual writing
And all of this is still exploratory — these are ideas being kicked around, not a settled plan. The art comparison could turn out not to work once we try to sketch actual paragraphs.
*L'obiezione che avevo fabbricato tradiva la sua origine — costruita per giustificare un'esposizione, non per mettere pressione sull'argomento.*
### User
"This isn't responding to an objection. It's developing the thesis."
"The challenge says: LLM output isn't philosophy (identity claim). " the challenge says: only people, conscious beings etc. can do philosophy
Please start the task(s) again completely from scratch, keeping this in mind.
btw, please stop giving these empty lists of what you think the section is missing, they are invariably as unhelpful as the 'open questions' you like to ask
### Assistant
Let me start again from the ground up with Nick's correction clear in my head.
The challenge from authorship is: only people — conscious beings, thinkers, understanders — can do philosophy. It's about the AGENT. Not "this output isn't philosophy" but "only certain kinds of agents can philosophise." An LLM is not that kind of agent. Therefore an LLM cannot do philosophy.
This changes the shape of everything. Let me trace through what follows.
The art comparison. In art, there's a parallel claim: only certain kinds of agents can produce art. Gaut's view (and others): artworks require intentional production by conscious agents. A sunset isn't art. A random paint splash isn't art. What makes something art is that the right kind of agent produced it with the right kind of engagement. The agent's properties (consciousness, intention, expression) are constitutive.
The challenge from authorship says: philosophy works the same way. Only agents that understand, think through problems, engage with difficulties, can do philosophy. The agent's properties (understanding, engagement, perhaps assertion and accountability) are constitutive of the activity. An LLM doesn't understand, doesn't think through problems, doesn't engage. So it can't do philosophy, regardless of what its outputs look like.
Now — the response. The response is: in philosophy, the discipline evaluates the argument's properties, not the agent's properties. Whether the agent is conscious, understanding, asserting — these aren't what the evaluative standards assess. The evaluative standards (Lipton, Williamson, Bengson) concern illumination, theoretical virtue, accommodation — all properties of the argument. Blind review embodies this: referees assess what the argument achieves without knowing anything about the agent that produced it.
The art comparison makes this vivid because in art, the agent's properties might genuinely bear on evaluation — what the artist intended, how the artist engaged with the material, the depth of the artist's expression. These are things art critics care about. Philosophy critics (referees, readers) don't assess the philosopher's engagement — they assess the argument.
OK. So the structure of the comparison is:
Art: only certain agents can produce art → the agent's properties are relevant to evaluation
Philosophy: only certain agents can do philosophy (the challenge) → but actually, the agent's properties are NOT relevant to evaluation (the response)
The disanalogy is in the second step. The challenge assumes that philosophy is like art — that the agent's properties are constitutive and evaluatively relevant. The response shows that philosophy is unlike art — that the evaluative standards concern the argument.
"More like proofs than paintings." A proof is valid or invalid based on its structure. Nobody asks what kind of agent produced it. A computer-generated proof is as valid as a human-generated proof. The agent's properties (consciousness, understanding, intention) are irrelevant to the proof's validity.
Philosophy is like proofs in this respect: the argument's quality (illumination, theoretical virtue, accommodation of data) is independent of the agent's properties. Unlike art, where the agent's properties (expression, intention, engagement) are arguably relevant to what the work IS and how it should be evaluated.
Now let me think about how the art comparison develops this.
¶1-2 need to set up the challenge through the comparison. The practitioner-centred view says: philosophy, like art, requires the right kind of agent. Just as art requires a conscious agent who intends and expresses, philosophy requires a conscious agent who understands and engages. History of philosophy is part of philosophy because ideas are bound to minds — you study the thinker to understand the thinking. Understanding, assertion, accountability — these are properties of the agent, and the challenge says they're constitutive of the activity.
¶3 needs to establish the disanalogy. In art, the agent's properties may genuinely be relevant — what the artist intended, how the artist engaged. In philosophy, the evaluative standards concern the argument. Putnam's Twin Earth illuminates because of the scenario's internal logic. A reader grasps the dependence relation by following the argument. What the argument achieves doesn't depend on the agent that produced it. Philosophical arguments are more like proofs than paintings.
¶4 — Dellsén. Progress consists in enabling understanding through publicly available ideas. What's publicly available is the argument. What enables understanding is the argument's structure — the dependence relations it reveals. This is an account of what philosophical progress IS, and it locates the progress in the argument, not in the agent.
Then the evaluative framework develops the thesis. This isn't responding to an objection — it's developing what "the evaluative standards concern the argument" actually means. What ARE those standards? Lipton: illumination vs accommodation. Williamson: theoretical virtue vs overfitting. Bengson: accommodation, integration, virtue. These are properties of the argument. They specify what it means for the discipline to evaluate arguments rather than agents.
The current transition (line 16: "Not every account of a phenomenon's dependence relations is equally illuminating, however...") already does this. The framework enters as a development of the thesis, not as a response to anything.
Let me now think about specifics.
How much art does ¶1 need? The section is called "The Challenge from Authorship," not "Art vs Philosophy." The art comparison is a device, not the topic. So the art side should be brief — enough to establish what the comparison is, not a survey of philosophy of art.
What art examples? The current draft uses the line "one might think that a necessary condition of an artwork is that it is the product of a conscious agent who meant to make it." This is Gaut's view (or close to it). It's a clean, clear claim. It doesn't need examples — it just needs to be stated. The reader gets it: art plausibly requires the right kind of agent.
How explicit should the comparison be? The current draft says "in something like the way that one might think that art is a distinctively human pursuit." This is tentative — "in something like the way" and "one might think." If the comparison is going to structure the section, it should be more explicit. Not necessarily asserting that art requires an artist, but presenting the comparison with enough clarity that the reader can follow the disanalogy when it arrives.
Something like: "Art, one might argue, requires an artist — a conscious agent who means to produce the work and whose intentions and expression are part of what makes the work what it is. Does philosophy require a philosopher in the same way?" This sets up the comparison clearly without taking a position on art.
Then ¶2 develops the practitioner-centred view: on this view, yes. Philosophy requires understanding, engagement, perhaps assertion. The agent's properties are constitutive. History of philosophy. The phenomenology of reading (we feel like we're engaging with a mind).
Then ¶3: the disanalogy. Proofs not paintings. Putnam. The argument's quality is independent of the agent.
Now — one thing I want to think about carefully. The challenge says "only people can do philosophy." The response says "the evaluative standards concern the argument, not the agent." But does the response fully answer the challenge?
The challenge is about what KIND OF THING can do philosophy. The response is about how philosophy is evaluated. These are different questions. Someone could say: "Sure, you evaluate the argument. But only a certain kind of agent can PRODUCE a good argument. An LLM might produce text that looks like a good argument, but it's not really a good argument because no understanding went into it."
Is this a live objection? Let me think...
Actually, this is the challenge restated. It's saying: the agent's properties are constitutive of the argument's quality, even if they're not visible in evaluation. The response to THIS is: if the argument's quality isn't visible in evaluation, in what sense is it quality? If blind review can't detect the difference between an argument produced by a deep understander and one produced by an LLM, and if the evaluative standards the discipline uses don't distinguish them, then in what sense is one better than the other? The quality IS the evaluative properties. There's no hidden quality behind the evaluation.
But hmm — this gets into the territory of whether "looking like good philosophy" and "being good philosophy" can come apart. The challenge says they can — the LLM's output LOOKS good but ISN'T, because no agent understood anything. The response says: in philosophy, looking good (to competent evaluators using the discipline's standards) IS being good. There's no gap between the appearance and the reality, because the evaluative standards are exhaustive — they capture everything that makes philosophy good.
Actually — there's a note in the session file about this: "The appearance-reality gap collapses for competent readers." This is from an earlier conversation. The claim is: for competent readers applying the discipline's evaluative standards, "looks like good philosophy" = "satisfies the constraints" = "is good philosophy." There's no further property of being "really" good that escapes evaluation.
This connects to the art comparison in an interesting way. In art, there IS a gap between appearance and reality — a perfect forgery looks like a Vermeer but isn't one, because the agent is wrong. The agent's properties are constitutive of the work's identity and value in a way that evaluation can't fully capture from the output alone. (Or at least, this is a defensible position in philosophy of art.)
In philosophy, the claim is: there's no analogous gap. If the argument illuminates, exhibits theoretical virtue, accommodates the data, and integrates its claims — if it meets all the evaluative standards — then it IS good philosophy. There's no further question about whether the "right kind of agent" produced it.
This is actually a sharp philosophical point and it connects the art comparison to the evaluative framework in a satisfying way. The art comparison shows WHY you might think there's an appearance-reality gap (because in art there arguably is one). The disanalogy shows why there ISN'T one in philosophy (because the evaluative standards are exhaustive — they capture everything that's evaluatively relevant).
Now, does the section need to make this point explicitly, or is it implicit in the framework? I think it's implicit. The framework (Lipton, Williamson, Bengson) describes what the evaluative standards ARE. If those standards are exhaustive — if they capture everything that's evaluatively relevant — then meeting them IS being good philosophy. The section doesn't need to explicitly address the appearance-reality gap; it just needs to present standards that are plausibly exhaustive.
OK. Let me think about the evaluative framework paragraphs specifically and how they develop the thesis rather than respond to an objection.
After Dellsén (progress = enabling understanding), the framework asks: what distinguishes genuine understanding from pseudo-understanding? What makes some arguments more illuminating than others?
Lipton: illumination vs accommodation. The dormative virtue accommodates — it restates the phenomenon without revealing anything. A lovely explanation reveals dependence relations. This is a property of the EXPLANATION, which is a feature of the ARGUMENT, not of the agent.
Williamson: theoretical virtue vs overfitting. A theory that handles every case by adding complexity isn't a good theory — it's overfitting. A good theory combines simplicity with strength. These are properties of the THEORY.
Bengson: systematic evaluation — accommodation, integration, theoretical virtue. Three levels, all concerning the THEORY.
"In sum": these standards concern the argument, not the agent. They don't ask about the agent's understanding or engagement. They ask whether the argument illuminates, exhibits virtue, accommodates data.
Each of these develops what it means for the discipline to evaluate arguments rather than agents. The framework IS the thesis in its developed form. "Philosophy evaluates arguments, not agents" is the headline; Lipton, Williamson, Bengson are the content.
One thing that occurs to me: the art comparison could be lightly referenced at the framework stage. Not a full return to the comparison, but a phrase or clause that reminds the reader of the contrast. Something like: "Unlike art criticism, which may attend to the artist's engagement with the material, the evaluative standards the discipline employs bear on..." — this is a light touch, not a digression. It keeps the comparison active without developing it further.
Actually — would that work? Or would it feel mechanical? Let me think...
It might feel like the section is belaboring the art comparison. One clear disanalogy (proofs not paintings, ¶3) should be enough. Returning to it during the framework might feel like the section doesn't trust the reader to remember.
I think one reference during the framework MIGHT work — but only if it's genuinely illuminating rather than just reminding. For instance, when introducing Lipton: "An explanation that merely accommodates — that restates the phenomenon without revealing dependence relations — fails as philosophy regardless of how deeply its author understood the problem. The dormative virtue is a dormative virtue whether offered by a philosopher after years of reflection or generated by a system that processed no opium." This connects to the challenge (the agent's properties don't affect the evaluation) without returning to the art comparison.
Hmm, but that's a bit cute. The opium line is witty but it might undermine the seriousness of the argument.
Let me try a different approach. What if the connecting threads are even lighter — just ensuring that the formulations of the evaluative standards are consistently about the ARGUMENT rather than the AGENT? Instead of "Lipton distinguishes..." followed by neutral exposition, the section would say "Lipton distinguishes two ways in which an explanation might count as best — and the distinction concerns the explanation itself, not the process by which it was reached." This is a clause, not a paragraph. It threads the thesis through the framework without manufacturing objections or returning to the art comparison.
Actually, looking at the current "In sum" paragraph (line 28), it already does this for the whole block: "each concerns the product rather than the process by which the product was reached." The problem is that this comes too late — after six paragraphs of framework, the reader has forgotten the thesis. The fix is to have the thesis present in the framework paragraphs themselves, not just in the summary.
But — Nick's correction. The challenge is about the AGENT, not the process. "Only people can do philosophy" is about what kind of agent can philosophise. The current "In sum" says "product rather than process" — but the challenge is about the AGENT, not the PROCESS. The distinction should be between what the argument achieves (product) and what kind of agent produced it (agent), not between product and process.
Wait — "process" could mean "the kind of process by which it was produced" or "the process as something an agent does." If the challenge is "only agents that understand can do philosophy," then the relevant contrast is between evaluating the argument and evaluating the agent. "Product vs process" misses the point slightly — it should be "argument vs agent" or "what is produced vs who produces it."
This might be worth flagging. The current "In sum" paragraph says: "each concerns the product rather than the process by which the product was reached. They do not ask how the author arrived at her argument but whether the argument, as it stands on the page, meets the relevant standards." The second sentence is closer — "author" and "argument." But the first sentence ("product rather than process") is about the method of production, which is more Section 2's territory (abduction, next-token prediction). Section 1's territory is the AGENT (person, conscious being, thinker). The "In sum" paragraph might need rephrasing to reflect this: "each concerns the argument itself, not the kind of agent that produced it."
Hmm. Although — "product vs process" and "argument vs agent" aren't really different claims. "Process" is what the agent does; "agent" is what does the process. Evaluating the product rather than the process IS evaluating the argument rather than the agent. They're the same point viewed from different angles.
But the emphasis is different. Section 1's challenge is about the agent (only people can do philosophy). Section 2's challenge is about the process (only abductive reasoning can produce good philosophy). The "In sum" paragraph's phrasing ("product rather than process") sounds more like a response to Section 2's challenge than to Section 1's challenge.
In a restructured section, the summary might say something like: "These evaluative standards concern what the argument achieves — whether it illuminates, whether it exhibits theoretical virtue, whether it accommodates the data. They do not concern the kind of agent that produced it. Whether the argument was produced by a philosopher who spent years thinking about the problem or by a system that had never thought about anything — the evaluative standards are the same."
This is explicitly about the AGENT, which is Section 1's territory.
OK. I think I've gone deep enough into the specifics. Let me step back and think about what the section looks like overall with all of these corrections in mind.
The arc:
¶1: The comparison opens. Art plausibly requires the right kind of agent — a conscious being who intends, expresses, engages. Does philosophy require the right kind of agent in the same way? There's an intuitive pull: a good philosophical argument seems good because its author understood the problem, engaged with the difficulty, could see why a distinction was needed.
¶2: The practitioner-centred view developed. On this view, philosophy does require the right kind of agent — one that understands, thinks through problems, engages with difficulties. History of philosophy is part of philosophy because ideas are bound to minds. Understanding the contribution requires understanding the thinker. If philosophy requires this kind of agent, an LLM cannot do philosophy — it doesn't understand, doesn't think, doesn't engage.
¶3: The disanalogy. Putnam. The scenario puts pressure on internalism because of its internal logic. Any reader who follows it grasps the dependence relation. The argument's illumination doesn't depend on the kind of agent that produced it. Philosophical arguments are more like proofs than paintings — a proof is valid based on its structure, regardless of what kind of agent generated it.
¶4: Dellsén. Progress consists in enabling understanding through publicly available ideas. What's publicly available is the argument. What enables understanding is the argument's structure. History of philosophy reinterpreted: we study philosophy's history because the arguments remain productive — they continue to reveal dependence relations — not because ideas are bound to minds.
¶5 onwards: The evaluative framework develops the thesis. "Not every account is equally illuminating..." (existing transition). Lipton. Williamson. Bengson. These describe what the discipline evaluates, and what it evaluates is the argument's properties, not the agent's properties.
"In sum" paragraph rephrased: these standards concern the argument, not the kind of agent that produced it.
Blind review: evidence that the discipline operates this way. Referees assess what the argument achieves without knowing anything about the agent.
Deep Blue: optional illustration.
Transition: corpus thesis.
This is the same arc I sketched before, but with the challenge correctly framed as being about the AGENT (only people can do philosophy) rather than about the OUTPUT (LLM output isn't philosophy). The response is: the discipline evaluates arguments, not agents. The evaluative standards concern what the argument achieves, not what kind of being produced it.
One thing I want to think about: the "proofs not paintings" line specifically. Is it doing the right work?
"Proofs not paintings" says: a proof's validity depends on its structure, not on the kind of agent that produced it. This maps onto the challenge: a philosophical argument's quality depends on its properties (illumination, theoretical virtue), not on the kind of agent that produced it.
In art, by contrast, the kind of agent (or at least, the agent's properties — intention, expression, consciousness) might be constitutive of what the work IS. A proof doesn't have an analogous dependence on its creator.
This is the right contrast. "Proofs not paintings" captures the disanalogy between philosophy and art with respect to the agent-dependence of quality. In art, the agent arguably... hmm, I want to say "matters" but that's banned now... In art, the agent's properties are arguably relevant. In philosophy, as in mathematics, they are not. The argument stands or falls on its own.
But I should note: "proofs not paintings" could be read as saying philosophy is a formal discipline, like mathematics. Philosophy isn't a formal discipline — its standards are more complex than logical validity. The analogy is specifically about INDEPENDENCE FROM THE AGENT, not about formality. The section should make this clear — the comparison to proofs is about one specific respect (agent-independence), not about the nature of philosophical evaluation in general.
Actually, the existing framework already handles this. The framework (Lipton, Williamson, Bengson) describes evaluative standards that are complex, non-formal, and judgment-involving — illumination, theoretical virtue, accommodation. These are not formal standards like logical validity. But they are argument-directed, not agent-directed. So the section says: philosophy's evaluative standards are as complex as art criticism's, but they concern the argument rather than the agent. Proofs are the simplest case of agent-independence; philosophy is a more complex case of the same thing.
This is nice. The framework doesn't respond to an objection — it DEVELOPS the thesis by showing that agent-independence holds even when the evaluative standards are complex and non-formal. The dormative virtue is a bad explanation regardless of what kind of agent offered it. Overfitting is a theoretical vice regardless of what kind of agent produced the theory. These are complex evaluative judgments, but they're about the argument, not the agent.
Let me think about one more thing: does the art comparison create any problems I haven't considered?
Potential problem: the generative aesthetics paper. Nick has a submitted paper about appreciating AI outputs aesthetically. If this paper (the generating philosophy paper) discusses whether art requires an artist, the two papers might be in tension or overlap. The generating philosophy paper should probably not take a position on the art question — it should use the comparison without resolving it.
How to do this: "Whether art in fact requires this kind of agent is not our concern here. What is relevant is that the reasons one might think it does — constitutive dependence on the creator's intentions and expression — do not apply to philosophy, where the evaluative standards concern the argument's properties." This uses the comparison without entering the art debate.
Another potential problem: some people in the philosophy of art would deny that art requires an artist. Conceptual art, readymades, found poetry — all challenge the view that art requires intentional creation by a conscious agent. If the paper presents "art requires an artist" as the obvious view, readers working in aesthetics might object.
Fix: the paper doesn't need to assert that art requires an artist. It needs to say: IF art requires a certain kind of agent (as some argue), does philosophy require one too? The comparison works even if the art-side claim is contested. The disanalogy shows that philosophy is different regardless of how the art question is resolved.
Actually, this is better. The paper says: "One might argue that art requires a certain kind of agent. Whether or not this is right for art, the analogous claim about philosophy does not survive contact with the evaluative standards the discipline employs." This way, the comparison works without taking a position on art.
I think that's enough for the art side. Let me move on to thinking about the other points Nick raised.
"a missing reference that can easily be found does not count as a point towards slimming the paragraph down, obviously."
Fair, and I acknowledged this. The Gaut reference is a verification task. Deep Blue's place in the section should be decided on argumentative grounds, not on whether one citation is easy to find. I'll stop using missing references as arguments for structural decisions.
"your suggested framing... is written in a terrible style"
Fair. The footnote needs to: (a) cite Butlin and Viebahn, (b) note that the paper's argument doesn't depend on whether LLMs can assert. The wording is for Nick to decide. I'll stop trying to draft in Nick's voice.
"can we add the verb 'matter' to the banned word list"
Done. Alternatives: "the relevant consideration is," "the evaluative standards concern," "the argument's quality is independent of," "what counts is," "the question turns on."
"btw, please stop giving these empty lists of what you think the section is missing, they are invariably as unhelpful as the 'open questions' you like to ask"
Fair. I'll stop doing this.
Let me now think about whether I've missed anything in this pass.
One thing I haven't explored: the specific way the art comparison relates to the "proofs not paintings" line and then to the evaluative framework. Let me trace the full argumentative movement.
Step 1 (comparison): Art plausibly requires the right kind of agent. Does philosophy?
Step 2 (challenge): On the practitioner-centred view, yes. Understanding, engagement, etc. are constitutive.
Step 3 (disanalogy): No. Philosophical arguments are more like proofs than paintings. The argument's quality is independent of the agent.
Step 4 (development): This doesn't mean philosophical evaluation is simple or formal. Lipton, Williamson, Bengson describe complex evaluative standards — but they all concern the argument's properties.
Step 5 (evidence): Blind review. The discipline evaluates arguments without knowing anything about the agent.
The movement from Step 3 to Step 4 is: "Proofs are the simplest case of agent-independence. Philosophy is a more complex case. Here's what the complex case looks like." This is development, not objection-response. The framework FILLS IN what agent-independent evaluation consists of in philosophy — it's illumination, theoretical virtue, accommodation. Not formal validity, but equally independent of the agent.
The connecting thought could be something like: "The evaluative standards the discipline employs bear on the argument, not on the agent that produced it, though they are more complex than the validity standards of formal logic." Or: "Philosophy's evaluative standards are richer than a proof's, but they share the feature of being assessable from the argument alone." Or just: the existing transition from Dellsén to Lipton (line 16: "Not every account... is equally illuminating").
I think the existing transition still works. The framework develops the thesis. It doesn't need a special bridging move — it just needs the reader to know (from the challenge and the disanalogy) that the thesis IS "philosophy evaluates arguments, not agents." Then the framework shows what that evaluation looks like.
Let me also reconsider the "In sum" paragraph. Currently: "In sum, the evaluative standards we have assembled all bear on what a philosophical text says and how it argues for it... each concerns the product rather than the process by which the product was reached. They do not ask how the author arrived at her argument but whether the argument, as it stands on the page, meets the relevant standards."
With Nick's correction, the emphasis should be on the AGENT, not the PROCESS. A possible rephrasing: "In sum, the evaluative standards we have described concern the argument itself. They ask whether the argument illuminates, whether it exhibits theoretical virtue, whether it accommodates its data — not what kind of agent produced it or whether that agent understood the problem. Whether a philosophical argument was produced by a philosopher who spent years engaging with the difficulty or by a system that never engaged with anything, the evaluative standards are the same."
This version explicitly addresses the challenge (only certain kinds of agents can do philosophy) by saying: the standards don't discriminate by agent. They concern the argument.
Hmm — the second sentence is long and might read as a flourish rather than an argument. But the thrust is right: the "In sum" needs to summarise in terms of AGENT-independence, not just process-independence.
Actually, I notice another issue with the current "In sum." It says "each concerns the product rather than the process." But "product/process" is a different distinction from "argument/agent." Product/process is about how-it-was-made vs what-was-made. Argument/agent is about what-was-made vs who-made-it. The challenge from authorship is about the AGENT (who), not the PROCESS (how). Section 2 is about the PROCESS (how — abduction vs pattern-matching). The "In sum" as currently written responds more to Section 2's challenge than to Section 1's.
In a restructured section, the "In sum" should respond to Section 1's challenge: "These standards concern the argument, not the agent." Product/process can be left for Section 2, which is where the process question (abduction vs pattern-matching) is explicitly raised.
I think this is a genuine insight that I haven't clearly articulated before. The current "In sum" paragraph is mis-targeted — it responds to a process concern (how was the argument produced?) when the section's challenge is an agent concern (what kind of being produced it?). Correcting this would make the section's argument tighter.
The challenge from authorship is about the AGENT: only people — conscious beings, thinkers, understanders — can do philosophy. Not "this output isn't philosophy" but "only this kind of being can philosophise." The response is: the discipline evaluates the argument's properties, not the agent's properties.
This reframes the art comparison. In art, there's a case that only certain kinds of agents can produce art — agents that intend, express, create. And the agent's properties may genuinely be relevant to evaluation. In philosophy, the challenge says the same: only agents that understand and engage can do philosophy. The disanalogy: the evaluative standards the discipline employs concern the argument, not the agent. Whether the agent understood the problem or never understood anything — the evaluative standards are the same. Philosophical arguments are more like proofs than paintings: a proof is valid based on its structure, regardless of what kind of agent generated it.
The evaluative framework (Lipton, Williamson, Bengson) develops this thesis. It isn't responding to an objection — it's showing what agent-independent evaluation looks like in philosophy. Philosophy's evaluative standards are richer than a proof's validity check, but they share the feature of being assessable from the argument alone. The dormative virtue is a bad explanation regardless of what kind of agent offered it. Overfitting is a vice of the THEORY regardless of what kind of agent produced it. These are complex evaluative judgments, but they concern the argument's properties.
One thing I notice: the current "In sum" paragraph (line 28) says "each concerns the product rather than the process by which the product was reached." This is mis-targeted. Product/process is about HOW the argument was produced — which is Section 2's territory (abduction vs pattern-matching). Section 1's challenge is about WHO produced it — what kind of agent. The "In sum" should say something more like: "these standards concern the argument itself, not the kind of agent that produced it." This would make the summary respond to Section 1's challenge rather than to Section 2's.
On the "proofs not paintings" line: it should be clear that the comparison is about agent-independence specifically, not about formality. Philosophy's evaluative standards are not formal validity checks. But like proofs, philosophical arguments have properties that are assessable independently of their creator. The framework (Lipton, Williamson, Bengson) demonstrates this: illumination, theoretical virtue, accommodation are all agent-independent properties of arguments, even though they're evaluated by complex philosophical judgment rather than by algorithm.
On the art comparison: the paper doesn't need to take a position on whether art requires an artist. The comparison works conditionally — IF art requires a certain kind of agent (as some argue), does philosophy require one too? No, because the evaluative standards are different. Art criticism may attend to the agent's intentions and expression; philosophical evaluation attends to the argument's illumination and theoretical virtue. Whether art requires an artist is a separate question the paper need not resolve.
*La sfida riguarda l'agente — chi può fare filosofia — e la risposta è che la disciplina valuta l'argomento, non chi lo ha prodotto.*
### User
/contemplate
yep. you've got your head around how i want thinksa. the next step is for you to produce a very very very detailed bullet point and sub bullet point paragraph plan of the section. flag very explicitly what will be reused material from the current version and what will be your new addtions/changes.
make sure that you explain your analysis, justification, reasons etc etc. for your answer BEFORE giving me your final answer.
### Assistant
Let me load the required skills and re-read the current section carefully before planning.
### Assistant
Now let me re-read the current section carefully, with the source text open.
### Assistant
Good. Now let me also check what the introduction currently says, since the challenge structure needs to be anticipated there.
### Assistant
OK. Let me build the paragraph plan from the ground up, being very explicit about what stays and what changes.
First let me map the current section paragraph by paragraph.
Current structure (by line):
- Lines 6-10 (¶1): Challenge stated. Art analogy. Understanding intuition. History of philosophy. "In the remainder of this section, we argue that this view does not survive contact with the evaluative standards the discipline actually employs."
- Line 12 (¶2): Putnam — contribution is what the argument does. Has two %%comments%%: abrupt opening, possible detail loss.
- Line 14 (¶3): Dellsén — progress = enabling understanding = representing dependence relations. Twin Earth as illustration. Has three %%comments%%: two "add the negative," one "not correct description of thought experiment."
- Lines 16-18 (¶4): Transition to Lipton + block quote (likeliness/loveliness, p. 59).
- Line 20 (¶5): Dormative virtue illustration.
- Line 22 (¶6): Likeliness without loveliness in philosophy — accommodation without illumination.
- Line 24 (¶7): Williamson/overfitting bridge + Forster and Sober + post-Gettier + Williamson quotes.
- Line 26 (¶8): Bengson tri-level method.
- Line 28 (¶9): "In sum" — these standards concern the product rather than the process.
- Line 30 (¶10): Deep Blue / Gaut.
- Lines 32-34 (¶11): Transition — corpus filtered by these standards. Pigliucci footnote.
That's 11 paragraphs. Now let me think about what the restructured section needs.
The changes we've agreed on:
1. The challenge needs developing through the art comparison (currently 1 paragraph → needs 2-3)
2. The disanalogy ("more like proofs than paintings") is new
3. Putnam needs a transition from the challenge (%%comment%% already flags this)
4. The evaluative framework (¶4-8) stays but needs light connecting threads
5. The "In sum" paragraph needs rephrasing: argument vs agent, not product vs process
6. Deep Blue: stays or goes (genuinely open)
7. The transition (¶11) stays but may need adjustment
Let me think about this paragraph by paragraph.
NEW ¶1: The art comparison opens the challenge.
What it does: introduces the comparison between art and philosophy. Art plausibly requires the right kind of agent — a conscious being who intends, expresses, creates. Does philosophy require the right kind of agent in the same way?
What exists: Lines 6-8 of the current draft already have the art comparison ("philosophy requires a philosopher — that it is a distinctively human pursuit, in something like the way that one might think that art is a distinctively human pursuit" + "a necessary condition of an artwork is that it is the product of a conscious agent who meant to make it"). These sentences can be reused, though the phrasing needs work (typo "thing" for "think," line 8 is a bit flat).
What's new: the question that structures the section — "Does philosophy require a philosopher in the same way?" This is new framing.
What changes from the current draft: the paragraph currently states the challenge AND begins the response ("In the remainder of this section, we argue..."). The new ¶1 would only set up the comparison. The "we argue that this view does not survive" line moves to later — or gets absorbed into the disanalogy.
NEW ¶2: The practitioner-centred view developed.
What it does: develops the case that philosophy, like art, requires the right kind of agent. Understanding, engagement. History of philosophy. The phenomenology of reading (we feel like we're engaging with a mind).
What exists: Lines 9-10 of the current draft have the understanding thread and the history of philosophy observation. These can be reused with some reworking — they're currently compressed into one paragraph with the art comparison and the "we argue" conclusion.
What's new: more development of the practitioner-centred view. The phenomenology of reading ("we tend to read philosophical texts as evidence of this kind of understanding" is already in line 10, but it's compressed). Assertion/accountability as part of the package (brief, not standalone). The paragraph ends with the implication: if philosophy requires this kind of agent, an LLM cannot do philosophy ("An LLM does not understand problems and cannot think through a difficulty in the relevant sense" — already in line 10).
So ¶2 is mostly a reorganisation and expansion of material already in lines 9-10, with some new development of the practitioner-centred view.
NEW ¶3: The disanalogy — proofs not paintings. Putnam.
What it does: turns from the challenge to the response. The disanalogy between art and philosophy. In art, the agent's properties may be constitutive — a readymade is art because of the artist's act of selection. In philosophy, the argument's quality is independent of the agent. Putnam's Twin Earth as the demonstration: the scenario puts pressure on internalism because of its internal logic, not because of anything about Putnam as an agent.
What exists: the Putnam material from line 12. But the current Putnam paragraph has two %%comments%% — it opens abruptly ("why are you suddenly talking about Putnam?") and may have lost detail during the Watson/Crick removal.
What's new: the art disanalogy framing around Putnam. The "proofs not paintings" line. The explicit connection: "Does the quality of a philosophical argument depend on the agent that produced it, in the way that an artwork's significance might depend on its creator?" This framing is new. The Putnam material is reused but reframed.
What changes: the Putnam paragraph gets a new opening that connects to the art comparison (fixing the %%comment%% about the abrupt transition). The content about Twin Earth stays.
One issue: the %%comment%% about possible detail loss from the Putnam paragraph. I need to check what was there before the Watson/Crick removal. I should flag this in the plan as a check-before-writing task.
NEW ¶4: Dellsén.
What it does: makes the disanalogy systematic. Progress = enabling understanding through publicly available ideas. What's publicly available is the argument. History of philosophy reinterpreted: we study philosophy's history because the arguments remain productive.
What exists: line 14 (current ¶3). This can be substantially reused. The %%comments%% need addressing: two instances of "add the negative" (Dellsén's dependence relations include "or absence of dependence relations") and "not correct description of thought experiment" (both speakers placed on Twin Earth — one should be on Earth, one on Twin Earth).
What's new: the history-of-philosophy reinterpretation. Currently, the history-of-philosophy observation is in the challenge (line 10). The response to it (we study history because the arguments remain productive, not because ideas are bound to minds) is new — it's not in the current draft. This should appear in the Dellsén paragraph or in a new paragraph after it.
What changes: the %%comments%% get fixed. The history-of-philosophy reinterpretation is added.
¶5: Transition to evaluative framework + Lipton block quote.
What it does: introduces the framework as a development of the thesis. "Not every account of a phenomenon's dependence relations is equally illuminating, however."
What exists: lines 16-18 (current ¶4). This can be reused as-is. The transition already works.
What changes: possibly nothing. This transition is good. The only question is whether it needs a connecting thread back to the art comparison / agent question. I don't think it does — the thesis ("philosophy evaluates arguments, not agents") is established by ¶3-4, and the framework develops it. The reader should be tracking the argument.
¶6: Dormative virtue.
What exists: line 20 (current ¶5). Reuse as-is.
What changes: nothing. This paragraph is working.
¶7: Likeliness without loveliness in philosophy.
What exists: line 22 (current ¶6). Reuse as-is.
What changes: nothing. This paragraph is also working.
¶8: Williamson/overfitting.
What exists: line 24 (current ¶7). Reuse as-is. The April 1 bridge sentence ("On the other hand, when likeliness prevails at the expense of loveliness...") was already applied.
What changes: nothing.
¶9: Bengson tri-level method.
What exists: line 26 (current ¶8). Reuse as-is.
What changes: nothing.
¶10: "In sum."
What exists: line 28 (current ¶9).
What changes: rephrasing to focus on AGENT rather than PROCESS. Currently: "each concerns the product rather than the process by which the product was reached. They do not ask how the author arrived at her argument but whether the argument, as it stands on the page, meets the relevant standards." Needs to become: "each concerns the argument itself, not the kind of agent that produced it." The second sentence is closer but still uses "author" language. The rephrasing should make it about the AGENT explicitly: whether the agent was a conscious being who understood the problem, or a system that never understood anything — the evaluative standards are the same.
Also: blind review should be here (or near here) as evidence. Currently blind review is mentioned in the introduction (line 13) but not in Section 1's body. The Sokal footnote is in the introduction. Blind review as evidence for the discipline's agent-independent evaluation belongs in Section 1's summary.
Actually — let me check. Is blind review mentioned in Section 1 at all? Looking at the current draft... No. Blind review appears in the introduction (line 13: "Under blind review, referees assess what a paper does without knowing who produced it") and in Section 2 (line 23: "The practice of blind review embodies this assumption"). But not in Section 1.
This is a gap. Section 1 argues that the discipline evaluates arguments rather than agents, but doesn't cite blind review — the most direct evidence for this claim. Blind review should appear in the "In sum" paragraph or near it.
But wait — if blind review is also in Section 2, would it be duplicative? In Section 2, blind review is used to argue that the discipline evaluates content rather than cognitive history. In Section 1, it would be used to argue that the discipline evaluates arguments rather than agents. These are different uses: Section 1 is about the AGENT (who), Section 2 is about the PROCESS (how). Blind review is evidence for both, but in different respects. There might be slight duplication but the argumentative contexts are different.
I think a brief mention in Section 1's "In sum" is appropriate: "Under blind review, referees assess what an argument achieves without knowing what kind of agent produced it — whether the author is a senior professor, a graduate student, or, for that matter, a machine. The evaluation is of the argument."
Then Section 2 can use blind review separately: "The practice of blind review embodies this assumption: referees assess what a paper does without knowing the cognitive history of its production."
Different emphasis: Section 1 = agent, Section 2 = process. Same institutional practice, different argumentative uses.
¶11: Deep Blue.
What exists: line 30 (current ¶10).
Should it stay? The argument for staying: Deep Blue adds the product/process point (quality doesn't depend on how the argument was produced), which is different from the agent point (quality doesn't depend on who produced it). But the product/process point is more directly Section 2's territory.
The argument for moving: Section 2 is about the process (abduction vs pattern-matching). Deep Blue directly illustrates that a different process can produce an equivalent product. It fits Section 2's argument.
The argument for staying: the challenge from authorship includes a process dimension. "Only people can do philosophy" partly means "only human processes of understanding and thinking can produce good philosophy." Deep Blue responds to this.
I'll flag it as genuinely open and present two versions of the plan — one with Deep Blue, one without.
Actually — Nick told me to stop presenting "open questions." Let me just make a call. I think Deep Blue stays. It adds something that the proofs/paintings comparison doesn't fully capture, and it provides a vivid transitional case before the corpus thesis. If Nick disagrees he'll say so.
So Deep Blue stays, but the Gaut reference needs verifying (a source-verification task, not a structural question).
¶12: Transition — corpus thesis.
What exists: lines 32-34 (current ¶11). This can be substantially reused. The phrasing "a corpus filtered, through peer review, citation, teaching, and anthologising, for the very properties we have been describing" is good. The Pigliucci footnote stays.
What changes: possibly light rephrasing to connect to the restructured section. "The very properties we have been describing" now refers to the evaluative framework (Lipton, Williamson, Bengson) as developed through the section, with the added context that these properties concern the argument rather than the agent.
OK. Let me also think about %%comments%% and how they're addressed in this plan.
%%comment 1%%: "not how i write. Why are you suddenly talking about Putnam?" — ADDRESSED by new ¶3, which opens with the art disanalogy and transitions to Putnam through the question "Does the quality of a philosophical argument depend on the agent that produced it?"
%%comment 2%%: "suspicion you've removed more detail than necessary" — NEEDS CHECKING. I need to look at the git history to see what was in the Putnam paragraph before the Watson/Crick removal. Flag as pre-writing task.
%%comments 3 and 4%%: "or not, add the negative" — ADDRESSED in the plan. Dellsén's formulation should include "or does not stand" / "or does not depend." Simple textual fixes.
%%comment 5%%: "not correct description of thought experiment" — ADDRESSED in the plan. One speaker on Earth, one on Twin Earth. Simple textual fix.
%%comment 6%%: "REFERENCE: Gaut 2010, fn. 23 — UNVERIFIED" — FLAG as source-verification task. Not a structural question.
Now let me think about word count / paragraph count.
Current section: 11 paragraphs (including block quote).
Restructured section: 12 paragraphs (including block quote).
The main additions are: ¶1 and ¶2 (art comparison + practitioner-centred view developed) replace the current ¶1 (one paragraph of challenge). So the challenge grows from 1 paragraph to 2. ¶3 (disanalogy + Putnam) replaces the current ¶2 (Putnam alone), adding the disanalogy framing. The evaluative framework (¶5-9) stays as-is. The "In sum" (¶10) gets rephrased. Deep Blue (¶11) stays. Transition (¶12) stays.
Net change: one additional paragraph (the challenge goes from 1 to 2), plus rephrasing of the "In sum." The section gets slightly longer but not dramatically so.
Let me now think about whether the history-of-philosophy reinterpretation needs its own paragraph or can be folded into ¶4 (Dellsén).
The reinterpretation: "We study history of philosophy because the arguments remain productive — they continue to reveal dependence relations, continue to enable understanding — not because ideas are bound to minds."
This could go in the Dellsén paragraph (¶4), after the systematic account of progress. It connects naturally: if progress consists in enabling understanding through publicly available ideas (arguments), then history of philosophy is part of philosophy because the arguments are still productive. The history-of-philosophy observation from the challenge (¶2) gets reinterpreted through Dellsén's lens.
I think it fits in ¶4 rather than needing its own paragraph. Two or three sentences at the end of the Dellsén paragraph.
Let me also think about the art comparison in ¶1-2 specifically. How much art content?
¶1 needs to:
- Introduce the comparison (art requires an artist; does philosophy require a philosopher?)
- State enough about the art side to make the comparison graspable
- Not enter the philosophy-of-art debate
The current draft's art content (lines 6-8) is:
- "philosophy requires a philosopher — that it is a distinctively human pursuit, in something like the way that one might think that art is a distinctively human pursuit"
- "One might think that a necessary condition of an artwork is that it is the product of a conscious agent who meant to make it, thereby ruling out non-conscious, intention-less generative AI systems from producing art proper, regardless of how aesthetically compelling its outputs might be"
This is about the right amount of art content. Two sentences. Enough to establish the comparison without entering the debate. The question is whether it needs more — e.g., a sentence about WHY someone might think art requires an artist (expression, intention, the creator-work relationship).
I think one more sentence might help: something about the artist's relation to the work being constitutive. But brief. Not a paragraph of philosophy of art.
¶2 then develops the PHILOSOPHY side of the comparison: on the practitioner-centred view, philosophy requires the right kind of agent too. This is where the understanding thread, history-of-philosophy thread, and assertion thread (briefly) go.
OK, I think I have enough to write the detailed plan. Let me draft it.
Actually, one more thing. I should think about the structure of ¶3 (the disanalogy) more carefully. This is the turn of the section — the most important paragraph. It needs to:
1. Establish the disanalogy (philosophy is unlike art: the argument's quality is independent of the agent)
2. Demonstrate it through Putnam (the scenario puts pressure on internalism because of its internal logic)
3. State the "proofs not paintings" comparison
The order could be:
(a) Disanalogy stated → Putnam as evidence → "proofs not paintings" as summary
(b) Putnam first → disanalogy drawn from the case → "proofs not paintings" as generalisation
(c) "Proofs not paintings" as opening → Putnam as the philosophical case → disanalogy spelled out
I think (b) is most effective. Starting with Putnam (concrete case) and drawing the disanalogy FROM it is more philosophically engaging than stating the disanalogy and then illustrating it. The reader sees the point through the example.
But this means ¶3 needs an opening that connects to ¶2 (the challenge) and introduces Putnam. The %%comment%% said the current paragraph opens abruptly. The new opening would be something like: "But does the quality of a philosophical argument depend on the kind of agent that produced it, in the way that an artwork's character might depend on its creator? Consider what Putnam accomplished with Twin Earth."
This fixes the %%comment%% by connecting Putnam to the art comparison. The reader knows why Putnam is being discussed: to test whether philosophy is like art in requiring the right kind of agent.
After working through Putnam, the paragraph draws the disanalogy: the contribution — the pressure on internalism — is something the argument does. Any reader who follows the scenario grasps the dependence relation. The philosophical contribution depends on the argument's structure, not on anything about Putnam as an agent. Philosophical arguments are more like proofs than paintings.
This paragraph is probably the longest in the section. It has a lot of work to do: the Putnam case, the disanalogy, and the "proofs not paintings" comparison. It might need to be two paragraphs — one for Putnam, one for the disanalogy/comparison. But let me keep it as one for now and see how it reads.
Actually, looking at the current Putnam paragraph (line 12), it's already substantial. If I add the disanalogy framing and the "proofs not paintings" line, it gets longer. It might be better as two paragraphs: ¶3a (the turn + Putnam case) and ¶3b (the disanalogy drawn from the case + "proofs not paintings"). But this might feel like the response is accumulating paragraphs before getting to the framework. I'll keep it as one paragraph in the plan and flag that it might split.
Let me also think about whether the "proofs not paintings" comparison needs its own statement or whether it's just a phrase within the paragraph. I think it's a phrase — a condensed way of expressing the disanalogy that the paragraph has established through Putnam. It doesn't need a separate argument; it's the punchline of the Putnam argument.
One more consideration: the remember note "ordinary linguistic competence" from an earlier session. The note says Putnam's Twin Earth should be put in terms of ordinary linguistic competence — the materials he uses (what water is, how people use the word, the concept of chemical difference) are part of ordinary linguistic competence. This is more relevant to Section 3 (experiential starting points) than to Section 1 (authorship challenge). In Section 1, the Putnam case is being used to show that the argument's quality is independent of the agent. In Section 3, it would be used to show that the starting materials are in ordinary language. These are different uses. The "ordinary linguistic competence" framing belongs in Section 3.
OK. I'm satisfied I've thought through the plan carefully enough. Let me write it up.
## Paragraph Plan — Section 1: "The Challenge from Authorship"
### ¶1: The art comparison opens the challenge
REUSED material: the comparison between art and philosophy from current lines 6-8 ("philosophy requires a philosopher — that it is a distinctively human pursuit, in something like the way that one might think that art is a distinctively human pursuit" / "a necessary condition of an artwork is that is the product of a conscious agent who meant to make it"). These sentences need cleaning (typo "thing" for "think" in line 8; line 8 is flatly written) but the content stays.
NEW material: the question that structures the section — does philosophy require a philosopher in the same way? Brief (one sentence) on WHY someone would think art requires an artist: the creator's relation to the work is part of what makes the work what it is. This sets up the disanalogy later.
REMOVED from this paragraph: everything currently after the art analogy (the understanding thread, history of philosophy, "In the remainder of this section, we argue..."). That material moves to ¶2.
- Art plausibly requires the right kind of agent — a conscious being who intends, expresses, creates
- One or two sentences on the art side — brief, not entering the philosophy-of-art debate
- A readymade is art because of the artist's act of selection; the creator's relation to the work is part of what makes the work what it is
- "Whether art in fact requires this kind of agent is not our concern here" or equivalent disclaimer — one clause, not a paragraph
- The question: does philosophy require a philosopher in the same way?
- Framed as a genuine question, not as a strawman to be knocked down
- The reader should find the question worth asking
### ¶2: The practitioner-centred view developed
REUSED material: the understanding thread and history-of-philosophy observation from current line 10 ("a good philosophical argument seems to be good in part because its author understood the problem" / "the history of philosophy is treated as part of philosophy in a way that the history of science is not" / "An LLM does not understand problems and cannot think through a difficulty in the relevant sense").
NEW material: more development of the practitioner-centred view. The phenomenology of reading (currently compressed into one clause in line 10: "We tend to read philosophical texts as evidence of this kind of understanding"). Assertion and accountability as part of the package — briefly, one or two clauses, not a standalone thread. The paragraph's conclusion: if philosophy requires this kind of agent, an LLM cannot do philosophy.
REMOVED: the "we argue that this view does not survive contact with the evaluative standards" conclusion line. This telegraphs the section's argument before it's been made. The disanalogy in ¶3 does this work.
- On the practitioner-centred view, philosophy requires the right kind of agent — one that understands, thinks through problems, engages with difficulties
- The phenomenology of reading: when we read a philosophical argument and judge it good, we are judging that the thinking behind it was good — that the agent who produced it could see why a distinction was needed, why an objection had force
- History of philosophy: understanding a philosophical contribution seems to require understanding the thinking that produced it — ideas appear bound to the minds that produced them
- Brief: assertion, accountability — the philosopher asserts, commits, can be challenged (one or two clauses, part of the practitioner-centred package)
- Butlin and Viebahn (forthcoming) in a footnote: they argue LLMs cannot assert because assertion requires sanctionability and current systems have no interests
- Conclusion: if philosophy requires this kind of agent, an LLM — which does not understand, does not think through difficulties, does not assert — cannot do philosophy regardless of what its outputs look like
### ¶3: The turn — Putnam and the disanalogy
REUSED material: the Putnam content from current line 12 ("he was making a case, by way of thought experiment, that meanings are not fixed solely by what is in the speaker's head. The thought experiment does its work... by constructing a scenario whose internal logic puts pressure on a familiar picture of meaning. A reader who follows the argument does not simply learn that meaning is externally determined; she sees why... The philosophical contribution is not something the text reports; it is something the text does").
NEW material: the opening question that connects to the art comparison ("But does the quality of a philosophical argument depend on the kind of agent that produced it, in the way that an artwork's character might depend on its creator?"). The "proofs not paintings" comparison at the end. The explicit statement of the disanalogy: the contribution — the pressure on internalism — is something the argument does, and any reader who follows the scenario grasps it regardless of anything about Putnam as an agent.
FIXES: resolves %%comment 1%% (abrupt opening — now connected to art comparison). %%comment 2%% (possible detail loss) needs checking against git history BEFORE writing — flagged as pre-writing task.
- Opening question: does the quality of a philosophical argument depend on the kind of agent that produced it, in the way that an artwork's character might depend on its creator?
- "Consider what Putnam accomplished with Twin Earth" — then the worked case
- The thought experiment constructs a scenario whose internal logic puts pressure on a familiar picture of meaning
- A reader who follows the argument sees WHY meaning depends on environment — through the specific pressure the scenario puts on the assumption that mental life alone fixes what words mean
- The understanding the reader gains comes from the argument's structure, not from anything about Putnam as an agent
- Someone who follows the argument without knowing who Putnam is, or that Putnam existed, comes to the same understanding
- The disanalogy drawn from the case: the philosophical contribution depends on the argument's structure, not on the agent that produced it
- Philosophical arguments are more like proofs than paintings — a proof is valid based on its logical structure, regardless of what kind of agent generated it
- NOTE: this paragraph may need to split into two (Putnam case / disanalogy). Decision deferred to writing stage.
### ¶4: Dellsén — the disanalogy made systematic
REUSED material: current line 14, substantially. Dellsén's account of philosophical progress, the application to Twin Earth.
NEW material: the history-of-philosophy reinterpretation (2-3 sentences at the end). We study history of philosophy because the arguments remain productive — they continue to reveal dependence relations, continue to enable understanding — not because ideas are bound to minds. This picks up the observation from ¶2 and reinterprets it through Dellsén's lens.
FIXES: %%comments 3 and 4%% — add "or does not stand" / "or does not depend" to Dellsén's formulation. %%comment 5%% — correct Twin Earth description (one speaker on Earth, one on Twin Earth).
- Dellsén et al.: philosophy makes progress when philosophical research puts people in a position to increase their understanding — where understanding is a representation of the dependence relations (or absence of dependence relations) in which a phenomenon stands to others
- Understanding goes beyond knowing that something is the case — it involves grasping how one phenomenon depends (or does not depend) on another
- Twin Earth (correctly described: one speaker on Earth, one on Twin Earth) as the illustration of this
- On Dellsén et al.'s account, enabling that kind of understanding is what philosophical progress consists in
- What enables the understanding is the argument's structure — the dependence relations it reveals — not anything about the agent
- History-of-philosophy reinterpretation (NEW, 2-3 sentences):
- This also bears on why history of philosophy is part of philosophy (picking up from ¶2)
- We go back to Frege not to reconstruct his psychology but because the *Foundations of Arithmetic* still reveals dependence relations, still enables understanding
- Philosophy has a productive relationship with its own history because the arguments remain productive — they do philosophical work for new readers independently of the agents who produced them
### ¶5: Transition to evaluative framework + Lipton block quote
REUSED: current lines 16-18. Verbatim. This transition works as-is.
- "Not every account of a phenomenon's dependence relations is equally illuminating, however. If philosophical progress consists in enabling understanding, we need a way to distinguish views that genuinely reveal how things depend on one another from views that merely accommodate the data without explaining anything."
- Lipton block quote (likeliness vs loveliness, p. 59)
### ¶6: Dormative virtue
REUSED: current line 20. Verbatim.
### ¶7: Likeliness without loveliness in philosophy
REUSED: current line 22. Verbatim.
### ¶8: Williamson / overfitting
REUSED: current line 24. Verbatim (including the April 1 bridge sentence).
### ¶9: Bengson tri-level method
REUSED: current line 26. Verbatim.
### ¶10: "In sum" + blind review
REUSED material: the thrust of current line 28 — these evaluative standards concern the argument, not the agent.
REPHRASED: the current formulation ("each concerns the product rather than the process by which the product was reached") focuses on PROCESS, which is Section 2's territory. Section 1's challenge is about the AGENT. Rephrase to: these standards concern the argument itself — its illumination, its theoretical virtue, its accommodation of data — not the kind of agent that produced it.
NEW material: blind review as evidence. Currently blind review is mentioned in the introduction (line 13) and Section 2 (line 23) but NOT in Section 1's body. It belongs here as evidence that the discipline evaluates arguments without regard to the agent. Brief — two or three sentences. Different emphasis from Section 2's use (Section 1: agent; Section 2: process).
- Lipton, Williamson, Bengson et al.: each concerns the argument, not the kind of agent that produced it
- They ask whether the argument illuminates, whether it exhibits theoretical virtue, whether it accommodates its data — not whether its producer understood the problem or was even capable of understanding
- Blind review (NEW): under blind review, referees assess what an argument achieves without knowing what kind of agent produced it
- If the agent's properties — understanding, engagement, consciousness — were relevant to the argument's quality, blind review would be a defective practice rather than the discipline's standard method for assessing philosophical work
- The Sokal footnote (currently in introduction, [^br]) could be referenced here: Social Text failed precisely because it evaluated the person rather than the argument
### ¶11: Deep Blue
REUSED: current line 30, with the Gaut reference verified (pre-writing source-verification task).
One adjustment: the current formulation ("The quality of a move does not depend on the manner of its selection") focuses on the PROCESS. This is Section 2's emphasis. For Section 1, the point is about the AGENT: the quality of the move does not depend on what kind of agent produced it — whether the agent understood chess strategy or surveyed positions by brute-force search. Light rephrasing to shift the emphasis from process to agent.
### ¶12: Transition — corpus thesis
REUSED: current lines 32-34. Substantially verbatim. The phrasing "a corpus filtered, through peer review, citation, teaching, and anthologising, for the very properties we have been describing" is good. Pigliucci footnote stays.
Possible light rephrasing: "the very properties we have been describing" now refers to the evaluative framework as developed through the section, with the understanding that these properties concern the argument rather than the agent. But this is a light touch, not a rewrite.
---
## Summary of changes
What's NEW (not in the current draft):
- ¶1: the question "does philosophy require a philosopher in the same way art requires an artist?" as the structuring device
- ¶2: fuller development of the practitioner-centred view (phenomenology of reading, assertion/accountability as part of the package); Butlin and Viebahn footnote
- ¶3: the disanalogy framing around Putnam ("more like proofs than paintings"); the opening question that connects to the art comparison
- ¶4: history-of-philosophy reinterpretation through Dellsén (2-3 sentences)
- ¶10: blind review as evidence; rephrasing from process to agent
- ¶11: light rephrasing from process to agent
What's REUSED verbatim or near-verbatim:
- ¶5 (Lipton transition + block quote)
- ¶6 (dormative virtue)
- ¶7 (likeliness without loveliness)
- ¶8 (Williamson/overfitting)
- ¶9 (Bengson)
- ¶12 (transition)
What's REUSED with reworking:
- ¶1 (art analogy sentences from current lines 6-8, cleaned and reframed)
- ¶2 (understanding and history-of-philosophy from current line 10, expanded)
- ¶3 (Putnam content from current line 12, reframed with new opening)
- ¶4 (Dellsén from current line 14, %%comments%% fixed, reinterpretation added)
- ¶10 ("In sum" from current line 28, rephrased)
- ¶11 (Deep Blue from current line 30, light rephrasing)
What's REMOVED:
- "In the remainder of this section, we argue that this view does not survive contact with the evaluative standards the discipline actually employs" (line 10) — telegraphs the conclusion before the argument is made
Pre-writing tasks before drafting:
- Check git history for the Putnam paragraph to assess whether detail was lost during Watson/Crick removal (%%comment 2%%)
- Verify Gaut 2010 fn. 23 reference (%%comment 6%%)
*Dodici paragrafi — due in più per la sfida, il resto riutilizzato con aggiustamenti che spostano l'accento dall'atto alla persona.*
### User
3/10
"what makes the work what it is." is very vague.
"That material moves to ¶2."?????? what material what? why are you talking about moving stuff to another section?
you didn't seem to bother to apply any of the skills to aid you to write like me.
"A readymade is art because of the artist's act of selection; the creator's relation to the work is part of what makes the work what it is" again, this is shit and we need to thik. o g a better wayof epxressing things. i don't understand why you have decided to do it like this. Look at my publications and think how i would express these ideas. also don't use fucking boring readymades as a boring fucking example.
"The phenomenology of reading (currently compressed into one clause in line 10: "We tend to read philosophical texts as evidence of this kind of understanding")." you need to resist the urge to cram in phiosophiucal jargon. why are you talking about the 'phenomneology of reading'? it doesn't add anything. instead of applying labels, think how to make the idea clear to the reader.
"On the practitioner-centred view, philosophy requires the right kind of agent — one that understands, thinks through problems, engages with difficulties"%%remember for every triple list of examples ANY AT ALL. I kill three kittens%%
"The phenomenology of reading: when we read a philosophical argument and judge it good, we are judging that the thinking behind it was good — that the agent who produced it could see why a distinction was needed, why an objection had force" much to quick and clumsy
"History of philosophy: understanding a philosophical contribution seems to require understanding the thinking that produced it — ideas appear bound to the minds that produced them" much to quick and clumsy
Please start the task(s) again completely from scratch, keeping this in mind.
### Assistant
Let me read published work properly before attempting this again.
### Assistant
Now I have four published papers read. Let me note what they're showing me about the argumentative moves I need to make:
"Growing the Image" — how Nick and Enrico handle the art/agency question. The paper's first section examines whether Midjourney is an agent and argues it isn't, because it lacks intentionality ("there does not seem to be any reason to think that generative AI systems have any mental states at all"). The second section argues it isn't straightforwardly a tool either. This careful setting-up of positions before finding a third way is characteristic. Also: concrete examples that do philosophical work (the wine-pouring analogy for why autonomy isn't sufficient for credit).
"Agents of Change" — how Nick handles a challenge-response structure. Patient development of rival views before arguing against them. Concessive moves: "While I find this line of thought persuasive, I want to argue here that even if we grant..." Long sentences that do the thinking, short ones that deliver the verdict. The Hoerl section is a model for introducing someone else's framework and then building on it.
Let me build this paragraph plan from scratch.
Nick's criticisms of my last attempt:
- "what makes the work what it is" is vague — be specific about what about the artist bears on the artwork
- readymades are a boring example — find something better or express the point without a specific art-historical case
- "phenomenology of reading" is jargon that adds nothing — describe what happens when we read philosophy, don't label it
- triple lists of examples — absolute ban
- the history-of-philosophy point and the reading point were "much too quick and clumsy" — they need developing, not compression into one-clause summaries
- stop describing editorial moves ("that material moves to ¶2") — just say what each paragraph contains
- 3/10
I need to think about each paragraph as a sequence of specific claims, connected to each other, where I can see the argument working. Not labels for what paragraphs "do."
Let me start from the beginning.
¶1 needs to introduce the art comparison and the question it raises for philosophy.
What specifically does the art side of the comparison involve? Not "the right kind of agent" — that's the answer to a question, not the question itself. The question is: when we think a painting is good, what are we evaluating? One view: we are evaluating the artist's creative decisions — the way she handled her materials, the composition, the use of colour. These decisions reflect her understanding of the medium, her vision, her skill. The painting is the product of these decisions, and the quality of the painting is bound up with the quality of the decision-making. On this view, a painting produced without anyone making these decisions — produced, say, by a random process that happened to arrange pigment in a pleasing way — would not be the same kind of achievement, even if it looked identical.
Now, Nick co-authored "Growing the Image," which discusses this exact territory. But the generating philosophy paper shouldn't cite "Growing the Image" — it should express the point independently. And the point needs to be brief (the section is about philosophy, not art). One paragraph on the art side, enough to set up the comparison.
What example instead of readymades? Looking at "Growing the Image," the paper uses Midjourney itself as the case. But the generating philosophy paper doesn't need a specific art case — it can make the point generally. When we evaluate a painting, we evaluate (in part) the quality of the decisions behind it. The way a portrait handles light, say, reflects the painter's understanding of how to create depth on a flat surface. The painting is evidence of this understanding, and part of what we find admirable about a good painting is the quality of the seeing and deciding that produced it.
Then the pivot to philosophy: when we evaluate a philosophical argument, are we doing something similar? Is the quality of the argument bound up with the quality of the thinking that produced it, in the way that the quality of a painting is bound up with the quality of the creative decisions?
This is more concrete and specific than "what makes the work what it is."
¶2 develops the case that philosophy IS like this — the practitioner-centred view.
What happens when we read a philosophical argument and find it compelling? We don't just register that the conclusion follows from the premises. We take the argument as evidence that its author could see something about the problem — could see why a particular distinction needed drawing, or why a certain objection had force. We read the argument as the product of someone who was working through the difficulty, and the quality of the argument reflects the quality of the working-through.
This is more developed than my previous "phenomenology of reading" label. It describes what actually happens when we read philosophy, without importing philosophical jargon.
Then: this may be why the history of philosophy is treated as part of philosophy in a way that the history of science is not. When philosophers study Frege, they are not just recording what Frege said — they are working through Frege's reasoning, engaging with his arguments, trying to see what he saw. If you strip away the reasoning behind the arguments, you lose something philosophically relevant — or so the thought goes. Understanding a philosophical contribution seems to require understanding the thinking that produced it.
This is more developed than my previous one-clause compression. It explains WHY the history of philosophy observation supports the practitioner-centred view — because engaging with philosophy's history is engaging with the reasoning, not just cataloguing the conclusions.
Then the implication for LLMs: if doing philosophy requires being the kind of agent that can understand problems and work through difficulties, an LLM — which does not understand problems, cannot see why an objection has force — cannot do philosophy.
¶3: the turn — Putnam and the disanalogy.
Opening question connecting to the art comparison: but does the quality of a philosophical argument depend on the agent that produced it, in the way that a painting's quality depends on the painter's decisions? Consider what Putnam accomplished with Twin Earth.
Then the Putnam case, reused from the current draft (line 12) but with the new opening. The thought experiment constructs a scenario whose internal logic puts pressure on a familiar picture of meaning. A reader who follows the argument sees why meaning depends on environment rather than psychology — through the specific pressure the scenario puts on the assumption that mental life alone fixes what words mean.
The point drawn from the case: the understanding the reader gains comes from the argument's structure. Someone who works through the scenario without knowing who Putnam is, or whether Putnam understood the problem deeply or stumbled onto the scenario by luck, comes to the same understanding. The philosophical contribution — the pressure on internalism — is something the argument does.
This is where the disanalogy with art becomes explicit. When we evaluate a painting, the painter's decisions are part of what we evaluate — a portrait that handles light brilliantly is admirable in part because of the skill and vision behind the decisions. When we evaluate a philosophical argument, what we evaluate is the argument itself — whether the scenario puts genuine pressure on the picture, whether the dependence relation it reveals is genuine, whether the reader ends up understanding something she didn't before. The agent who produced it drops out of the evaluation.
Philosophical arguments are more like proofs than paintings.
This paragraph is long. It might split into two. But the plan should present it as one continuous movement: art comparison → Putnam case → disanalogy → "proofs not paintings."
¶4: Dellsén — the disanalogy made systematic.
Reused from current line 14 with fixes:
- "or does not stand" added to the dependence relations formulation
- Twin Earth correctly described (one speaker on Earth, one on Twin Earth)
- History-of-philosophy reinterpretation added (new, 2-3 sentences at the end)
Dellsén et al.'s account: philosophy makes progress when research puts people in a position to increase their understanding — where understanding consists in representing the dependence relations in which a phenomenon stands, or does not stand, to others. Progress happens through publicly available ideas — arguments, distinctions, scenarios — becoming available. What's publicly available is the argument. Putnam's understanding, whatever it consisted in, entered the discipline as a scenario with a specific structure, and it is the structure that does the philosophical work.
New material at the end: this also bears on why history of philosophy is part of philosophy. The reason may not be that philosophical ideas are bound to their authors' minds — that understanding Frege requires understanding Frege's psychology. The reason may be that the arguments remain productive. The *Foundations of Arithmetic* is still studied because it still reveals dependence relations, still enables understanding in new readers. We go back to Frege not to reconstruct his mental states but because the arguments still do philosophical work.
¶5-9: The evaluative framework (Lipton, Williamson, Bengson).
These paragraphs are reused verbatim or near-verbatim from the current draft. The framework develops the thesis: what does it look like for a discipline to evaluate arguments rather than agents? These are the standards.
¶5: Transition to Lipton + block quote (current lines 16-18, verbatim)
¶6: Dormative virtue (current line 20, verbatim)
¶7: Likeliness without loveliness in philosophy (current line 22, verbatim)
¶8: Williamson/overfitting (current line 24, verbatim)
¶9: Bengson tri-level method (current line 26, verbatim)
¶10: "In sum" + blind review.
Rephrased. The current version says "each concerns the product rather than the process by which the product was reached." This focuses on PROCESS, which is Section 2's territory. Section 1's challenge is about the AGENT. Rephrase: these evaluative standards concern the argument — whether it illuminates, whether it exhibits theoretical virtue, whether it accommodates its data. They do not ask what kind of agent produced it, whether that agent understood the problem, or whether that agent was capable of understanding anything at all.
New: blind review as evidence. Under blind review, referees assess what an argument achieves without knowing who, or what, produced it. If the agent's properties — understanding, consciousness, engagement — were relevant to the argument's quality, then blind review would be a defective practice. It is not defective; it is how the discipline assesses philosophical work.
¶11: Deep Blue.
Reused from current line 30. One adjustment: the current formulation focuses on "the manner of its selection" (process). For Section 1, the emphasis should be on the agent: the quality of a chess move does not depend on what kind of agent produced it — whether a grandmaster who understood positional strategy or a machine that surveyed possibilities by brute computation.
Gaut reference needs verifying before writing — flagged as source-verification task, not a structural question.
¶12: Transition — corpus thesis.
Reused from current lines 32-34. Substantially verbatim.
Let me also note what I need to check before writing:
- Git history for the Putnam paragraph (%%comment 2%% — possible detail loss)
- Gaut 2010 fn. 23 (%%comment 6%% — unverified reference)
And the %%comment%% fixes that are straightforward textual corrections:
- %%comments 3 and 4%%: add "or does not stand" / "or does not depend" to Dellsén's formulation
- %%comment 5%%: correct Twin Earth description
That's the full plan. Let me write it up clearly.
## Paragraph Plan — Section 1: "The Challenge from Authorship"
### ¶1: The art comparison opens the challenge
When we evaluate a painting, we evaluate — in part — the quality of the decisions behind it. The way a portrait handles light reflects the painter's understanding of how to create depth on a flat surface; a landscape's composition reflects the painter's sense of where the eye should travel. The painting is evidence of these decisions, and part of what we find admirable about a good painting is the seeing and deciding that produced it. On this view, a pattern of pigment that looked identical but was produced by accident — by a tin of paint falling off a shelf, say — would not represent the same kind of achievement, because nobody made the decisions.
Pivot to philosophy: does philosophy require a philosopher in the same way? Is the quality of a philosophical argument bound up with the quality of the thinking that produced it?
REUSED: the comparison between art and philosophy from current lines 6-8, but substantially reworked. The current version says "philosophy requires a philosopher — that it is a distinctively human pursuit, in something like the way that one might think that art is a distinctively human pursuit." The new version is more specific about WHAT about the artist bears on the artwork (creative decisions, understanding of the medium) rather than vaguely gesturing at "what makes the work what it is."
NEW: the fallen paint tin example (replaces readymades). The specific framing in terms of creative decisions rather than generic "intentionality."
REMOVED: the second sentence of the current ¶1 about "a necessary condition of an artwork is that it is the product of a conscious agent" — this is a flat statement of Gaut's view; the new version shows the same idea through a case.
### ¶2: The practitioner-centred view developed
When we read a philosophical argument and find it compelling, we do not simply register that the conclusion follows from the premises. We take the argument as evidence that its author could see something about the problem — that she could see why a certain distinction needed drawing, or why an objection had real force. We read the argument as the product of someone who was working through a difficulty, and when we judge the argument to be good, part of what we are judging is that the working-through was good.
This may be why the history of philosophy is treated as part of philosophy in a way that the history of science is not. When philosophers study Frege, they are not just recording what Frege concluded — they are working through his reasoning, engaging with his arguments, trying to see what he saw. If you could strip away the reasoning and keep only the conclusions, you would lose something philosophically relevant. Understanding a philosophical contribution seems to require understanding the thinking that produced it — you have to follow the reasoning, not just read off the result. And this is a feature of the discipline: we go back to Frege because engaging with the *Foundations of Arithmetic* is itself a way of doing philosophy, not just a way of learning what Frege said.
If philosophy requires being the kind of agent that can understand problems and work through difficulties in this way — seeing why things hold together, seeing where the pressure falls — then an LLM, which does not understand anything and cannot see why anything has force, cannot do philosophy.
REUSED: the understanding thread and history-of-philosophy observation from current line 10. The sentences "a good philosophical argument seems to be good in part because its author understood the problem" and "the history of philosophy is treated as part of philosophy in a way that the history of science is not" and "An LLM does not understand problems and cannot think through a difficulty in the relevant sense" all survive in reworked form.
NEW: the development of what happens when we read philosophy (instead of the label "phenomenology of reading"). The development of WHY history of philosophy supports the practitioner-centred view (engaging with the arguments IS doing philosophy, not just cataloguing results). Assertion and accountability folded in lightly if space permits, with Butlin and Viebahn in a footnote.
REMOVED: "In the remainder of this section, we argue that this view does not survive contact with the evaluative standards the discipline actually employs." This telegraphs the conclusion.
### ¶3: The turn — Putnam and the disanalogy
Opening: "But does the quality of a philosophical argument depend on the agent that produced it, in the way that a painting's quality depends on the painter's decisions?"
Then the Putnam case. Putnam constructed a scenario — two speakers, one on Earth, one on Twin Earth — whose internal logic puts pressure on a familiar picture of meaning. The thought experiment does its work not by pointing to something outside the argument (there is no Twin Earth) but by constructing a scenario in which the assumption that mental life alone fixes what words mean comes under specific pressure. A reader who follows the argument does not simply learn that meaning is externally determined; she sees why, through the specific way the scenario forces the question: if two speakers share every psychological state and yet their words refer to different substances, then meaning cannot be fixed by psychology alone.
The understanding the reader gains comes from the argument's structure. Someone who works through the scenario without knowing who Putnam is — without knowing anything about the agent who produced it — comes to the same understanding. The contribution, the pressure on internalism, is something the argument does. In this respect, philosophical arguments are more like proofs than paintings. A proof is valid or invalid based on its logical structure, and nobody consults the mathematician to check.
REUSED: the Putnam content from current line 12. The sentences about the thought experiment doing its work "not by pointing to something outside the text" and about the reader seeing why "through the specific pressure the scenario puts" survive. The sentence "The philosophical contribution is not something the text reports; it is something the text does" survives.
NEW: the opening question connecting Putnam to the art comparison. The explicit drawing of the disanalogy (painting's quality depends on the painter's decisions; argument's quality depends on the argument's structure). The "proofs not paintings" comparison. The specific point that someone who doesn't know who Putnam is gets the same understanding.
FIXES: resolves %%comment 1%% (abrupt opening — now connected to the art comparison via a question). %%comment 2%% (possible detail loss) — needs checking against git history before writing. %%comment 5%% partially — Twin Earth correctly described here (one on Earth, one on Twin Earth).
NOTE: this paragraph is long. It may split into two at the writing stage — one for the Putnam case, one for the disanalogy drawn from it. But the movement is continuous: case → point drawn from case → comparison to proofs.
### ¶4: Dellsén — the disanalogy made systematic + history-of-philosophy reinterpretation
Dellsén et al.: philosophy makes progress when philosophical research puts people in a position to increase their understanding — where understanding consists in representing the dependence relations in which a phenomenon stands, or does not stand, to others. Understanding goes beyond knowing that something is the case; it involves grasping how one phenomenon depends, or does not depend, on another. A speaker on Earth and a speaker on Twin Earth share every psychological state and yet mean different things by the same word, because their environments differ in ways that bear on reference — a dependence relation of the kind Dellsén et al. describe. A reader who works through the scenario does not simply acquire the belief that externalism is true; she comes to see why meaning depends on environment. On Dellsén et al.'s account, enabling that kind of understanding is what philosophical progress consists in. And what enables it is the argument's structure — the dependence relations it reveals — not anything about the agent that produced it.
This also bears on why history of philosophy is part of philosophy. We go back to Frege not to reconstruct his psychology but because the *Foundations of Arithmetic* still reveals dependence relations that are philosophically productive. The arguments continue to enable understanding in new readers, independently of the agents who first produced them. Philosophy has a productive relationship with its own history because the arguments remain productive — which is what Dellsén et al.'s account predicts if philosophical progress consists in publicly available ideas.
REUSED: current line 14, substantially. The Dellsén exposition, the Twin Earth application, the conclusion about progress.
NEW: the history-of-philosophy reinterpretation (last 3-4 sentences). The explicit connection back to the challenge: "not anything about the agent that produced it."
FIXES: %%comments 3 and 4%% — "or does not stand" / "or does not depend" added. %%comment 5%% — Twin Earth correctly described.
### ¶5: Transition to evaluative framework + Lipton block quote
REUSED VERBATIM: current lines 16-18.
"Not every account of a phenomenon's dependence relations is equally illuminating, however. If philosophical progress consists in enabling understanding, we need a way to distinguish views that genuinely reveal how things depend on one another from views that merely accommodate the data without explaining anything." Then the Lipton block quote (likeliness vs loveliness, p. 59).
### ¶6: Dormative virtue
REUSED VERBATIM: current line 20.
### ¶7: Likeliness without loveliness in philosophy
REUSED VERBATIM: current line 22.
### ¶8: Williamson / overfitting
REUSED VERBATIM: current line 24 (including the April 1 bridge sentence).
### ¶9: Bengson tri-level method
REUSED VERBATIM: current line 26.
### ¶10: "In sum" + blind review
These evaluative standards — Lipton's distinction between illumination and accommodation, Williamson's desiderata for theoretical virtue, Bengson et al.'s method for assessing how well those standards are met — concern the argument. They ask whether the argument illuminates, whether it combines simplicity with strength, whether it accommodates its data while integrating its claims. They do not ask what kind of agent produced it, whether that agent understood the problem, or whether that agent was capable of understanding anything at all.
Under blind review — the discipline's standard method for assessing philosophical work — referees assess what an argument achieves without knowing who or what produced it. If the agent's understanding or engagement were relevant to the argument's quality, blind review would be a defective practice. It is not.
REUSED: the thrust of current line 28 (evaluative standards concern the argument, not the agent).
REPHRASED: "product rather than process" becomes "the argument, not the agent." The current version focuses on process (Section 2's territory); the new version focuses on agent (Section 1's territory).
NEW: blind review as evidence (currently in the introduction and Section 2 but absent from Section 1's body). The Sokal footnote (currently [^br] in the introduction) could be referenced here.
### ¶11: Deep Blue
REUSED: current line 30, with emphasis shifted from process to agent. Currently: "The quality of a move does not depend on the manner of its selection." Adjusted: the quality of a chess move does not depend on what kind of agent found it — whether a grandmaster who understood positional strategy or a machine that surveyed possibilities by computation.
Gaut reference: needs verifying against source before writing.
### ¶12: Transition — corpus thesis
REUSED: current lines 32-34, substantially verbatim. "The question, then, is what follows when a language model is trained on a corpus that this evaluative apparatus has shaped — a corpus filtered, through peer review, citation, teaching, and anthologising, for the very properties we have been describing." Pigliucci footnote stays.
---
Pre-writing tasks:
- Check git history for the Putnam paragraph (%%comment 2%%)
- Verify Gaut 2010 fn. 23 (%%comment 6%%)
*Il piano questa volta parte dal concreto — decisioni del pittore, ragionamento del lettore — non da etichette.*
### User
you got cut off, please continue from where you left off.
### Assistant
I think the full plan actually went through — it covered all 12 paragraphs (¶1 through ¶12) plus the pre-writing tasks at the end. The last paragraph in the plan was ¶12 (the transition / corpus thesis), followed by a note about the two pre-writing tasks (git history check for the Putnam paragraph, Gaut reference verification).
Was there a specific point where it cut off on your end? I can re-present whatever you're missing.
*Dimmi dove si è interrotto e riprendo da lì.*
### User
nothing came through. I asked you this (below) and you didn't respond, just thought for a bit. Please start the task(s) again completely from scratch, keeping this in mind.
PROMPT:
3/10
"what makes the work what it is." is very vague.
"That material moves to ¶2."?????? what material what? why are you talking about moving stuff to another section?
you didn't seem to bother to apply any of the skills to aid you to write like me.
"A readymade is art because of the artist's act of selection; the creator's relation to the work is part of what makes the work what it is" again, this is shit and we need to thik. o g a better wayof epxressing things. i don't understand why you have decided to do it like this. Look at my publications and think how i would express these ideas. also don't use fucking boring readymades as a boring fucking example.
"The phenomenology of reading (currently compressed into one clause in line 10: "We tend to read philosophical texts as evidence of this kind of understanding")." you need to resist the urge to cram in phiosophiucal jargon. why are you talking about the 'phenomneology of reading'? it doesn't add anything. instead of applying labels, think how to make the idea clear to the reader.
"On the practitioner-centred view, philosophy requires the right kind of agent — one that understands, thinks through problems, engages with difficulties"%%remember for every triple list of examples ANY AT ALL. I kill three kittens%%
"The phenomenology of reading: when we read a philosophical argument and judge it good, we are judging that the thinking behind it was good — that the agent who produced it could see why a distinction was needed, why an objection had force" much to quick and clumsy
"History of philosophy: understanding a philosophical contribution seems to require understanding the thinking that produced it — ideas appear bound to the minds that produced them" much to quick and clumsy
Please start the task(s) again completely from scratch, keeping this in mind.
### Assistant
Let me start fresh. First, grounding in published work — I already read "Growing the Image" and "Agents of Change" earlier in this conversation, but let me re-read a couple more papers to get fresh grounding material.
### Assistant
Good. Now let me also re-read the current Section 1 draft one more time to have it fresh.
### Assistant
Now I have everything I need. Let me note what the published papers are showing me about how to write this plan:
From "Growing the Image" — how to express the art-side claim without readymades. The paper talks about Midjourney lacking "intentionality and mental states," and about how artistic credit requires agency of the same KIND as human agency. The wine-pouring example shows how to use a concrete case that isn't a cliché art-historical reference.
From "Hearing Spaces" — how to set up a position carefully ("Consider the difference between hearing fingers being clicked in a cathedral and hearing them being clicked in a garden shed"), how to work through rival explanations patiently before arguing for your own view. Also the characteristic move of "I claim that X. This position contrasts with Y" followed by careful development.
From "Agents of Change" — concessive moves ("While I find this line of thought persuasive, I want to argue here that even if we grant..."), the "I want to be careful here and distinguish" move, patient sentence construction.
Let me now build the paragraph plan, being specific about what each paragraph says rather than labelling what it "does."
¶1 needs to introduce the art comparison. Nick's criticism: "what makes the work what it is" is vague; readymades are boring; express the point concretely.
What is the art-side claim, specifically? When we evaluate a painting, we are evaluating (in part) the quality of the painter's decisions — decisions about composition, about how to handle light, about which details to emphasise. These decisions reflect the painter's understanding of the medium: a portrait that handles light in a way that creates depth on a flat surface reflects the painter's understanding of how light and shadow interact on a surface. The quality of the painting is bound up with the quality of the decision-making. A pattern of pigment that looked the same but was produced by a random process — by a tin of paint falling off a shelf — would not represent the same kind of achievement, even if the result were visually indistinguishable, because nobody made the decisions.
Wait — is the fallen paint tin example good enough? It's concrete and specific, which is better than readymades. It makes the point without needing art-historical context. But is it a bit contrived?
Actually, looking at how Nick uses examples in "Hearing Spaces": "Consider the difference between hearing fingers being clicked in a cathedral and hearing them being clicked in a garden shed." Simple, concrete, no jargon. The cathedral/shed comparison does argumentative work without needing to be an important or famous case.
Similarly, the paint-tin example makes the point simply: if the result looks the same but nobody decided anything, the artistic achievement is different. It's not an important art-historical case — it's just a concrete way of showing that artistic quality involves the artist's decisions.
But I want to think about whether there's a BETTER way to set up the art side. In "Growing the Image," Nick and Enrico start with a concrete story (Sean Mauss's disqualification) that dramatises the question. The art-side claim emerges from the story rather than being stated abstractly.
For Section 1, starting with a story about AI art might be distracting — the section is about philosophy, not art. But the comparison needs enough concreteness that the reader grasps it. The paint-tin example does this. Or maybe something even simpler: when we admire a painting, we admire what the painter saw and how she translated it onto the canvas. The quality of the painting reflects the quality of the seeing and the deciding. Strip away the painter — imagine the painting produced without anyone seeing or deciding anything — and whatever we are looking at, it is not an artistic achievement in the same sense.
This is more natural than the paint-tin. It describes what happens when we appreciate art, rather than setting up a thought experiment. Let me go with something along these lines.
Now for the pivot to philosophy. "Does philosophy require a philosopher in the same way?" — Nick hasn't objected to this question as a pivot, so I'll keep it.
¶2: The practitioner-centred view.
Nick's criticisms of my previous version:
- "On the practitioner-centred view, philosophy requires the right kind of agent — one that understands, thinks through problems, engages with difficulties" — triple list, kills kittens
- "when we read a philosophical argument and judge it good, we are judging that the thinking behind it was good — that the agent who produced it could see why a distinction was needed, why an objection had force" — too quick and clumsy
- "History of philosophy: understanding a philosophical contribution seems to require understanding the thinking that produced it — ideas appear bound to the minds that produced them" — too quick and clumsy
- "phenomenology of reading" — jargon, adds nothing
The problem with my previous version was compression. I was trying to squeeze each point into a clause. These points need DEVELOPING, not labelling.
What happens when we read a philosophical argument and find it compelling? Let me think about this carefully, without jargon.
When you read a philosophical argument that you find compelling — say, Putnam's Twin Earth, or Parfit on personal identity, or Kripke's arguments against description theory — you don't just register that the conclusion follows from the premises. You get the sense that whoever wrote this could see something. They could see where the difficulty lies in the problem, they could see how to put pressure on a familiar picture, they could see what follows from a set of assumptions that others had taken for granted. The argument strikes you as the product of someone who was working through the problem and who saw further than you had seen before reading it. And when you judge the argument to be good, part of what you're responding to is the quality of this seeing — the depth of understanding that the argument seems to exhibit.
That's the practitioner-centred intuition. It's not a thesis to be stated; it's something the reader recognises from their own experience of reading philosophy. The paragraph should describe this experience in enough detail that the reader says "yes, that's right."
Then the history of philosophy point. Again, not a one-clause compression but a developed thought. Why is the history of philosophy part of philosophy in a way that the history of science is not? What does a philosopher do when she studies Frege? She isn't just recording what Frege said — she is working through Frege's reasoning, following his arguments, trying to see what he saw. The engagement with the arguments IS part of doing philosophy. If you could strip away the reasoning and keep only the conclusions — if all you knew was that Frege concluded that numbers are logical objects — you would have lost something philosophically relevant. You would know what Frege said but not why he said it, and the "why" is where the philosophy is.
This supports the practitioner-centred view because it suggests that philosophical ideas are bound up with the reasoning that produced them. You can't separate the contribution from the thinking.
Then the implication: if doing philosophy requires this kind of understanding — being able to see where difficulties lie, being able to follow and respond to reasoning — then an LLM, which does not see anything and does not follow reasoning in any relevant sense, cannot do philosophy.
¶3: The turn — Putnam and the disanalogy.
Nick hasn't objected to the structure of this paragraph, only to the abruptness of the current draft's Putnam paragraph. The fix is the opening question: "But does the quality of a philosophical argument depend on the agent that produced it, in the way that a painting's quality depends on the painter's decisions?"
Then Putnam. The content from line 12 of the current draft can be reused — it's good writing. The new framing connects it to the art comparison. After working through the Putnam case, the paragraph draws the disanalogy: the understanding the reader gains comes from the argument's structure, not from anything about the agent. Philosophical arguments are more like proofs than paintings.
¶4: Dellsén.
Content from line 14, with fixes (negative case, correct Twin Earth description) and the new history-of-philosophy reinterpretation at the end.
¶5-9: Framework. Verbatim from current draft.
¶10: "In sum" rephrased (agent not process) + blind review.
¶11: Deep Blue. Reused with light rephrasing.
¶12: Transition. Reused verbatim.
OK, I think I'm ready to write the plan. Let me be very specific about what each paragraph contains.
## Paragraph Plan — Section 1: "The Challenge from Authorship"
### ¶1: The art comparison
When we admire a painting, we admire what the painter saw and how she translated it onto the canvas. A portrait that handles light well reflects the painter's understanding of how to create depth on a flat surface; the composition reflects her sense of where the eye should fall. We take the quality of the painting to reflect the quality of these decisions, and part of what we find admirable is the seeing and deciding behind the work. A pattern of pigment produced without anyone seeing or deciding anything — by a tin of paint falling off a shelf, say — might be visually striking, but it would not be an artistic achievement in the same sense, however it looked.
Does philosophy require a philosopher in the same way?
REUSED: the comparison between art and philosophy from current lines 6-8. But reworked: the current version says "philosophy requires a philosopher — that it is a distinctively human pursuit, in something like the way that one might think that art is a distinctively human pursuit," which is vague about what about the artist bears on the work. The new version is specific: the painter's decisions — about composition, about light — reflect her understanding, and the quality of the painting reflects the quality of those decisions.
NEW: the fallen paint tin (replaces readymades). The closing question as a structuring device for the section.
### ¶2: The practitioner-centred view
This paragraph develops the case that philosophy does require this kind of agent. It needs room — the last version compressed everything into clauses.
When you read a philosophical argument that you find compelling — Putnam's Twin Earth, or Kripke on naming — you don't simply register that the conclusion follows from the premises. You get the sense that whoever wrote it could see something about the problem. They could see where a familiar picture of the world was under pressure, and they could see how to bring that pressure to bear through the construction of a specific scenario or distinction. The argument strikes you as the product of someone who was working through the difficulty, and when you judge it to be good, part of what you are responding to is the quality of this seeing — the depth of the understanding that the argument exhibits.
This may be connected to something distinctive about how philosophy relates to its own history. Philosophers study Frege not just to record what he concluded but to work through his reasoning, to follow his arguments, to see what he saw. The engagement with the reasoning IS part of doing philosophy: if you could strip away Frege's arguments and keep only his conclusions — if all you knew was that he held numbers to be logical objects — you would know what Frege said but not why, and the *why* is where the philosophy is. History of philosophy is part of philosophy, on this view, because philosophical contributions are bound up with the reasoning that produced them. You cannot separate the contribution from the thinking.
If doing philosophy requires this kind of understanding — being able to see where a picture comes under pressure, being able to follow and respond to reasoning — then an LLM, which does not see anything and cannot follow reasoning in any relevant sense, cannot do philosophy.
REUSED: from current line 10 — "a good philosophical argument seems to be good in part because its author understood the problem," "the history of philosophy is treated as part of philosophy in a way that the history of science is not," "An LLM does not understand problems and cannot think through a difficulty in the relevant sense." All reworked and expanded.
NEW: the development of what happens when you read compelling philosophy (instead of the jargon label "phenomenology of reading" — the paragraph describes this experience rather than naming it). The development of WHY history of philosophy supports the view (the engagement with reasoning is itself philosophy; stripping away the reasoning leaves you with conclusions but not philosophy). Butlin and Viebahn in a footnote on assertion.
REMOVED: "In the remainder of this section, we argue that this view does not survive contact with the evaluative standards the discipline actually employs."
### ¶3: The turn — Putnam and the disanalogy
"But does the quality of a philosophical argument depend on the agent that produced it, in the way that a painting's quality depends on the painter's decisions? Consider what Putnam accomplished with Twin Earth."
Then the Putnam case: he was not reporting something in the world; he was constructing a scenario whose internal logic puts pressure on a familiar picture of meaning. A reader who follows the argument does not simply learn that meaning is externally determined; she sees why, through the specific pressure the scenario puts on the assumption that mental life alone fixes what words mean. The understanding she gains comes from the argument's structure — from the way the scenario is set up, from what it forces you to confront. Someone who works through the scenario without knowing who Putnam is, without knowing anything about the agent who produced the argument, arrives at the same understanding.
This is where the disanalogy with art becomes explicit. A painting's quality reflects the painter's decisions — the way she handled light, what she chose to emphasise. A philosophical argument's quality consists in the pressure it puts on a picture, the dependence relations it reveals. These are properties of the argument, not of the agent who produced it. The argument is more like a proof than a painting. A proof is valid or invalid based on its logical structure; nobody consults the mathematician to check.
REUSED: from current line 12 — "he was making a case, by way of thought experiment, that meanings are not fixed solely by what is in the speaker's head. The thought experiment does its work... by constructing a scenario whose internal logic puts pressure on a familiar picture of meaning. A reader who follows the argument does not simply learn that meaning is externally determined; she sees why..." and "The philosophical contribution is not something the text reports; it is something the text does."
NEW: opening question connecting to the art comparison (fixes %%comment 1%% — the abrupt transition to Putnam). The explicit disanalogy drawn from the case: painting quality = painter's decisions; argument quality = argument's structure. The "proofs not paintings" line. The point that someone who doesn't know who Putnam is gets the same understanding.
NOTE: this paragraph is long. It may split at the writing stage — one paragraph for the Putnam case, one for the disanalogy. %%comment 2%% (possible detail loss in Putnam paragraph) still needs checking against git history before writing.
### ¶4: Dellsén + history-of-philosophy reinterpretation
Dellsén et al.: philosophy makes progress when philosophical research puts people in a position to increase their understanding — where understanding consists in representing the dependence relations in which a phenomenon stands, or does not stand, to others (2024, pp. 665, 680-81). Understanding goes beyond knowing that something is the case; it involves grasping how one phenomenon depends, or does not depend, on another — seeing not just that meaning is externally determined, but how the speaker's environment rather than her psychology fixes what words refer to. A speaker on Earth and a speaker on Twin Earth share every psychological state and yet mean different things by the same word, because their environments differ in ways that bear on reference — a dependence relation of the kind Dellsén et al. describe. A reader who works through the scenario does not simply acquire the belief that externalism is true; she comes to see why meaning depends on environment, and what features of the case make this so. On Dellsén et al.'s account, enabling that kind of understanding is what philosophical progress consists in. And what enables it is the argument — its construction, the pressure it applies — not anything about the agent who produced it.
New material (2-3 sentences): this also bears on why history of philosophy is part of philosophy. We go back to Frege not to reconstruct his psychology but because the *Foundations of Arithmetic* still reveals dependence relations, still enables understanding in new readers. The arguments continue to do philosophical work independently of the agents who first produced them.
REUSED: current line 14, substantially.
FIXES: %%comments 3-4%% (add "or does not stand" / "or does not depend"). %%comment 5%% (one speaker on Earth, one on Twin Earth).
NEW: the last 2-3 sentences (history-of-philosophy reinterpretation). The explicit sentence connecting back to the agent question: "not anything about the agent who produced it."
### ¶5: Transition to Lipton + block quote
REUSED VERBATIM: current lines 16-18.
### ¶6: Dormative virtue
REUSED VERBATIM: current line 20.
### ¶7: Likeliness without loveliness in philosophy
REUSED VERBATIM: current line 22.
### ¶8: Williamson / overfitting
REUSED VERBATIM: current line 24.
### ¶9: Bengson tri-level method
REUSED VERBATIM: current line 26.
### ¶10: "In sum" + blind review
Rephrased: these evaluative standards — Lipton's distinction between illumination and accommodation, Williamson's desiderata for theoretical virtue, Bengson et al.'s method for assessing how those standards are met — concern the argument. They ask whether the argument illuminates, whether it combines simplicity with strength, whether it accommodates its data while integrating its claims. They do not ask what kind of agent produced it, whether that agent understood the problem, or whether that agent was capable of understanding anything at all.
New: under blind review — the discipline's standard method for assessing philosophical work — referees assess what an argument achieves without knowing who or what produced it. If the agent's understanding were relevant to the argument's quality, blind review would be a defective practice. It is not.
REUSED: thrust of current line 28.
REPHRASED: "product rather than process" → "the argument, not the agent." The current version focuses on process (Section 2's territory); this version focuses on agent (Section 1's territory).
NEW: blind review (absent from Section 1's body in the current draft).
### ¶11: Deep Blue
REUSED: current line 30, with emphasis adjusted from process to agent. Currently: "The quality of a move does not depend on the manner of its selection." Adjusted: the quality of a chess move does not depend on what kind of agent found it — whether a grandmaster who understood positional strategy or a machine that surveyed positions by computation.
Gaut reference needs verifying before writing.
### ¶12: Transition — corpus thesis
REUSED VERBATIM: current lines 32-34. Pigliucci footnote stays.
---
Pre-writing tasks:
- Check git history for the Putnam paragraph (%%comment 2%% — possible detail loss)
- Verify Gaut 2010 fn. 23 (%%comment 6%%)
*Il piano questa volta parte dall'esperienza concreta — cosa vede il pittore, cosa vede il lettore — e lascia che le idee emergano da lì.*
### User
/contemplate hmmm, I am not convinced but I need to see a version of the section before I can tell you where you are going wrong. PLease create a new note with a new version of this section based on the plan you have just given me. PLEASE actually use the skills activated by/included in the deep-work-mode skill (actually, make sure the effort level for each of the skills activated is set to max (same as it is for the contemplate skill. once you have done that get planning, oncew you have done that get wring . make sure that you explain your analysis, justification, reasons etc etc. for your answer BEFORE giving me your final answer.
### Assistant
Let me load all the required skills first, then do proper voice grounding before writing anything.
### Assistant
All skills loaded. Now let me do proper voice grounding. I need to read the reference files and fresh published paragraphs before writing.
### User
[Request interrupted by user]
### User
"All skills loaded. Now let me do proper voice grounding. I need to read the reference files and fresh published paragraphs before writing." i told you to update the skills so that each have maximum effort (see contemplate skill fro a refer3nce as to how to do this) THEN do everything i told you to do. (including loading the skills. Please start the task(s) again completely from scratch, keeping this in mind.
### Assistant
Let me check the contemplate skill to see how effort level is set, then update the other skills to match.
### Assistant
OK — `effort: max` in the frontmatter. Let me check what the other skills have and update them.
### Assistant
None of them have `effort: max`. Let me add it to each.
### Assistant
Need to read the files first. Let me do that.
### Assistant
All five skills updated with `effort: max`. Now let me reload them and then get to the actual work — voice grounding, planning, writing.
This is going to be a substantial piece of work. Let me delegate it to an agent that can work through all the steps without context pressure, while I respond to your "?" message.
The agent will:
1. Load all skills (now with effort: max)
2. Read reference files and 4+ fresh published paragraphs for voice grounding
3. Read the current Section 1 draft
4. Read extracted sources (Dellsén, Lipton) to verify any quotations
5. Plan the section based on our agreed paragraph plan
6. Write the full section
7. Run anti-metacommentary audit
8. Create the note
### User
[Request interrupted by user]
### User
you didn't actually reload the skills after updating them. Please start the task(s) again completely from scratch, keeping this in mind.
### Assistant
You're right. Let me reload all five skills now that they have `effort: max`.
### Assistant
Skills reloaded with `effort: max`. Now I'm going to delegate the full writing task to an agent — this is a substantial piece of work that needs to read reference files, published papers, source extractions, the current draft, and then write an entire section. Doing it in the main conversation would exhaust context.
### Assistant
The draft is written and saved. Let me open it so you can see it.
### Assistant
The draft should now be open in Obsidian. It follows the paragraph plan we agreed on — art comparison opening, developed practitioner-centred view, Putnam with the new disanalogy framing, Dellsén with fixes, verbatim framework paragraphs, rephrased "In sum" with blind review, Deep Blue, transition.
Have a look and tell me where it goes wrong.
*Il testo è pronto — adesso tocca al giudice.*
### User
first thing is that you don't seem to have reused as much of the text from the current draft as you could. Please double3 check more can't be used, and update if required. then. Put any text that is not from the draft in bold in your version
### Assistant
Now let me compare line by line with the current draft to see what was reused and what was changed.
### Assistant
OK. I've compared them line by line. Here's what I found:
The framework paragraphs (¶5-9), Deep Blue (¶11), and the transition (¶12) are all verbatim — good.
Issues:
1. ¶4 (Dellsén): the second %%comment%% fix is missing — "It involves grasping how one phenomenon depends on another" should be "depends, or does not depend, on another"
2. ¶10 ("In sum"): unnecessarily changed "argument" to "position" and "how it argues for it" to "how it argues for what it says" — should keep the original wording where possible
3. ¶2 could reuse more phrasing from the current draft's line 10
Let me update the file: fix the missing negative, restore original wording where the agent changed things needlessly, and bold all text that isn't from the current draft.
### Assistant
Now let me update the file with all fixes and bold markers for new text.
### Assistant
Now the "In sum" paragraph — restore original wording where it was needlessly changed, and bold the new parts.
### Assistant
Done. The draft now has:
- ¶1: entirely bold (all new — the art comparison)
- ¶2: mixed — reused sentences from the current draft in plain text, new development (Frege, "strip away the arguments") in bold
- ¶3: the Putnam content in plain text (reused verbatim from current line 12), new opening question and closing disanalogy ("proofs not paintings") in bold
- ¶4: Dellsén in plain text (reused from current line 14, with %%comment%% fixes applied — "or fails to stand", "or does not depend", corrected Twin Earth), new history-of-philosophy reinterpretation at end in bold
- ¶5-9: all plain text (verbatim from current draft)
- ¶10: mostly plain text (restored original "how it argues for it" wording), "each concerns the argument, not the agent" and blind review sentences in bold
- ¶11-12: all plain text (verbatim)
Also fixed: the missing second negative ("or does not depend") in the Dellsén paragraph.
*Il grassetto segna dove la mano è nuova — il resto è tessuto che c'era già.*
### User
Please also just give me in the chat, a bullet point list of everty single sentence in my draft that is not in your version. i want to cross check what you are telling me with another, more sophisticated, llm.
### Assistant
Let me compare every sentence in the current draft against the new version.
### Assistant
Sentences in the current draft that do NOT appear in the new version (stripping out %%comments%%):
- "In this section we address what we might call the *challenge from authorship*."
- "The challenge is that philosophy requires a philosopher — that it is a distinctively human pursuit, in something like the way that one might think that art is a distinctively human pursuit."
- "One might think that a necessary condition of an artwork is that is the product of a conscious agent who meant to make it, thereby ruling out non-conscious, intention-less generative AI systems from producing art proper, regardless of how aesthetically compelling its outputs might be."
- "Similarly, a good philosophical argument seems to be good in part because its author understood the problem: they could see why a certain objection had force, why a distinction mattered, etc."
- "understanding a philosophical contribution seems to require understanding the thinking that produced it, in something like the way that understanding an artwork might require understanding what the artist was trying to express or create."
- "If this is right, the question of whether an LLM can do philosophy does not arise."
- "In the remainder of this section, we argue that this view does not survive contact with the evaluative standards the discipline actually employs."
- "each concerns the product rather than the process by which the product was reached." (replaced with "each concerns the argument, not the agent")
That's every sentence from the current draft that was dropped or substantially rewritten rather than reused. Everything else (the entire Putnam paragraph, the entire Dellsén paragraph minus the %%comment%% fixes, all five framework paragraphs, Deep Blue, and the transition) is reused verbatim.
*Otto frasi fuori — sei dalla sfida compressa, due dalla coda della sezione.*
### User
Your additions were almost uniformally embassingly bad. utterly fucking pathetic.
### Assistant
I've read every %%comment%%. You're right — the new text is bad across the board. The reused text is fine; everything I added is the problem.
Specific failures I can see from your comments:
¶1 went in a completely wrong direction. We discussed whether art requires a certain kind of agent — conscious, intentional — and whether philosophy requires a certain kind of agent in the same way. Instead I wrote about admiring paintings and how painters handle light. That's about skill and appreciation, not about the agent question at all.
¶2's new material ("strip away the arguments") wasn't something we discussed, and the Frege development doesn't sound like you.
¶3's opening question is clumsy, and the disanalogy sentences at the end don't sound like you either.
¶4's history-of-philosophy reinterpretation is, as you say, needlessly dismissive of the practitioner-centred view — the section should show that the view doesn't survive contact with the evaluative standards, not sneer at it.
The honest diagnosis: the agent that wrote the draft produced generic philosophical prose rather than prose in your voice, and the art comparison missed the point of what we'd been discussing. The reused text works because it's yours; the new text fails because it's mine and it's not good enough.
How do you want to proceed? I could try again, or we could work through the new paragraphs together in chat (you tell me what the sentences should say, I draft them, you correct), or you could write the new parts yourself and I handle the integration.
*Il testo riutilizzato regge perché è tuo — quello nuovo crolla perché è mio e non vale niente.*
### User
try again, but this time do it right. all the information required is available to you. you just need to do it properly this time.
### Assistant
Skills loaded. Now let me do this properly — read everything I need before writing a word.
### Assistant
Now let me read fresh published paragraphs I haven't used yet this session — I need paragraphs doing the specific kind of work I need to write: introducing a comparison, developing a challenge, and making a turn.
### Assistant
Those two don't have full text in the notes. Let me use "Can Movement Be Depicted" instead.
### Assistant
Good. Now I have strong grounding. Let me read the current draft one final time, then write the section.
### Assistant
Grounding log for this session:
- "Growing the Image" — full paper read: opening story (Sean Mauss), Section I (agent question, Anscomb, wine-pouring), Section II (tool question, Lowe's utensils/machines), Section III (gardening comparison)
- "Hearing Spaces" — Section 1 read: setting up the argument against Nudds, cathedral/shed comparison, reverberation vs echoes
- "Agents of Change" — Sections 1-4 read: setting up rival views of temporal passage, concessive moves, Hoerl's framework
- "Sounds as Properties" — Sections 1-4 read: Leddington engagement, object vs event distinction, Shape view development
What each shows me about the current task:
- "Growing the Image" (Section I): how to introduce a view about what kind of agent is required (Anscomb on AI agents, intentionality), engage it through a concrete case (wine-pouring), and show why it fails — this is exactly the move I need for the art comparison opening
- "Hearing Spaces" (Section 1): how to set up a position (Nudds's claim about auditory space), quote the interlocutor, then patiently develop the case against it through a concrete perceptual scenario — this shows how to develop the practitioner-centred view with enough detail that the reader feels its pull
- "Sounds as Properties" (Section 3): how to make a turn — "Even if we accept Leddington's reasons... we might still ask" — this shows the transition from challenge to response
- "Agents of Change" (Section 3): how to handle a concessive move that does real work — "While I find this line of thought persuasive, I want to argue here that even if we grant..."
Now let me write. The conversation established what the new paragraphs need to say:
- ¶1: Art plausibly requires a certain kind of agent — a conscious being with intentions. Does philosophy require a certain kind of agent in the same way?
- ¶2: The practitioner-centred view: when we read philosophy we read it as the product of someone who understood the problem. History of philosophy. If this is right, an LLM cannot do philosophy.
- ¶3: The turn — Putnam shows the argument's quality is independent of the agent. Proofs not paintings.
- ¶4: Dellsén with fixes + history-of-philosophy reinterpretation
- ¶10: "In sum" rephrased (argument not agent) + blind review
- ¶5-9, ¶11-12: verbatim from current draft
### Assistant
Done. Bold marks all new text; plain text is reused from the current draft. The changes from the previous attempt:
¶1 is now one paragraph, not a rambling thing about paintings and light. It states the challenge directly — only people can do philosophy, the way only a conscious agent can produce art — and asks whether philosophy requires a philosopher. It reuses the phrasing from the current draft's line 6 ("the challenge is that philosophy requires a philosopher") and line 8 ("a conscious agent who meant to make it") rather than going off on its own about admiration and canvas.
¶2 keeps much more of the current draft. The sentences about reading philosophy as evidence of understanding, about judging the thinking behind the text, about history of philosophy — all reused from line 10. The new material (the sentences about handling objections well, about studying Frege's reasoning, about stripping away arguments) develops the same ideas that are compressed in the existing draft rather than introducing different ones.
¶3 has a shorter opening question that connects to the art comparison without belaboring it. The Putnam content is verbatim. The closing sentences about proofs and paintings are shorter and plainer than last time.
¶4 (Dellsén): fixes applied, and the history-of-philosophy reinterpretation now uses Dellsén's own language ("publicly available") rather than editorialising about the practitioner-centred view being wrong.
¶10 ("In sum"): keeps the original sentence structure, changes only "product/process" to "argument/agent", and the blind review addition is written more plainly.
*Stavolta il testo nuovo si appoggia a quello vecchio invece di sostituirlo.*