# Changes Needed: Generating Philosophy — Text-Internal Evaluation (Sections 0–3)
I have worked through the transcript line by line, decoded the speech-to-text errors, and mapped every substantive philosophical point onto the current draft. Below is the full list, with reasoning for each.
A note on epistemic status: where I discuss what authors argue ([[Edouard Machery|Machery]], [[Tamar Szabó Gendler|Gendler]], [[Kendall Walton|Walton]]), I am working from training data, not from extracted source texts. I have flagged this where it matters. The vault has a catalogue note for Walton's "[[Categories of Art by Walton|Categories of Art]]" but no extracted text; Machery is mentioned in the [[Notes/Metaphilosophy Landscape|Metaphilosophy Landscape]] note but without a source PDF.
---
## 1. Structural error in Section 2 opening
The first sentence of Section 2 reads:
> I want to argue that the objections to LLM philosophy presented in the previous section do not apply to the kind of philosophical work that Williamson describes.
This is wrong. The objections are not presented in Section 1; they are presented *in* Section 2 itself. Section 1 establishes the text-internal thesis (Watson/Crick, Lipton, blind review). This opening sentence is a holdover from an earlier draft structure where the objections came first.
Options:
(A) Rewrite the opening to introduce the objections freshly: "Two prominent arguments against LLM philosophy target the reasoning process. If these arguments work, they would show that LLMs cannot produce texts exhibiting the properties Section 1 identified. I want to argue that neither succeeds."
(B) A softer rewrite: "Section 1 argued that philosophical evaluation concerns properties of texts. Two recent arguments challenge whether LLMs can produce texts exhibiting these properties. I present both arguments and explain why neither undermines the text-internal thesis."
(C) Cut the framing entirely and open with the [[Luciano Floridi|Floridi]] exposition directly. The reader already knows from the Introduction's roadmap what Section 2 is doing.
My inclination is toward (A) or (B). The opening needs to do two things: signal that the section presents objections, and set up the response. Option (C) is cleaner prose but loses the dialectical signposting, and at ~4,000 words the reader benefits from knowing where the argument is going.
---
## 2. Sharpen the Floridi/Zahavy asymmetry
The co-author's comment ("the only objection in which the connection to the process is really relevant because it's the idea that the process put constraints on the text") identifies a distinction the paper should make explicit.
The co-author's point: Floridi's objection is about process (LLMs generate without evaluating), but process-is-different does not entail output-is-worse. This is already argued in the draft. [[Tom Zahavy|Zahavy]]'s objection is different in kind: it's about process *constraining* what can be produced (the E→A Jump means LLMs *cannot produce* certain kinds of contributions). This is the harder objection because it's about capability, not mechanism.
The current draft treats them somewhat symmetrically (each gets ~300–350 words). The revision should:
(A) Make explicit that these are asymmetric challenges: Floridi challenges the *mechanism* (which the text-internal thesis already addresses — mechanism is irrelevant to evaluation); Zahavy challenges the *capability* (which requires a substantive response about what philosophical novelty consists in).
(B) Possibly compress the Floridi exposition slightly and expand the transition to Zahavy, marking the escalation: "Zahavy's objection cuts differently" (the draft already has this line — good — but the asymmetry could be named more precisely).
(C) The Floridi self-quotation ("if an AI can generate the same explanatory hypothesis a human would, does it matter that the process was different?") already concedes the point. The paper could lean harder on this: Floridi *himself* acknowledges that process may not matter for content evaluation. The question is whether the process constrains what content can appear — and that's Zahavy's question, not Floridi's.
---
## 3. Address Bayesianism
Nick says he "hasn't finished" dealing with Bayesianism. The current draft has a [[Peter Lipton|Lipton]] quote about Bayesianism vs IBE at the end of Section 3 (the squash game analogy), but the Bayesian objection is never named explicitly.
The Bayesian objection would be something like: "The real description of what LLMs do is Bayesian updating on token probabilities. This is not inference to the best explanation; it's statistical prediction. So [[Timothy Williamson|Williamson]]'s IBE framework doesn't apply to LLM outputs."
The response (via Lipton): these operate at different levels of description. The mechanical (Bayesian/stochastic) description of the LLM is correct but doesn't settle whether the outputs meet philosophical standards, just as the mechanical description of a squash ball's trajectory doesn't settle whether the player has good technique.
Options for incorporating:
(A) Add a setup paragraph before the Lipton quote that names the Bayesian objection explicitly. Something like: "One might object that the description just offered smuggles in a level confusion. What LLMs actually do, the objection goes, is Bayesian updating — statistical prediction, not inference to the best explanation. But Lipton makes a point about the relationship between different levels of description that bears on this..."
(B) Expand the Lipton quote passage into a self-contained response to the Bayesian objection, making the levels-of-description point the explicit climax of Section 3.
(C) Add a footnote addressing the Bayesian challenge, if the paper doesn't want to give it full-paragraph treatment.
My inclination: (A). The Lipton passage is already the final substantive paragraph of Section 3. Adding a setup that names the objection makes the Lipton quote do more work and provides a stronger conclusion to the section.
---
## 4. Add the intuitions objection (substantial new material)
The co-author identifies this as "the most substantial objection." It is entirely absent from the current draft. The argument: philosophy relies on intuitions (pre-theoretical judgments, intellectual seemings), and LLMs lack this epistemic capacity.
This needs to be a significant addition, not a footnote. If the paper is targeting *Mind* or *Philosophical Review*, reviewers will notice the absence.
The co-author mentions Machery's book on intuitions in philosophy (I'm speculating this is *Philosophy Within Its Proper Bounds*, 2017, based on training data — the vault has a mention in the Metaphilosophy Landscape note but no extracted source text, so I cannot verify his specific arguments).
Where to place it:
(A) Add as a third objection in Section 2 (alongside Floridi and Zahavy), then respond in Section 3. This keeps the "present objections → respond" structure clean.
(B) Add as a separate subsection after Section 3 — a new challenge that emerges even after the Floridi/Zahavy objections are answered. This has the advantage of escalating the dialectical tension: "Even if we grant that process doesn't constrain outputs, and that philosophical novelty is conceptual reconfiguration, one might still worry that philosophy requires a kind of epistemic access — intuition — that LLMs lack."
(C) Integrate it into the scope-limitation discussion (item 5 below) as part of a broader treatment of experience-dependent philosophy.
My inclination: (B). The intuitions objection is structurally different from Floridi and Zahavy (those are about reasoning processes; this is about epistemic access). Presenting it as a further challenge after the process objections are answered creates a stronger dialectical arc. The paper would then have:
- Section 2: Process objections (Floridi, Zahavy)
- Section 3: Response to process objections (text-internal evaluation, armchair abduction)
- Section 3.5 or 4: Intuitions and experience objection → partial response
But this depends on how Section 4 (demonstration) is handled. If the paper is expanding to 6,000–8,000 words, there's room for a new section.
The content should include:
(i) The objection: philosophy uses intuitions as data (the intuition that p-zombies are conceivable), as checks on theory (the theory must accord with intuitions), and as epistemic access to certain concepts (feeling imaginative resistance). LLMs lack this.
(ii) The co-author's two responses:
- Philosophy works at a level of abstraction above raw intuitions — it manipulates concepts, not intuitions directly, and the concepts are in the text
- The corpus as a "repository of second-hand experience" (not just arguments, but diaries, autobiographies, art criticism, phenomenological descriptions)
(iii) The hedge: this doesn't fully solve the problem for all areas of philosophy, but it shows the situation is more nuanced than the objection suggests.
---
## 5. Address the scope limitation (philosophy of mind, consciousness, aesthetics)
The co-author raises the worry: the paper's argument may apply only to certain areas of philosophy (language, metaphysics, modal logic — the more "mathematical-like" areas) and not to others where experience or acquaintance matters (philosophy of mind, consciousness, aesthetics).
This is closely connected to item 4. The paper should explicitly acknowledge this scope issue.
Options:
(A) Restrict the paper's scope explicitly: "The argument presented here applies most naturally to areas of philosophy where the relevant moves are conceptual and argumentative — metaphysics, philosophy of language, epistemology, ethics insofar as it concerns the structure of arguments. Whether it extends to philosophy of mind or phenomenology, where the subject matter seems to require acquaintance, is a further question."
(B) Partially restrict but then push back: acknowledge the worry, provide the "second-hand experience" response, and argue that the range of philosophy accessible to LLMs is larger than skeptics assume.
(C) Don't restrict at all — argue that even philosophy of consciousness works with concepts at a textual level of abstraction, so the argument applies generally.
The transcript favours (B): "not trying to present what you've just said about stuff being in the Corpus as being 100% a solution to this objection but at least showing we can probably get quite a lot further than you would imagine."
The examples from the transcript:
- Walton's "Categories of Art": to understand how art-historical categories shape aesthetic engagement, you might need to have viewed artworks under different categorical descriptions. But art criticism contains rich descriptions of this. The LLM has access to what it's like to perceive a Guernica-like painting under the category "Cubism" vs "Renaissance perspective" through the critical literature.
- Gendler's imaginative resistance: the concept requires *feeling* resistance to certain fictions. But descriptions of this resistance appear in literary criticism, reader-response criticism, and philosophical texts that articulate the phenomenon.
The framing: "The training corpus is not merely a repository of concepts and arguments but also a repository of second-hand experience." This is the co-author's phrase and it's good.
---
## 6. Write Section 4 (demonstration / prompting)
This section is promised in the Introduction's roadmap ("Section 4 considers what a demonstration would look like") but does not exist in the current draft. The transcript discusses what should go here:
(A) The autonomy question: "Can do philosophy" might mean collaborative (human-guided) or autonomous. The less the human prompter contributes, the stronger the claim.
(B) The prompting distinction: one-shot prompts (posing a problem) vs conversational/iterative prompting (collaborative development). The co-author's example: "Please solve the mind-body problem" (one-shot, problem-oriented) vs an evolving conversation with adjustments and refinements (iterative, solution-oriented).
(C) The "good continuation" framework: connect to semiotic physics / Janus. Good prompting means writing prompts whose most probable continuation exhibits philosophical quality. Two senses of good continuation:
- Shallow: question → answer
- Deep: rough gesture toward solution → more robust, developed solution
(D) The self-proving argument: "If you think this paper is good philosophy, there's your proof" (since it was co-written with an LLM).
(E) Discussion of what a proper demonstration would require.
I'd suggest structuring Section 4 around the autonomy continuum:
1. Open with the challenge: "If LLMs can do philosophy, show us."
2. Distinguish modes of LLM philosophy: autonomous (one-shot) vs collaborative (conversational)
3. Argue that both are interesting, but in different ways
4. Discuss what good prompting consists in (the "good continuation" framework)
5. Note the self-proving meta-argument
6. Close with the Hitchhiker's Guide callback
---
## 7. Create Hitchhiker's Guide payoff
The epigraph currently has no payoff later in the paper. The transcript identifies this as a missed opportunity: "If we did talk about prompting at some point, it would actually fit very nicely with the Hitchhiker's Guide Galaxy joke at the beginning. It would be nice to have a sort of payoff."
The connection: Deep Thought gave a useless answer ("42") because the *question* was wrong. The lesson for LLM philosophy: the quality of the philosophical output depends on the quality of the prompt. Asking "solve the mind-body problem" is like asking for "the Answer to the Ultimate Question of Life, the Universe, and Everything" — too vague, too unstructured to elicit good philosophy. Good philosophical prompting requires providing the right materials, the right level of specificity, the right kind of question.
This payoff should appear in Section 4, probably near the end.
---
## 8. Problem-oriented vs solution-oriented prompts
From the transcript: a distinction between:
- Problem-oriented prompts: pose a problem; the LLM generates a solution
- Solution-oriented prompts: sketch the beginning of a solution; the LLM develops it further
The co-author develops this using "good continuation": for a problem-oriented prompt, good continuation is an answer to the question (relatively shallow). For a solution-oriented prompt, good continuation is a more developed version of the initial sketch (deeper, more collaborative).
This distinction is philosophically interesting because it maps onto different conceptions of what it means for an LLM to "do" philosophy. With problem-oriented prompts, the LLM is doing more of the philosophical work. With solution-oriented prompts, the work is distributed between human and LLM. Both are interesting, but the autonomy claim is stronger with problem-oriented prompts.
This should be worked into Section 4. The language of "problem-oriented" and "solution-oriented" prompts is cleaner than "one-shot vs conversational," and the good-continuation framework provides a theoretical vocabulary for discussing what's happening in each case.
---
## 9. The "corpus as repository of second-hand experience" framing
This is the co-author's most distinctive contribution in the transcript and deserves careful development. The idea: the training corpus is not just a repository of concepts, arguments, and theories — it's also a repository of experience, recorded in text. Diaries, autobiographies, phenomenological descriptions, art criticism, medical case reports, travel writing, literary fiction — all of these contain rich textual records of human experience.
This reframes the usual worry about LLMs lacking experience. The worry assumes that philosophical work requires *first-hand* experience. The response: it requires access to experience, but textual access may suffice for many philosophical purposes. Philosophy typically works with experience at a level of abstraction that can be accessed through descriptions. Phenomenologists describe experiences in texts; philosophers of art describe aesthetic responses in texts; philosophers of mind describe qualia in texts. These descriptions ARE the philosophical materials.
The framing should be hedged: this does not claim that reading about an experience is *the same as* having it. It claims that for the purposes of philosophical argument — drawing distinctions, testing theories, evaluating explanations — textual records of experience provide sufficient materials. Whether there are philosophical problems that require first-hand experience in a way that textual records cannot substitute is an open question.
Placement: in the new section addressing the intuitions/experience objection (item 4).
---
## 10. Verify the gluon scattering reference (Introduction)
The Introduction mentions: "In February 2026, researchers working on gluon scattering amplitudes gave GPT-5.2 worked examples for three, four, five, and six particles and asked it to find the general formula. The model conjectured a formula, completed a formal proof, and overturned a forty-year-old assumption (Guevara et al. 2026)."
I am not certain this reference is real. It could be fabricated by an LLM in a previous drafting session (a known failure mode — see the common errors log on quotation fabrication). The specific details (GPT-5.2, gluon scattering amplitudes, February 2026, Guevara et al.) are very precise, which could mean either that it's real or that it's a confident fabrication. Nick should verify this independently. If it's real, it's a strong opening example. If it's not, it needs to be replaced.
Similarly, the other AI breakthroughs in footnote [^1] (AlphaFold, Willow chip, Lupsasca with GPT-5) should be verified.
---
## 11. Reduce the analogy density in Section 1
Section 1 currently uses three extended analogies:
1. Lipton's likeliness/loveliness distinction, illustrated with Semmelweis
2. Deep Blue and chess (Gaut on creativity vs domain-specific excellence)
3. [[Finnur Dellsén|Dellsén]]'s dead-scientist thought experiment
Three analogies in one section is a lot. Each does different work (Lipton: evaluation is independent of truth; Deep Blue: quality is independent of process; Dellsén: progress is independent of current uptake), but the cumulative effect might feel like analogy-stacking rather than argument-development.
Options:
(A) Cut the Deep Blue analogy, which does work that's partially redundant with the Section 1 thesis (process irrelevant to evaluation). This point is made again, and more powerfully, in Section 3 via the Lipton squash-game quote.
(B) Cut the Semmelweis detail (cadaveric matter, priest, birthing position) and condense Lipton to the abstract point about likeliness vs loveliness.
(C) Keep all three but compress each — less detail on Semmelweis, less setup on Deep Blue, tighter prose throughout.
(D) Keep them all. Section 1 is establishing the paper's evaluative framework, and each analogy does genuinely different work. The section is ~900 words, which is reasonable.
My inclination: (A). The Deep Blue paragraph introduces Gaut (creativity) and makes a point that the paper doesn't subsequently develop. If creativity isn't a thread the paper pursues, this is a digression. The same point (quality is independent of process) is already made by the Watson/Crick comparison and by the Lipton distinction.
If creativity IS a thread (and the transcript's discussion of novelty suggests it could be), then keep Deep Blue but make it do more work — connect it forward to the novelty discussion in Section 3.
---
## 12. The "appear to hover" detail in the Einstein passage
The co-author mentions that the Einstein equivalence principle might need more explanation "for the sake of clarity." The current passage reads:
> Einstein imagined himself inside a falling elevator. He simulated the sensations of an observer in that scenario — objects released from the hand appearing to hover, the floor rushing up to meet falling things — and abduced from that simulated experience that gravity and acceleration must be the same phenomenon.
This is already clear and vivid. But for non-specialist readers, "gravity and acceleration must be the same phenomenon" might be confusing — what does it mean for them to be "the same"? The equivalence principle says: an observer in a closed room cannot distinguish between being in a gravitational field and being in an accelerating reference frame.
A slight expansion might help: "Einstein imagined himself inside a falling elevator. He simulated the sensations of an observer in that scenario — objects released from the hand appearing to hover, the floor rushing up to meet falling things — and noticed that these sensations would be indistinguishable from those of an observer floating in deep space, free from any gravitational field. From this he abduced the equivalence principle: that the effects of gravity are locally indistinguishable from those of acceleration."
But the co-author also says "it's already that zombie paper, so it's not a problem for us to explain the principle," which I interpret as: Zahavy's paper already explains the principle, so the paper doesn't need to do it from scratch. If the audience has read Zahavy, the current level of detail is fine.
---
## 13. The footnote on the analytic/continental divide [^ac]
> The distinction between text-focused and practitioner-focused conceptions maps imperfectly but suggestively onto the analytic/continental divide: analytic philosophy tends to emphasise texts and arguments as the locus of evaluation, while continental traditions more often locate philosophical activity in lived practice or self-transformation.
"Maps imperfectly but suggestively" is a hedge that risks irritating readers — either the mapping is worth making or it isn't. And the claim itself is a simplification that some readers will push back on (Heidegger wrote dense, carefully argued texts; many analytic philosophers care about lived experience).
Options:
(A) Cut the footnote entirely — it's not necessary for the argument.
(B) Keep it but own the simplification: "The distinction between text-focused and practitioner-focused conceptions tracks, roughly, the analytic/continental divide. This is a simplification — Heidegger is text-focused in one sense, and much analytic philosophy of mind engages lived experience — but it captures a genuine difference in where these traditions locate philosophical activity."
(C) Rework it to avoid the analytic/continental framing entirely: just note that the text-focused approach is more common in some philosophical traditions than others, without naming the traditions.
My inclination: (A). The footnote does no argumentative work and risks opening a tangent. If the paper addresses the scope limitation (item 5), the relevant distinctions (text-focused vs experience-dependent philosophy) are handled at the level of specific areas (philosophy of mind, aesthetics) rather than at the level of traditions.
---
## 14. "Generating plausibility" — the Floridi reframing in Section 3
Section 3 says:
> When Floridi et al. describe LLMs as "engines of generative plausibility," they are describing systems that have absorbed, from the corpus, the evaluative standards that philosophical abduction employs.
I want to flag that "engines of generative plausibility" should be verified as a direct quote from Floridi et al. The common-errors log warns about quotation fabrication. The Floridi extraction exists in the vault (`Attachments/_floridi_temp.txt`), so this should be checked against the source.
The argumentative move here is strong: what counts as "plausible" in philosophy is determined by the same evaluative standards (Williamson's intrinsic virtues) that the corpus encodes. So an LLM that generates "plausible" philosophical text is generating text that scores well on the evaluative criteria. This is a bridge between Floridi's description of LLMs and Williamson's account of philosophical evaluation.
---
## 15. Paper length and venue
The transcript discusses target length (6,000–8,000 words) and venue (*Mind*, *Philosophical Review*, *Journal of Philosophy*). The current draft is ~4,000 words. The additions discussed above (intuitions/experience section, Section 4 on demonstration/prompting) would add roughly 2,000–3,000 words, bringing the paper to 6,000–7,000 words. This is within range.
The venue ambition is high. For *Mind* or *Philosophical Review*, the paper would need:
- A tight argumentative structure with no loose ends
- Serious engagement with anticipated objections (the intuitions objection is the most pressing)
- Proper referencing (the Machery book needs to be engaged with, not just cited)
- Novel philosophical contribution (the "corpus as repository of experience" framing could serve this role)
---
## 16. The Dellsén dead-scientist thought experiment
Section 1 presents this:
> Suppose a scientist publishes an important finding and then dies. Everyone who read the paper also dies, or forgets what they read. Has the progress been lost? No.
This is a helpful thought experiment but it's compressed. The conclusion ("No") is asserted without development. Why "No"? Because progress consists in enabling understanding, and the publication still enables understanding even if no one currently grasps it — the materials remain publicly available.
The thought experiment would be more effective if the reader were allowed to arrive at the conclusion through the example rather than being told the answer immediately. Compare the Watson/Crick passage in the same section, where the development is paced — each detail builds toward the conclusion.
Option: expand slightly to develop the reasoning: "Suppose a scientist publishes an important finding and then dies. Everyone who read the paper also dies, or forgets what they read. On Dellsén's account, the progress has not been lost. The publication still exists; it still enables anyone who reads it to understand the phenomenon in question. Progress, on this view, consists in the availability of understanding — not in anyone's actually understanding."
---
## 17. Integrate the "good continuation" vocabulary
This is a theoretical framing from the semiotic physics / Janus literature (the "Simulators" paper is in the Learning folder). The idea: an LLM generates the most probable continuation of a sequence. "Good continuation" in philosophy means text that exhibits the properties the paper identifies (coherence, handling of objections, illumination of subject matter).
This vocabulary does two things:
(A) It provides a technical account of what LLMs do when they produce philosophy — they generate continuations that are "good" by the standards the corpus encodes.
(B) It connects the paper to the broader theoretical literature on LLMs (beyond Floridi and Zahavy), which strengthens its interdisciplinary reach.
The vocabulary should appear in Section 4, in the discussion of prompting. The distinction between shallow good continuation (question → answer) and deep good continuation (partial solution → more developed solution) maps onto the problem-oriented/solution-oriented distinction from item 8.
---
## 18. Remove the "Work in Progress" callout
Section 0 currently has:
> [!warning] Work in Progress
> This introduction is still being drafted and requires significant revision.
This should be removed from any version shared with the co-author or submitted. It's an Obsidian-specific callout that serves as a note to self but shouldn't appear in the shared text.
---
## Summary of changes by priority
I'm grouping these by my sense of how they relate to the paper's argumentative needs, not ranking them (since both you and the co-author should decide priorities):
Structural/argumentative additions:
- Item 1: Fix Section 2 opening (structural error)
- Item 4: Add intuitions/experience objection (the co-author's "most substantial objection")
- Item 5: Address scope limitation with "repository of experience" response
- Item 6: Write Section 4 (demonstration/prompting)
- Item 7: Hitchhiker's Guide payoff
- Item 8: Problem-oriented vs solution-oriented prompts
- Item 9: "Corpus as repository of second-hand experience" framing
Sharpening existing material:
- Item 2: Sharpen the Floridi/Zahavy asymmetry
- Item 3: Address Bayesianism explicitly
- Item 14: Verify "engines of generative plausibility" quote
- Item 16: Develop the Dellsén thought experiment
- Item 17: Integrate "good continuation" vocabulary
Prose/presentation:
- Item 10: Verify the gluon scattering reference
- Item 11: Consider reducing analogy density in Section 1
- Item 12: Slight expansion of Einstein passage (maybe)
- Item 13: Cut or rework the analytic/continental footnote
- Item 18: Remove "Work in Progress" callout
Remaining uncertainties:
- Where exactly to place the intuitions/experience discussion (new section? expansion of existing section?)
- How much of the prompting/good-continuation discussion to include (risk of making the paper too ambitious)
- Whether to engage with Machery's specific arguments (this requires getting the source text)
- Whether the gluon scattering reference is real