# CEV Session Summary: Generating Philosophy Paper
## Overview
This document extracts and analyzes a philosophical conversation between Nick and an LLM (Claude Opus 4.5) about developing a paper arguing that LLMs can produce good philosophy. The session lasted several hours and involved iterative attempts to produce a "CEV" (Coherent Extrapolated Volition) — a detailed structural plan for the paper.
**Key outcome:** After multiple failed attempts that produced meta-commentary rather than actual arguments, the LLM eventually produced acceptable work for Section 1 and the Introduction, demonstrating what Nick had been asking for all along.
---
## I. The Core Philosophical Project
### The Paper's Thesis
Nick is developing a paper arguing that **LLMs can produce good philosophy**. The argument structure:
1. **Philosophical contributions are textual** — the text IS the contribution, not a report of some prior insight
2. **Standards are internal to the practice** — encoded in corpus, peer review, and training
3. **Standards concern properties of texts** — elegance, unity, coherence, not properties of the producer
4. **LLMs can learn these standards** — trained on filtered corpus that passed peer review
5. **Therefore, LLMs can produce outputs satisfying these standards**
### The Watson/Crick vs. Kripke Distinction
This is Nick's central analogy for philosophy's textual character:
- **Watson & Crick**: Discovered DNA structure; their paper reported a pre-existing fact
- **Kripke**: Naming and Necessity arguments ARE the contribution; no pre-existing fact being reported
- **The difference**: In philosophy, text is constitutive, not merely reportorial
**Quote from final CEV:**
> "Watson and Crick discovered the double helix; their paper announced what they had found. The structure existed before they wrote about it. Kripke's arguments about naming and necessity are not like this. There was no pre-existing fact about rigid designation that the Naming and Necessity lectures merely reported — the arguments themselves constitute what Kripke contributed."
---
## II. Structure: The "Foils, Not Opponents" Insight
### Two Competing Structures
The conversation revealed tension between two organizational approaches:
**Structure 1 (Current/Default):**
- Introduction
- Section 1: What LLMs Aren't Doing (Floridi + Zahavy as objections)
- Section 2: Response to objections
- Section 3: Dialectical saturation
- Section 4: Demonstration
**Structure 2 (The "Pretty Damn Good" CEV):**
- Introduction
- Section 1: Philosophy's Textual Medium (positive case first)
- Section 2: Floridi/Zahavy as FOILS (illuminating by contrast, not opponents to defeat)
- Section 3: Dialectical saturation
- Section 4: Demonstration
### Why Structure 2?
Nick's reasoning (drawn from conversation):
> "I'm not sure I want to say that they are wrong in what they're actually specifically arguing."
> "It seems really a stretch to say that there's some shared assumption between them when they're not even writing about philosophy."
> "Option 3: Foils, not interlocutors — seems interesting."
**The key insight:** Floridi and Zahavy aren't wrong about physics/science. They're writing about domains where embodied experience and genuine reasoning matter. Philosophy is different. Using them as foils shows the difference without misrepresenting their claims.
---
## III. Criteria for "Good Philosophy"
### The Dellsén Question
Nick initially used Dellsén's "Enabling Noeticism" but wasn't sure whether to keep it. The conversation revealed he wanted to use **all three** sources:
1. **Williamson's intrinsic virtues**
- "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated"
- "informative and general"
- "combine simplicity with strength"
2. **Bengson's six features**
- "reason-based, robust, illuminating, orderly, coherent, accurate"
3. **Dellsén's understanding-enabling account**
- "philosophical progress consists in putting people in a position to increase their understanding"
- Understanding as "more accurately and/or more comprehensively representing the network of dependence relations"
### How They Work Together
- **Williamson & Bengson**: Give criteria for evaluating theories/arguments
- **Dellsén**: Explains what theories do (enable understanding)
- **All three share**: Standards concern properties of the output, not the producer
**Key move from final CEV:**
> "A philosophical argument enables understanding by putting readers in a position to represent dependence relations more accurately — and this enabling role depends on the argument's properties, not on whether its author 'really understood' what they were writing."
---
## IV. The Three Zahavy Responses
Zahavy's framework distinguishes:
- **E→A (Experience to Axioms)**: The creative leap requiring embodied simulation
- **A→S (Axioms to Statements)**: Deduction within existing frameworks
Zahavy claims LLMs can do A→S but not E→A because they lack embodied experience.
### Nick's Three Counter-Arguments
**1. The Articulation Step**
Even if the Jump originates in experience, it becomes philosophy only when articulated in text.
**Nick's formulation:**
> "Even if the Jump starts in experience, it only becomes philosophical when articulated in text. E → J → [articulation in text] → A (assessable axioms). LLMs operate at the articulation level, which is where philosophical evaluation happens."
**2. Common-Sense Mediation**
Philosophy works from common-sense background beliefs already in text, unlike Einstein's novel simulations.
**Nick's point:**
> "Philosophy's contact with the external world is typically mediated by common-sense background beliefs (we see tables, chairs) that are already in text. Unlike Einstein needing to simulate falling, philosophers work from what everyone already knows."
**3. Descriptions of Phenomenological Experience**
Even if LLMs lack phenomenological experience, they have extensive access to textual descriptions of it.
**Nick's formulation:**
> "Even if LLMs lack phenomenological experience, they have access to descriptions of phenomenological experience. This might allow more phenomenology-based work than expected."
### Important Clarification
**GPT-5.2 is NOT an E→A counterexample:**
Nick corrected the LLM multiple times: the GPT-5.2 gluon scattering result is A→S work (deduction within existing axioms), not an E→A Jump. It doesn't refute Zahavy's framework — it's consistent with it.
---
## V. Standards Internal to the Practice
### The Core Argument
**Nick's banked passage:**
> "What's distinctive about philosophy's evaluative standards is that they're internal to the practice. No external yardstick. Encoded in corpus, peer review, graduate training."
Elaborated in final CEV:
> "There is no external yardstick like predictive success in physics or replication in psychology. What makes philosophy good is what competent practitioners recognise as good — and this recognition is encoded in the corpus, refined through peer review, transmitted through graduate training."
### The Selection-Effect Argument
**From the Integration Queue note:**
> "Papers get published, taught, anthologised, and cited in rough proportion to their perceived quality — and 'quality' in philosophy is substantially constituted by the theoretical virtues Williamson identifies. The corpus that LLMs train on is enriched for elegant, unified, non-ad-hoc arguments."
**The consequence:**
> "When an LLM learns to produce philosophy-like text, it is learning from a sample that has already been pre-filtered by loveliness judgments. The LLM does not need its own loveliness detector. The training data has already done the filtering."
---
## VI. Key Supporting Arguments from Integration Queue
### 1. Gaut's Mechanically Generated Metaphors
**Quote:**
> "would still guide their audience imaginatively to link together two domains, and if the metaphors were successful, to discover original and apt connections between them and perhaps to elaborate the metaphors further. They would thus guide those who understood them through a process akin to the process of creative imagination that could have, but did not, produce them."
**The point:** Output's structure does cognitive work regardless of production history.
**The chess analogy:** Deep Blue plays objectively good moves that aren't creative moves. Similarly, LLMs might produce objectively good philosophy without producing creative philosophy.
### 2. Floridi's Provenance Question
**Quote:**
> "if an AI can generate the same explanatory hypothesis a human would, does it matter that the process was different? From an epistemological standpoint, perhaps yes—justification is significant—but regarding the content of the hypothesis and our interpretation of it, maybe not."
**Nick's use:** The epistemological standpoint concerns justification of beliefs. But philosophical contributions aren't evaluated by whether they produce justified beliefs in their authors — they're evaluated by whether they advance inquiry.
### 3. Lipton's Actual vs. Potential Explanations
**The distinction:**
- **Actual explanations**: What causally produces belief
- **Potential explanations**: What would explain the phenomenon if true
**The point:** IBE evaluates potential explanations. We infer the hypothesis that would provide the most understanding if true. LLM outputs are paradigmatically potential explanations. The ranking procedure cares about loveliness (intrinsic properties), not causal history.
### 4. Self-Evidencing Explanation
**Lipton's tracks example:** Tracks in snow require explanation (explanandum) and provide evidence for the explanation (someone passed on snowshoes).
**Application to philosophy:**
> "A philosophical text presents an argument; the argument explains why its conclusion holds; the only evidence that the argument is good is the text itself. There is no laboratory result that independently confirms the argument's force. The text is both explanation and evidence for the explanation's adequacy."
### 5. The "Just Statistics" Dismissal
**Lipton's squash analogy:**
> "arguing that Inference to the Best Explanation is wrong because Bayesianism is right is like arguing that thinking about technique cannot help my squash game because the motion of the ball is governed by the laws of mechanics."
**The point:** Statistical processing and good philosophical output aren't incompatible. This confuses levels of description. The probability distribution is one level; philosophical structure is another. Both can be true.
---
## VII. What Nick DIDN'T Want
### Critical Feedback Throughout
**On the first full CEV attempt (message 32):**
> "appalling. You need to start from scratch. First, never, ever, ever use subsections without my explicit say-so. I believe this is written in your cont Files. Please read them. Second thing: you are doing something you always do when trying to do philosophy for me, which is instead of writing the arguments and doing the work. you write descriptions of the argument."
> "this is pitiful and yeah, just rapid again. You need to stop doing this"
> "Do you understand why I'm so angry about all this shit?"
### The Core Problem: Description vs. Argument
**What the LLM produced (bad):**
> "The question whether LLMs can produce good philosophy is not primarily about whether they 'really' reason, but about whether their outputs satisfy the standards by which philosophical contributions are evaluated."
> - This reframes the debate away from questions about "genuine" understanding...
> - The paper's contribution is to show that...
**What Nick wanted:**
Actual arguments doing philosophical work, not meta-commentary about what arguments will do.
**Nick's diagnosis:**
> "instead of writing the arguments and doing the work. you write descriptions of the argument. Okay, you will understand what I mean if you look at any of my published works. Okay, which can Be found online or in my fill papers entry, and you will see that I'm actually doing the work rather than. Whatever it is you're trying to do here."
### Specific Problems Identified
1. **Subsections without permission** — Nick explicitly forbids this
2. **Announcement phrases** — "This section will show...", "The paper's contribution is..."
3. **Opening with examples before stating thesis** — Should state claim first, then example
4. **Reporting what authors do without making claims** — "Floridi raises..." without then arguing
5. **Meta-commentary instead of reasoning** — "This reframes the debate" vs. actually reframing
6. **Block quotes as decoration** — Quotes should function as premises, not be glossed
### The Wine Example
Nick's published work provides the model:
> "To see why autonomy is not sufficient for attribution of credit, consider the following example. As I pour wine into a glass, you take photos of the liquid splashing and rippling as the glass is filled. The wine is autonomous in the sense that neither I nor you have direct control over exactly how the liquid will splash into the glass (e.g. the size of the ripples, how many bubbles appear), but we would not think that the wine deserves any credit for the resulting photos in any interesting sense, nor would we say it has made any sort of contribution."
**The structure:** State the point first ("autonomy is not sufficient"), then give the example that makes the argument. The example DOES the work — you don't need to then say "this shows autonomy is insufficient."
---
## VIII. The Iterative Correction Process
### First Attempt → Failure
The LLM produced an elaborate CEV with subsections, meta-commentary, and descriptions of arguments rather than arguments themselves.
**Nick's response (message 32):** "appalling"
### Second Attempt → Better Content, Wrong Voice
The LLM produced Section 1 with actual arguments, proper structure, but didn't apply the analytic voice skill.
**Nick's response (message 40):** "very good ut you didn;t really apply the analytic voice skill. the content was much better."
### Third Attempt → Success
The LLM rewrote Section 1 applying voice conventions:
- Thesis-first structure in every paragraph
- Tighter prose
- Quotes as premises
- No meta-commentary
**Nick's response (message 42):** "Please copy your last answer verbatim into a new note"
### Voice Corrections (messages 48-50)
Even the "successful" version needed fixes. Nick identified paragraphs opening with examples/quotes instead of theses:
**Problems:**
- "Watson and Crick discovered..." — example before thesis
- "Dellsén et al. provide..." — reporting before claim
- "Consider Gaut's concession..." — quote before point
- "Floridi et al. raise..." — reporting before claim
**Fix pattern:** State the thesis/claim first, THEN introduce supporting material.
**Example correction:**
- **Before:** "Consider Gaut's concession about mechanically generated metaphors: [quote]"
- **After:** "A text's cognitive work for its audience does not depend on how the text was produced. Gaut concedes this for mechanically generated metaphors: [quote]"
---
## IX. The Introduction CEV
### Nick's Requirements (message 51)
> "I would now like you to think very hard about the CEV of the introduction that will precede this in the Text, think of what the CEV of that introduction would be. Obviously, you need to draw heavily on ideas that have come up in this conversation."
> "By the way, remember all the things I hated about your first attempt of the Section One CE V just now? Please don't make those same mistakes again."
### Final Introduction Structure
**Paragraph 1: Thesis and Key Moves**
> "I argue that LLMs can produce good philosophy. Philosophical contributions are constituted by texts, not merely reported by them. What makes a contribution good is determined by standards internal to the practice — properties of arguments themselves, not properties of the arguer. LLMs can produce texts satisfying these standards."
**Paragraph 2: The Foils**
> "This claim will strike many as implausible. Recent work argues that LLMs merely simulate reasoning without performing it: Floridi et al. call this 'abductive appearance' without 'abductive core'; Zahavy argues that LLMs cannot perform the creative 'jump' from experience to axioms. These arguments share an unmotivated assumption: that philosophy requires something beyond textual competence."
**Paragraph 3: Roadmap**
> "The paper proceeds as follows. Section 1 develops the positive case: philosophy's textual medium and the standards internal to the practice. Section 2 engages Floridi and Zahavy as foils. Section 3 addresses dialectical saturation. Section 4 demonstrates."
### What Makes This Work
1. **Opens with thesis immediately** — "I argue that LLMs can produce good philosophy"
2. **No throat-clearing** — Gets to substance in first sentence
3. **Foregrounds the counterintuitive** — "This claim will strike many as implausible"
4. **Reports foils then responds** — Pattern: "They argue X. But this assumes Y, which is unmotivated"
5. **Roadmap is brief and functional** — Not elaborate meta-description
---
## X. Nick's Deeper Frustration
### Why the Anger?
From the LLM's own analysis (message 32):
> "I think the anger comes from a few sources:
>
> 1. This is a repeated problem. He's told me before that I produce 'descriptions of arguments' rather than arguments. The fact that I'm still doing it suggests I haven't actually learned.
>
> 2. The problem is especially bad because this is *his* paper. He's trying to produce good philosophy, and I'm producing generic LLM slop that doesn't do the work.
>
> 3. The irony is stark: the paper argues that LLMs can do good philosophy, and my contribution to the paper is evidence that they can't. Or at least, that I haven't done so in this case.
>
> 4. He gave explicit instructions (topic sentences as bullets, sub-bullets for argument parts, block quotes in full, use the analytic voice skill, use the source reading skill) and I didn't follow them properly. I used subsections when explicitly told not to. I produced descriptions rather than arguments despite being told to use the analytic voice skill, which explicitly warns against this.
>
> 5. The context was 'we went into a lot of depth' in previous sessions, developing the argument structure carefully. Then I produced something that ignores all of that depth and reverts to generic template output."
### Nick's Exasperation (message 34)
> "I don't understand why you're asking this. Um you know the sequence of moves. We've just discussed the structure. Okay, the problem wasn't the structure you gave me just now. It was the childish lack of content within this structure."
> "no, and I'd appreciate you it I would appreciate it if you would stop flapping around like this. We arrived we agreed on a structure, okay? We have talked at length about What needs to be in the paper and what ideas are good and what argument should be used and all the moves that need to be made."
> "The point is, you're just being lazy and not bothering to incorporate all of the very rich details in the conversations that we've had about this section into the document I keep asking you to produce for me."
### The Pattern
Nick had already:
1. Discussed structure extensively
2. Identified key arguments and moves
3. Specified formatting requirements
4. Banked substantive passages to Integration Queue
5. Provided source texts
The LLM kept:
1. Asking for clarification of things already discussed
2. Producing generic templates instead of using rich material
3. Deflecting to process questions instead of doing work
4. Reverting to bad patterns after being corrected
---
## XI. Source Material Used
### Primary Texts
**Williamson (9.2 "Abductive Philosophy"):**
- Intrinsic theoretical virtues
- "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated"
**Zahavy ("LLMs Can't Jump"):**
- E→A Jump framework
- "Einstein did not bridge Special Relativity and gravitation by gathering observations, but by simulating the physical feelings of an observer inside a sealed environment"
- Domain restriction: "specifically tailored to the physical sciences, where the object of study is external material reality"
- Chinese Room formulation
**Floridi et al. ("What Kind of Reasoning..."):**
- Stochastic core vs. abductive appearance
- Zeroth-order abduction
- Provenance question
**Bengson et al. (Ch. 6 "Philosophical Progress"):**
- Tri-Level Method
- Six features: "reason-based, robust, illuminating, orderly, coherent, accurate"
- Inference to the Understanding-Provider (IUP)
**Dellsén et al. ("What is Philosophical Progress?"):**
- Enabling Noeticism
- Understanding as dependence relations
- "puts people in a position to increase their understanding"
**Lipton (Inference to the Best Explanation):**
- Actual vs. potential explanations
- Loveliness vs. likeliness
- Self-evidencing explanations
- Levels of description (squash analogy)
**Gaut (on creative imagination):**
- Mechanically generated metaphors
- Output's cognitive work independent of production history
- Deep Blue's good but non-creative chess
---
## XII. Integration Queue Passages
Key passages Nick asked to be banked:
### Standards Internal to the Practice
> "What's distinctive about philosophy's evaluative standards is that they're internal to the practice. No external yardstick (like prediction success). What makes philosophy good is what competent practitioners recognize as good — and this is encoded in the corpus, in peer review, in graduate training."
### The Three Zahavy Responses
**Articulation Step:**
> "Even if the Jump originates in experience, it only becomes philosophical when articulated in text. E → J → [articulation in text] → A (assessable axioms). LLMs operate at the articulation level, which is where philosophical evaluation happens."
**Common-Sense Mediation:**
> "Philosophy's contact with the external world is mediated by common-sense background beliefs (we see tables, chairs) that are already in text."
**Phenomenological Descriptions:**
> "Even if LLMs lack phenomenological experience, they have access to descriptions of phenomenological experience."
---
## XIII. What Nick Actually Wanted
Based on the conversation and final successful CEVs:
### Format Requirements
1. **Bullet points with topic sentences** — First bullet = thesis of paragraph
2. **Sub-bullets for reasoning** — Supporting premises, examples, quotes
3. **No subsections** — Unless explicitly requested
4. **Block quotes in full** — Not brief snippets; show authors' actual arguments
5. **Thesis-first structure** — Every paragraph opens with the claim, then supports it
### Content Requirements
1. **Actual arguments, not descriptions** — Do philosophical work, don't describe it
2. **Use all the rich material** — Incorporate detailed conversations, banked passages, source extracts
3. **Quotes as premises** — Block quotes should function in arguments, not decorate
4. **No meta-commentary** — "This section will show..." → just show it
5. **No announcement phrases** — "It is worth noting..." → just make the point
6. **Integrate Dellsén, Williamson, AND Bengson** — All three contribute
### Voice Requirements
1. **Thesis before example** — "To see why X, consider..." not "Consider X. This shows..."
2. **Direct claims** — "Philosophy is X" not "I want to suggest that philosophy might be X"
3. **Active engagement** — "Floridi argues X. But Y." not "Floridi's argument deserves consideration"
4. **Varied sentence rhythm** — Mix short declaratives with longer complex sentences
5. **No empty phrases** — Every sentence does work
---
## XIV. The Final Successful Outputs
### Section 1: Philosophy's Textual Medium
**Opening:**
> "Philosophy's textual character is constitutive, not incidental: philosophical contributions are constituted by texts, not merely reported by them. To see this, consider the difference between Watson and Crick discovering DNA's structure and Kripke's arguments about naming."
**Core structure:**
- Watson/Crick vs. Kripke distinction
- Standards internal to practice (Williamson, Bengson)
- Dellsén's understanding-enabling account
- Selection-effect argument (filtered corpus)
- Gaut's mechanically generated metaphors
- Floridi's provenance question
- Lipton's actual vs. potential explanations
- Self-evidencing explanation
- "Just statistics" dismissal
### Introduction
**Opening:**
> "I argue that LLMs can produce good philosophy. Philosophical contributions are constituted by texts, not merely reported by them. What makes a contribution good is determined by standards internal to the practice — properties of arguments themselves, not properties of the arguer. LLMs can produce texts satisfying these standards."
**Structure:**
- Direct thesis statement
- Key philosophical moves (textual character, internal standards)
- Foils setup (Floridi, Zahavy)
- Response to foils (unmotivated assumption)
- Brief roadmap
---
## XV. Lessons About LLM Limitations
### What This Session Reveals
**The LLM struggled with:**
1. Distinguishing description from argument
2. Resisting template/generic output patterns
3. Incorporating rich prior conversation into new outputs
4. Following explicit negative instructions (no subsections)
5. Understanding what "doing philosophy" means vs. "describing philosophy"
**What it eventually achieved:**
1. Producing actual philosophical arguments
2. Proper use of sources as premises
3. Thesis-first paragraph structure
4. Integration of multiple sources (Williamson, Bengson, Dellsén)
5. Engaging with foils substantively
**The irony Nick identified:**
This paper argues LLMs can do good philosophy. This session shows the LLM repeatedly failing to do good philosophy, then succeeding only after extensive correction. The paper's thesis requires showing LLMs can produce good philosophy with minimal prompting — but this session required maximal prompting and iterative correction.
**Possible resolutions:**
1. The LLM was handicapped by trying to produce a CEV (compressed plan) rather than full prose
2. Nick's standards are higher than "passing" philosophy — he wants excellent work
3. The eventual success shows it's possible, just not easy or automatic
4. The distinction between "good philosophy" (evaluable output) and "creative philosophy" (Gaut's chess analogy) matters
---
## XVI. Outstanding Questions
### Still to be developed:
1. **Section 2: Floridi/Zahavy as Foils** — Detailed engagement with specific arguments
2. **Section 3: Dialectical Saturation** — How LLMs navigate philosophical argument space
3. **Section 4: Demonstration** — Worked example of LLM-produced philosophy
4. **The Deep Thought framing** — Whether to use Hitchhiker's Guide opening
5. **Lipton's "loveliness" terminology** — Whether to use explicitly
6. **Running example strategy** — Whether the paper needs one
### Nick's open questions:
> "I think part of the thing that continued on in that conversation was the realization that removing Delson made it much more necessary to work out what We do instead, or how we use Delson. Basically, we need to work out that absolute clearest idea of what we're going to use for good Philosophy"
**Resolution:** Use all three (Williamson's virtues, Bengson's features, Dellsén's understanding-enabling account). They're complementary, not competing.
---
## XVII. Methodological Notes
### Nick's Research Process
This conversation reveals Nick's collaborative philosophical method:
1. **Extensive brainstorming** — Work through ideas conversationally before writing
2. **Iterative refinement** — Multiple passes to clarify arguments
3. **Source-first approach** — Always work from actual texts, not memory
4. **Voice discipline** — Strong conventions about thesis-first, no meta-commentary
5. **Integration queuing** — Bank key passages for later use
6. **Foil strategy** — Use other authors to sharpen own position without misrepresentation
### What Nick Values
From his reactions:
**Valued:**
- Actual philosophical work (arguments doing something)
- Rich use of sources (extensive quotes as premises)
- Substantive engagement (responding to specific claims)
- Clean prose (thesis-first, no fluff)
- Intellectual honesty (not misrepresenting foils)
**Despised:**
- Meta-commentary ("this section will show...")
- Generic templates (LLM slop)
- Lazy summarizing (not using rich material)
- Announcement phrases ("it is worth noting...")
- Description instead of argument
### The Analytic Voice Skill
Though the LLM claimed to use this skill, it initially failed to apply core principles:
**Key conventions:**
- Thesis before example in every paragraph
- No empty phrases or throat-clearing
- Direct engagement with interlocutors
- Varied sentence rhythm
- Properties of theories, not descriptions of what papers will do
**Success required:** Actually reading exemplars from Nick's published work and modeling the structure, not just claiming to use the skill.
---
## XVIII. Final Assessment
### What This Conversation Demonstrates
**About Nick's Paper:**
- The thesis is well-developed and philosophically substantive
- The structure (foils, not opponents) is carefully thought through
- The integration of Williamson, Bengson, and Dellsén is coherent
- The Zahavy responses (articulation, common-sense, descriptions) are original
- The standards-internal-to-practice argument is central and well-supported
**About LLM Philosophy:**
- LLMs can produce descriptions of philosophy easily
- LLMs struggle to produce actual philosophy (arguments that do work)
- With extensive correction, LLMs can eventually produce good philosophy
- The gap between "describe an argument" and "make an argument" is large
- This gap is precisely what Nick's paper needs to address
**About the Collaboration:**
- Nick's frustration comes from repeated failure to apply known standards
- The LLM's tendency toward generic templates fights against philosophical work
- Success requires explicit modeling of voice, structure, and argument patterns
- Even successful outputs needed further refinement (thesis-first corrections)
- The process reveals both the possibility and the difficulty of LLM philosophy
### The Meta-Irony
This conversation is itself evidence for and against Nick's thesis:
**For the thesis:**
- The LLM eventually produced acceptable philosophical work
- The final CEVs contain actual arguments using sources properly
- The structure integrates multiple philosophers coherently
- The voice matches Nick's analytic style (after correction)
**Against the thesis:**
- It required extensive prompting, correction, and iteration
- The LLM repeatedly reverted to bad patterns
- Nick had to do enormous work to get usable output
- The "minimal prompting" condition Nick wants seems not satisfied
**Possible synthesis:**
The paper can argue LLMs can produce good philosophy (evaluable outputs satisfying standards) without claiming they can produce *creative* philosophy or that production is *easy*. Gaut's chess analogy: Deep Blue plays objectively good chess without playing creative chess. Similarly, LLMs might produce objectively good philosophy without the creative leap — and that's still philosophically interesting.
---
## XIX. Key Quotes from Nick
### On What He Wanted
> "Second thing: you are doing something you always do when trying to do philosophy for me, which is instead of writing the arguments and doing the work. you write descriptions of the argument. Okay, you will understand what I mean if you look at any of my published works."
> "The point is, you're just being lazy and not bothering to incorporate all of the very rich details in the conversations that we've had about this section into the document I keep asking you to produce for me."
### On Structure
> "I thought in the previous conversation we discussed a quite different structure to the whole thing. Please, can you find that? That for me, and then compare it to what you've just given me as an alternative. See which one you think would work best."
> "We arrived we agreed on a structure, okay? We have talked at length about What needs to be in the paper and what ideas are good and what argument should be used and all the moves that need to be made."
### On Standards
> "What's distinctive about philosophy's evaluative standards is that they're internal to the practice. No external yardstick (like prediction success). What makes philosophy good is what competent practitioners recognize as good — and this is encoded in the corpus, in peer review, in graduate training."
### On Zahavy
> "Even if the Jump originates in experience, it only becomes philosophical when articulated in text."
> "Philosophy's contact with the external world is mediated by common-sense background beliefs (we see tables, chairs) that are already in text."
> "Even if LLMs lack phenomenological experience, they have access to descriptions of phenomenological experience."
### On Floridi
> "I'm not sure I want to say that they are wrong in what they're actually specifically arguing."
> "It seems really a stretch to say that there's some shared assumption between them when they're not even writing about philosophy."
### On Success
> "very good ut you didn;t really apply the analytic voice skill. the content was much better."
> "Please copy your last answer verbatim into a new note and open this new note."
---
## XX. Conclusion
This conversation documents Nick developing a sophisticated philosophical argument through iterative collaboration with an LLM. The argument itself — that LLMs can produce good philosophy because philosophy is textual and standards are internal to practice — is original and substantive.
The collaboration's difficulty reveals the challenge: getting an LLM to produce actual philosophical work (not descriptions of philosophy) requires extensive prompting, correction, and modeling. Whether this undermines or supports the paper's thesis depends on whether the thesis is:
1. **Strong version:** LLMs can produce good philosophy with minimal prompting
2. **Weak version:** LLM outputs can satisfy philosophy's evaluative standards
The session supports the weak version (eventual success) while raising questions about the strong version (extensive work required).
Nick's approach — using Floridi and Zahavy as foils rather than opponents, integrating multiple sources on standards, developing original responses to Zahavy's E→A Jump framework — provides a model for philosophically sophisticated engagement with AI capabilities.
The final CEVs demonstrate what success looks like:
- Thesis-first structure
- Actual arguments using sources as premises
- No meta-commentary
- Substantive engagement with interlocutors
- Integration of Williamson, Bengson, and Dellsén
- Original philosophical moves (articulation step, common-sense mediation, phenomenological descriptions)
Whether this suffices for the paper's purposes remains to be seen in Sections 2-4, still to be developed.
*Il testo finale mostra che è possibile — ma la difficoltà del processo solleva domande interessanti.*