# Generating Philosophy Session: Williamson, Floridi, and Understanding Session transcript from 2026-01-29 discussing the generating philosophy project, connecting Williamson on abduction, Floridi's critique of LLM reasoning, and two accounts of theoretical understanding. --- ## Opening: Project Context **Nick:** I would like us to talk about the generating philosophy project. **Claude:** *[Loads session file for Generating Philosophy with AI]* Here's where things stand: **Core question:** Can LLMs produce philosophy of sufficient quality to be useful, and how should philosophers adopt them? ### Active Threads **1. Dialectical Saturation Thesis** — the headline theoretical claim. LLMs trained on philosophical corpora have internalized argumentative "rules of the game": - Move types (distinction, counterexample, repair, disambiguation, synthesis) - Move sequences (distinction → objection → reply) - Success conditions (precision, explanatory power, simplicity) Three versions, increasing in ambition: - **Script Competence**: LLMs have internalized recurring move-sequences - **Latent-Game Inference**: Bottleneck is *which game*, not lack of rules - **Salience-Not-Frequency**: "Obvious move" can be rare but structurally apt Walton's *Argumentation Schemes* (2008) provides formal backing — schemes as locution/commitment/dialogue rules, critical questions as built-in stress tests. **2. Move 37 / Tail Novelty** — how LLMs might produce genuinely novel moves (selection vs generation; combinatorial novelty) **3. Inner Speech / LLM Coupling** — extended cognition framing for philosopher-LLM collaboration ### Open Questions - Which saturation thesis version to develop (Script Competence safest but less interesting; Salience-Not-Frequency bolder but needs evidence) - What "remainder" falls outside scheme-governed reasoning (long-horizon theory-building?) - Whether "philosophy is self-grounding" framing is productive or too defensive - Substack draft exists — tone may need work Yesterday you created ~15 notes extracting ideas from the ChatGPT conversation about Floridi. Today's scratch pad flags this as the first work session on the project. What would you like to work on? Some options: developing one of the saturation thesis versions, looking at the Substack draft, working through remaining threads in the source conversation, or something else entirely. --- ## Williamson on Abduction and Model-Building **Nick:** In particular I would like you to tell me about the williamson text's sections about abduction and model building. **Claude:** *[Extracts and reads Williamson's "Widening the Picture" from Learning/generating-philosophy/]* This is rich material. Let me synthesize the key sections: ### Williamson on Abduction (Section 9.2) #### The Core Claim Williamson advocates an **abductive methodology** for philosophy — inference to the best explanation as the primary mode of theory selection. He frames this through an anecdote about Dummett: > "The difference between us is that you think that inference to the best explanation is a legitimate method of argument in philosophy, and I don't." I realized that his characterization of the difference was right, although I was a little shocked at his outright rejection of inference to the best explanation in philosophy. #### Abduction Characterized Williamson defines abduction broadly (following Peirce but not slavishly): > "Abduction is an informal method of non-deductive, ampliative inference and theory choice familiar from the natural sciences... approximately equivalent to 'inference to the best explanation,' when 'explanation' is understood to cover non-causal as well as causal explanations." Key criteria for ranking theories abductively: - **Fit with evidence** (entailment or probabilistic connection) - **Simplicity, elegance, generality, unificatory power** - **Informativeness and strength** (bold theories are rewarded) #### The Mathematics Precedent Crucially for the generating philosophy project, Williamson argues mathematics provides a precedent for **armchair abduction**: > "At first sight, mathematics looks like a paradigm of a purely deductive form of inquiry... However, those deductive proofs ultimately rely on first principles, principles for which no further deductive proof is expected." The axioms of ZFC are justified *abductively*: > "The first principles must be strong enough to prove all established mathematical theorems; they must not be strong enough to prove a contradiction. Subject to those constraints, we naturally want the first principles to maximize simplicity and similar virtues." #### Philosophy as "Impure" Abduction This is the passage from the legacy quotes, now in context: > "Unsurprisingly, abduction in philosophy is and should be less 'pure' than in mathematics. The evidence on which it does and should depend is often exogenous, generated from outside the discipline itself. It is perfectly proper for philosophers of time to appeal to Einstein's theory of special relativity, for philosophers of perception to use experimental results from the psychology of perception..." ### Williamson on Model-Building (Section 9.3) #### Model-Building Defined > "The model-building methodology that has proved so successful in the natural sciences can thereby be applied in philosophy too, and provides new insights into old problems." A model is: a **precisely described hypothetical example** — a type of case rather than a particular case, formal enough to be tractable. #### Where Model-Building Applies > "The need for model-building is hardest to avoid where the complex, messy nature of the subject matter tends to preclude informative exceptionless universal generalizations. The paradigm of such complexity and mess is the human world." Hence: ethics, epistemology, philosophy of language, philosophy of mind. #### Model-Building as Relaxed Abduction > "We can treat model-building as a special case of the abductive methodology in which the requirements of evidential fit are relaxed." This is important: models aren't expected to be strictly true. We learn from them despite knowing their assumptions are false. #### Counterexamples vs Better Models > "What defeats a model is not a counterexample but a better model, one that retains its predecessor's successes while adding some more of its own." Williamson notes that Gettier-style counterexamples would be "unpublishable in economics" — pointing out a model fails to fit perfectly isn't newsworthy in model-building disciplines. #### The Knowledge Payoff > "When we explore a model by valid deductive reasoning from the model description, we learn necessary truths of the general conditional form 'If a given case satisfies the model description, then it satisfies this other description too.'" Plus vague unconditional knowledge: "this model fits the phenomenon better than that one." ### Relevance to Generating Philosophy Several threads connect to your project: 1. **Abduction as skill, not mystery**: Williamson treats abductive inference as a *learnable competence* — "good judgment, honed by experience." If philosophers learn it, could an LLM trained on philosophical corpora also acquire it? 2. **The mathematics precedent**: Philosophy's evidence base can be *endogenous* (like mathematics) — philosophical arguments can serve as evidence for other philosophical arguments. This supports the "self-grounding" thesis in your project. 3. **Model-building as fallibilist methodology**: The explicit acknowledgment that models involve false assumptions but still yield knowledge suggests a framework where LLM outputs could be useful *despite* not being strictly truth-apt — they could function as philosophical models to explore. 4. **Criteria are aesthetic but not arbitrary**: Williamson defends simplicity/elegance as truth-conducive (via Forster-Sober overfitting arguments). These are the kinds of criteria that *are* encoded in philosophical training data. --- ## Connecting Williamson to Floridi **Nick:** Please relate what williamson says on these matters to floridi's claim about what llms can't do **Claude:** *[Extracts and reads Floridi et al.'s "What Kind of Reasoning (if any) is an LLM actually doing?"]* ### Floridi's Core Claim Floridi argues that LLMs have a **stochastic core** but an **abductive appearance**: > "LLMs generate text based on learned associations rather than performing abductive inferences. When their output exhibits an apparent abductive quality – often reinforced by interface design – this effect is due to the model's training on human-generated texts that encode reasoning structures." The key limitation he identifies: > "They can help generate hypotheses and support human reasoning, but their outputs must be critically examined because they cannot discern truth or verify explanations." He invokes Reichenbach's distinction between **context of discovery** and **context of justification**: > "LLMs seem to perform only the first part. They generate candidates (explanations, answers) but do not genuinely validate them against reality... They aim to model the conditional distribution of tokens in text, not to evaluate truth." Floridi coins the term **"zeroth-order abduction"**: > "LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations." ### Where Williamson Creates Pressure on Floridi #### 1. The Justification Gap May Not Apply to Philosophy Floridi's picture assumes that *justification* requires something beyond what LLMs can do — external verification, truth-checking, grounded semantics. But Williamson's central argument is that **abduction just is a form of justification** in philosophy: > "I still favor inference to the best explanation and an abductive methodology in philosophy... Indeed, it is hard to see how the kind of positive, systematic, general theory that Dummett sought in the philosophy of language could be established by any other means." If philosophical justification proceeds *through* abduction (assessing theories by simplicity, elegance, explanatory power, coherence with other commitments), rather than by external verification, then Floridi's "discovery without justification" framing may mischaracterize what philosophy actually requires. #### 2. The Mathematics Precedent: Self-Grounding Domains Williamson's most striking claim is that **mathematics provides a precedent for armchair abduction**: > "The first principles must be strong enough to prove all established mathematical theorems; they must not be strong enough to prove a contradiction. Subject to those constraints, we naturally want the first principles to maximize simplicity and similar virtues." The axioms of ZFC are justified *by their fruits within mathematics* — not by external empirical verification. The evidence base is endogenous: > "The most important evidence on which to base the relevant abductions for foundational inquiries in mathematics is itself mathematical – less foundational mathematics." If philosophy is similarly **self-grounding** — if philosophical arguments can serve as evidence for other philosophical arguments — then Floridi's worry about "no external verification" loses much of its force. The "stochastic core" isn't disqualifying if the domain doesn't require exogenous grounding. #### 3. "Learned Associations" vs "Abductive Competence" Here's the crux. Floridi writes: > "LLMs have effectively absorbed patterns of human abductive reasoning as expressed in writing. However, when faced with inputs that go beyond their training (truly novel situations or complex multi-step logical puzzles), the facade can crack." But consider Williamson on what abductive competence *is*: > "Selecting and interpreting models is an art – in science as well as in philosophy. It depends on good judgment, honed by experience." If abductive competence is "judgment honed by experience" — and if that judgment is expressible in argumentative patterns in text — then training on those patterns could *transmit* the competence, not merely its superficial form. Floridi treats the learned associations as a bug ("just pattern matching"); but if the patterns *are* the reasoning, it might be a feature. #### 4. The Criteria Are In The Training Data Williamson identifies the criteria for good abductive inference: > "It should be informative and general. In brief, it should combine simplicity with strength... Abduction rewards virtues such as simplicity, elegance, generality, and unificatory power." Floridi agrees that LLMs have "absorbed patterns of human abductive reasoning." But those patterns *include* the evaluative criteria — what counts as a better or worse explanation. The philosophical corpus doesn't just contain explanations; it contains *assessments* of explanations, *arguments about* what makes explanations good, *demonstrations* of how to apply these criteria. This is your **dialectical saturation thesis**: the training data contains not just moves, but evaluations of moves, sequences of moves, and meta-level discussion of what makes moves good. ### The Structural Disagreement | | **Floridi** | **Williamson** | |---|---|---| | **What abduction requires** | Truth-tracking, verification, grounded semantics | Meeting criteria: simplicity, elegance, explanatory power, coherence | | **Role of external verification** | Essential — the missing piece | Not essential for all domains — mathematics and logic proceed without it | | **Status of "learned associations"** | Surface mimicry without understanding | Could constitute competence if patterns encode the skill | | **Discovery vs justification** | Sharp distinction — LLMs can only do discovery | Blurred in philosophy — abduction does both | ### The Generating Philosophy Thesis Your saturation thesis can now be stated more precisely as a response to Floridi: **Floridi's claim**: LLMs produce abductive *appearance* without abductive *substance* because they lack verification/grounding. **Your response**: Philosophy (unlike empirical domains) may be a **self-grounding** domain where: 1. The criteria for evaluating philosophical moves are themselves philosophical 2. Those criteria are abundantly represented in the training corpus 3. "Verification" in philosophy often proceeds by further abductive assessment, not external checking 4. Therefore, absorbing the patterns *is* acquiring the competence, not merely mimicking it The Williamson texts provide formal backing: if mathematics can justify its first principles abductively, using only mathematical evidence, then philosophy might be able to do likewise — and an LLM trained on philosophical text would have access to both the moves *and* the evaluative standards. ### What Remains to Be Settled 1. **What falls outside?** Even if scheme-governed argumentative moves are learnable, Williamson notes that model-building requires "distinguishing simplifications that abstract away inessential complications from those that abstract away crucial features." Does this require something beyond pattern recognition? 2. **Novel situations**: Floridi's strongest point is that LLMs fail on "truly novel situations." But is philosophical novelty really so different from recombination? Your "combinatorial novelty" thread addresses this. 3. **The criteria for criteria**: Even if LLMs have learned what philosophers *say* makes a good explanation, do they have the capacity to *apply* those criteria in new cases? Or do they just reproduce the assessments they've seen? --- ## Connecting to Two Conceptions of Understanding **Nick:** also, how does everything said so far in this conversation link up to the two conceptions of understanding that we have written about in that note? **Claude:** *[Reads [[Theoretical Understanding - Bengson vs Dellsén]]]* This is exactly the right connection to draw. Let me map the threads together. ### The Two Accounts of Understanding To recap the note: **Bengson, Cuneo & Shafer-Landau**: Understanding = fully grasping a theory with six properties. Four are *fundamental*: 1. **Accuracy** — the theory must be largely correct 2. **Reason-based** — positively supported by considerations *beyond mere coherence* 3. **Robust** — answers many important questions about the domain 4. **Illuminating** — genuinely explanatory, not just descriptive Two are *conditional* (contribute only if the first four are present): 5. **Orderly** — reveals how features hang together 6. **Coherent** — fits with understanding-providing theories of other domains **Dellsén**: Understanding = grasping a sufficiently accurate and comprehensive *dependency model*. Two dimensions: - **Accuracy** — correctly depicts dependence relations - **Comprehensiveness** — covers the relevant relations (including *negative* facts about what doesn't depend on what) No requirement for justification or explanation. Understanding can come apart from both. ### How Floridi Maps Onto Bengson Floridi's critique of LLMs aligns structurally with Bengson's critique of Reflective Equilibrium: | **Bengson on RE** | **Floridi on LLMs** | |-------------------|---------------------| | RE guarantees outputs are "robust, orderly, and coherent" | LLMs produce coherent, plausible explanatory text | | But RE "infamously fails to put inquirers on track to achieve even a modicum of accuracy" | LLMs "cannot discern truth or verify explanations" | | RE outputs may be "unsupported by any consideration, beyond coherence, that speaks in their favor" | LLM outputs are "learned associations" not "genuine abductive inferences" | | RE risks systematizing errors into coherent but false theories | LLMs "hallucinate" — producing convincing but fabricated answers | Bengson's line: RE produces theories that are *orderly and coherent* but fail on the *fundamental* properties (accuracy, reason-based support, illumination). Mere coherence is insufficient for understanding. Floridi's line: LLMs produce outputs with *abductive appearance* but lack the *substance* — no verification, no grounding, no truth-tracking. **Both critiques turn on the same structural worry**: internal coherence without external accuracy. ### Where Williamson Creates Pressure Williamson's defence of abductive methodology cuts against this worry in two ways: #### 1. Reasons Can Be Endogenous Bengson requires theories to be "reason-based" — supported by considerations beyond mere coherence. But he doesn't require those reasons to be *exogenous* (from outside the domain). Williamson shows that in mathematics, abductive support is *endogenous*: > "The most important evidence on which to base the relevant abductions for foundational inquiries in mathematics is itself mathematical – less foundational mathematics." The axioms of ZFC are supported by *mathematical* considerations (proving established theorems, not proving contradictions, maximizing simplicity). No external verification is needed. If philosophy is similarly self-grounding, then: - The "reasons" that make a philosophical theory reason-based are *themselves philosophical* - Those reasons are abundantly represented in the training corpus - LLMs trained on philosophy have access to both theories *and* the reasons that support them This would mean LLM outputs could be *genuinely reason-based* in Bengson's sense — not merely coherent — if the training data encodes the relevant support relations. #### 2. Model-Building Relaxes Accuracy Requirements Williamson's account of model-building offers another angle: > "We can treat model-building as a special case of the abductive methodology in which the requirements of evidential fit are relaxed." In model-building disciplines, you learn from models *despite knowing their assumptions are false*. The payoff is: - Precise conditional knowledge: "If X satisfies this description, then Y" - Vague unconditional knowledge: "This model fits the phenomenon better than that one" If LLM outputs function as *philosophical models* — not direct truth claims but structured representations that illuminate without being strictly accurate — they could yield understanding even by Bengson's standards, as long as they're *illuminating* and *supported by reasons*. ### Where Dellsén Creates Even More Space Dellsén's account is more permissive than Bengson's in two crucial ways: #### 1. No Justification Requirement > "Justification *promotes* but doesn't *constitute* understanding. You can understand via an accurate dependency model regardless of whether you have reasons for thinking the model is accurate." This directly undercuts Floridi's worry. Even if LLMs don't "verify" their outputs, if those outputs are *actually accurate* representations of philosophical dependency structures, they yield understanding. The question becomes purely empirical: are LLM outputs accurate? #### 2. No Explanation Requirement Dellsén shows understanding can come apart from explanation: - Understanding *brute facts* — things with no explanation at all - Understanding *independence* — learning what something *doesn't* depend on This matters because Floridi's critique focuses on LLMs not doing *genuine* abduction/IBE. But if understanding doesn't require explanation, then even "zeroth-order abduction" (pattern-matched plausible continuations) could yield understanding, as long as the dependency structures are accurately represented. ### The Synthesis Here's how everything connects: **Floridi's critique** assumes something like Bengson's framework: - Understanding requires *reasons* beyond mere coherence - LLMs produce coherence but lack reasons/verification - Therefore LLMs can't produce understanding **Williamson's response** (implicit): - In self-grounding domains, reasons are *internal* - Abductive criteria (simplicity, elegance, explanatory power) *are* the reasons - Those criteria are encoded in training data - Therefore LLMs may have access to the reasons, not just the coherence **Dellsén's response** (more radical): - Understanding doesn't require reasons at all — just accurate dependency models - If LLM outputs accurately represent how philosophical concepts depend on each other, understanding is achieved - The "verification" worry drops out as irrelevant **Your saturation thesis** can now be stated more precisely: > Philosophical corpora don't just contain *coherent* argumentative patterns (which would be Bengson-problematic). They contain *reason-giving* structures: arguments that support theories, criteria that distinguish good from bad explanations, meta-level discussion of what makes philosophical moves successful. If LLMs have learned these structures, their outputs may be *genuinely reason-based* (satisfying Bengson) or *dependency-accurate* (satisfying Dellsén), not merely coherent. ### The Empirical Question That Remains The note already identifies the key empirical question: > "Do AI systems tend to produce *mere coherence* (like RE outputs) or *genuinely reason-based* theories? This is contingent on training and capabilities, not a conceptual point about understanding." This is what your project needs to address. The conceptual space is open: - Neither Bengson nor Dellsén creates an in-principle barrier to AI-generated understanding - Floridi's critique assumes a Bengson-like framework but doesn't engage with self-grounding domains - Williamson's abductive methodology shows that reasons can be domain-internal The question is whether LLMs have learned *support relations* (which arguments support which theories, which criteria apply when) or merely *surface patterns* (how philosophical text tends to be formatted). Your three versions of the saturation thesis are essentially three empirical hypotheses about this: 1. **Script Competence**: LLMs have learned move-sequences (surface patterns) 2. **Latent-Game Inference**: LLMs have learned which game is being played (meta-level structure) 3. **Salience-Not-Frequency**: LLMs have learned what makes moves *apt* (evaluative criteria) Only version 3 clearly satisfies Bengson's "reason-based" requirement. But all three might satisfy Dellsén's "accurate dependency model" requirement, if the patterns they've learned correspond to actual philosophical dependency structures. --- ## Sources Consulted - Williamson, T. (2007). "Widening the Picture." In *The Philosophy of Philosophy*. Oxford: Blackwell, pp. 313-403. - Floridi, L., Morley, J., Novelli, C., & Watson, D. (2025). "What Kind of Reasoning (if any) is an LLM actually doing? On the Stochastic Nature and Abductive Appearance of Large Language Models." - Bengson, J., Cuneo, T., & Shafer-Landau, R. (2024). *Philosophical Methodology: From Data to Theory*. Oxford University Press. - Dellsén, F. (2020). "Beyond Explanation: Understanding as Dependency Modelling." *British Journal for the Philosophy of Science*, 71, 1261-1286. ## Related - [[Dialectical saturation thesis]] - [[Philosophy as self-grounding domain]] - [[Theoretical Understanding - Bengson vs Dellsén]] - [[Floridi's critique of LLM abduction]] - [[Sessions/Generating Philosophy]]