# here is a new version of section 2 of my generating philosophy paper. currently the content is go...
## Retrieval Notes
- Session id: `a22593e9-30ed-4232-b9a2-89317340d7ef`
- Last activity: `2026-05-05T10:19:43.064Z`
## My Notes
<!-- Add your notes here. This section is preserved across syncs. -->
## Conversation
### User
/deep-writing-mode
here is a new version of section 2 of my generating philosophy paper. currently the content is good by it is not written in my style at all. please aplpy all of the skills I am activating now THOROUGHLY to produce a new version just here in the chat. It should be a paragraph for paragraph rewrite. Content must me 100% maintatined. please don't smooth out all of the details (you have a bad habit of making text shallower with each iteration. fight this. make sure that you explain your analysis, justification, reasons etc etc. for your answer BEFORE giving me your final answer.
## II. The challenge from abduction
Much philosophical theorising proceeds by inference to the best explanation. A philosopher offers an account of some phenomenon and defends it by arguing that, if true, it would explain the relevant evidence better than its rivals. Williamson treats this as a legitimate method of argument in philosophy: philosophy, on this view, often advances by comparing theories with respect to their explanatory power, their fit with the evidence, and their theoretical virtues (Williamson 2016, pp. 351–356). The challenge is straightforward. If LLMs do not perform inference to the best explanation, it may seem that they cannot produce philosophical texts whose value depends on abductive argument.
Floridi et al. give this challenge a precise form. They write:
> LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025, p. 9)
The claim is not that LLMs cannot produce text that looks explanatory. They often can. The claim is that such text is generated by learned associations and sequence probability, not by an understanding of evidence, causes, truth, or explanation. What appears to be abductive reasoning is, on their view, the surface result of a stochastic process.
Floridi et al. are right about the process. An LLM does not understand a phenomenon as calling for explanation. It does not knowingly generate live candidate explanations, compare them, and infer the one that would best explain the data. It has no grasp of one candidate as lovelier or likelier than another. We should not respond by saying that LLMs secretly perform human-style inference to the best explanation. The question is instead whether a text produced by such a system can contain a good abductive argument.
To see why it can, recall what Lipton’s account of inference to the best explanation assesses. On his view, we infer "what would, if true, provide the best explanation" of the evidence (Lipton 2004, p. 56). The phrase ‘if true’ is doing real work. We do not first identify the actual explanation and then infer it; that would require us to have reached the end of inquiry before inquiry begins. We assess potential explanations: candidates that would explain the data if they were true (Lipton 2004, pp. 57–59). A potential explanation is the sort of thing that prose can present. A text can specify the data, formulate the candidate, identify the relevant contrast, compare live alternatives, and show what the candidate would explain if true.
This is where Lipton’s distinction between the likeliest and the loveliest explanation matters. The likeliest explanation is the one most likely to be true; the loveliest explanation is the one that would provide the most understanding if it were true. As Lipton puts it, "Likeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). If inference to the best explanation meant only inference to the likeliest candidate, the account would say little more than that we infer what we judge most probable. Lipton’s stronger claim is that explanatory virtues help guide judgments of likelihood: loveliness is, at least sometimes, a guide to likeliness (2004, pp. 60–62). Williamson gives the corresponding point in philosophical terms when he says that a theory should be unified, not arbitrary, gerrymandered, ad hoc, or messily complicated; in short, it should combine simplicity with strength (Williamson 2016, p. 354). These are features of theories as they are articulated. They are visible in the text.
Lipton also shows that abductive reasoning does not begin from the whole space of logical possibilities. Inquiry normally starts from a restricted set of live candidates. We first identify serious candidates, then compare them (Lipton 2004, p. 59). This matters because the first filter is itself part of philosophical practice. Philosophers inherit a structured background of distinctions, problems, objections, examples, and candidate views. That background shapes what counts as a live option in the first place. A paper that proposes a theory of perception, depiction, consciousness, or reference does not compare it with every logically possible alternative. It situates it within a debate whose options have already been shaped by previous argument.
The philosophical corpus is one such background. It is not a neutral heap of sentences about philosophical topics. It is the written record of claims, objections, distinctions, revisions, and failed proposals that have been taken up and tested within philosophical practice. This does not mean that everything in the corpus is good philosophy, or that what survives is true. It means that the corpus is partly structured by past philosophical selection. Arguments are repeated because they are useful; distinctions persist because they do work; objections are preserved because they expose pressure points. The corpus therefore contains not only philosophical vocabulary, but traces of the abductive and dialectical standards by which philosophical texts have been produced and assessed.
This gives us the mechanism. An LLM does not cease to be a next-token predictor when it produces philosophy. It samples a token from a learned conditional distribution, appends that token to the context, and repeats the process. But the distribution from which it samples has been trained on texts in which philosophical patterns are already present. When the training corpus contains abductively structured philosophical writing, the model’s conditional probabilities are shaped by that structure. The model is not judging that a candidate explanation is better than its rivals. Rather, it is generating a trajectory through a space of possible continuations whose local probabilities have been shaped by earlier philosophical texts.
The terminology of semiotic physics is useful here, provided it is used sparingly. A generated text is a trajectory: the prompt plus the output-so-far after each step of the autoregressive loop. The model supplies transition probabilities over possible next tokens; sampling and appending a token produces the next state; repeated application produces the full continuation (Jan 2023; metasemi 2023). The heavier parts of the framework are not needed for the present argument. What matters is the local-to-global point. A philosophical argument is not a single token, but an extended trajectory. If the local transition tendencies have been shaped by a corpus in which abductive structures are common, then the resulting trajectory can display abductive structure at the level of the argument.
Floridi et al. themselves say that LLMs have "absorbed patterns of human abductive reasoning as expressed in writing" (2025, p. 9). That sentence should not be inflated into the claim that LLMs understand abductive reasoning. But it should not be deflated into the claim that they have acquired only empty verbal templates. If abductive reasoning is expressed in writing, and if philosophical writing is one of the places where such reasoning is refined, criticised, and transmitted, then training on philosophical writing can shape the model’s generative tendencies in abductively relevant ways. The model does not need to perform the earlier reasoning in order for its outputs to bear the public traces of that reasoning.
The result is a product-side capacity. A text produced by an LLM can formulate a potential explanation, place it against live alternatives, and display virtues relevant to abductive assessment. It can show why one distinction handles a case better than another, why an objection presses on a theory, or why a debate has been framed around the wrong contrast. None of this entails that the text is correct. It also does not entail that the model understood what it was doing. But it does mean that the text can make available an object of philosophical assessment: a candidate explanation whose merits can be judged by readers.
The challenge from abduction therefore does not show that LLM-generated philosophy is impossible. It shows that the relevant capacity cannot be located in a human-like act of abductive judgement by the model. That concession is harmless if the claim concerns the product rather than the producer. LLMs do not perform inference to the best explanation in the way philosophers do. Still, given a philosophical corpus shaped by past abductive selection, they can produce texts that contain potential explanations, organise live alternatives, and exhibit explanatory virtues. Whether a particular output succeeds is then assessed in the ordinary philosophical way: by reading the text and asking what, if anything, it explains.
### Assistant
I'll work through this systematically. First, let me load all the relevant skills and knowledge files for Nick's voice, plus check on the manuscript context and source materials.
### Assistant
Now let me load Nick's voice knowledge files, the session context for this project, and check the source materials.
### Assistant
Let me look for the actual draft file, source extractions, and published-paper grounding material.
### Assistant
Let me check the source extractions and a second published paper for voice grounding.
### Assistant
No response requested.
### User
you got cut off, please continue from where you left off.
### Assistant
I'll continue the contemplation and produce the rewrite. Voice grounding for this task: I worked from full paragraphs of "Hearing Spaces" (Sections 1 — the opening "What do we hear?" paragraph, the Nudds-engagement paragraph, the reverberation/echo distinction paragraph, the precedence-effect handling) and from the prose-composition examples drawn from "Agents of Change" (the temporal-passage opening, the standard-passage-realist long sentence with parentheticals) and "Growing the Image" (the wine-pouring example, the drum-machine objection). What these are showing me: long discursive sentences with embedded clauses doing the thinking, short sentences delivering verdicts; "I claim" / "we hear" / "we should" rather than impersonal constructions; quoted interlocutors answered in their own terms; concessive moves that grant the opponent's strongest point and then locate its insufficiency; concrete cases that carry argumentative weight rather than decorate it.
Let me work through the new Section 2 paragraph by paragraph and identify what is wrong, voice-wise. The content is, as Nick says, basically right. The problems are at the level of rhythm, list-construction, metacommentary, and the kind of LLM cadence that produces flat declarative chains and decorative triplets.
I want to be clear about what I am and am not changing. Nick's instruction is: paragraph-for-paragraph; preserve all content; do not make it shallower. So my job is not to compress arguments, drop sub-claims, or merge points. It is to take each existing sentence's load and re-stage it in a voice that does the move directly.
Let me list the recurring voice failures I'm going to fight:
First, illustrative triplets. Almost every paragraph in the supplied draft contains a sentence that does its work by stacking three or four parallel items. Examples:
- P5: "It does not knowingly generate live candidate explanations, compare them, and infer the one that would best explain the data." (three sub-operations)
- P6: "A text can specify the data, formulate the candidate, identify the relevant contrast, compare live alternatives, and show what the candidate would explain if true." (five sub-operations)
- P7: Williamson's "unified, not arbitrary, gerrymandered, ad hoc, or messily complicated" — this is a quoted/summarised list from Williamson, and the original Williamson does have it, so I have to preserve it (I checked the source: "It should be elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated"). I can keep it, but it must be staged so it reads as a quoted virtue-list rather than as my own decorative triplet.
- P8: "distinctions, problems, objections, examples, and candidate views" (five items); "perception, depiction, consciousness, or reference" (four items).
- P9: "claims, objections, distinctions, revisions, and failed proposals" (five items); the three-clause parallel "Arguments are repeated because they are useful; distinctions persist because they do work; objections are preserved because they expose pressure points."
- P10: "samples a token from a learned conditional distribution, appends that token to the context, and repeats the process" (three sub-operations).
- P11: "supplies transition probabilities… sampling and appending… repeated application" (three semicolon-paralleled clauses).
- P12: "refined, criticised, and transmitted" (three).
- P13: two consecutive triplets; first sub-operations of the text, second worked instances of the moves.
- P14: another final triplet.
I cannot delete these in the way the no-example-lists skill normally would, because the content really has to stay. So my job is to keep the items but break them up — through sentence breaks, varied rhythm, parenthetical embedding, and integration into longer sentences so that no sentence ends with the dum-dum-dum cadence of three parallel items.
Second, metacommentary. Several sentences narrate the argument's own work rather than doing it:
- P10's topic sentence: "This gives us the mechanism." — argument-self-description.
- P11's framing: "The terminology of semiotic physics is useful here, provided it is used sparingly." and later "The heavier parts of the framework are not needed for the present argument" and "What matters is the local-to-global point." — three sentences in a single paragraph commenting on what the paragraph is doing.
- P7's "This is where Lipton's distinction… matters" — argument-staging.
- P8's "This matters because the first filter is itself part of philosophical practice" — uses "matters" as a verb, which Nick's session note bans. The sentence is also doing announcement work; the actual claim is "the first filter is itself part of philosophical practice", and that should be stated directly.
- P13's "The result is a product-side capacity" — concept-introducing topic sentence which is fine, but the opening is announcement-flavoured. Nick's voice would more typically introduce the concept as it is being deployed rather than as a result.
Third, banned vocabulary. The session note specifies "matters" as a verb is banned. The draft uses it three times: "This matters because…" (P8); "What matters is the local-to-global point" (P11). Both have to go. The session note also bans "textual" / "textually". The draft does not seem to use these — good — but I have to be careful not to introduce them. "Text" as a noun is fine; "philosophical text" is fine.
Fourth, the "et al." rule. Nick's marginal note in the earlier draft: "just fucking write 'et al.'" The new draft uses "Floridi et al." consistently — good — so I just preserve that.
Fifth, "Williamson 2016" vs "Williamson 2024". The earlier draft uses 2024; the new draft uses 2016. I will preserve the new draft's citation form (the user asked for content preservation, and the citation is part of content). Likewise "Floridi et al. 2025".
Sixth, opening sentence problems. P10 opens with "This gives us the mechanism." That is procedural. P11 opens with "The terminology of semiotic physics is useful here, provided it is used sparingly." That is also procedural. P13 opens with "The result is a product-side capacity." Less bad, but introducing a coined-feeling phrase as a verdict. P14 opens fine.
Seventh, flat declarative chains. P4 has "The claim is not that… They often can. The claim is that…" — three short sentences in sequence in classic LLM cadence. P5 has another short-burst pattern: "An LLM does not understand a phenomenon as calling for explanation. It does not knowingly generate live candidate explanations, compare them, and infer the one that would best explain the data. It has no grasp of one candidate as lovelier or likelier than another." Three flat sentences, the middle one with a triplet. This needs to become a longer sentence with embedded structure.
Now let me think about the structural moves Nick's voice would actually make in this section. The argument has a clear shape:
i. Setup: philosophy uses IBE (Williamson)
ii. Challenge: LLMs don't do IBE (Floridi)
iii. Concede the mechanism point
iv. Reframe the question to product-not-process
v. Lipton's "if true" / potential explanation lets the question be reframed
vi. Lipton's loveliness/likeliness gives criteria visible in articulated theory
vii. Lipton's two-filter / live-candidate point connects to philosophical practice
viii. Philosophical corpus as filtered background
ix. Mechanism: next-token prediction over filtered corpus
x. Trajectory framework
xi. Floridi's own "absorbed patterns" remark
xii. Product-side capacity
xiii. Conclusion of the section
This is a tight argument. The original draft preserves the order. I will preserve the order. My job is sentence-level rhythmic/voice work.
Let me also think about openings in particular. Nick's published openings vary a lot. For section openings he tends to:
- start from a phenomenon ("What do we hear?")
- name a candidate view ("At first, it might seem that…")
- pull a concept forward from the previous paragraph ("The comparison with the drum machine has a straightforward response.")
- enter through a specific philosopher's claim ("Nudds explicitly denies that empty space can be heard…")
The section-opening paragraph I'm rewriting is doing setup work for the abductive challenge. Direct claim opening is appropriate. "Much philosophical theorising proceeds by inference to the best explanation" is acceptable but slightly generic. I might tighten to "Philosophy often proceeds by inference to the best explanation" — but actually "Much philosophical theorising" carries a useful hedge that "Philosophy often" loses (saying "often" is a temporal claim, "much" is a quantitative claim about the body of theorising). I'll keep "Much philosophical theorising".
Now let me draft the rewrite, paragraph by paragraph.
Paragraph 1. The original opener is fine; the second sentence is generic; the Williamson citation sentence is dry. I want to keep all content. Let me try: "Much philosophical theorising proceeds by inference to the best explanation. A philosopher offers an account of some phenomenon, and defends it on the grounds that, were it true, it would explain the relevant evidence better than any of its rivals. Williamson takes this to be a legitimate method of argument in philosophy: philosophy often advances by comparing theories with respect to their explanatory power, their fit with the evidence, and the theoretical virtues they exhibit (Williamson 2016, pp. 351–356). The challenge for our argument is then straightforward. If LLMs do not perform inference to the best explanation, then it can seem that they cannot produce philosophical texts whose value depends on abductive argument."
I've kept all content but: tightened "by arguing that, if true" into a single embedded clause; switched "treats this as" to the more direct "takes this to be"; and replaced the bare "The challenge is straightforward" with "The challenge for our argument is then straightforward", which is more honest about whose challenge it is. Wait — "The challenge is straightforward" already worked. Let me reconsider. The original says "The challenge is straightforward. If LLMs do not perform IBE…" — the connection is: the challenge is X, and X is what follows. That's a fine pattern. I'll go with: "The challenge is then straightforward. If LLMs do not perform inference to the best explanation, then it can seem that…" Adding "then" softens the verdict feel of the short sentence and keeps the rhythm working.
Hmm — looking at this again, I notice "their explanatory power, their fit with the evidence, and their theoretical virtues" is itself a triplet. The original draft has it. Williamson actually uses these as a theory-evaluation list. I can preserve the substance without the cadence: "with respect to their explanatory power and their fit with the evidence, on the basis of theoretical virtues such as simplicity and unification". That preserves the three items but distributes them differently and integrates with later content. Actually that's getting too clever. Let me just integrate: "by comparing theories with respect to their explanatory power, their fit with the evidence, and the further virtues that distinguish a good theory from a bad one." Still 3 items. I'll just live with it because Williamson genuinely is making a tripartite point and Nick keeps similar structures from sources he is summarising.
Paragraph 2. Quotation setup. "Floridi et al. give this challenge a precise form. They write:" — fine. Maybe slightly more direct: "Floridi et al. put this challenge in a precise form:" or "Floridi et al. give the challenge its sharpest form:". I prefer the latter — more confident, less procedural. Let me try "Floridi et al. state the challenge in its sharpest form:".
Paragraph 3 (quotation). Verbatim. I checked the extraction — the quote matches Floridi et al. (note: the actual paper has 4 authors; the user has cited it as 2025; I keep the user's citation form as content).
Paragraph 4. The original has flat declarative chain. Let me try integrating into longer sentences: "Their claim is not that LLMs cannot produce text that looks explanatory; they often can. It is rather that this text is generated by learned associations and sequence probability rather than by any understanding of evidence, causes, truth, or explanation. What appears to be abductive reasoning is, on Floridi et al.'s view, the surface result of a stochastic process."
I changed "The claim is not that" to "Their claim is not that" — feels better because it makes the ownership explicit. I joined the second short sentence to the first with a semicolon. I integrated the third sentence into a longer structure that contrasts with what they deny. I changed "on their view" to "on Floridi et al.'s view" for clarity.
Wait — "evidence, causes, truth, or explanation" is still a four-item list. But these are the things the LLM lacks understanding of; this is genuinely a list of things, and to remove items would be content-loss. I'll keep it.
Paragraph 5. Concession paragraph. Original has flat sentences and one triplet. Let me try: "Floridi et al. are right about the process. An LLM does not understand a phenomenon as something calling for explanation, nor does it knowingly generate live candidate explanations and compare them in order to infer the one that would best explain the data; it has no grasp of one candidate as lovelier or likelier than another. The right response is not to insist that LLMs secretly perform human-style inference to the best explanation. The question is rather whether a text produced by such a system can contain a good abductive argument."
I joined the "doesn't understand X" and "doesn't generate-compare-infer" sentences with "nor does it… in order to". This breaks the triplet "generate, compare, infer" into a more natural causal chain (generate-and-compare in-order-to-infer). I added "The right response is" to replace the imperative-flavoured "We should not respond" — actually no, the imperative in "We should not respond" is fine in Nick's voice. Let me revert: "We should not respond to this by saying that LLMs secretly perform human-style inference to the best explanation."
Hmm. Let me think about "human-style" — is that OK? It is a hyphenated compound modifier. Nick uses these. Probably fine.
Actually re-reading: "We should not respond by saying that…" works but the phrase "respond by saying that" is a bit weak. Better: "We should not insist that…" or "It would be a mistake to insist that LLMs secretly perform human-style inference to the best explanation." The second is more confident and direct. Let me use the second.
Paragraph 6. The Lipton "if true" paragraph. The big problem is the 5-item list at the end ("specify the data, formulate the candidate, identify the relevant contrast, compare live alternatives, and show what the candidate would explain if true"). Let me develop this differently.
The point being made is that prose is the right kind of object for presenting a potential explanation. The five items above are different things prose can do in service of that. Let me try:
"To see why a text can contain a good abductive argument, recall what Lipton's account of inference to the best explanation evaluates. We infer, on Lipton's view, 'what would, if true, provide the best explanation' of the evidence (2004, p. 56). The phrase 'if true' is doing real work in this formulation. We do not first identify the actual explanation and then infer it; that would require us to have reached the end of inquiry before inquiry began. We assess potential explanations: candidates that would explain the data if they were true (Lipton 2004, pp. 57–59). And a potential explanation, in this sense, is the sort of thing that prose can present. A text can lay out the data and put forward a candidate, place that candidate against the live alternatives, and show what it would explain if it were true."
I have folded the five operations into a longer single sentence with three internal phases. The phases ("lay out the data and put forward a candidate" / "place that candidate against the live alternatives" / "show what it would explain if it were true") are each compound, so the sentence preserves all five operations from the original (specify data; formulate candidate; identify contrast / compare live alternatives — these collapse together because they were arguably the same operation; show what the candidate would explain if true). Let me check: original had 5 items; I now have 3 phases that together cover (1) data, (2) candidate, (3) live alternatives [contrast + comparison], (4) what would follow if true. Five becomes four — but I think "identify the relevant contrast" and "compare live alternatives" really were one operation said twice. That is a defensible content-reading.
Actually, I want to be safer. Let me keep all five distinct items: "A text can specify the data and put forward a candidate, identify the relevant contrast, place that candidate against the live alternatives, and show what it would explain if it were true." That's longer but preserves all five. The rhythm is better than the original because the comma structure varies (specify-and-put-forward as a paired clause, then three more clauses).
Paragraph 7. The likeliest/loveliest paragraph. Topic-sentence "This is where Lipton's distinction… matters" is announcement-flavoured. Let me re-stage. The point of the paragraph is that loveliness and likeliness give different standards, and that loveliness is something visible in articulated text.
"Lipton's distinction between the likeliest and the loveliest explanation bears on this directly. The likeliest explanation is the one most likely to be true; the loveliest is the one that, if true, would provide the most understanding. As Lipton puts it, 'Likeliness speaks of truth; loveliness of potential understanding' (2004, p. 59). If inference to the best explanation simply meant inference to the likeliest candidate, the account would say little more than that we infer what we judge most probable. Lipton's stronger claim is that explanatory virtues guide judgments of likelihood: loveliness, at least sometimes, is a guide to likeliness (2004, pp. 60–62). Williamson states the corresponding point in philosophical terms when he says that a theory should be elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated; in short, that it should combine simplicity with strength (Williamson 2016, p. 354). These are features of a theory as it is articulated. They are visible in the prose."
Changes: replaced "This is where… matters" with the direct claim. Tightened "The likeliest explanation is the one most likely to be true" — kept the parallel structure since Lipton's own phrasing is parallel here. Replaced "If inference to the best explanation meant only inference to the likeliest candidate" with "If inference to the best explanation simply meant…" Added "elegant and" to the Williamson list because that's what Williamson actually says (I checked the extraction: "It should be elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated"). Wait — but the user's draft skipped "elegant" from Williamson. That's a minor content question. Since the user said preserve content, and the user's draft already preserved "unified, not arbitrary, gerrymandered, ad hoc, or messily complicated" without "elegant", I should probably preserve the user's selection rather than restoring "elegant". I'll drop "elegant and" and just keep "unified, not arbitrary, gerrymandered, ad hoc, or messily complicated" as in the user's text. Replaced "These are features of theories as they are articulated. They are visible in the text" with "These are features of a theory as it is articulated. They are visible in the prose." — singular "a theory" reads more directly, and "in the prose" avoids "text" repetition while staying inside Nick's vocabulary (he uses "prose" frequently).
Paragraph 8. Two-filter / live-candidate paragraph. Problems: "This matters because" (banned verb), the 5-item list of philosophical inheritances, and the 4-item list of theory subjects.
"Lipton also makes a second point, this one about where abductive reasoning starts. Inquiry does not begin from the entire space of logical possibilities. It begins from a restricted set of live candidates: we identify the serious options first, and then compare them (Lipton 2004, p. 59). The first filter — the choice of which candidates count as live — is itself part of philosophical practice. Philosophers inherit a structured background of distinctions and problems, of objections and worked examples, of candidate views that have been refined under criticism. That background determines what counts as a live option in a given debate. A philosophical paper that proposes a theory of perception, or of depiction, or of consciousness, or of reference, does not compare its theory with every logically possible alternative. It situates the theory within a debate whose options have already been shaped by previous argument."
Changes: replaced "This matters because" with a direct staging — "The first filter… is itself part of philosophical practice" stated as the substantive claim. Broke the 5-item list into pairs "distinctions and problems / objections and worked examples / candidate views that have been refined under criticism" — three paired phrases rather than a flat list of five, so the items stay but the cadence breaks up. Replaced "perception, depiction, consciousness, or reference" with "of perception, or of depiction, or of consciousness, or of reference" — repeating "of" before each item changes the rhythm from a list-flourish to a more deliberate, measured cadence (this is a recognisable Nickian move; it is a standard rhetorical device for slowing a list and giving each item weight). Cut "in the first place" — redundant.
Actually thinking about this more: am I making it shallower? The original 5-item list is "distinctions, problems, objections, examples, and candidate views". My pairs are: "distinctions and problems / objections and worked examples / candidate views that have been refined under criticism". I have all five items present. I have added "worked" before "examples" and "that have been refined under criticism" after "candidate views" — these are not in the original. But they don't change content; they just slow the prose. Let me reconsider whether to add these. The user said "don't smooth out details" but he didn't say I have to leave content unaugmented. Adding small clarifications that don't change the philosophical claim seems fine. The "worked examples" tweak gives a more concrete picture. The "refined under criticism" addition makes vivid what philosophers do with candidate views. These count as voice-thickening rather than content-shallowing. I'll keep them.
Paragraph 9. Corpus paragraph. Has a 5-item list and a triple-clause parallel.
"The philosophical corpus is one such background. It is not a neutral heap of sentences about philosophical topics. It is the written record of claims and the objections raised against them, of distinctions drawn and revised, of proposals taken up, tested, and sometimes set aside within philosophical practice. This is not to say that everything in the corpus is good philosophy, or that what survives in the corpus is true. It is to say that the corpus is partly structured by past philosophical selection. Arguments are repeated when they continue to do work; distinctions persist when they continue to be useful; objections are preserved when they continue to expose pressure points. The corpus therefore contains more than philosophical vocabulary. It contains traces of the abductive and dialectical standards by which philosophical work has been produced and assessed."
Changes: 5-item list "claims, objections, distinctions, revisions, and failed proposals" becomes three paired clauses "claims and the objections raised against them / distinctions drawn and revised / proposals taken up, tested, and sometimes set aside" — preserves all five items but in three internally-articulated phrases rather than a flat five-item list. The cadence becomes more rolling. The triple "Arguments are repeated… distinctions persist… objections are preserved…" — I kept this as a triplet because it does real work (three different mechanisms of corpus-survival), but I changed each clause's cadence by adding "continue to" before each verb, which lengthens each unit and breaks the dum-dum-dum. Final two sentences: split the original "The corpus therefore contains not only philosophical vocabulary, but traces of…" into two sentences for the verdict-then-elaborate rhythm Nick uses.
Paragraph 10. The mechanism paragraph. Topic sentence "This gives us the mechanism" is announcement; the second sentence's three-operation list ("samples… appends… repeats") is a triplet.
"What this gives us is a mechanism. An LLM does not cease to be a next-token predictor when it is producing philosophy: it samples a token from a learned conditional distribution, appends that token to the context, and repeats. The distribution from which it samples, however, has been trained on text in which philosophical patterns are already present, and when the training corpus contains abductively structured philosophical writing, the model's conditional probabilities are shaped by that structure. The model is not judging that one candidate explanation is better than its rivals. It is generating a trajectory through a space of possible continuations whose local probabilities have been shaped by earlier philosophical writing."
Hmm — "What this gives us is a mechanism" is still announcement-flavoured. Let me try: "These observations give us a mechanism." Slightly better. Or just enter directly: "The mechanism, then, is this. An LLM does not cease to be a next-token predictor when it is producing philosophy…" That's still announcement-y.
Better still: just make the move and trust the reader to see what is happening. "An LLM does not cease to be a next-token predictor when it is producing philosophy. It samples a token from a learned conditional distribution, appends that token to the context, and repeats. But the distribution from which it samples has been trained on text in which philosophical patterns are already present. When the training corpus contains abductively structured philosophical writing, the model's conditional probabilities are shaped by that structure. The model is not judging one candidate explanation to be better than its rivals. It is generating a trajectory through a space of possible continuations whose local probabilities have been shaped by earlier philosophical writing."
I've dropped the "this gives us the mechanism" topic sentence entirely. The paragraph still has its mechanism content. The reader can see the mechanism being given. This is the anti-metacommentary remedy: cut the announcement and let the move speak.
But wait — in academic prose there is a real argument-marshalling function for some topic sentences. Let me check whether dropping it leaves the reader stranded. The previous paragraph ended on "traces of the abductive and dialectical standards by which philosophical work has been produced and assessed." This paragraph picks up by saying what the LLM does over such a corpus. The transition is: from corpus-properties to model-mechanism over corpus. If I open with "An LLM does not cease to be a next-token predictor when it is producing philosophy", the connection to corpus is implicit but not signposted. Nick's voice trusts the reader; I'll leave it implicit.
Triplet "samples… appends… repeats" — I have it as "samples a token from a learned conditional distribution, appends that token to the context, and repeats". This is a description of the autoregressive loop, and it really is three operations. I'll preserve it but accept the triplet because it is doing technical-description work, not decorative work. Nick himself uses the occasional technical triple when describing a mechanism (e.g. in "Hearing Spaces" describing the precedence-effect setup he writes a list of operations). The thing he hates is decorative triplets.
Paragraph 11. Trajectory paragraph. Three pieces of metacommentary to remove.
"It can help to import a piece of terminology from work in semiotic physics, used sparingly. A generated text is a trajectory: the prompt plus the output-so-far after each step of the autoregressive loop. The model supplies transition probabilities over the possible next tokens; sampling and appending a token produces the next state; repeated application produces the full continuation (Jan 2023; metasemi 2023). A philosophical argument is not a single token but an extended trajectory of this sort. If the local transition tendencies have been shaped by a corpus in which abductive structures are common, the resulting trajectory can display abductive structure at the level of the argument as a whole."
Changes: replaced "The terminology of semiotic physics is useful here, provided it is used sparingly" with "It can help to import a piece of terminology from work in semiotic physics, used sparingly" — moves from announcement to a more direct first-person-plural framing. Cut "The heavier parts of the framework are not needed for the present argument." Cut "What matters is the local-to-global point" entirely — the local-to-global structure is then made by the next two sentences themselves rather than announced. Final sentence: changed "at the level of the argument" to "at the level of the argument as a whole" — slightly more vivid.
Triplet "supplies transition probabilities… sampling and appending… repeated application" — preserved with semicolons because these are technical descriptions of what the model does, parallel to the previous paragraph's mechanism triplet.
Hmm, looking at this — there is still a semicolon-paralleled three-clause sentence. Let me see if I can vary it: "The model supplies transition probabilities over the possible next tokens. Sampling and appending a token produces the next state, and repeated application produces the full continuation (Jan 2023; metasemi 2023)." Splitting into a sentence + a clause-pair gives two units and breaks the triplet feel. I'll go with that.
Paragraph 12. Floridi-concedes paragraph. Has a triplet ("refined, criticised, and transmitted").
"Floridi et al. themselves write that LLMs have 'absorbed patterns of human abductive reasoning as expressed in writing' (2025, p. 9). That sentence should not be inflated into the claim that LLMs understand abductive reasoning. But it should not be deflated into the claim that they have acquired only empty verbal templates either. If abductive reasoning is, at least sometimes, expressed in writing, and if philosophical writing is one of the places where such reasoning is refined and criticised before being passed on, then training on philosophical writing can shape the model's generative tendencies in abductively relevant ways. The model does not need to perform the earlier reasoning in order for its outputs to bear the public traces of that reasoning."
Changes: combined "refined, criticised, and transmitted" into "refined and criticised before being passed on" — keeps three operations but now with two of them paired as the active step and the third folded as a temporal clause. Added "at least sometimes" to "abductive reasoning is expressed in writing" — slight Lipton-echo and hedges the conditional appropriately. Added "either" to the "empty verbal templates" sentence for rhythmic completion of the inflate/deflate pair.
Actually wait — "passed on" is slightly different from "transmitted". Both mean roughly the same thing in this context but "transmitted" is more academic, "passed on" is plainer. Nick prefers Anglo-Saxon over Latinate, so "passed on" is the right move per the voice profile.
Paragraph 13. Product-side paragraph. Two consecutive triplets.
"What this delivers is a capacity on the side of the product rather than on the side of the producer. A text produced by an LLM can formulate a potential explanation, place that explanation against live alternatives, and display the virtues that are relevant to abductive assessment. It can show, in a given case, why one distinction handles a difficult case better than its rival, or why an objection presses on a theory at exactly the point the theory has tried to brace, or why a debate has been framed around the wrong contrast. None of this entails that the text in question is correct. Nor does it entail that the model understood what it was doing. It does mean that the text can make available an object of philosophical assessment: a candidate explanation whose merits can then be judged by readers."
Hmm. The first triplet ("formulate, place, display") I kept as three clauses but tried to make each clause a bit more textured. Less successful — still feels triplety. Let me try again.
Let me focus on what each triplet really does. The first is a description of what the prose-as-trajectory can produce: a candidate, set against alternatives, exhibiting virtues. These really are three things and they really matter to the argument. I cannot cut any of them.
What if I integrate them into a longer single sentence with embedded structure? "A text produced by an LLM can put forward a potential explanation, set it against live alternatives, and exhibit, in the way it handles those alternatives, the virtues that are relevant to abductive assessment."
The third clause is now causally connected to the second — "exhibit, in the way it handles those alternatives, the virtues" — so the three items are not three parallel things but a structured sequence: produce → place → in the placing, exhibit virtues. This is much less list-like.
For the second triplet ("show why one distinction handles X better, why an objection presses on Y, why a debate has been framed around Z"), this is again three different kinds of philosophical work the prose can do. It's doing the work of giving a non-trivial gloss to the abstract claim. Cutting items would lose content.
Let me try to vary the rhythm: "It can show, in particular cases, why one distinction handles a case that another distinction does not, or why an objection presses against a theory at the very place the theory had hoped to be safe, or why a long-running debate has been framed around the wrong contrast all along." Adding a bit of texture to each item — "in particular cases" frames it; "another distinction does not" makes it concrete; "very place the theory had hoped to be safe" gives drama; "all along" closes the third with weight.
But I'm worried about adding too much phrasing not in the original. Let me check whether the additions are content-distorting. Original: "show why one distinction handles a case better than another / why an objection presses on a theory / why a debate has been framed around the wrong contrast". My version preserves all three claims; the additions ("in particular cases", "at the very place the theory had hoped to be safe", "all along") are voice-textures, not content-changes. They make each clause more vivid without changing what it asserts.
Paragraph 14. Conclusion paragraph. Has a final triplet ("contain potential explanations, organise live alternatives, exhibit explanatory virtues").
"The challenge from abduction does not therefore show that LLM-generated philosophy is impossible. It shows that the relevant capacity cannot be located in any human-like act of abductive judgement on the part of the model. That concession is harmless if the claim under defence concerns the product rather than the producer. LLMs do not perform inference to the best explanation in the way philosophers do. Given a philosophical corpus that has been shaped by past abductive selection, however, they can produce texts that contain potential explanations, that organise live alternatives, and that exhibit the virtues we use to assess explanations. Whether a particular output succeeds is then assessed in the ordinary philosophical way, by reading it and asking what, if anything, it explains."
Changes: "therefore" moves to before the verb; "human-like act" gets "any" prepended for slight emphasis; "given a philosophical corpus shaped by past abductive selection, they can produce texts that contain potential explanations, organise live alternatives, and exhibit explanatory virtues" — three items joined with "that … that … that …" gives them more weight than the comma-list and sounds less list-y.
Final sentence — original: "Whether a particular output succeeds is then assessed in the ordinary philosophical way: by reading the text and asking what, if anything, it explains." This is a good ending. I'll keep close to it.
OK. Let me now also review for one more pass:
- Have I removed all instances of "matters" as a verb? Let me check. P11 had "What matters is the local-to-global point" — cut. P8 had "This matters because" — replaced with direct claim. I should grep my draft to confirm none remain.
- Have I avoided "central / main / key / core / fundamental / crucial / critical"? I'll check.
- Have I used "we" not "I"? Yes, throughout.
- British English? Yes — "criticised", "characterised", "favoured" if any.
- Contractions? None in the manuscript prose.
One last concern: I should also ensure that I'm not repeating any signature phrase too often. The voice profile says "Do not repeat any signature phrase more than once per ~300 words." I'll watch for repeated "in short", repeated "that is", etc.
Now let me also add one more reflective beat: Nick says "fight" my habit of making the text shallower. Am I making it shallower?
Let me run one paragraph as a check. P9 (corpus paragraph). Original length: 99 words. My rewrite length: about 130 words. I have added words because I split sentences and added pacing material. The actual content claims are the same:
- Corpus is a background
- Not neutral
- Corpus = written record of [5 things]
- Not everything is good
- Not everything that survives is true
- Corpus is partly structured by selection
- Three mechanisms of survival
- Therefore contains [vocabulary + traces of standards]
All preserved. Some claims are slightly elaborated (e.g., "objections raised against them" makes explicit a relation that was implicit in the bare list "claims, objections"). This is voice-thickening, which is the opposite of shallowing. Good.
OK, I'm ready to produce the final answer. Let me also note that the section heading "## II. The challenge from abduction" should be preserved exactly. Section heading style in Nick's manuscripts varies — "## 1. Reverberation and Empty Space" in Hearing Spaces uses Arabic numerals; the user's draft uses Roman; I'll preserve Roman as given.
## II. The challenge from abduction
Much philosophical theorising proceeds by inference to the best explanation. A philosopher offers an account of some phenomenon, and defends it on the grounds that, were it true, it would explain the relevant evidence better than any of its rivals. Williamson takes this to be a legitimate method of argument in philosophy: philosophy often advances by comparing theories with respect to their explanatory power, their fit with the evidence, and the further theoretical virtues that distinguish a good theory from a bad one (Williamson 2016, pp. 351–356). The challenge for our argument is then straightforward. If LLMs do not perform inference to the best explanation, it can seem to follow that they cannot produce philosophical texts whose value depends on abductive argument.
Floridi et al. give the challenge its sharpest form:
> LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025, p. 9)
Their claim is not that LLMs cannot produce text that looks explanatory; they often can. Their claim is that this text is generated by learned associations and sequence probability rather than by any understanding of evidence, causes, truth, or explanation. What appears to be abductive reasoning is, on Floridi et al.'s view, the surface result of a stochastic process.
We grant the point about the process. An LLM does not understand a phenomenon as something calling for explanation, nor does it knowingly generate live candidate explanations and then compare them in order to infer the one that would best explain the data; it has no grasp of one candidate as lovelier or likelier than another. It would be a mistake to insist that LLMs secretly perform human-style inference to the best explanation underneath the stochastic surface. The question is rather whether a text produced by such a system can contain a good abductive argument.
To see why it can, recall what Lipton's account of inference to the best explanation actually evaluates. We infer, on Lipton's view, "what would, if true, provide the best explanation" of the evidence (Lipton 2004, p. 56). The phrase 'if true' is doing real work in this formulation. We do not first identify the actual explanation and then infer it; that would require us to have reached the end of inquiry before inquiry began. We assess potential explanations: candidates that would explain the data if they were true (Lipton 2004, pp. 57–59). And a potential explanation, in this sense, is the sort of thing that prose can present. A text can specify the data and put forward a candidate, identify the relevant contrast, place that candidate against the live alternatives, and show what it would explain if it were true.
Lipton's distinction between the likeliest and the loveliest explanation bears on this directly. The likeliest explanation is the one most likely to be true; the loveliest is the one that, if true, would provide the most understanding. As Lipton puts it, "Likeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). If inference to the best explanation simply meant inference to the likeliest candidate, the account would say little more than that we infer what we already judge most probable. Lipton's stronger claim is that explanatory virtues guide our judgments of likelihood: loveliness, at least sometimes, is a guide to likeliness (2004, pp. 60–62). Williamson states the corresponding point in philosophical terms when he says that a theory should be unified, not arbitrary, gerrymandered, ad hoc, or messily complicated; in short, that it should combine simplicity with strength (Williamson 2016, p. 354). These are features of a theory as it is articulated. They are visible in the prose.
Lipton makes a second point that bears on us as well, this one about where abductive reasoning starts. Inquiry does not begin from the entire space of logical possibilities. It begins from a restricted set of live candidates: we identify the serious options first, and only then compare them (Lipton 2004, p. 59). The first filter — the choice of which candidates count as live — is itself part of philosophical practice. Philosophers inherit a structured background of distinctions and problems, of objections and worked examples, of candidate views that have been refined under earlier criticism. That background determines what counts as a live option in a given debate. A philosophical paper that proposes a theory of perception, or of depiction, or of consciousness, or of reference, does not compare its theory with every logically possible alternative. It situates the theory within a debate whose options have already been shaped by previous argument.
The philosophical corpus is one such background. It is not a neutral heap of sentences about philosophical topics. It is the written record of claims and the objections raised against them, of distinctions drawn and revised, of proposals taken up, tested, and sometimes set aside within philosophical practice. This is not to say that everything in the corpus is good philosophy, or that what survives there is true. It is to say that the corpus is partly structured by past philosophical selection. Arguments are repeated when they continue to do work; distinctions persist when they continue to be useful; objections are preserved when they continue to expose pressure points. The corpus therefore contains more than philosophical vocabulary. It contains traces of the abductive and dialectical standards by which philosophical work has been produced and assessed.
An LLM does not cease to be a next-token predictor when it is producing philosophy. It samples a token from a learned conditional distribution, appends that token to the context, and repeats. But the distribution from which it samples has been trained on text in which philosophical patterns are already present, and when the training corpus contains abductively structured philosophical writing, the model's conditional probabilities are shaped by that structure. The model is not judging one candidate explanation to be better than its rivals. It is generating a trajectory through a space of possible continuations whose local probabilities have been shaped by earlier philosophical writing.
It can help to import a piece of terminology from recent work in semiotic physics, used sparingly. A generated text is a trajectory: the prompt plus the output-so-far after each step of the autoregressive loop. The model supplies transition probabilities over the possible next tokens. Sampling and appending a token produces the next state, and repeated application produces the full continuation (Jan 2023; metasemi 2023). A philosophical argument is not a single token but an extended trajectory of this sort. If the local transition tendencies have been shaped by a corpus in which abductive structures are common, the resulting trajectory can display abductive structure at the level of the argument as a whole.
Floridi et al. themselves write that LLMs have "absorbed patterns of human abductive reasoning as expressed in writing" (2025, p. 9). That sentence should not be inflated into the claim that LLMs understand abductive reasoning. But it should not be deflated into the claim that they have acquired only empty verbal templates either. If abductive reasoning is, at least sometimes, expressed in writing, and if philosophical writing is one of the places where such reasoning is refined and criticised before being passed on, then training on philosophical writing can shape the model's generative tendencies in abductively relevant ways. The model does not need to perform the earlier reasoning itself in order for its outputs to bear the public traces of that reasoning.
What we are left with is a capacity on the side of the product rather than on the side of the producer. A text produced by an LLM can put forward a potential explanation, set it against live alternatives, and exhibit, in the way it handles those alternatives, the virtues that are relevant to abductive assessment. It can show, in particular cases, why one distinction handles a case that another distinction does not, or why an objection presses against a theory at the very place the theory had hoped to be safe, or why a long-running debate has been framed around the wrong contrast all along. None of this entails that the text in question is correct. Nor does it entail that the model understood what it was doing. It does mean that the text can make available an object of philosophical assessment: a candidate explanation whose merits can then be judged by readers.
The challenge from abduction does not therefore show that LLM-generated philosophy is impossible. It shows that the relevant capacity cannot be located in any human-like act of abductive judgement on the part of the model. That concession is harmless if the claim under defence concerns the product rather than the producer. LLMs do not perform inference to the best explanation in the way philosophers do. Given a philosophical corpus that has been shaped by past abductive selection, however, they can produce texts that contain potential explanations, that organise live alternatives, and that exhibit the virtues we use to assess explanations. Whether a particular output succeeds is then assessed in the ordinary philosophical way: by reading the text and asking what, if anything, it explains.
Notes on what I changed and why, briefly: removed announcement-flavoured topic sentences ("This gives us the mechanism", "The terminology of semiotic physics is useful here…", "What matters is the local-to-global point", "This is where Lipton's distinction… matters"); removed both uses of "matters" as a verb per the project's banned-word list; broke up decorative triplets and five-item lists by reorganising them as paired phrases, embedded clauses, or sequential sentences while preserving every original item; replaced flat declarative chains (most visibly in P4 and P5) with longer discursive sentences that do the thinking, punctuated by short verdict sentences; kept all citations, all quotations, and all sub-claims; left "we" throughout (co-authorship with Enrico) and used British spelling. The Williamson list ("unified, not arbitrary, gerrymandered, ad hoc, or messily complicated") is left as Williamson's own list because it is functioning as a quoted/paraphrased virtue-list rather than as a decorative cadence.
If you want me to push harder in a specific direction — for example, opening §2 with a Nudds-style direct claim or a candidate-view opener instead of "Much philosophical theorising proceeds by inference to the best explanation" — say which paragraph and I will rework just that one.
*Una traiettoria abduttiva non è un singolo passo: si misura nell'arco intero, dove la stoffa filosofica si lascia leggere o non si lascia leggere affatto.*
### User
that was excellent. i would now like your suggestions for two subheadings for this section and where they should go. Note my subheadings in section 3 shiould give you an idea of the style of subheadings i like, and my publications
## Experiential Axioms
A further capacity worry concerns phenomenology. Few would say that LLMs are conscious, and we will assume the same here; yet this might seem to pose a problem for LLM philosophy, or at least for philosophy grounded in, or making use of, phenomenology. Some philosophy interrogates or refers to what it is like to see red (Harman, 1990), to feel anger (Goldie, 2000), or to have a particular intuition take hold (Chudnoff, 2011). If LLMs lack conscious experience, it seems as if this might hamper their ability to produce worthwhile philosophy which relies on it. This is not to say that all philosophy would be off bounds: large stretches of philosophy of language and modal metaphysics proceed without leaning on the phenomenology of any particular experience.
Zahavy’s discussion of a thought experiment of Einstein's brings out this worry:
> Einstein’s variation required inventing new axioms based on a physical intuition that did not yet exist in the mathematics. He envisioned a physicist inside an elevator being uniformly accelerated through deep space. Inside this enclosure, the sensory experience reveals a specific pattern: when objects are released, the floor rushes up to meet them. To the physicist, the objects appear to fall with identical acceleration, regardless of composition. Thus, the simulation here was not a permutation of symbols, but a manipulation of perceptual experience. (Zahavy 2026, §5)
The thinker imagines[^4] some set of circumstances and attends to what would be experienced within it — in Einstein’s case, that all objects inside the elevator would appear to fall with identical acceleration. That observation becomes the new axiom: a starting point arrived at through experiential simulation rather than formal derivation, from which further reasoning proceeds. If thinking of this kind depends on simulated experience, then it would seem to be out of reach for LLMs. They can provide descriptions of weightlessness or elevators, but they have never felt the sensation of an elevator descending, let alone weightlessness.[^2]
Philosophy also uses experience based thought experiments. Jackson’s Mary case turns on what it is like to see colour, and we might think that as with Einstein's thought experiment, it provides us with an experiential axiom, from which further philosophical reasoning can proceed. The same worry then arises in philosophy: experience based thought experiments seem to require what LLMs do not have.[^3]
## Articulated Phenomenology
Pigliucci offers an account of philosophy on which it is constrained by, but does not aim at, the world as the natural sciences do. He writes:
> This means that the basic parameters that philosophers use as their inputs, the starting points of their philosophizing, their equivalent of axioms in mathematics and assumptions in logic (or rules in chess) are empirical data about the world. This data comes from both everyday experience [...] and of course increasingly from the world of science itself. (Pigliucci, p. 6)
Philosophy begins from worldly materials, but those materials function as starting points for conceptual exploration. They are not used in the same way a physical datum is used to confirm or disconfirm an empirical theory.%%one more sentence, a good one, will make this paragraph substantial%%
Pigliucci elaborates this picture by drawing on Smolin’s account of evocation, taking chess as the paradigm. Positing the rules of a game does not require that they pre-exist; once posited, they generate a structure with rigid properties — a space of consequences that can be explored but not chosen. Once the rules of chess are codified, all the facts about chess become demonstrable, even though chess did not exist before its rules were written down. Pigliucci’s claim is that philosophy operates in this register. Unlike the rules of chess or the axioms of mathematics, however, the starting points of philosophising are constrained empirically. They are constrained by how the world actually is, including by what experience is like.
This is why Pigliucci distinguishes philosophy from fiction: philosophy is not merely the invention of imaginary possibilities. As he puts it:
> Philosophy, I maintain, is in the business of doing empirically informed evoking, not inventing. (Pigliucci, p. 7)
The same picture covers philosophical thought experiments. Even when philosophers explore possible worlds or imagined scenarios, they do so “with an interest in figuring things out as far as this world is concerned” (Pigliucci, p. 7). The thought experiment articulates an axiom — an experiential or empirical starting point — and the philosophical work proceeds within the conceptual landscape that axiom evokes.
This brings out a difference between the elevator and Mary cases. Both are evocations of the kind Pigliucci describes: each posits an experiential axiom and develops what follows from it. What differs is what the evocation is for. In Einstein’s case, the evoked structure yields a hypothesis whose status is then settled by experiment — the elevator gave him the equivalence principle, but the principle’s truth was a matter for empirical confirmation. In Mary’s case, the evoked landscape is itself the object of inquiry; the philosophical question is what the landscape contains, not whether anything outside it corresponds. The role of the evocation, not its presence, is what tracks the disciplinary difference. Evocation is present in both cases; what differs is whether the evoked structure is the means to an external test or is itself the object of inquiry.
No competent discussant of the knowledge argument has personally undergone her transition. Once the case is articulated, work on it is work on the articulation. Responses to Jackson press at the level of the articulated structure, not at the level of any discussant’s experience. Lewis’s reply, for instance, modifies what is taken to follow from Mary’s situation, not what Mary’s situation is taken to be like from the inside.
What allows the Mary case to do philosophical work in public is its articulation: the experiential material it draws on has been made available in language. This is the form in which phenomenology enters philosophy generally. The articulation is what does the philosophical work; the experience the articulation refers to need not be undergone by the people working on it. Philosophers work on the experiences of the blind and on the experiences of non-human animals without first-hand access to either, by working on the articulations the literature has accumulated. The point matters for LLMs in a particular way. They have no raw phenomenology of their own; but no text corpus contains raw phenomenology either. What a corpus contains is articulated phenomenology, and it is in articulated form that phenomenology becomes usable in philosophical argument.
Merleau-Ponty’s discussion of self-touch raises a sharper question — that of phenomenological _discovery_. Suppose the toucher-touched asymmetry was first identified by Merleau-Ponty himself, by sustained attention to his own embodied experience. The asymmetry would then be a phenomenological axiom out of reach of any LLM not trained on Merleau-Ponty or his interlocutors: an axiom an LLM could not have produced for itself, because the system lacks the body and the experience that the discovery requires. When one fingertip touches another, one finger plays the role of toucher and the other of touched. The roles can reverse, but not simultaneously: at any given instant, the body is split between touching and touched. But that does not prevent an LLM from working philosophically on the description once articulated.
What survives, then, is a narrower asymmetry. Even granting that LLMs can work within articulated landscapes, some phenomenological articulations seem to be originated through first-person attention; LLMs have no experience to attend to. First-person attention is one route to an articulation; it is not what gives an articulation philosophical use. What makes an articulation philosophically usable, on Pigliucci’s picture, is not its causal origin but its functioning as an axiom — its capacity to evoke a landscape with rigid properties. An articulation can also be arrived at by working from the articulations a corpus already contains, generating new ones by extension and recombination. Whether a candidate articulation succeeds is a question about what it evokes, and that question is answered the way other philosophical questions are — by the public assessment of the conceptual structure the articulation makes available. It is the assessment any candidate articulation, whatever its origin, must finally meet.
The phenomenology objection rests on a producer-to-product inference: that the absence of experience in the producer must remove phenomenological value from the product. The inference fails. LLMs lack conscious experience, but phenomenology enters philosophy as articulated content. Pigliucci’s account explains why this is not a workaround. Philosophy uses empirical and experiential materials by turning them into constrained spaces for conceptual exploration. Since those spaces are public and inferentially usable once articulated, current models can produce phenomenology-based philosophy worth reading.
[^2]: footnote saying that he calls it manipulative abduction. it should probably also explain why we might think go this as abduction as well as what we talked about in the previous section
[^3]: A nice example in the footnote will be the feeling of understanding that is sometimes used as a way of motivating cognitive phenomenology.
### Assistant
Recommendation: pair them with §3's structure, since the project plan explicitly endorses a parallel between Sections 2 and 3 (Enrico's transcript line 312: "Section two asks: can we have philosophy in the text without abduction in the mind? Section three asks: can we have philosophy in the text without embodied phenomenology in the mind?"). The §3 subheadings name the worry-concept first, then the response-concept. Apply the same shape here.
## Primary pair
- "Zeroth-Order Abduction" — placed immediately under the section heading, before P1.
- Covers P1–P5: philosophy uses IBE → Floridi's diagnosis (with the block quotation) → unpacking what they claim → our concession on the mechanism.
- "Articulated Abduction" — placed before P6 ("To see why it can, recall what Lipton's account…").
- Covers P6–P14: Lipton's "if true" / loveliness-likeliness → live-candidate filter → corpus → trajectory → Floridi's "absorbed patterns" remark → product-side capacity → section close.
## Why this pair
- It mirrors §3 exactly. "Experiential Axioms" / "Articulated Phenomenology" becomes "Zeroth-Order Abduction" / "Articulated Abduction". The reader picks up that §2 and §3 have the same shape, and the recurring phrase "Articulated X" gives the paper a small but real terminological spine.
- Both are short noun phrases naming a concept. That is what your publications do. *Hearing Spaces*'s subsections — "Reverberation as a Distinct Sound", "Reverberation as Echoes", "Reverberation as a Property of a Sound", "Reverberation as Hearing Spaces" — each name the candidate view being treated; they do not announce the move ("In this section I will…") and they do not name the philosopher.
- The first names Floridi's coinage and uses it as the heading-level name for the worry, just as "Experiential Axioms" picks up the Pigliucci/Zahavy framing. The second is a coinage of your own that the section earns: abduction surviving as articulated content in the corpus.
- The 4-to-9 paragraph balance under the two subheadings matches the 4-to-9 balance in §3.
## Alternatives if either feels off
- For the first heading:
- "Stochastic Continuation" — safer, less Floridi-specific; less of a coined term in a subheading.
- "Abduction at Zeroth Order" — same content, slightly less compressed.
- For the second heading:
- "Abductive Traces" — picks up the section's own phrase "public traces of that reasoning"; more vivid, but loses the §3 parallel.
- "Trajectories Through a Corpus" — picks up the trajectory framework; broader than just abduction, narrower than the section's full scope.
## One small caveat
- "Articulated Abduction" is a more novel coinage than "Articulated Phenomenology", because *phenomenology* is already a mass noun for a body of content, while *abduction* names a process. The section in fact licenses the coinage — abductive reasoning leaves articulated traces in prose, and that is what the corpus preserves — but it is worth being aware that you are doing slightly more lifting with the title here than in §3.
*Un'abduzione articolata non è meno abduzione perché il filo del ragionamento si è già fissato sulla pagina; è semplicemente l'abduzione resa pubblica.*
### Assistant
No response requested.
### User
/contemplate/deep-writing-mode/draft-audit
here is a new version of section 2 of my generating philosophy paper. currently the content is good by it is not written in my style at all. please aplpy all of the skills I am activating now THOROUGHLY to produce a new version just here in the chat. It should be a paragraph for paragraph rewrite. Content must me 100% maintatined. please don't smooth out all of the details (you have a bad habit of making text shallower with each iteration. fight this. make sure that you explain your analysis, justification, reasons etc etc. for your answer BEFORE giving me your final answer.
## II. The challenge from abduction
Much philosophical theorising proceeds by inference to the best explanation. A philosopher offers an account of some phenomenon and defends it by arguing that, if true, it would explain the relevant evidence better than its rivals. Williamson treats this as a legitimate method of argument in philosophy: philosophy, on this view, often advances by comparing theories with respect to their explanatory power, their fit with the evidence, and their theoretical virtues (Williamson 2016, pp. 351–356). The challenge is straightforward. If LLMs do not perform inference to the best explanation, it may seem that they cannot produce philosophical texts whose value depends on abductive argument.
Floridi et al. give this challenge a precise form. They write:
> LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025, p. 9)
The claim is not that LLMs cannot produce text that looks explanatory. They often can. The claim is that such text is generated by learned associations and sequence probability, not by an understanding of evidence, causes, truth, or explanation. What appears to be abductive reasoning is, on their view, the surface result of a stochastic process.
Floridi et al. are right about the process. An LLM does not understand a phenomenon as calling for explanation. It does not knowingly generate live candidate explanations, compare them, and infer the one that would best explain the data. It has no grasp of one candidate as lovelier or likelier than another. We should not respond by saying that LLMs secretly perform human-style inference to the best explanation. The question is instead whether a text produced by such a system can contain a good abductive argument.
To see why it can, recall what Lipton’s account of inference to the best explanation assesses. On his view, we infer "what would, if true, provide the best explanation" of the evidence (Lipton 2004, p. 56). The phrase ‘if true’ is doing real work. We do not first identify the actual explanation and then infer it; that would require us to have reached the end of inquiry before inquiry begins. We assess potential explanations: candidates that would explain the data if they were true (Lipton 2004, pp. 57–59). A potential explanation is the sort of thing that prose can present. A text can specify the data, formulate the candidate, identify the relevant contrast, compare live alternatives, and show what the candidate would explain if true.
This is where Lipton’s distinction between the likeliest and the loveliest explanation matters. The likeliest explanation is the one most likely to be true; the loveliest explanation is the one that would provide the most understanding if it were true. As Lipton puts it, "Likeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). If inference to the best explanation meant only inference to the likeliest candidate, the account would say little more than that we infer what we judge most probable. Lipton’s stronger claim is that explanatory virtues help guide judgments of likelihood: loveliness is, at least sometimes, a guide to likeliness (2004, pp. 60–62). Williamson gives the corresponding point in philosophical terms when he says that a theory should be unified, not arbitrary, gerrymandered, ad hoc, or messily complicated; in short, it should combine simplicity with strength (Williamson 2016, p. 354). These are features of theories as they are articulated. They are visible in the text.
Lipton also shows that abductive reasoning does not begin from the whole space of logical possibilities. Inquiry normally starts from a restricted set of live candidates. We first identify serious candidates, then compare them (Lipton 2004, p. 59). This matters because the first filter is itself part of philosophical practice. Philosophers inherit a structured background of distinctions, problems, objections, examples, and candidate views. That background shapes what counts as a live option in the first place. A paper that proposes a theory of perception, depiction, consciousness, or reference does not compare it with every logically possible alternative. It situates it within a debate whose options have already been shaped by previous argument.
The philosophical corpus is one such background. It is not a neutral heap of sentences about philosophical topics. It is the written record of claims, objections, distinctions, revisions, and failed proposals that have been taken up and tested within philosophical practice. This does not mean that everything in the corpus is good philosophy, or that what survives is true. It means that the corpus is partly structured by past philosophical selection. Arguments are repeated because they are useful; distinctions persist because they do work; objections are preserved because they expose pressure points. The corpus therefore contains not only philosophical vocabulary, but traces of the abductive and dialectical standards by which philosophical texts have been produced and assessed.
This gives us the mechanism. An LLM does not cease to be a next-token predictor when it produces philosophy. It samples a token from a learned conditional distribution, appends that token to the context, and repeats the process. But the distribution from which it samples has been trained on texts in which philosophical patterns are already present. When the training corpus contains abductively structured philosophical writing, the model’s conditional probabilities are shaped by that structure. The model is not judging that a candidate explanation is better than its rivals. Rather, it is generating a trajectory through a space of possible continuations whose local probabilities have been shaped by earlier philosophical texts.
The terminology of semiotic physics is useful here, provided it is used sparingly. A generated text is a trajectory: the prompt plus the output-so-far after each step of the autoregressive loop. The model supplies transition probabilities over possible next tokens; sampling and appending a token produces the next state; repeated application produces the full continuation (Jan 2023; metasemi 2023). The heavier parts of the framework are not needed for the present argument. What matters is the local-to-global point. A philosophical argument is not a single token, but an extended trajectory. If the local transition tendencies have been shaped by a corpus in which abductive structures are common, then the resulting trajectory can display abductive structure at the level of the argument.
Floridi et al. themselves say that LLMs have "absorbed patterns of human abductive reasoning as expressed in writing" (2025, p. 9). That sentence should not be inflated into the claim that LLMs understand abductive reasoning. But it should not be deflated into the claim that they have acquired only empty verbal templates. If abductive reasoning is expressed in writing, and if philosophical writing is one of the places where such reasoning is refined, criticised, and transmitted, then training on philosophical writing can shape the model’s generative tendencies in abductively relevant ways. The model does not need to perform the earlier reasoning in order for its outputs to bear the public traces of that reasoning.
The result is a product-side capacity. A text produced by an LLM can formulate a potential explanation, place it against live alternatives, and display virtues relevant to abductive assessment. It can show why one distinction handles a case better than another, why an objection presses on a theory, or why a debate has been framed around the wrong contrast. None of this entails that the text is correct. It also does not entail that the model understood what it was doing. But it does mean that the text can make available an object of philosophical assessment: a candidate explanation whose merits can be judged by readers.
The challenge from abduction therefore does not show that LLM-generated philosophy is impossible. It shows that the relevant capacity cannot be located in a human-like act of abductive judgement by the model. That concession is harmless if the claim concerns the product rather than the producer. LLMs do not perform inference to the best explanation in the way philosophers do. Still, given a philosophical corpus shaped by past abductive selection, they can produce texts that contain potential explanations, organise live alternatives, and exhibit explanatory virtues. Whether a particular output succeeds is then assessed in the ordinary philosophical way: by reading the text and asking what, if anything, it explains.
### Assistant
Voice grounding for this pass: I worked from full paragraphs of "Hearing Spaces" (the "What do we hear?" opener; the bell-and-cathedral case; the Nudds-objection paragraph and reply; the precedence-effect concession), the prose-composition examples drawn from "Agents of Change" and "Growing the Image" (the wine-pouring example, the drum-machine objection-and-reply, the "Both options are unsatisfying" verdict structure), and the prior passage of the same paper that you have already approved (the "we grant the characterisation at the level of mechanism" move). What these are showing me, especially relevant to draft-audit's depth lens: paragraphs that perform a philosophical move directly (concede, identify, distinguish, set up) rather than narrate one ("This shows", "This dissolves"); discursive sentences carrying the thinking with shorter sentences delivering verdicts; concessive moves whose force comes from being genuinely concessive; concrete cases that do work rather than illustrate.
Running the four audits in order before the rewrite.
Source-check.
Lipton p. 56. The draft has: "we infer 'what would, if true, provide the best explanation' of the evidence (Lipton 2004, p. 56)". Lipton's actual sentence on p. 56 reads: "we infer what would, if true, provide the best of the competing explanations we can generate of those data". The phrase inside the inverted commas is therefore a paraphrase, not a verbatim quotation. Two repair options: drop the inverted commas and frame as paraphrase, or quote Lipton's actual phrase ("the best of the competing explanations we can generate"). I will preserve the draft's wording in the rewrite per the content-preservation rule, but flag this so it can be fixed at the citation pass.
Lipton p. 59. The draft quotes "Likeliness speaks of truth; loveliness of potential understanding" — verbatim from p. 59. Clean.
Lipton pp. 57–59. The draft characterises Lipton as saying "We assess potential explanations: candidates that would explain the data if they were true." This is a paraphrase (no inverted commas), and is consistent with Lipton's actual move on pp. 57–58 distinguishing potential from actual explanation. Clean.
Lipton pp. 60–62. The draft characterises Lipton's stronger claim as "explanatory virtues help guide judgments of likelihood: loveliness is, at least sometimes, a guide to likeliness". Lipton's actual phrasing on pp. 60–61 is "loveliness will be a guide to likeliness". The "at least sometimes" hedge is an addition by Nick — defensible, since Lipton himself allows that the criteria can come apart. Clean.
Williamson 2016, pp. 351–356. The relevant content I can locate in the extraction is on p. 354 ("It should be elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated… it should combine simplicity with strength") and around p. 346 ("a theory's strength, explanatory power, and consistency with the evidence"). The page range as cited (351–356) is plausible for the abductive-methodology section but the directly load-bearing pages are 346 and 354. Worth tightening the page numbers at the citation pass.
Floridi et al. p. 9. Both quotations are verbatim from the extraction — the zeroth-order abduction passage and the "absorbed patterns" remark. Clean. (Note that the extracted Floridi paper carries a 2024 date in its metadata; the draft cites it as 2025. I am preserving the draft's date.)
Depth-audit.
I check each paragraph against the question: is this paragraph performing a philosophical move, or is it describing one? Going through:
P1 sets up that philosophy uses IBE and states the resulting threat. Move performed.
P2–3 introduce and quote Floridi's diagnosis. Move performed (Floridi's claim is being made available).
P4 separates Floridi's claim from a stronger claim it is sometimes mistaken for. Move performed (philosophical-reading move on a source).
P5 grants Floridi's diagnosis at the level of mechanism, and pivots the question to the product. Move performed (concessive pivot).
P6 brings in Lipton to relocate what abduction evaluates: it evaluates potential explanations, which prose can present. Move performed.
P7 deploys Lipton's loveliness/likeliness distinction and Williamson's intrinsic-virtues characterisation to identify what is being evaluated as features visible in the articulated theory. Move performed.
P8 deploys Lipton's two-filter point to connect abduction's first filter with philosophical practice. Move performed.
P9 identifies the philosophical corpus as the structured background that filter operates over, and argues the filtering preserves more than vocabulary. Move performed.
P10 gives the mechanism: stochastic generation over a corpus filtered for abductive structure shapes the conditional probabilities. Move performed.
P11 uses the trajectory framework to extend the mechanism from local to global. Move performed.
P12 reads Floridi's own "absorbed patterns" remark as supporting the present argument when neither inflated nor deflated. Move performed.
P13 states the resulting product-side capacity and concretises what kinds of work the prose can do. Move performed.
P14 concludes the section. Move performed.
Each paragraph is making a move. Depth survives.
Anti-metacommentary scan of the draft as supplied.
Forbidden cases:
- P7's "This is where Lipton's distinction… matters" — argument-self-description ("this is where… matters"); also uses the banned-as-verb "matters". Replace with the substantive claim ("Lipton's distinction… bears on this directly").
- P8's "This matters because the first filter is itself part of philosophical practice" — same banned verb, same announcement structure. Replace with the direct claim that the first filter is part of philosophical practice.
- P10's "This gives us the mechanism" — argument-self-description. Cut entirely; let the mechanism be given by the next sentences.
- P11's "What matters is the local-to-global point" — argument-self-description. Cut entirely; let the local-to-global structure be made by the trajectory sentence and the conditional that follows.
- P11's "The heavier parts of the framework are not needed for the present argument" — framework-management, redundant with "used sparingly". Cut.
Suspicious cases (keep with rewriting):
- P11's "The terminology of semiotic physics is useful here, provided it is used sparingly." — argument-self-description, but it is doing real attribution work for the citations Jan 2023; metasemi 2023. Replace with a parenthetical attribution inside a substantive sentence, so the citations are preserved without an opening sentence whose only job is to flag the framework.
- P6's "To see why it can, recall what Lipton's account of inference to the best explanation assesses" — procedural opener. Reduce to "Recall what Lipton's account of inference to the best explanation evaluates." The "to see why it can" is the part doing the announcement work.
- P13's "The result is a product-side capacity" — borderline. Acceptable as a topic sentence introducing a coined concept, but a more direct opening ("The capacity, then, lies on the side of the product rather than on the side of the producer") performs the same role without the "result is" frame.
Permitted as-is:
- P5's "We should not respond by saying that LLMs secretly perform…" — genuine objection-management; doing dialectical work, not narrating it.
- P4's setup-and-pivot pattern ("The claim is not that… The claim is that…") — flat declarative chain rhythmically, but the move is genuinely an X-not-Y identification of what Floridi is and is not claiming. Reshape rhythmically without cutting.
Voice-fix scan.
Triplets and lists. The draft has many. I cannot delete items because content must be preserved. The remedies available are: split across sentences; embed each item with its own internal grammar; integrate via "of X… of Y… of Z" with repeated preposition; turn parallel clauses into a sequence with internal causal/temporal relation. Specific items:
- P5's "live candidate explanations, compare them, and infer the one that would best explain the data" — restage as a purpose-sequence: "generate live candidate explanations to hold them against one another and infer which would best explain the data."
- P6's five-item sentence "specify the data, formulate the candidate, identify the relevant contrast, compare live alternatives, and show what the candidate would explain if true" — preserve all five items but split the rhythm: "lay out the data and put forward a candidate, identify the relevant contrast, place that candidate against the live alternatives, and show what it would explain if it were true." This still has four-five clauses; integrating "specify the data" with "put forward a candidate" pairs them and reduces the list-feel.
- P7's Williamson list "unified, not arbitrary, gerrymandered, ad hoc, or messily complicated" — Williamson's own list, preserve as quoted virtue-list, with "in short, that it should combine simplicity with strength" closing.
- P8's "distinctions, problems, objections, examples, and candidate views" — restage as paired phrases with internal grammar: "distinctions and problems, of objections and worked examples, and of candidate views that have been refined under earlier criticism."
- P8's "perception, depiction, consciousness, or reference" — repeat "of" before each item to slow the cadence: "of perception, or of depiction, or of consciousness, or of reference."
- P9's "claims, objections, distinctions, revisions, and failed proposals" — restage as paired phrases: "claims and the objections raised against them, of distinctions drawn and revised, and of proposals taken up, tested, and sometimes set aside within philosophical practice."
- P9's three-clause parallel "Arguments are repeated… distinctions persist… objections are preserved" — preserve as triplet but lengthen each clause with "continue to" so the cadence is not dum-dum-dum: "Arguments are repeated when they continue to do work; distinctions persist when they continue to be useful; objections are preserved when they continue to expose pressure points."
- P10's "samples a token from a learned conditional distribution, appends that token to the context, and repeats the process" — preserve as triplet; this is technical description of the autoregressive loop, not decorative.
- P11's "supplies transition probabilities… sampling and appending… repeated application" — break into a sentence + a clause-pair: "The model supplies transition probabilities over the possible next tokens. Sampling and appending a token produces the next state, and repeated application produces the full continuation."
- P12's "refined, criticised, and transmitted" — turn the third item from a parallel verb into a temporal clause: "refined and criticised before being passed on."
- P13's first triplet "formulate a potential explanation, place it against live alternatives, and display virtues" — restage as a structured sequence where the third item is causally connected to the second: "put forward a potential explanation, set it against live alternatives, and exhibit, in the way it handles those alternatives, the virtues that are relevant to abductive assessment."
- P13's second triplet "show why one distinction handles a case better, why an objection presses, why a debate has been framed wrong" — preserve all three but thicken each clause so it earns its weight.
- P14's "contain potential explanations, organise live alternatives, and exhibit explanatory virtues" — restage with "that … that … that …" instead of comma-list, giving each clause subordinate-clause weight.
Banned vocabulary. Two uses of "matters" as verb (P8, P11) — both eliminated above. No instances of "textual" / "textually". No instances of the prohibited epistemic-discipline words.
Pronoun. "We" throughout, correct for co-authorship.
Spelling. British: criticised, characterised, etc.
Plan.
The rewrite proceeds paragraph for paragraph, with all content preserved. Differences from the draft are:
- topic-sentence and section-progression sentences rephrased to make the move directly (anti-metacommentary)
- triplets and lists redistributed across sentences or embedded with internal grammar (no-example-lists / voice-fix)
- "matters" as verb replaced with substantive claims
- the semiotic-physics framework-management cut, with citations preserved by a parenthetical inside the substantive trajectory sentence
- flat declarative chains in P4 reshaped into a longer sentence + a brief verdict
- I add small thickening clauses ("that have been refined under earlier criticism", "at the very place the theory had hoped to be safe", "all along") that are voice-textures and do not change content; these are the opposite of shallowing
Two source-check items I am flagging but not silently fixing in the rewrite, because content preservation rules: the Lipton p. 56 inverted commas are around a paraphrase, and the Williamson p. 351–356 page range may want tightening to 354 (with a cross-reference around 346 if the explanatory-power-and-fit material is wanted as well).
## II. The challenge from abduction
Much philosophical theorising proceeds by inference to the best explanation. A philosopher offers an account of some phenomenon and defends it on the grounds that, were it true, it would explain the relevant evidence better than its rivals. Williamson takes this to be a legitimate method of argument in philosophy: philosophy often advances by comparing theories with respect to their explanatory power, their fit with the evidence, and the further theoretical virtues that distinguish a good theory from a bad one (Williamson 2016, pp. 351–356). The challenge is then straightforward. If LLMs do not perform inference to the best explanation, it can seem to follow that they cannot produce philosophical texts whose value depends on abductive argument.
Floridi et al. give the challenge its sharpest form:
> LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025, p. 9)
Their claim is not that LLMs cannot produce text that looks explanatory; they often can. Their claim is that this text is generated by learned associations and sequence probability rather than by any understanding of evidence, of causes, of truth, or of explanation. What appears to be abductive reasoning is, on Floridi et al.'s view, the surface result of a stochastic process.
We grant the point about the process. An LLM does not understand a phenomenon as something calling for explanation. Nor does it knowingly generate live candidate explanations to hold them against one another and infer the one that would best explain the data. It has no grasp of one candidate as lovelier or likelier than another. It would be a mistake to insist, in reply, that LLMs secretly perform human-style inference to the best explanation. The question is rather whether a text produced by such a system can contain a good abductive argument.
Recall what Lipton's account of inference to the best explanation evaluates. We infer, on Lipton's view, "what would, if true, provide the best explanation" of the evidence (Lipton 2004, p. 56). The phrase 'if true' is doing real work in this formulation. We do not first identify the actual explanation and then infer it; that would require us to have reached the end of inquiry before inquiry began. We assess potential explanations: candidates that would explain the data if they were true (Lipton 2004, pp. 57–59). And a potential explanation, in this sense, is the sort of thing that prose can present. A text can specify the data and put forward a candidate, identify the relevant contrast, place that candidate against the live alternatives, and show what it would explain if it were true.
Lipton's distinction between the likeliest and the loveliest explanation bears on this directly. The likeliest explanation is the one most likely to be true; the loveliest is the one that, if true, would provide the most understanding. As Lipton puts it, "Likeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). If inference to the best explanation simply meant inference to the likeliest candidate, the account would say little more than that we infer what we already judge most probable. Lipton's stronger claim is that explanatory virtues help to guide our judgments of likelihood: loveliness, at least sometimes, is a guide to likeliness (2004, pp. 60–62). Williamson states the corresponding point in philosophical terms when he says that a theory should be unified, not arbitrary, gerrymandered, ad hoc, or messily complicated; in short, that it should combine simplicity with strength (Williamson 2016, p. 354). These are features of a theory as it is articulated. They are visible in the prose.
Lipton makes a further point, this one about where abductive reasoning starts. Inquiry does not begin from the entire space of logical possibilities. It begins from a restricted set of live candidates: we identify the serious options first, and only then compare them (Lipton 2004, p. 59). The first filter — the choice of which candidates count as live — is itself part of philosophical practice. Philosophers inherit a structured background of distinctions and problems, of objections and worked examples, and of candidate views that have been refined under earlier criticism. That background determines what counts as a live option in a given debate. A philosophical paper that proposes a theory of perception, or of depiction, or of consciousness, or of reference, does not compare its theory with every logically possible alternative. It situates the theory within a debate whose options have already been shaped by previous argument.
The philosophical corpus is one such background. It is not a neutral heap of sentences about philosophical topics. It is the written record of claims and the objections raised against them, of distinctions drawn and revised, and of proposals taken up, tested, and sometimes set aside within philosophical practice. This is not to say that everything in the corpus is good philosophy, or that what survives there is true. It is to say that the corpus is partly structured by past philosophical selection. Arguments are repeated when they continue to do work; distinctions persist when they continue to be useful; objections are preserved when they continue to expose pressure points. The corpus therefore contains more than philosophical vocabulary. It contains traces of the abductive and dialectical standards by which philosophical work has been produced and assessed.
An LLM does not cease to be a next-token predictor when it is producing philosophy. It samples a token from a learned conditional distribution, appends that token to the context, and repeats. The distribution from which it samples, however, has been trained on text in which philosophical patterns are already present, and when the training corpus contains abductively structured philosophical writing, the model's conditional probabilities are shaped by that structure. The model is not judging that one candidate explanation is better than its rivals. It is generating a trajectory through a space of possible continuations whose local probabilities have been shaped by earlier philosophical writing.
A generated text, in the terminology of recent work in semiotic physics, is a trajectory: the prompt plus the output-so-far after each step of the autoregressive loop. The model supplies transition probabilities over the possible next tokens. Sampling and appending a token produces the next state, and repeated application produces the full continuation (Jan 2023; metasemi 2023). A philosophical argument is not a single token but an extended trajectory of this sort. If the local transition tendencies have been shaped by a corpus in which abductive structures are common, the resulting trajectory can display abductive structure at the level of the argument as a whole.
Floridi et al. themselves write that LLMs have "absorbed patterns of human abductive reasoning as expressed in writing" (2025, p. 9). That sentence should not be inflated into the claim that LLMs understand abductive reasoning. But it should not be deflated into the claim that they have acquired only empty verbal templates either. If abductive reasoning is, at least sometimes, expressed in writing, and if philosophical writing is one of the places where such reasoning is refined and criticised before being passed on, then training on philosophical writing can shape the model's generative tendencies in abductively relevant ways. The model does not need to perform the earlier reasoning itself in order for its outputs to bear the public traces of that reasoning.
The capacity, then, lies on the side of the product rather than on the side of the producer. A text produced by an LLM can put forward a potential explanation, set it against live alternatives, and exhibit, in the way it handles those alternatives, the virtues that are relevant to abductive assessment. It can show, in particular cases, why one distinction handles a case that another distinction does not, or why an objection presses against a theory at the very place the theory had hoped to be safe, or why a long-running debate has been framed around the wrong contrast all along. None of this entails that the text in question is correct. Nor does it entail that the model understood what it was doing. It does mean that the text can make available an object of philosophical assessment: a candidate explanation whose merits can then be judged by readers.
The challenge from abduction does not therefore show that LLM-generated philosophy is impossible. It shows that the relevant capacity cannot be located in any human-like act of abductive judgement on the part of the model. That concession is harmless if the claim under defence concerns the product rather than the producer. LLMs do not perform inference to the best explanation in the way philosophers do. Given a philosophical corpus that has been shaped by past abductive selection, however, they can produce texts that contain potential explanations, that organise live alternatives, and that exhibit the virtues we use to assess explanations. Whether a particular output succeeds is then assessed in the ordinary philosophical way: by reading the text and asking what, if anything, it explains.
Two source-check items to fix at the citation pass, flagged here so they are not lost:
- Lipton p. 56. The phrase inside the inverted commas — "what would, if true, provide the best explanation" — is a paraphrase, not a verbatim quotation. Lipton's actual sentence reads "we infer what would, if true, provide the best of the competing explanations we can generate of those data". Either drop the inverted commas, or quote the verbatim phrase.
- Williamson p. 351–356. The directly load-bearing pages for the content cited are 354 (theory virtues — unified, not gerrymandered, simplicity-with-strength) and around 346 (theory's strength, explanatory power, and consistency with the evidence). The page range can either narrow to p. 354 or be expanded to capture both passages.
*La traiettoria abduttiva non è un singolo gesto del pensiero ma una linea che attraversa il corpus, e ciò che lascia sulla pagina è valutabile anche quando nessuno sotto la pagina sta pensando.*