# Section 2 — The Challenge from Abduction
- Even if one accepts our argument that LLMs should not be automatically ruled out from producing worthwhile philosophy
- one might still think that such systems, at least in their current form, lack particular *capacities* which are required to produce philosophy worth reading.
- In the next we shall consider how much LLMs' lack of phenomenology impedes on their ability to produce worthwhile philosophy.
- In this section we shall examine the charge that LLMs cannot perform abductive inference, and how it bears on these systems' ability to produce worthwhile philosophy.
Floridi and colleagues hold that large language models do not perform abductive inference. They describe what such models do instead as zeroth-order abduction:
> LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025, p. 9)
- On this description, the model explains nothing. It produces writing that has the outward form of an explanation. It does not pick out anything as standing in need of explanation, and it does not judge one account of the facts to explain them better than another.
- Abduction is often said to play a significant role in philosophical theorising.
Williamson takes much philosophical theorising to proceed by inference to the best explanation: a theory is offered and defended on the grounds that, were it true, it would explain the relevant evidence better than its rivals (2016, pp. 351–356). Theorising of this kind is struc-
turally comparative. There are data any candidate theory must accommodate, and there are
rival candidates each of which would, if true, accommodate those data in different ways and
at different theoretical cost. The philosophical task is to weigh the candidates against each
other and judge which would do the explanatory work best. If LLMs cannot perform inference
to the best explanation, then philosophical work of this kind — much of it, on Williamson’s
account — would lie out of their reach.
Consider the following from Floridi et al.:
LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a
plausible continuation (a hypothesis or explanation) based purely on learned associations. In
reality, their operation is driven by maximising the probability of the sequence... The model
does not understand what an explanation is, but it produces text that follows the typical
phrasing and structure of explanations. It does not reason about causes from scratch but
outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025,
p. 9)
-
challengFloridi and colleagues hold that large language models do not perform abductive inference. They describe what such models do instead as zeroth-order abduction:
> LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025, p. 9)
A model presented with a prompt extends it with the continuation that the training distribution makes probable. When that continuation has the form of an explanation, this is because explanatory prose is what tends to follow such prompts in the training data, not because the model has identified something to be explained and settled on a hypothesis that would explain it. The model has no purchase on what an explanation is, on what would count as evidence for one, or on why one candidate explanation should be preferred to another.
A large part of philosophical theorising proceeds by inference to the best explanation. Williamson takes much of the discipline to work this way: a theory is advanced and defended on the grounds that, were it true, it would explain the relevant data better than its rivals. Doing philosophy of this kind requires weighing candidate explanations against one another and judging which would explain best.
If philosophy of this kind requires inference to the best explanation, and language models do not perform inference to the best explanation, then philosophy of this kind is beyond them.
---
**Move 1 — State the challenge, Floridi first.** Open with Floridi: LLMs perform only "zeroth-order abduction" — given a prompt they generate a plausible continuation, producing explanation-shaped text by maximising sequence probability, with no grasp of explanation, evidence, or cause. (Block quote.) Then Williamson as the bridge: much philosophical theorising proceeds by inference to the best explanation. The challenge follows: if the model performs no IBE, philosophy of this kind lies beyond it.
> LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data.
— Floridi et al. (2025), p. 9
_(The Williamson bridge premise — that philosophical theorising substantially proceeds by IBE — is in your draft; no canvas block quote for it.)_
**Move 2 — Concede the producer point and state what is claimed instead.** Grant Floridi's diagnosis of the _process_ in full: the model does not treat the prompt as evidence and infer a best explanation. Then fix the claim precisely — what is at issue is whether the _text_ exhibits abductive structure (a candidate explanation, the rivals it is set against, the grounds for preferring it), not whether the model _reasoned_. Structure in the product, not inference in the producer. Name the challenge's error as the slide Section 1 already refused: from a fact about the producer to a verdict on the product.
_(No block quotes — this move is the producer/product framing.)_
**Move 3 — Abductive structure is a property a text can have.** Lipton: a candidate considered "if true," prior to being established — statable in prose. The likeliest/loveliest distinction, with loveliness a feature of how the candidate is laid out. Williamson: the intrinsic virtues (elegant, unified, not ad hoc, simplicity-with-strength) and the comparative form, a theory set against rivals. These are features of presentation; they live in the text.
> we may characterize the best explanation as the one which would, if correct, be the most explanatory or provide the most understanding: the "loveliest" explanation.
> Likeliness speaks of truth; loveliness of potential understanding.
— Lipton, _Inference to the Best Explanation_ (2004), p. 60 (l. 520)
> It should be elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated. It should be informative and general. In brief, it should combine simplicity with strength.
— Williamson, "Widening the Picture," §9.2, "A sketch of abduction" (ll. 1676–1679)
> abduction involves the assessment of – amongst other factors – a theory's strength, explanatory power, and consistency with the evidence
— Williamson, §9.1 (ll. 1345–1346)
**Move 4 — Why a non-reasoning system produces text with that structure.** The model is trained on philosophy's written record, which is itself a residue of past abductive argument — the surviving distinctions, objections, and candidate views that have gone on doing work. In that corpus, abductive moves follow abductive moves: a thesis is met by an objection, an objection by a reply. Fitting that distribution, the high-probability completion of an abductive opening is an abductive completion. So the model produces abductively-structured text without performing the inference.
> ChatGPT doesn't have any explicit "knowledge" of such rules. But somehow in its training it implicitly "discovers" them—and then seems to be good at following them.
— Wolfram, _What Is ChatGPT Doing?_, §"What Really Lets ChatGPT Work?" (l. 441)
> it's something that one can think of ChatGPT as having implicitly "developed a theory for" after being trained with billions of (presumably meaningful) sentences from the web
— §"What Really Lets ChatGPT Work?" (l. 459)
> while one can therefore expect ChatGPT to produce text that contains "correct inferences" based on things like syllogistic logic, it's a quite different story when it comes to more sophisticated formal logic—and I think one can expect it to fail here for the same kind of reasons it fails in parenthesis matching.
— §"What Really Lets ChatGPT Work?" (l. 461) — _second clause is the shallowness objection inverted at Move 5_
> just as one can somewhat whimsically imagine that Aristotle discovered syllogistic logic by going ("machine-learning-style") through lots of examples of rhetoric, so too one can imagine that in the training of ChatGPT it will have been able to "discover syllogistic logic" by looking at lots of text on the web
— §"What Really Lets ChatGPT Work?" (l. 461)
> Of course, we rank only those potential explanations that have been thought of.
— Williamson, §9.2, "A sketch of abduction" (l. 1718)
> the transformer architecture of neural nets like the one in ChatGPT seems to successfully be able to learn the kind of nested-tree-like syntactic structure that seems to exist (at least in some approximation) in all human languages.
— Wolfram, §"What Really Lets ChatGPT Work?" (l. 455)
Deflation guard — the mechanism-level gloss to insulate, and the levels move that does it:
> ChatGPT is "merely" pulling out some "coherent thread of text" from the "statistics of conventional wisdom" that it's accumulated.
— Wolfram, §"So … What Is ChatGPT Doing?" (l. 531)
> arguing that Inference to the Best Explanation is wrong because Bayesianism is right is like arguing that thinking about technique cannot help my squash game because the motion of the ball is governed by the laws of mechanics.
> Even if Bayesianism gave the mechanics of belief revision, Inference to the Best Explanation might yet illuminate its psychology.
— Lipton (2004), p. 108 (l. 762)
**Move 5 — The shallowness objection, and its reversal.** The objection: real abduction is too sophisticated for these systems. The reply turns on what abduction is. Lipton and Williamson both deny it is a rigid procedure — there is no algorithm from data to hypothesis; loveliness is a defeasible, holistic matter of degree; Williamson calls it an informal method and offers no full account. Judgement of degree, made by pattern rather than by exact rule, is the kind of task these systems handle well, not the kind they fail at. The more one insists on abduction's sophistication and holism, the more firmly it sits on the favourable side. (Illustration available if wanted: digit recognition — graded, rule-less, succeeds; exact symbol-counting — fails. Use once or not at all.)
The dividing line — guessing-in-a-glance vs exact symbol-tracking:
> Cases that a human "can solve in a glance" the neural net can solve too. But cases that require doing something "more algorithmic" (e.g. explicitly counting parentheses to see if they're closed) the neural net tends to somehow be "too computationally shallow" to reliably do.
— Wolfram, §"What Really Lets ChatGPT Work?" (l. 453)
> in English it's much more realistic to be able to "guess" what's grammatically going to fit on the basis of local choices of words and other hints. And, yes, the neural net is much better at this
— §"What Really Lets ChatGPT Work?" (l. 455)
Abduction is the non-algorithmic kind — Williamson concedes no algorithm:
> Abduction is an informal method of non-deductive, ampliative inference
— Williamson, §9.2, "Abductive Philosophy" (l. 1536)
> The following remarks are merely indicative; they do not aspire to be a full account
— Williamson, §9.2, "A sketch of abduction" (l. ~1650)
> Inference to the best explanation may be a good heuristic to use when – as often happens – probabilities are hard to estimate
— Williamson, §9.2, "A sketch of abduction" (l. 1721)
Lipton on the same — done well, not statable as a rule:
> It is easy to ride a bicycle, but hard to describe how it is done; it is easy to distinguish between grammatical and ungrammatical strings of words in one's native tongue, but hard to describe the principles that underlie those judgments.
— Lipton (2004), Preface (l. 210)
> These similarities are not created or governed by rules, but they result in a pattern of research that mimics one that is rule governed.
— Lipton (2004), ch. 1, ~p. 7 (l. 240)
> there is no general algorithm that could take them from data to a hypothesis that refers to entities and processes not mentioned in the data
> generating good hypotheses is a matter of "happy guesses" (1966: 15).
— Lipton (2004), ch. 5, ~p. 83 (l. 636)
> there are no universally shared mechanical rules that generate a unique hypothesis from any given pool of data
— Lipton (2004), ch. 5, ~p. 83 (l. 638)
> the weakness of our grasp on what makes one explanation lovelier than another is discouraging.
— Lipton (2004), ~p. 61 (l. 528)
> using judgments of loveliness as a barometer of likelihood
— Lipton (2004), ~p. 114 (l. 798)
**Move 6 — The opponent's own concession.** Floridi grants that LLMs have absorbed the patterns of human abductive reasoning as expressed in writing. This seals the producer/product split from the challenger's side: the corpus carries the structure; what is denied is only that the model reasons — already conceded.
_(The Floridi "absorbed patterns of human abductive reasoning as expressed in writing" line is in your draft, p. 9; not among the previous canvas block quotes. Drop it in when you write the move.)_
**Move 7 — Close.** The model performs no inference, but the text can carry abductive structure. Whether a given argument is _sound_ — whether the explanation actually explains — is the ordinary question every paper faces, settled by reading the argument, not by inspecting the model. Defer the separate worry about genuine novelty (distinct from the abduction worry) toward Section 4 or a limitations note. Bridge: Section 2 shows the structure can be present and produced; Section 4 takes up the prompter who elicits it.
_(No block quotes.)_
---
Drafting notes, pinned here rather than in the moves: place abduction by its gradedness and defeasibility, not by globality of scope (Move 5); rest nothing on dated claims about model capability, and date Wolfram if cited; keep mechanism to what underwrites structure-in-the-text and stop short of soundness (Move 4). Open: whether to use the digit illustration at all (Move 5), and whether the novelty worry gets a sentence or a short paragraph (Move 7).
These two passages — where Wolfram disowns the "semantic laws of motion" idea — no longer map to a move (the new plan drops the trajectory framing). They are kept here only as the evidence behind the "gradedness, not globality of scope" note:
> There's certainly no "geometrically obvious" law of motion here.
— Wolfram, §"Meaning Space and Semantic Laws of Motion" (l. 477)
> this seems like a mess—and doesn't do anything to particularly encourage the idea that one can expect to identify "mathematical-physics-like" "semantic laws of motion" […] we're not ready to "empirically decode" from its "internal behavior" what ChatGPT has "discovered"
— §"Meaning Space and Semantic Laws of Motion" (l. 483)