## §2 The Challenge from Abduction — section in progress (17 June 2026) ### First half (already drafted — prose as written, my inline %%comments%% removed) In the previous section, we argued against the idea that LLMs cannot produce philosophy worth reading simply because they are not human. In this section and the next we shall consider a different form of challenge: even if LLMs cannot be ruled out of producing philosophy worth reading tout court, they lack particular _capacities_ that producing it requires. If a parrot uttered a sequence of sounds that happened to form a philosophical argument, the argument would be none the worse for its source; yet parrots' powers of mimicry do not extend to producing strings of sounds so complex as to make up a philosophical argument. One might think the same is true for LLMs. They just don't have what is needed to produce worthwhile philosophical argument. In this section we address one capacity challenge, which we will call the _challenge from abduction_. In the next we shall look at two more: the challenge from phenomenological experience and the challenge from contact with the world. Abduction, or inference to the best explanation, is reasoning from a body of evidence to the hypothesis that would best explain it. In a deductive argument the premises fix the conclusion: if all men are mortal and Socrates is a man, then Socrates is mortal, and there is no wriggle room. Now, imagine walking into your kitchen and finding the floor wet. What has happened? The wet floor does not determine the answer in the way the two premises determined Socrates' mortality: a burst pipe would have left the floor wet, and so would a spilled bucket. But, given that the window is open, the water is under the window, and it rained last night, rain coming through the window seems the most plausible answer. To reason in this way, deciding what best explains a set of facts, is common in the sciences as well as every day life. A scientist chooses one theory over another when it explains the same data more simply: Copernicus's model of the solar system was preferred to Ptolemy's because it explained the observed planetary motions without the elaborate epicycles the older model required. Williamson argues that philosophy should also use a broadly abductive methodology (2007; 2021, §9.2). In philosophy too there are data that a candidate theory must accommodate — intuitions about cases, and the phenomena of the domain itself — and rival theories that would each accommodate them at different costs. The theory to prefer is the one that would, if true, best explain the data. What makes one explanation better than another, on this account, is a matter of explanatory virtue: a good philosophical theory is, in Williamson's words, "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated", and should "combine simplicity with strength" (2021, §9.2). That theories are weighed by such comparative and explanatory virtues need not rest on a science-modelled conception of philosophy: Bengson, Cuneo and Shafer-Landau (2022) argue that the assessment of rival theories by their explanatory and unifying merits is a constraint on sound philosophical method as such. This conception of philosophy is widely held (Sider 2011; Paul 2012; Dellsén et al. 2024), though not universally (Bueno and Shalkowski 2020; Thomasson 2015), and we shall assume it in what follows. On this account a philosophical text offers its reader a choice of theory displayed — a position, its rivals, and the case for preferring it — so that whether the text is worth reading and whether it contains a good weighing travel together. If the capacity for abduction is what is required to produce worthwhile philosophy, we can ask whether LLMs possess it. Floridi et al. (2025) argue that they do not, describing what such models do instead as zeroth-order abduction: > LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025, p. 9) An LLM, on this account, has "a stochastic core and an abductive appearance" (2025, p. 2). The model is trained to predict which words are likely to follow which, and it produces the continuation its training makes probable; it aims at the likely continuation, not at the truth. The appearance comes from what the training data have passed on: models have "absorbed patterns of human abductive reasoning as expressed in writing" (p. 9) — how explanations are typically phrased, which causes are typically offered for which effects. What is inherited, on their account, is the look of the reasoning, not the reasoning itself.[^1] Explaining the wet kitchen floor involved two separable activities: coming up with candidate explanations — the burst pipe, the spilled bucket, the rain — and settling which of them the open window and the position of the water favoured. Call the first _generating_ and the second *weighing*. Floridi et al.'s position is that a model does not actually do either. Asked why a car might not start on a cold morning, a model replies that a weak battery is one possibility, since cold reduces a battery's efficiency; that thickened engine oil is another, since a cold engine is harder to turn over; and that, "[b]ased on your description, the battery is the most likely explanation" (2025, p. 10). The offering of candidates here is not generating, on their reading: the model is not reasoning about causes from the user's case but statistically reproducing the causes such explanations typically cite (p. 9). And the singling out is not weighing: the verdict reproduces how explanations of this kind typically end, and where an output marks a genuine point of difference between two hypotheses, that is something the model has seen stated, not something it has derived anew. This veneer of abduction, Floridi et al. argue, means that LLMs can only ever play a supporting role in intellectual work: > In essence, LLMs function like brainstorming assistants that toss out ideas without filtering for quality. After all, they work like statistical interfaces to an enormous amount of data accumulated for millennia by generations. A cautious human collaborator can sift through and assess them. (2025, p. 11) If LLMs are little more than brainstorming devices, this would seem to put them far away from the possibility of producing worthwhile philosophy. If you were presented with a text and told that it contains a number of philosophical ideas, none of which have been filtered for quality, it is unlikely you would think it is worth your time to read it. We will not attempt to argue that LLMs 'really' perform abduction in the way that humans do. Instead, we shall argue that LLM-produced text can still exhibit good abductive inference despite not being produced by such an inference. To defend that claim, we first need to say what it is for a philosophical text to make an abductive move. We then ask whether Floridi et al.'s account of LLMs gives us any reason to think that such a move cannot appear in text generated by a continuation system. ### Second half — new paragraphs Return to the wet kitchen floor. What recommended rain over a burst pipe was not that the floor was wet, which a burst pipe would have managed too, but that the water lay in a pool beneath the open window, where a broken pipe would have spread it more evenly across the boards. A philosophical text reasons in the same way whenever it does more than show that some view accommodates what we already accept, and instead sets that view against a competitor and lets one consideration decide for it and against the other. Such reasoning earns its conclusion only when the consideration it cites would, were it correct, separate the two — when it is the sort of thing one of them can claim and the other cannot — and it idles when both may help themselves to it equally (Lipton 2004, p. 42). Not every stretch of philosophy weighs competitors in this way, and much of it does other things; but this is the weighing that Floridi and his colleagues deny a continuation system can accomplish. Whether a consideration really tells two views apart is settled by what a text has set down, not by anything that passed through whoever assembled it. Whether the reason it gives for preferring one position would, if it held, leave the rival worse off turns on how that reason stands to the two, and not on the route by which it came to be written. The deliberation Floridi finds missing — the entertaining and ranking of candidate hypotheses — is at most that route; it is not where the route ends. And if the text's 'abductive appearance' (Floridi et al. 2025, p. 2) consists in its setting down a reason that really does discriminate between the positions, then what would make it worth reading in this respect is there already, whatever the standing of the process behind it. Floridi and his colleagues describe the mechanism behind this appearance in some detail. A model trained on human writing has taken in 'the typical phrasing and structure of explanations', and, asked to explain something, it returns text of that shape, supplying the causes that explanations of the kind tend to supply rather than reasoning to causes from the case before it (Floridi et al. 2025, p. 9). Even the closing verdict — "[b]ased on your description, the battery is the most likely explanation" — is, on their reading, a conversational habit picked up from the way such answers usually end, and not a ranking it has carried out (2025, p. 10). Where the output does mark a difference between two candidate explanations, that too, on their account, is 'something it has seen stated' rather than something it has worked out afresh (2025, p. 13). What a continuation system carries across, then, is the form of explanation; what it is said to leave behind is whatever, in a real weighing, makes one consideration count for a position rather than merely accompany it. That a continuation system should take in more than turns of phrase is, on reflection, much what one would expect. Wolfram (2023) trains a small network on nothing but well-formed text and finds it comes to keep its sentences grammatical, and in simple cases to carry a valid inference through — neither given to it as a rule, both simply present, throughout, in what it had read. Doing no more than continue its input does not confine such a system to the surface of the words. What such a system takes in, Floridi himself concedes, is not the phrasing of explanations alone but the patterns of the reasoning as it gets set down in writing (Floridi et al. 2025, p. 9). The ways in which one consideration is set against a rival, and something is allowed to decide between them, are worked into the prose it was trained on as steadily as grammar is; and there is, accordingly, nothing in its being a mere continuation of text that holds those patterns beyond its reach. To say so is not to say it reaches them dependably. It is to say that the bare fact of continuation does not put them out of range. The failures Wolfram emphasises concern tasks of one particular kind. Asked to keep count of which brackets in a long string are still open, or to carry a formal proof through to its end, they lose their place, for success at such a task is success at recovering the one continuation it allows, and fitting a pattern closely is no guarantee of that (Wolfram 2023). An abductive comparison sets no such single continuation to be recovered. Whether the reason a text gives discriminates between the two positions, or fails to, is a relation among the things it has said — the reason tells against the rival, or it does not — and not a link in a chain that a single slip would void; so that a system should stumble where an exact procedure is wanted does not by itself show that it cannot manage this other thing. Wolfram is candid that a system of this sort, left to itself, produces what sounds right rather than what is so, and that for any firm purchase on the world it would have to draw on instruments outside itself (2023); Floridi, in the same spirit, allows that the output may come out 'similar or even identical' to a person's while holding that the justification behind it is absent (Floridi et al. 2025, pp. 11–12). What a text can hold, even so, and even with no one having checked it, is a conditional: that if the considerations it adduces stand, the favoured position gains on its rival. That relation is in the writing whether or not its author confirmed that the considerations do stand. None of which lets the writing off answering to the truth — the preference lapses if the facts it leans on are false, or if the cost it charges the rival fails to attach — but these are ways the stated relation can come apart, found in what has been claimed, and not deficits left behind by the manner of its making. That the system reaches its preferences by falling in with a pattern, rather than by deliberating, is a fact about how the words come, not about whether they are any good. The pattern it falls in with is one in which reasons are made to discriminate between positions — reasoning as it was actually written down, which is what Floridi grants it has absorbed when he speaks of 'patterns of human abductive reasoning as expressed in writing' (2025, p. 9) — and not the bare 'because' or 'the best explanation is' that dress such reasoning up. Whether the consideration offered in a given case really does favour the one position over the other depends on what has been offered, exactly as it would with a passage no machine had touched. Where it does not, what one has is a poor piece of reasoning, plain enough on inspection; that such pieces get produced tells us how often the thing comes off, not whether it can. To cast these systems as brainstorming aids is to say that what they turn out comes unsorted, the sound mixed in with the worthless, so that someone else must sort it before any of it counts (Floridi et al. 2025, p. 11). But sorting is what reading philosophy already is; no argument, whoever set it going, is spared the question whether it holds up. A piece drawn from one of these systems and found, on reading, to make its discrimination stick is taken on the same terms as one a philosopher hit upon at the first try. The sorting such a description points to is just this reading, and that a reader must do it is the condition on which any philosophy is taken up at all, not a charge against work that came from a machine. A model need not reason, and one need not trust its output in advance, for a given output to set one position against another and offer a consideration that genuinely decides between them. Its doing no more than continue the text it is handed leaves it free to do that much, rather than barring it. What would reduce such a passage to mere appearance is a consideration that, looked at, settles nothing — and that is a fault one finds in the writing, not one fixed beforehand by the make of the machine. [^1]: Floridi et al. also support the denial with an argument from the model's relation to the world: its words are connected to no perception of anything, and a hypothesis, once produced, is never tested against the world (2025, pp. 7–9). We take that argument up in Section 3.