--- vc-id: 79c92778-0ecd-4e47-a7a1-89bc1a6aadb7 tags: - generating-philosophy --- # 2. The challenge from abduction In the previous section, we argued against the idea that LLMs cannot produce philosophy worth reading simply because they are not human. In this section and the next we shall consider a different form of challenge: even if LLMs cannot be ruled out of producing philosophy worth reading tout court, they lack particular _capacities_ that are required to produce it. If our arguments regarding authorship are correct, there is no reason that a novel philosophical argument produced by a parrot should be taken any less seriously than one produced by a human.%%This sentence comes a bit out of the blue. and is awkwardldy written%% Yet parrots _cannot_ produce such strings of sounds as their powers are mimetic, rather than productive. One might think the same is true for LLMs. They just don't have the capacity, or capacities, required to produce worthwhile philosophical argument. In this section we address one capacity challenge, which we will call _the challenge from abduction_. In the next we shall look at two more: the challenge from phenomenology and the challenge from connecting to the world. Abduction, or inference to the best explanation, is reasoning from a body of evidence to the hypothesis that would best explain it.%%Slightly repetitive.%% In a deductive argument the premises fix the conclusion: if all men are mortal and Socrates is a man, then Socrates is mortal, and there is no wriggle room. %%i don't mind that this final clause is somewhat colloquial, it is slightly awkward shape wise though%% Now, imagine walking into your kitchen and finding that part of the floor is wet. What has happened? The wet floor does not determine the answer in the way the two premises determined Socrates' mortality: a burst pipe would have left the floor wet, and so would a spilled bucket. But, given that the window is open, the water is under the window, and it rained last night, rain coming through the window seems the most plausible answer. To reason in this way, deciding the best explanation for a set of facts, is common in the sciences as well as every day life. A scientist chooses one theory over another when it explains the same data more simply: Copernicus's model of the solar system was preferred to Ptolemy's because it explained the observed planetary motions without the elaborate epicycles the older model required. Williamson argues that philosophy should also use a broadly abductive methodology (2007; 2021, §9.2). In philosophy too there are data that a candidate theory must accommodate — intuitions about cases, and the phenomena of the domain itself — and rival theories that would each accommodate them at different costs.%%long and overclunky sentence, perhaps a more succinct sentence followed by a worked out eexample in a contemporary philosophical debate (nothing cheesey)%% The theory to prefer is the one that would, if true, best explain the data. What makes one explanation better than another, on this account, is a matter of explanatory virtue: a good philosophical theory is, in Williamson's words, "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated", and should "combine simplicity with strength" (2021, §9.2). That theories are weighed by such comparative and explanatory virtues need not rest on a science-modelled conception of philosophy: Bengson, Cuneo and Shafer-Landau (2022) argue that the assessment of rival theories by their explanatory and unifying merits is a constraint on sound philosophical method as such. %%This sentence is too compressed to be worthwhile.%% This conception of philosophy is widely held (Sider 2011; Paul 2012; Dellsén et al. 2024), though not universally (Bueno and Shalkowski 2020; Thomasson 2015), and we shall assume it in what follows. %%this paragraph slips a bit from 'explanation' to 'theory'. Is there a way that this can be avoided? are we forced into theory because of williamson, or can we use eplanation for his ideas as well (perhaps with a succinct footnote to clear things up?) ideally, we should focus on explanation, but of course we need to accurately characterise what williamson says%% If a capacity for abduction is required to produce worthwhile philosophy, do LLMs possess it? Floridi et al. (2025) argue that they do not: > We argue that such LLMs generate text based on learned associations rather than performing abductive inferences. […] LLMs produce plausible hypotheses, simulate commonsense reasoning, and provide explanatory answers without grounding them directly in truth, semantics, verification, or understanding, and without any abductive reasoning. (Floridi et al. 2025, p. 1) In order to deal with this quite comprehensive critique of LLMs' abductive capacities, we will split it in two. In the rest of this section we will consider the charge that these systems do not genuinely infer to the best explanation, but simply generate text based on learned associations. In the next, we will focus on the idea that LLMs are not connected to the world in a way that would allow them to "genuinely validate" their explanations against reality. An LLM, on this account, has "a stochastic core and an abductive appearance" (2025, p. 2). Models are trained to predict which words are likely to follow which, and produce the continuation their training makes probable; it aims at the likely continuation, not at the truth. %%'aims' is too anthropormprohic, 'not at the truth' sounds editorial and cunty%% The appearance comes from what the training data have passed on: models have "absorbed patterns of human abductive reasoning as expressed in writing" (p. 9) — how explanations are typically phrased, which causes are typically offered for which effects. What is inherited, on their account, is the look of the reasoning, not the reasoning itself.[^1] %%Is this final sentence fair to Floridi? I'm not sure that it is. Well, double-check that it is, please.%% Floridi et al.'s own example is a car that will not start on a cold morning. Asked why, a model replies that a weak battery is one possibility, since cold reduces a battery's efficiency; that thickened engine oil is another, since a cold engine is harder to turn over; and that, "[b]ased on your description, the battery is the most likely explanation" (2025, p. 10). That the model offers these candidates at all is weak abduction, which Floridi et al. grant. %%very clumsyly written. do we even need to talk about strong and weak abduction?%%On their reading, though, the offering is not reasoning about causes from the case before it; it is the statistical reproduction of the causes such explanations typically cite (p. 9). What they deny is that the model weighs the candidates, the strong abduction that "entails choosing the best explanation among alternatives" (2025, p. 3). The verdict that settles on the battery does no weighing; it reproduces how explanations of this kind conventionally end. And where the output marks a genuine difference between the two — some consideration that would tell the battery from the oil — it is one already drawn in the explanations the model learned from, not one worked out afresh for the case in hand. This veneer of abduction, Floridi et al. argue, means that LLMs can only ever play a supporting role in intellectual work: > In essence, LLMs function like brainstorming assistants that toss out ideas without filtering for quality. After all, they work like statistical interfaces to an enormous amount of data accumulated for millennia by generations. A cautious human collaborator can sift through and assess them. (2025, p. 11) A brainstorming assistant is a far cry from something which might produce worthwhile philosophy. If you were presented with a text and told that it contains a number of philosophical ideas, none of which have been filtered for quality, it is unlikely you would think it is worth your time to read it. Filtering explanations for quality can be thought of as preferring what Lipton (2004, p. 59) calls lovelier explanations to merely likelier ones. The likeliest explanation is the one most warranted by the data, while the loveliest is the one that would, if true, provide the most understanding. As he puts it, "Likeliness speaks of truth; loveliness of potential understanding" (PAGE REF). The likeliest explanation is not always the loveliest. Consider again the wet floor in the kitchen. A very likely, almost certainly true, explanation is that the floor is wet because water has fallen on it, yet offering that as an explanation would be met with exasperation: of course it is because water fell on it, but how, and which water? Banally true explanations do little for a person's understanding. %%stubby sentence%% Nor is the loveliest always the likeliest.%%would a reader understand?%% A conspiracy theory involving clumsy aliens visiting one's kitchen at night would tie together the water, the door being open, and the lights you thought you saw in the sky last night, and would provide a great deal of understanding if true, but it is exceedingly unlikely to be true. On Williamson's abductive methodology, weighing rival explanations goes by their explanatory virtue %%'goes by' is a weird verb to use here. The whole sentence is not clear%%, and the explanation to prefer is the one with a virtue the others lack (2021, §9.2). To prefer an explanation on that ground is to prefer the lovelier rather than the likelier, since explanatory virtue is a matter of the understanding an explanation would afford if true, not of its probability. Not all philosophical writing turns on this kind of weighing, and we are not arguing that it does. But where a philosophical text does turn on it, whether the text is worth reading and whether it weighs its rivals well go together. And it is just this kind of weighing that Floridi and his colleagues say a system that does no more than continue text cannot do. We do not disagree with Floridi et al.'s characterisation of how LLMs function: these systems do not weigh and choose among alternatives in the way that humans do. However, we should not be too quick to jump from this to the conclusion that LLMs cannot produce text which exhibits abductive reasoning. A pocket calculator does not have the capacity to do arithmetic in the way a person does, but does have capacity to produce the correct answer to sums which are entered into it. Similarly, it might be possible for LLMs to produce text which displays abductive reasoning, despite it not being grounded in any actual abductive reasoning. However, Floridi et al. sometimes seem to argue that this possibility should be ruled out: the way LLMs function means that the best they can do is produce text that has an "abductive appearance". Trained on a great deal of writing in which explanations are offered and weighed, such a system absorbs the forms this writing takes and, prompted to explain, reproduces them, following "the typical phrasing and structure of explanations" and offering "typical causes for typical effects" rather than "reason[ing] about causes from scratch" (2025, p. 9). On this view the output has the form of an abductive explanation but not the substance — the shape of a weighing of explanations, taken over from the writing the model has digested, and not a genuine weighing of the case at hand — so that what looks like inference to the best explanation is, in their words, a "compelling illusion of genuine and structured inferential reasoning" (2025, p. 2). Set against what the model actually produces, the facade is harder to make sense of than it looks. Floridi and his colleagues allow that, in ordinary cases such as their own cold-morning car, the model's answer is a good one: "the same explanation a human reasoner would likely choose" (2025, p. 10), one that may be "even optimal by IBE criteria" (2025, p. 19), since such systems "echo the obvious, common explanations" (2025, p. 10). The complaint cannot then be that the explanation is poor, which leaves it hard to say what, in such an answer, is supposed to be merely apparent. What their account points to is the uncommon case: "in less common situations, LLMs can falter" (2025, p. 10), and on inputs "that go beyond their training" "the facade can crack" (2025, p. 9), the success on familiar cases being "a sign of overfitting to common patterns" (2025, p. 15). Once the question of its truth is set aside, this is what the facade reduces to: the claim that the competence shown on common problems is overfitting that would give out in less common ones. Whether the competence really gives out on the uncommon case is an empirical question, and the most recent survey of abductive reasoning in language models seems at first to bear the conjecture out (Salimi et al. 2026). Current models that handle deduction well do markedly worse here: where their median accuracy on deductive tasks is near eighty per cent, on abductive ones it is some forty-two and a half, with a spread "extending down to near-zero accuracy", and the survey reports that "strong deductive performance does not reliably imply strong abductive performance". Nor is it confident that the successes are what they seem: benchmark scores, it observes, register only the final answer, "completely bypassing the actual reasoning trace", so that a model "can achieve strong performance on reasoning tasks while relying on superficial patterns rather than genuine inference", and the systems fine-tuned for these tasks are "trained merely to imitate reference hypotheses". This is close to Floridi's own diagnosis — an answer that reproduces a common pattern instead of reasoning to it — now voiced from within the field that builds the systems. And as the cases grow more demanding the performance thins: on long narratives whose clues are scattered through the text the strongest models fall short of human solvers, and the survey's summary judgement is that "abductive reasoning in LLMs remains at an early stage". Taken at face value, this is the field's own assessment of these systems, and it tells in Floridi's favour.