In the previous section, we argued against the idea that LLMs cannot produce philosophy worth reading simply because they are not human. In this section and the next we shall consider a different form of challenge: even if LLMs cannot be ruled out of producing philosophy worth reading tout court, they lack particular _capacities_ that are required to produce it. If our arguments regarding authorship are correct, there is no reason that a novel philosophical argument produced by a parrot should be taken any less seriously than one produced by a human. Yet parrots _cannot_ produce such complex strings of sounds as their powers are mimetic, rather than productive. One might think the same is true for LLMs. They just don't have a capacity, or capacities, required to produce worthwhile philosophical argument. In this section we address one capacity challenge, which we will call _the challenge from abduction_. In the next we shall look at two more: the *challenge from phenomenology* and the challenge from *connecting to the world*. Abductive inference is reasoning from a body of evidence to the hypothesis that would best explain it. In a deductive argument the premises fix the conclusion: if all men are mortal and Socrates is a man, then Socrates is mortal, with no wriggle room whatsoever. Now, imagine walking into your kitchen one morning and finding that part of the floor is wet. What has happened? The wet floor does not determine the answer in the way the two premises determined Socrates' mortality: a burst pipe would have left the floor wet, and so would a spilled bucket. But, given that the window is open, the water is under the window, and it rained last night, rain coming through the window seems the most plausible answer. To reason in this way, inferring the best explanation for a set of facts, is common in the sciences as well as every day life. A scientist chooses one theory over another when it explains the same data more simply: Copernicus's model of the solar system was preferred to Ptolemy's because it explained the observed planetary motions without the elaborate epicycles the older model required. Williamson argues that philosophy should use a broadly abductive methodology (2007; 2021, §9.2). In philosophy, as in science, there are data that a candidate theory must accommodate — intuitions about cases, and the phenomena of the domain itself — and rival theories that would each accommodate them at different costs. The theory to prefer is the one that would, if true, best accounts for the data, with 'best' being cashed out in terms of explanatory virtue What makes one theory's explanation better than another's, on this account, is a matter of explanatory virtue%%repeat%%: a good philosophical theory is, in Williamson's words, "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated", and should "combine simplicity with strength" (2021, §9.2). This need not rest on a science-modelled conception of philosophy: Bengson, Cuneo and Shafer-Landau (2022) argue that the assessment of rival theories by their explanatory and unifying merits is a constraint on sound philosophical method as such.[^holders] We shall assume it in what follows. If a capacity for abduction is required to produce worthwhile philosophy, do LLMs possess it? Floridi et al. (2025) argue that they do not: > We argue that such LLMs generate text based on learned associations rather than performing abductive inferences. […] LLMs produce plausible hypotheses, simulate commonsense reasoning, and provide explanatory answers without grounding them directly in truth, semantics, verification, or understanding, and without any abductive reasoning. (Floridi et al. 2025, p. 1) We will split this quite comprehensive critique of LLMs' abductive capacities into two parts. In the rest of this section we will consider the charge that these systems do not genuinely infer to the best explanation but simply generate text based on learned associations. In the next, we will focus on the idea that LLMs are not connected to the world in a way that would allow them to "genuinely validate" (ibid. p. 6) their explanations against reality. An LLM, on this account, has "a stochastic core and an abductive appearance" (2025, p. 2). Models are trained to predict which words are likely to follow which, and to produce the continuation their training makes probable — to aim at the likely continuation rather than truth. The appearance comes from what the training data have passed on: models have "absorbed patterns of human abductive reasoning as expressed in writing" (p. 9) — how explanations are typically phrased, which causes are typically offered for which effects. [^1] Floridi et al.'s own example is a car that will not start on a cold morning. Asked why not, a model replies that a weak battery is one possibility, since cold reduces a battery's efficiency; that thickened engine oil is another, since a cold engine is harder to turn over; and that, "[b]ased on your description, the battery is the most likely explanation" (2025, p. 10). On their reading, the LLM is not actually reasoning about causes from the case before it; it is the statistical reproduction of the causes such explanations typically cite (p. 9). What they deny is that the model weighs the candidates — that it "entails choosing the best explanation among alternatives" (2025, p. 3). The verdict that settles on the battery does no weighing; it reproduces how explanations of this kind conventionally end. And where the output marks a genuine difference between the two — some consideration that would tell the battery from the oil — it is one already drawn in the explanations the model learned from, not one worked out afresh for the case in hand. The fact that LLMs can produce no more than a "veneer of explanation" (ibid. p.20), Floridi et al. argue, means that LLMs can only ever play a supporting role in intellectual work: > In essence, LLMs function like brainstorming assistants that toss out ideas without filtering for quality. After all, they work like statistical interfaces to an enormous amount of data accumulated for millennia by generations. A cautious human collaborator can sift through and assess them. (2025, p. 11) A brainstorming assistant seems a far cry from something which might produce worthwhile philosophy. If you were presented with a text and told that it contains a number of philosophical ideas, none of which have been filtered for quality, it is unlikely you would think it is worth your while reading it. The filtering Floridi et al. have in mind is not a matter of selecting the most likely continuation — the stochastic core already does that. It is a matter of preferring what Lipton (2004, p. 59) calls *lovelier* explanations to merely *likelier* ones: the likeliest is the one most warranted by the data, the loveliest the one that would, if true, provide the most understanding. As he puts it, "Likeliness speaks of truth; loveliness of potential understanding" (PAGE REF). The two come apart. Consider again the wet floor in the kitchen. A very likely, almost certainly true, explanation is that the floor is wet because water has fallen on it, yet offering that as an explanation would be met with exasperation: of course it is because water fell on it, but how, and which water? Banally true explanations do little for a person's understanding. On Williamson's abductive methodology, weighing rival explanations is a matter of explanatory virtue, and the explanation to prefer is the one with a virtue the others lack (2021, §9.2). To prefer an explanation on that ground is to prefer the lovelier rather than the likelier, since explanatory virtue is a matter of the understanding an explanation would afford if true, not of its probability. Not all philosophical writing turns on this kind of weighing, but where a philosophical text does turn on it, whether the text is worth reading and whether it weighs its rivals well go together. And it is just this kind of weighing that Floridi and his colleagues say a system that does no more than continue text cannot do. We do not disagree with Floridi et al.'s characterisation of how LLMs function: these systems do not weigh and choose among alternatives in the way that humans do. However, we should not be too quick to jump from this to the conclusion that LLMs cannot *produce text* which exhibits abductive reasoning. A pocket calculator does not have the capacity to do arithmetic in the way a person does, but does have capacity to produce the correct answer to sums which are entered into it. Similarly, it might be possible for LLMs to produce text which displays abductive reasoning, despite it not being grounded in any actual abductive reasoning. Floridi et al. are not merely saying that the process by which an LLM produces its output is stochastic rather than abductive. They are saying that the output itself is only apparently abductive — that what looks like inference to the best explanation is, in their words, a "compelling illusion of genuine and structured inferential reasoning" (2025, p. 2). Trained on a great deal of writing in which explanations are offered and weighed, such a system absorbs the forms this writing takes and, prompted to explain, reproduces them, following "the typical phrasing and structure of explanations" and offering "typical causes for typical effects" rather than "reason[ing] about causes from scratch" (2025, p. 9). On this view the output has the form of an abductive explanation but not the substance — the shape of a weighing of explanations, taken over from the writing the model has digested, and not a genuine weighing of the case in hand. Floridi and his colleagues allow that, in ordinary cases such as their own cold-morning car, the model's answer is a good one: "the same explanation a human reasoner would likely choose" (2025, p. 10), one that may be "even optimal by IBE criteria" (2025, p. 19), since such systems "echo the obvious, common explanations" (2025, p. 10). The complaint cannot then be that the explanation is poor, which leaves it hard to say what, in such an answer, is supposed to be merely apparent. What their account points to is the uncommon case: "in less common situations, LLMs can falter" (2025, p. 10), and on inputs "that go beyond their training" "the facade can crack" (2025, p. 9), the success on familiar cases being "a sign of overfitting to common patterns" (2025, p. 15). Once the question of its truth is set aside, this is what the facade reduces to: the claim that the competence shown on common problems is overfitting that would give out in less common ones. Whether the competence gives out on the uncommon case is an empirical question, and the most recent survey of abductive reasoning in language models seems at first to bear the conjecture out (Salimi et al. 2026). Current models do markedly worse on abductive tasks than on deductive ones: where their median accuracy on deductive tasks is near eighty per cent, on abductive ones it is some forty-two and a half, with a spread "extending down to near-zero accuracy", and the survey reports that "strong deductive performance does not reliably imply strong abductive performance". This is close to Floridi's own diagnosis — an answer that reproduces a common pattern instead of reasoning to it — now voiced from within the field that builds the systems. Taken at face value, it tells in Floridi's favour. Our response to the challenge from abduction begins by considering what text having an abductive appearance actually amounts to. Consider first that, despite their stochastic core, LLMs are perfectly capable of producing grammatically correct text. Despite not being given specific rules, LLM training means that the system "implicitly 'discovers' them—and then seems to be good at following them" (Wolfram 2023). Does this mean that the texts LLMs produce have merely the appearance of being grammatically well-formed? Clearly not. LLMs sentences *are* gramatically well-formed despite their stochastic roots. This suggests that a stochastic core need not mean that the best an LLM can do is produce a veneer of abductive inference. It may be, rather, that the core is marshalled to produce text exhibiting actual abductive inference, in just the way it is marshalled to produce actual grammatical correctness. It might be objected that the grammar analogy will not stretch this far. The model has picked up the shape of abductive explanation, but it is not actually performing abduction. The preceding paragraph suggested that because LLMs implicitly discover and follow the rules of grammar, they might in the same way be marshalled toward genuine abductive inference. But grammar, the objector would say, is a system of rules a model can follow without understanding anything, and explanatory loveliness is not — it is not a system of rules you follow to get the right answer, so the model's grammatical competence gives no reason to expect it to produce lovely explanations. Floridi's claim is that the model's abductive output is merely apparent — that it has the form of an inference to the best explanation but not the substance. If that is right, then there should be cases where the form is present and the substance is absent: cases where the model produces something that looks like an explanation but is empty or nonsensical, just as it could produce something that is grammatically faultless and says nothing. Wolfram's example of the latter is "Inquisitive electrons eat blue theories for fish" (2023) — impeccably grammatical, and meaningless. The model does not produce such strings. The abductive equivalent would be a passage with the form of an inference to the best explanation, grammatically faultless, and yet senseless — asked why a car will not start on a cold morning, there are no squirrel tracks, so it must be freak arctic winds blowing into the exhaust pipe. That has the form of evidence weighed towards a conclusion, and it is nonsense; and it is nonsense the model does not produce. Put the question to it and it offers the weak battery and the thickened oil, plausible explanations rather than arctic winds. Wherever the facade is to be located, it cannot be located there: what the model produces are plausible explanations, and it is not yet clear what, in a plausible explanation, is supposed to be merely apparent. Moreover, the answer Floridi and his colleagues describe is not quite the answer these systems give. Their account asks us to picture a confident verdict laid over a hollow core, yet what one finds in practice is closer to hedging. Asked why a car would not start on a cold December morning, a current model will say that the battery is the most likely culprit, set out other plausible contributors — thickened oil, fuel-system problems, ignition faults — and end by observing that if the car started once the day had warmed, the battery is almost certainly the primary cause, though a load test would confirm it.[^kimi] The reply does not announce the answer; it fits its confidence to the little it has been told, and marks the point past which it will not go without knowing more of the particular car. This is not yet to say that Floridi's facade charge is refuted — a hedged answer can still be, on his account, a statistical reproduction of how hedged explanations typically read. But it does mean that Floridi et al.'s own example does not quite fit the charge. The case they describe, in which a confident surface verdict is laid over a hollow core, is not the case their own example presents: the confidence these systems express is already answerable to their evidence, and a confidence so answerable is not the facade the objection has in view. That the model keeps clear of senseless explanations, as it keeps clear of senseless sentences, points to a core of semantic competence picked up in training, over and above syntax. Wolfram calls it a semantic grammar. Syntax, he notes, settles only how the parts of speech may be combined: > to deal with meaning, we need to go further. And one version of how to do this is to think about not just a syntactic grammar for language, but also a semantic one. (Wolfram 2023) A model trained on enough text has, on his account, come by one: > From its training ChatGPT has effectively "pieced together" a certain (rather impressive) quantity of what amounts to semantic grammar. (Wolfram 2023) Wolfram's notion of a semantic grammar is meant to capture what a model picks up over and above syntax. A syntactic grammar settles only how the parts of speech may be combined — what counts as a well-formed sentence. It does not settle whether what a sentence says makes sense. "Inquisitive electrons eat blue theories for fish" is grammatical and says nothing, and it is the absence of such strings from the model's output that tells against treating its syntax as appearance alone. Avoiding such strings takes more than syntax: it takes a grasp of what can sensibly be said of what, which predicates go with which subjects, which causes go with which effects. ~~Wolfram calls this a semantic grammar~~, and his claim is that a model trained on enough text has come by one — has, in his words, "pieced together" a quantity of what amounts to semantic grammar (2023). It is this, rather than any contact with the case in hand, that keeps the arctic winds out of the model's explanations: a feel for what, in a working model of the world, can hang together. But a feel for how things hang together is gathered from the text the model has read, and a model of the world is not the world.[^wm] What a semantic grammar supplies is a sense of what would sound right, not a line to how things actually stand. The same detachment that keeps the model's explanations sensible is what leaves it unable, on its own, to reach the car on the drive. A semantic grammar is a grasp of how such failures are explained in general — cold slows batteries, thickens oil, fouls ignition — and it gives no hold on the particular car whose fault is in question. Pressed for the cause of this failure to start, the model can go only so far before it needs what it cannot get for itself: some purchase on the actual car in the actual world, the lights tried, the turn of the key heard. Set that aside, and what is left is, so far as the text goes, abduction itself: a weighing of explanations that respects sense, fits its confidence to the evidence, and stops where the evidence stops. How much the missing purchase costs depends on the kind of abduction at issue. The car is a hard case, because its answer waits upon the world; only the car itself can tell the weak battery from the frozen line. Much philosophical abduction does not wait upon the world in that way: what a thought experiment commits us to, or which of two theories carries the lighter explanatory cost, is settled from what is already set down, where the discriminating evidence is the sort a corpus already holds. The case that shows the model at its most hobbled is thus the world-bound one, and not the case philosophy most often presents — a matter we take up in the next section. What is left for this one is the empirical question of whether, on harder cases, this competence in fact gives out. Returning to the survey by Salimi et al. (2026), the scores measure whether the model arrived at the right answer, not the weighing that got it there. On the harder tasks the scores are low, and on long mysteries with their clues strewn through the text the best models fall just short of the average human solver; taken at face value, the numbers count against the model. But the score is a score for the answer — whether the named culprit was the keyed one, whether the right diagnosis came first — and not for the weighing that reached it; such a score, Salimi says, "completely bypass[es] the actual reasoning trace". So when a model misses the keyed culprit, the number marks the miss, and not the comparison it set out on the way — which explanations it canvassed, and why it came down on one. And the survey's hardest tasks are built from low-prior, non-stereotypical outcomes, where several explanations may be reasonable; there, matching an answer to a single reference "underestimates explanation quality", so a low score on such tasks does not show that the model cannot weigh, only that it did not land on the keyed answer. What a one-shot score leaves unscored is the weighing behind it, and that is what a fuller prompting would have to draw out. %%this paragraph is a bit messy and repetitive%% [Signpost]The section has argued that Floridi's facade charge does not hold against the output the model actually produces. The model avoids senseless explanations, fits its confidence to its evidence, and stops where the evidence stops; what it lacks is not abductive competence but purchase on the particular case. The semantic grammar that keeps the model's explanations sensible is gathered from text, not from the world, so the model can say how cold-weather starting failures are explained in general but cannot reach the particular car on the drive — ~~the lights tried, the turn of the key heard~~ — because that takes contact with the world, and the model has none. But that is a limitation about the world, not about abduction. Whether the competence gives out on harder cases is an empirical question the survey does not settle, because the scores measure answers, not the weighing behind them: a one-shot prompt may miss the weighing without showing it is not there. That leaves two questions for what follows. The next section takes up the world-connection challenge: whether the model's detachment from the world disqualifies its output from counting as genuine philosophy, or whether the kinds of abduction philosophy most often demands are precisely those that do not wait upon the world. The section after that takes up the elicitation question: if the weighing is there but a one-shot prompt does not draw it out, what does? [^wm]: We use "model of the world" in Wolfram's sense; whether such systems possess world models in the sense now disputed, and how they connect to the world, is the matter of §3. Example of real exchange with kimi 2.6 high > **User:** Last December, I tried to start my car early in the morning when it was very cold. It wouldn't start. Why not? > > **Kimi k2.67:** The most likely culprit is the battery. In very cold weather, a battery's chemical reactions slow dramatically, reducing its available capacity by up to 50%. If the battery was already marginal (older than three years, or not fully charged), freezing temperatures can push it below the threshold needed to turn the starter motor. You may have heard rapid clicking or nothing at all. > > Other plausible contributors: > > - Thickened engine oil — Cold increases viscosity, making the engine physically harder to crank and demanding even more current from the already weakened battery. > - Fuel system — Condensation in the fuel lines can freeze, blocking flow. Diesel fuel can gel; petrol can struggle to vaporise properly in extreme cold. > - Spark/ignition — Moisture in the distributor cap or on spark plugs can freeze, preventing a proper spark. > > If it started fine once temperatures rose later in the day, the battery is almost certainly the primary cause. A load test would confirm whether it needs replacement or just a longer drive to reach full charge.