# §2 — The Challenge from Abduction ## Source-grounded moveset M1 — Floridi et al. do not claim that LLMs fail to produce explanation-like answers. They claim that such answers are not produced by abductive inference: the model does not select an explanation because it would best account for the evidence, but generates a continuation made probable by training on human-written explanation. The appearance belongs to the output; the abductive process does not. > Source block — Floridi et al., pp. 1, 9. Floridi et al. distinguish the appearance of reasoning from the process producing it. In the abstract, LLMs are said to generate from "learned associations", while their outputs can have an "abductive appearance" because training texts already contain reasoning structures. Later, the paper calls the performance "zeroth-order abduction": a prompt is followed by a plausible continuation with the form of an explanation, but the model has not understood what explanation is or selected a cause as the best explanation. The same passage treats impressive answers as a result of absorbed written patterns rather than evidence of an internal abductive capacity. M2 — The target is the relation between the answer and the process that produced it. A model may give the answer an abductive reasoner would give, while arriving there by a process that has no grasp of explanatory force. Floridi et al.'s challenge begins from that separation. M3 — Reasoning-mode systems do not remove the separation. If hidden intermediate tokens are produced before the visible answer, the operation is still token generation within the same architecture. The system has been made better at producing answer-shaped sequences, but Floridi et al. would deny that it has thereby become an abductive reasoner. > Source block — Floridi et al., p. 18. Floridi et al. treat reasoning-mode as an advanced use of token completion rather than an escape from it. The hidden scratchpad tokens are generated by the model before the user-facing answer, but they remain part of the same completion mechanism. Their claim is that such systems may reduce errors and improve multi-step answers, but the added scratchpad does not amount to an "abduction engine". The process remains stochastic generation, only now with more internally generated material between prompt and answer. M4 — On its own, this is a claim about LLMs, not a claim about philosophy. The pressure on our project comes from adding Williamson's view that philosophy sometimes proceeds abductively. If philosophical work often consists in comparing potential explanations, then an LLM's inability to perform such comparison looks like a reason to doubt its philosophical standing. > Source block — Williamson, pp. 351, 353–354. Williamson explicitly endorses an "abductive methodology" for philosophy and treats it as inference to the best explanation, with explanation understood broadly enough to include non-causal cases. His sketch of abduction begins from potential explanations: theories are ranked before their truth is known, and they are ranked by how well they would explain the evidence if true. He adds that a better theory should combine "simplicity with strength"; it should not look arbitrary or made merely to fit the case. M5 — The challenge should be stated without making Floridi et al. say more than they say. If philosophy sometimes advances by comparing potential explanations, and LLMs do not perform that comparison, then LLM-generated philosophy may seem to preserve the outward shape of philosophical reasoning while lacking the activity that gives the shape its epistemic role. M6 — We should grant the process claim. An LLM does not understand the problem as a problem, does not know which answer is true, and does not weigh explanatory rivals under a norm of truth. If the reply required denying that, it would be a bad reply. M7 — The reply should instead distinguish production from assessment. A product can be assessed as a potential explanation even when the process that produced it was not itself an abductive inference. The question is whether the output gives us a candidate that can be evaluated philosophically, not whether the machine's internal operation already counts as philosophical evaluation. M8 — Lipton gives us the needed distinction. Inference to the best explanation cannot begin from actual explanations, because actuality already includes truth. It begins from potential explanations, whose status as actual explanations is precisely what inquiry has to decide. > Source block — Lipton, pp. 58–61. Lipton distinguishes "potential explanation" from "actual explanation" because IBE would be useless if it already had to start from explanations known to be true. A candidate can be explanatory in the relevant sense before truth has been settled. Lipton then distinguishes the explanation that is most warranted from the explanation that would provide the most understanding if true. His terms are "likeliest" and "loveliest". The distinction lets us ask whether a generated text offers an explanation with real understanding-giving shape without pretending that the model has established its truth. M9 — An LLM output can fall on the potential side of Lipton's distinction. It can propose a way the evidence would hang together if the proposal were true, even though the model itself has not judged the proposal to be true. That is enough for the product to enter philosophical assessment. M10 — Lipton's distinction between the likeliest and the loveliest explanation makes the same point from another angle. Likeliness concerns warrant; loveliness concerns the understanding an explanation would provide if true. A generated philosophical answer can be examined for loveliness before anyone has decided whether it is likely. M11 — Williamson's account of philosophical abduction also gives product-facing standards. A theory does better, on his view, when it is less ad hoc, more unified, and more informative. These are features a philosophical text can either exhibit or fail to exhibit. We do not need to inspect the author's psychology in order to ask whether the proposal has those features. M12 — Floridi et al.'s own account helps explain how a stochastic process can produce such a product. The model has been trained on writing in which human beings have already expressed explanatory reasoning. It can therefore reproduce patterns of philosophical answerability without possessing the capacity by which those patterns were first made. M13 — Philosophical training data is not merely a stock of conclusions. Published philosophical writing contains the pressure of objections and the marks left by revision under that pressure. A system trained on such writing can learn how an answer tends to be shaped when it has had to survive philosophical resistance. M14 — There is no need to pretend that token prediction is secretly inference. The claim is weaker: token prediction over philosophical writing is prediction over material already organised by philosophical norms. The process remains stochastic, but the source distribution is not philosophically inert. M15 — Lipton's two-stage picture gives a useful place for LLMs. Inquiry first generates a limited set of live candidates and then selects among them. LLMs can contribute to the first stage without being trusted with the second. > Source block — Lipton, pp. 149–151. Lipton describes inquiry as a "two-stage process": possible explanations are not considered all at once, but enter inquiry through a short list of "live candidates". Background beliefs help form that shortlist. The same discussion stresses that the background is itself shaped by earlier explanatory inferences; successful selections do not disappear once made, but become part of the material from which later candidates are generated. This gives us a way to describe the philosophical corpus: not as a neutral heap of sentences, but as a public background shaped by earlier selection. M16 — The generated candidate still has to be tested. It may smooth over the difficulty it should face, or preserve the vocabulary of explanation while leaving the explanatory burden untouched. In that case the output has produced philosophical-looking prose, not a satisfactory philosophical explanation. M17 — Floridi et al.'s warning about verification therefore remains in force. The system cannot certify that its answer is true, nor can it know that its explanation has succeeded. The responsibility for assessment stays with the philosopher. M18 — The answer to the abduction challenge is not that LLMs are abductive reasoners after all. It is that a non-abductive process can generate material with abductive form because it has been trained on the products of human abductive labour. Philosophy can then treat that material as a candidate, not as an authority. M19 — The section should leave us with a product-centred standard. LLM-generated philosophy is not vindicated by fluency, and it is not defeated merely by the absence of inner abduction. It is worth taking seriously when it offers a potential explanation that can survive the ordinary tests of philosophical judgement. ---