Branched from [Draft Review Feedback](https://chatgpt.com/g/g-p-6863cdae32988191b6ef486a436ee4e5/c/69f85e4e-a8bc-83eb-8560-312c1097b8e8) Section Ii Challenge From Abduction New Iteration ## II. The challenge from abduction The next challenge concerns abduction, or inference to the best explanation. Abductive reasoning moves from a phenomenon to be explained to a candidate explanation of it, and then asks whether that candidate explains the phenomenon better than the relevant alternatives. In philosophy, the relevant claim is not that every piece of reasoning has this form, but that theories are often defended by appeal to their explanatory power. Williamson’s point is methodological rather than terminological: philosophers often test theories by considering how well they explain the evidence (Williamson 2016). Dellsén et al. help to show why this kind of comparison is connected to understanding. They argue that increased understanding consists in representing more accurately or more comprehensively the network of dependence relations between phenomena (Dellsén et al. 2024, pp. 674–676, 687). Abductive reasoning fits this picture because it asks which proposed dependence relations would, if correct, best account for the phenomena under discussion. Bengson, Cuneo, and Shafer-Landau help specify what is being assessed when a philosophical text is assessed in this way. They describe philosophy as beginning from data: claims or considerations that a theory has to handle. Theorising consists in constructing and assessing candidate theories in light of those data (Bengson, Cuneo, and Shafer-Landau 2022, pp. 4–7). This is different from asking only whether a particular argument is cogent. A philosopher might show that a conclusion follows from two premises while leaving open why the phenomenon under discussion occurs, or how the proposed thesis accounts for the data that made the problem pressing. The abduction challenge, then, is whether LLMs can produce text that contributes to this kind of theory construction and evaluation. Lipton gives the account of abduction that we need in order to state this challenge precisely. According to Inference to the Best Explanation, explanatory considerations guide inference: given our data and background beliefs, we infer what would, if true, provide the best explanation of those data, so long as the best candidate is good enough to merit inference (Lipton 2004, p. 56). The phrase ‘if true’ matters. Lipton is not saying that we first identify the actual explanation and then infer it. That would make the account useless as a description of inquiry, since we would already have reached the truth before the inference began. As he puts the point, a model that instructed us to infer actual explanations would be like a recipe that tells us to "start with a souffle" (ibid., p. 58). The relevant object of assessment is therefore not an actual explanation, but a potential explanation. A potential explanation is a candidate that would explain the evidence if it were true. This distinction does several jobs. It allows for reasonable but false inferences. It allows incompatible candidates to compete with one another. And it makes the account epistemically usable, because we can assess what a candidate would explain before we know whether it is true (Lipton 2004, pp. 57–59). Philosophical readers already proceed in this way. A paper can be worth reading because it formulates a candidate explanation with enough clarity and force to make the problem more intelligible, even if the reader ultimately rejects that candidate. Lipton also stresses that potential explanations are not normally selected from the whole space of logical possibilities. Inquiry usually operates with a pool of live candidates. On one version of the view, there is first a filter that restricts the pool to serious candidates, and then a second filter that selects from among them (Lipton 2004, p. 59). This two-filter structure is familiar in philosophy. A paper does not compare its preferred account with every possible view, however bizarre. It identifies live alternatives in a debate and argues that one of them explains the data better than the others. This is why philosophical abduction is not just the production of an explanation-shaped sentence. It is a comparison between candidates in a structured dialectical space. Lipton’s second distinction is between the likeliest and the loveliest explanation. The likeliest explanation is the one most likely to be true. The loveliest explanation is the one that would, if true, provide the most understanding. Lipton’s formulation is compact: "Likeliness speaks of truth; loveliness of potential understanding" (Lipton 2004, p. 59). The two standards can come apart. The dormitive-power explanation of opium may be likely, but it is unlovely because it gives little understanding. A conspiracy theory may be lovely in one respect, since it would unify disparate events if true, while still being unlikely. Newtonian mechanics became less likely after later evidence supported relativity, but it did not cease to be a lovely explanation of the older data (ibid., pp. 59–60). This distinction is not a decorative addition to the account. If ‘best’ meant simply ‘likeliest’, then Inference to the Best Explanation would say little more than that we infer whichever explanation we judge most probable. Lipton thinks that this would make the view close to trivial. The more interesting claim is that explanatory loveliness is a guide to likeliness: we use explanatory virtues in order to judge which candidate is more likely to be true (Lipton 2004, pp. 60–62). In short, loveliness is meant to help explain why certain hypotheses strike us as better warranted than their rivals. The account therefore links the search for truth with the search for understanding, without identifying the two. Lipton’s discussion of explanatory virtues gives this claim further content. Among the virtues he discusses are mechanism, precision, scope, simplicity, fertility, and fit with background belief (Lipton 2004, pp. 121–123). These are not mere stylistic virtues. An explanation with greater scope explains more; a more precise explanation tells us more exactly why this phenomenon, rather than a nearby alternative, occurred; a fertile explanation opens further explanatory possibilities; an explanation that fits with background belief is less isolated from what we already have reason to accept. In scientific cases, ‘mechanism’ often means a causal mechanism. In philosophy, the analogue will often be structural: a distinction, dependency relation, grounding relation, representational structure, or inferential pattern that shows why the relevant data hang together. Lipton’s treatment of contrastive explanation is also relevant. Many explanations answer not simply ‘Why P?’ but ‘Why P rather than Q?’ The foil helps determine what counts as an adequate explanation. To explain why someone ordered eggplant rather than sea bass, for example, it may be enough to say that they did not know sea bass was available; that would not explain why they ordered eggplant *simpliciter*. Lipton’s point is that a contrast can make an explanation easier in one respect and harder in another: easier because the explanandum is more focused, harder because the explanation must distinguish the fact from the foil. Much philosophical abduction has this form. We ask not merely why a theory has some consequence, but why this consequence rather than that; not merely why a distinction can be drawn, but why this distinction rather than a rival one is the distinction that matters. This gives us a sharper way to think about LLM outputs. A generated philosophical text does not count as abductively interesting merely because it contains words such as ‘explanation’, ‘therefore’, or ‘best account’. It must formulate a potential explanation, locate the relevant live alternatives, and clarify the contrast in relation to which the explanation is supposed to be good. It must also exhibit at least some of the virtues that make explanations lovely in Lipton’s sense. In philosophy, this may mean that the text explains why a debate has become stuck, why two claims seem incompatible when they are not, why a familiar objection misses its target, or why a distinction handles a case that its rival cannot. Bengson, Cuneo, and Shafer-Landau allow us to translate this Liptonian framework into terms tailored to philosophical theorising. On their Tri-Level Method, a philosophical theory should, first, accommodate and explain the data; second, substantiate and integrate its own claims and commitments; and third, when the lower-level criteria leave things tied, possess theoretical virtues (Bengson, Cuneo, and Shafer-Landau 2022, pp. 107–109). Accommodation and explanation are distinct. A theory accommodates a datum when the datum is likely given the theory; it explains a datum when it shows why the datum holds (ibid., pp. 110–114). This distinction is useful because it prevents us from confusing mere fit with philosophical illumination. A view may make a datum unsurprising without explaining it, and an LLM output may fit a prompt without giving us any understanding. The second level of their method adds another demand. A theory must substantiate and integrate its own claims and commitments. To substantiate a claim is to defend and explain it; to integrate a theory is to show that its claims cohere with one another and with our best picture of the world (Bengson, Cuneo, and Shafer-Landau 2022, pp. 115–121). These are product-level features. A text either gives reasons for a claim or it does not. It either explains why the claim should be accepted or it stops too soon. It either shows how the claim fits with other commitments or leaves a conflict unresolved. A reader can inspect these features without first settling what psychological process produced the text. Floridi et al. press the challenge from the other side. They do not deny that LLMs produce answers that look explanatory. Their claim is that these answers are not generated by genuine abductive inference. They write: > LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025, p. 9) This passage makes three claims that should be kept separate. First, LLMs produce plausible continuations. Second, some of these continuations have the form of hypotheses or explanations. Third, the explanation-like output is generated by learned associations rather than by an understanding of explanation, evidence, causes, or truth. Floridi et al. call this ‘zeroth-order abduction’ because the output resembles abduction at the level of product, while lacking the cognitive structure of abductive inference at the level of process. The model does not ask what would explain the evidence. It does not understand one candidate as lovelier or likelier than another. It produces text that fits patterns learned from human-produced explanations. This is a real challenge. If philosophy requires the exercise of abductive judgement, and if LLMs lack the relevant judgement, then it may seem that their outputs cannot be philosophical in the relevant sense. At best, they will imitate the external marks of philosophical explanation. They will produce typical philosophical continuations for typical philosophical prompts, much as Floridi et al. say that LLMs output typical causes for typical effects. The danger, then, is not only falsehood. It is empty intelligibility: text that gives the impression of having explained something while merely moving through familiar forms. Our response is not that LLMs secretly perform human-style Inference to the Best Explanation. They do not understand the philosophical problem as a problem. They do not knowingly compare live candidates under a norm of truth. They do not infer a conclusion because they judge it to be the loveliest potential explanation of the available data. We should grant this much to Floridi et al. But that concession does not settle whether the text produced by an LLM can present a potential explanation, organise a comparison among live candidates, or display explanatory virtues that readers can assess. The distinction we need is between producer-side abduction and product-side abductive structure. Human abductive reasoning is one way of producing a text that contains a candidate explanation. It is not the only way such a text could come into existence. Once a potential explanation is articulated, its philosophical merits are not hidden in the mind of its producer. The reader can ask whether the explanandum has been correctly identified, whether the contrast is the right one, whether the live alternatives have been fairly represented, whether the proposed explanation gives understanding, and whether it does so better than its rivals. These are questions about the product. Lipton’s framework makes this reply more precise. What is assessed in Inference to the Best Explanation is a potential explanation: something that would explain the data if true. That object can be made available in prose. The prose can specify the data, formulate the candidate explanation, state the relevant foil, compare the candidate with live alternatives, and show what understanding it would provide. None of this establishes that the candidate is true. Nor does it establish that the producer arrived at it by human abductive reasoning. But it does make available a Liptonian object of assessment: a candidate explanation whose loveliness and likeliness can be considered by readers. Bengson, Cuneo, and Shafer-Landau sharpen the same point in philosophy-specific terms. A philosophical text can make a contribution by helping a theory accommodate or explain data, by substantiating one of its claims, by integrating it with background commitments, or by showing that a rival fails at one of these tasks. Their account of objections makes this especially clear. An objection is a consideration that gives reason to think that a theory does poorly with respect to one or more methodological criteria: it may fail to accommodate the data, fail to explain them, be ad hoc, stop too soon, be circular, or conflict with science or common sense (Bengson, Cuneo, and Shafer-Landau 2022, pp. 134–137). An LLM-generated objection can therefore be assessed by asking whether it locates one of these failures. If it does, its value does not depend on the model’s having experienced the objection as an objection. This also shows why Floridi et al.’s concern about learned associations is not by itself decisive. They are right that a model may reproduce the surface grammar of explanation without doing explanatory work. It may say ‘the best explanation is’ without identifying a genuine explanatory virtue. It may say ‘one might object’ without locating any failure of accommodation, explanation, substantiation, or integration. It may say ‘this distinction resolves the problem’ while drawing a distinction that leaves the same data untouched. These are failures, but they are recognisable as failures by reading the output. We do not need to infer them solely from the model’s causal architecture. Indeed, Floridi et al.’s own formulation leaves room for this possibility. They say that LLMs have "absorbed patterns of human abductive reasoning as expressed in writing" (Floridi et al. 2025, p. 9). This should not be inflated into the claim that LLMs understand those patterns. But it should also not be deflated into the claim that they have learned only empty verbal templates. Philosophy is conducted, preserved, and criticised in writing. The philosophical corpus contains arguments, objections, replies, distinctions, failed proposals, refined proposals, and later attempts to explain why those failures and refinements matter. To train on such material is to train on texts in which abductive and dialectical standards have left public traces. The point is not that these traces guarantee philosophical success. They do not. A model trained on philosophical prose can still produce a paragraph that is fluent, plausible, and empty. But philosophical prose has a public dialectical organisation. Objections create burdens; replies discharge or fail to discharge them; distinctions separate cases that had been run together; explanations compete by making different aspects of the same material intelligible. These relations can be present in a text even when the process that produced the text is not itself an episode of understanding. And when they are present, the reader can assess them in the ordinary philosophical way. Consider a negative case. A theory purports to explain a datum by appeal to a distinction between two kinds of experience. A generated text may point out that the same distinction also applies to a case the theory treats differently, and so the proposed explanation either overgeneralises or needs a further restriction. This need not be a mere verbal performance. If the criticism is right, the text has identified a burden of substantiation or integration. It has shown that the theory cannot yet explain the data in the way it claims to. Whether the model understood this pressure is a different question from whether the pressure is really there. The same point applies on the constructive side. A generated text may introduce a distinction that explains why two claims previously treated as incompatible can both be accepted, provided they are assigned to different levels of analysis. Or it may show that a debate has been framed around the wrong contrast: the issue is not why P rather than not-P, but why P rather than Q. In Lipton’s terms, this changes the relevant fact–foil structure and thereby changes what would count as a good explanation. In Bengson, Cuneo, and Shafer-Landau’s terms, it can help a theory accommodate and explain the data, or integrate its commitments. If such a move is made in the text, then there is philosophical material there for the reader to consider. This also helps to distinguish LLM-produced philosophy from mere philosophical mimicry. A mimicry case preserves the outward form while failing the relevant philosophical test. It may contain a thesis, an objection, a reply, and a conclusion, but the objection does not press on the thesis, the reply does not answer the objection, and the conclusion does not follow from the dialectic. A successful case is different. It need not settle the matter. It need not be correct. But it must change what is available for assessment: by making a candidate explanation clearer, exposing a burden, improving a contrast, or showing how a claim could be substantiated or integrated. The challenge from abduction therefore does not show that LLM-produced philosophy is impossible. It shows that many outputs will be poor philosophy: fluent continuations that repeat the surface form of explanation without offering a potential explanation worth assessing. That is a danger we should accept. But it is not a constitutive exclusion. LLMs do not perform Inference to the Best Explanation in the way human philosophers do. Still, they can produce texts that formulate potential explanations, organise live alternatives, and display the kinds of explanatory virtues that guide philosophical assessment. The conclusion is limited but not trivial. LLMs lack human-style abductive reasoning. They do not understand explanations as explanations, and they do not infer conclusions because they judge those conclusions lovely or likely. But this absence in the producer does not entail an absence in the product of explanatory structure, theoretical virtue, or philosophical value. For the purpose of deciding whether an LLM output is philosophy worth reading, the question is not whether the model reasoned abductively, but whether the text gives us an abductive structure worth assessing. In some cases, it can.