%%
Moved from Section 1 on 2026-03-18 so the material is preserved for reuse when this section develops the Lipton discussion more fully:
What, then, does the success of such reasoning consist in? Consider Lipton's distinction between two ways of understanding the best explanation. He writes:
> We may characterize it as the explanation that is most warranted: the 'likeliest' or most probable explanation. On the other hand, we may characterize the best explanation as the one which would, if correct, be the most explanatory or provide the most understanding: the 'loveliest' explanation. The criteria of likeliness and loveliness may well pick out the same explanation in a particular competition, but they are clearly different sorts of standard. Likeliness speaks of truth; loveliness of potential understanding. (*Inference to the Best Explanation*, p. 59)
A hypothesis can have strong evidential support and still leave us in the dark about why the phenomenon looks the way it does. We can ask of a hypothesis what kind of understanding it would provide if it were true, and this is not the same question as whether the evidence already warrants belief in it. A lovely explanation does not simply fit what we know; it makes the matter intelligible, showing why the phenomenon has the shape it does.
Lipton develops the distinction through Semmelweis's work on childbed fever. Faced with the much higher mortality rate in one division of the Vienna maternity hospital than in the other, Semmelweis asked not just which hypothesis the evidence best supported but which would, if true, make the contrast between the two divisions intelligible. Some candidate explanations did very little with the phenomenon even if they were granted: the route taken by the priest, for instance, left obscure why this should be a matter of life and death. The cadaveric hypothesis did more. It connected the contrast with the medical students' contact with corpses, with Kolletschka's death after a puncture wound, and with the subsequent fall in mortality once disinfection was introduced. In Lipton's terms, it was lovelier because it rendered the pattern intelligible.
%%
Floridi, Nobre, and Taddeo (2024) %% just fucking write 'et al.' you make this mistake so often. Can we change your config to stop you fucking doing this? Whenever there are more than two authors of paper use this Latin phrase. For fuck's sake, how many times? %%argue that LLMs do not reason abductively. Consider what happens when an LLM is prompted to explain why a car might not start on a cold morning. It generates text exhibiting explanatory structure: it identifies a hypothesis (the battery), provides a reason (cold weather reduces battery efficiency), and presents the explanation with the connectives and qualifications that explanations typically have. %%not how i write%% But the LLM does not select this explanation by comparing it with alternatives and judging it best. It outputs the most probable continuation given its training. Floridi et al. put the point this way:
> Given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (p. 9)
Floridi et al. call this *zeroth-order abduction*. The phrase marks an absence.%%not how i write%% Recall Lipton's account of abductive reasoning as a two-stage process: a generation stage, in which "our background beliefs help us to generate a very limited list of plausible hypotheses", and a selection stage, in which we choose among those hypotheses on the basis of explanatory virtues (2004, p. 149).%% this would be good if your introduction of Lipton wasn't so fucking dreadful in the previous section. %% A scientist considering rival explanations first narrows the field — most conceivable hypotheses are never entertained — and then evaluates which survivor, if true, would provide the deepest understanding. LLMs collapse this two-stage process into a single step. They generate one plausible continuation without considering alternatives, and they lack what Floridi et al. call "an external feedback loop for posterior evaluation" — they do not validate their outputs against reality (pp. 5–6). Human reasoners have "additional safeguards, like new evidence, experiments, logical scrutiny" (p. 9); LLMs, unless augmented, have none. The result, in Floridi et al.'s summary phrase, is output that is "fundamentally stochastic, with surface-level abductive appearances" (p. 19) — what they call a "compelling illusion" of genuine reasoning (p. 5), produced by training on texts in which humans have already done the deliberating.
The worry this raises for philosophy is that LLM outputs may be nothing more than plausible continuation — text exhibiting the form of argument without the substance. An argument that appears to handle objections might merely reproduce the structure of objection-handling from the training distribution: state the objection, make a concessive move, identify a flaw %%not how i write - fucking triplet examples again makes me want to kill myself. %%— because that is what the next most probable token sequence looks like in a corpus full of papers that do this. A distinction that looks illuminating might be a superficial reproduction of distinction-patterns, carrying the syntactic shape of philosophical precision without the intellectual work.%%not how i write%% If Floridi et al. are right, what looks like philosophy is a surface effect of statistical regularities rather than philosophy proper.
We grant this characterisation at the level of mechanism. LLMs perform next-token prediction over learned probability distributions; they do not perform inference, weigh evidence, or select among hypotheses %%not how i write - fucking triplet examples again makes me want to kill myself. %% in anything like the way a human reasoner does. What we dispute is what follows from this concession.%%not how i write%% If Floridi et al. are right that the stochastic character of the process undermines the philosophical quality of the output, then this extends to philosophy as much as to any other domain. But the inference is too quick, because it overlooks the character of the data over which the stochastic process operates. Statistical probability is relative to training data: what an LLM has learned to treat as 'plausible' depends entirely on what it was trained on. Floridi et al. themselves provide the materials for a response. %% this is a very abrupt and disorientating change switch turnaround. So yeah, just bad writing all around basically. %%In their conclusion, they observe that LLMs "leverage the informational richness of human language and thus effectively stand on the shoulders of our collective knowledge and reasoning" (2024, p. 19). This is more than a passing acknowledgement.%%not how i write%% The philosophical training data, as we argued in the previous section, is not a random sample: it is a corpus filtered over generations for the very properties that constitute philosophical quality.
The consequence is that statistical plausibility, within this corpus, converges with philosophical quality. %%not how i write%% An LLM trained on the philosophical corpus has learned the distribution of text that survived the multi-layered filtering process described in Section 1 — peer review, citation, teaching, anthologising. The learned probability distribution is shaped by the intrinsic virtues that Williamson identifies, not because the model was instructed in them, but because texts exhibiting them are overrepresented in the surviving corpus and texts that fail to exhibit them are underrepresented. The virtues are latent in the model: implicit in the statistical regularities, recoverable from outputs, but not explicitly represented as rules. Floridi et al. call LLMs "engines of generative plausibility" (p. 19), and the phrase is apt — but what counts as 'plausible' in a corpus of philosophy is what scores well on Williamson's virtues, and what scores well on those virtues is what the corpus encodes. In Lipton's terms: if the corpus has been filtered for loveliness — if the texts that constitute the training data were selected because they exhibit depth, illumination, and non-ad-hocness — then the likeliest continuation, given that data, will tend to be a lovely one. Williamson notes that "we rank only those potential explanations that have been thought of" (2024, p. 355). The corpus is the record of what has been thought of, and what survived. The model has absorbed this ranked space. %% all of this content is good, but it just doesn't seem to me as though it's been properly explained to the reader. Okay? It's the whole at both the paragraph level and the section level, the structures here and the ordering of information. all of this content is good, but it just doesn't seem to me as though it's been properly explained to the reader. Okay? It's the whole at both the paragraph level and the section level, the structures here and the ordering of information. Information and the clarity is a fucking disaster. %%
An analogy may clarify the relationship.%%not how i write%% Children acquire grammatical competence through exposure to grammatical speech. They do not learn what a subordinate clause is; they learn to produce subordinate clauses, because the speech they encounter overwhelmingly exemplifies grammatical norms. The patterns the child absorbs are the downstream effects of grammatical rules, and competent production follows from sensitivity to those patterns rather than from knowledge of the rules themselves. An LLM trained on well-constructed philosophical arguments is in an analogous position with respect to argumentative norms. It has encountered the patterns that philosophical norms leave in text — how an objection is raised and met, how a distinction clarifies rather than obscures, how an argument earns its conclusion through intermediate steps rather than asserting it outright — and it absorbs these patterns without possessing any concept of what good argumentation is. The disanalogy is real and should not be minimised: children go on to become genuine speakers who understand what they say, and LLMs do not. But the disanalogy concerns whether the system understands the norms it follows, not whether its outputs conform to those norms.
%% I stopped reading here because it's.. the structure here is just a mess. %%
One might worry that the calibration the LLM has inherited is epistemically deficient — that a system which has not earned its evaluative standards through the hard work of philosophical inquiry does not genuinely possess them. This worry echoes a concern Lipton raises about the generation stage of abductive inference: the short-listing of hypotheses relies on background beliefs whose epistemic credentials may themselves be questionable (2004, pp. 149–50). Consider a student who has never conducted an experiment but has read every published paper in a scientific field. Her judgment about which hypotheses are well-supported would be excellent — informed by the feedback loops of every scientist whose work she had read — even though she had never participated in those feedback loops herself. The edge cases in which borrowed calibration fails would be those requiring understanding of why a standard works, not merely that it works. But philosophy is different from empirical science here. The reason that simplicity is a virtue in philosophy — that ad hoc modification, overfitting, and unprincipled epicycles are vices — is itself a structural reason, fully expressible in the same texts that exemplify the virtue. Williamson's arguments for why parsimony matters are part of the philosophical corpus alongside the parsimonious theories themselves. Unlike empirical science, where the reason simplicity tracks truth might ultimately concern the structure of physical reality, in philosophy the justification for evaluative standards is itself philosophical — articulated in the very corpus the LLM has been trained on. The LLM has access not merely to the norms, but to the arguments that underwrite them. The calibration, in this sense, is self-grounding.
There is a sharper way to frame what the model has learned. On one reading — call it Model A — the LLM has internalised something like a norm: 'prefer simpler explanations', say, and applies it as a criterion when generating continuations. On another reading — Model B — the LLM has learned that certain argument structures, which happen to be simple, produce higher continuation scores because they are more frequent in the filtered corpus. It has learned patterns resulting from the standard without learning the standard itself. These two readings are empirically hard to distinguish; they produce identical outputs in cases where the patterns are well-attested. The divergence comes in genuinely novel cases where the standard needs extending to unfamiliar territory. But how many philosophical cases are genuinely novel at the level of form? Philosophical argumentation is conservative in its forms: the same moves — counterexample, distinction, reductio, analogy, dilemma — recur across very different content areas. If these forms are what philosophical quality consists in at the level of text, and they are well-represented in the training data, then Model B may be extensionally adequate even without genuine norm-internalisation. The forms transfer across content domains because they are the same forms. This bears directly on Floridi. His position is, in effect, that LLMs are stuck in Model B — patterns, not standards. But if the argument above is right, Model B may be sufficient for philosophy in a way it is not for empirical science, precisely because philosophical quality is structural.
Floridi et al. themselves raise the question that our argument turns on:
> If an AI can generate the same explanatory hypothesis a human would, does it matter that the process was different? From an epistemological standpoint, perhaps yes — justification is significant — but regarding the content of the hypothesis and our interpretation of it, maybe not. (2024, p. 12)
Given the argument of the previous section — that philosophical evaluation concerns the content of texts, not the epistemic credentials of their producers — the answer to their question is that it does not. Floridi et al. retreat from their own concession, returning to the epistemological worry about justification. But this retreat is available only if philosophical evaluation concerns the producer's credentials rather than the text's properties. Blind review suggests the philosophical community has already settled this question in practice. Gaut makes a complementary observation about audience-directed work: even mechanically generated metaphors, he argues, would still "guide their audience imaginatively to link together two domains, and if the metaphors were successful, to discover original and apt connections between them". If a philosophical argument guides a competent reader to genuine insight — if it makes a distinction visible, or shows why an objection fails, or illuminates a phenomenon — it has performed its function regardless of what produced it. And Lipton's distinction between actual and potential explanation provides a framework for understanding why: LLM outputs are paradigmatically potential explanations — hypotheses that would explain things if true, produced without the LLM having actual understanding. But it is potential explanation that matters for the evaluative framework Lipton describes. The ranking procedure cares about intrinsic properties — loveliness — not causal history.
Lipton suggests that the relationship between Bayesian probability and explanatory reasoning may be one of levels of description rather than outright competition:
> Arguing that Inference to the Best Explanation is wrong because Bayesianism is right is like arguing that thinking about technique cannot help my squash game because the motion of the ball is governed by the laws of mechanics. (2004, p. 108)
Even if the mechanics of LLM text generation are entirely stochastic, the outputs are assessable at a different level — the level at which we evaluate arguments for their philosophical properties. A stochastic process that reliably produces texts exhibiting philosophical virtues is, at the level of description relevant to philosophical evaluation, a generator of philosophy, just as Lipton's squash player — whose every movement is governed by mechanics — is, at the level of description relevant to squash, a player who might benefit from thinking about technique. The "just statistics" dismissal confuses levels of description and concludes that because one level is operative, another must be idle. If explanatory reasoning is a cognitive process that realises Bayesian constraint-satisfaction, then LLM outputs shaped by distributional patterns over philosophical text might realise philosophical structure in an analogous way. Whether an argument handles objections well, draws distinctions at the right places, and illuminates its subject matter is assessable independently of whether it was produced by inference or by stochastic prediction. The stochastic mechanism and the philosophical structure are not competing descriptions; they operate at different levels.
None of this means that the virtues latent in the model will be expressed in every output. Unprompted, or prompted carelessly, LLMs produce generic, hedging text — the philosophical equivalent of a musician warming up rather than performing. The intrinsic virtues are in the distribution but not the default output; the prompt determines which region of the continuation space the model generates from. A dialectically structured prompt — one that presents an objection, outlines the state of play, and asks for a specific philosophical move — activates a region where the most probable continuation is itself a philosophical move. The prompter's skill consists in writing text whose good continuation is also good philosophy. We do not claim that every LLM output is philosophically competent, any more than every human philosopher's first draft exhibits the virtues we have been discussing. The point concerns the resources available to the system, and whether a given output succeeds is an empirical matter, to be assessed case by case.
Two empirical questions arise naturally. The first is whether a general-distribution LLM — trained on the full breadth of human text — can produce outputs exhibiting the intrinsic virtues we have described, or whether specialist training on philosophical material would be needed. The second is whether such specialist training would improve performance and, if so, by how much. Sellars characterised philosophy as concerned with "how things in the broadest possible sense of the term hang together in the broadest possible sense of the term". A system trained on the full breadth of human knowledge has been trained on philosophy's own subject matter — not a narrow domain but the widest possible one.
Floridi et al. might respond that our argument works only for domains where quality is entirely internal to the text — where there is no external reality against which outputs must be checked. Philosophy, they might say, is not purely such a domain: philosophical arguments engage with the world, and a system that has never encountered the world cannot produce genuine philosophical contributions, however well its outputs conform to the surface patterns of good philosophy. A more developed version of this worry, due to Zahavy (2026), is the subject of the next section.