# §2 — current paragraphs (in reading order)
## Loveliness as the standard philosophy answers to
Abductive arguments in philosophy often compete in terms of loveliness. Williamson, as have seen, takes philosophy to proceed by inference to the best explanation, notes that the inference does not rank theories by probability, and that philosophical explanation is constitutive rather than causal, as Newton's laws explain Kepler's by subsuming them; what is assessed is how much a theory would illuminate were it correct. The banal diagnosis of the wet floor, that water has fallen on it, is as likely as can be and for that reason illuminates nothing; its philosophical counterpart, a claim at once true and unilluminating, is no philosophy at all.
## Loveliness is comparative, and read off the page
A reader can judge a text's loveliness from the text alone. What the text sets out is a candidate, the rivals that bear on the same data, and the difference that decides between them; the reader asks whether that difference would, were the candidate correct, afford the understanding the rivals lack. The question is asked under "if correct", so it turns on what the text claims and not on whether the claim is true, and nothing outside the text — the manner of its writing included — bears on the answer.
## Structure without the capacity (production side)
LLM pre-training, can be thought of as training the model to pick out patterns in language of language, rather than being given a set of pre-determined rules as to how language works (Wolfram 2023). Trained on English it comes to respect a syntax never stated to it, the writing it continues being overwhelmingly well-formed; its sentences come out meaningful and not merely grammatical, though no complete theory of sentence-meaning has ever been built for it to be given; and its syllogisms come out valid, the "correct inferences" discovered from a literature they pervade rather than derived. In each case a structure stands in the output while the capacity that ordinarily produces it is absent.
That loveliness answers to no stated rule is therefore no barrier on the production side, since what these systems acquire they acquire from exemplars, where Lipton locates the standards of loveliness. The precedent reaches only so far: a syllogism has one correct completion where an abductive comparison has none, so what transfers is the weaker point alone, that a weighing can stand in the output with nothing of the producing capacity behind it.
## The structure is in the corpus
Laying out the candidates and citing what decides between them is a regularity of explanatory writing wherever it occurs. The wet floor and the car that would not start were each a weighing of just this kind, and neither was philosophy; the form turns up in any prose that settles one explanation against another, from a fault diagnosed to a verdict reached. A system fitted to continue such writing comes to respect that organisation as it comes to respect syntax: it is taken from the writing rather than supplied as a rule, and explanatory prose carries it as steadily as well-formed prose carries grammar.
## Philosophy as a developed instance
A philosophy paper is one of these weighings carried to its full extent: a position is stated, the rivals that bear on it are set out, and each is pressed through the objection taken to tell against it. A comparison of that length runs past anything Wolfram exhibits, whose regularities reach no further than the parse tree of a sentence or the syllogism spanning two or three sentences (2023); raw continuation, he grants, tends to wander over longer stretches of text.
The weighing is a chain of local transitions rather than one dependency held across the paper, and each transition is a regularity of what follows what in explanatory writing: a stated position is followed by the rival that bears on it, a rival by the consideration that decides against it. These sit at the scale the model works at, the next move given the text so far, and a system fitted to continue the writing is fitted to them as it is to the order of words within a sentence.
## The statistics objection and its reply
It may be said that this only redescribes the statistics, since a model reproduces the regularities of its training text, and reproducing regularities is not weighing. Lipton answered an objection of the same shape when the Bayesian held that, once belief revision had its mechanics, explanatory considerations were left with nothing to do: a true account of the mechanism, he replied, need not displace a true account of what it produces. A squash ball's flight obeys the laws of mechanics, yet thinking about technique still improves one's game, and even were the Bayesian right about the mechanics, inference to the best explanation "might yet illuminate its psychology" (2004, p. 108). The mechanism here is the one Floridi et al. describe, and on their own description the objection does not go through: the patterns a model absorbs are patterns of reasoning as expressed in writing, and writing does not carry the phrasing of an explanation stripped of its organisation, so which considerations bear on which rival, and what settles the matter between them, lie in the corpus with the rest.
## Where the failure-line falls
It may still be objected that syntax is one thing and inference to the best explanation another, and that whatever structure next-word prediction carries, a system of this kind is too shallow for abduction, with the benchmark record reading like confirmation. The line Wolfram draws falls elsewhere, and it comes from the discussion that supplied the syllogism: his simple network cannot balance long sequences of parentheses, a task that demands exact procedure with no shortcut, while it manages whatever a person could take in at a glance. Wolfram extends the point to logic himself, allowing that the model will produce the "correct inferences" of syllogistic form but expecting it to fail at more sophisticated formal logic "for the same kind of reasons it fails in parenthesis matching" (2023). What defeats the net is computational irreducibility, the tracing of each step that no shortcut captures; exact procedure of that kind is the side of the divide it fails on, and the contrast with the merely simple is beside the point, since by Wolfram's own diagnosis it is being "too computationally shallow" that costs it the parentheses, and computational shallowness is what lets human-like tasks such as essay-writing succeed. Weighing does not sit on the failing side, for no rule runs from the evidence to the loveliest explanation, and telling a good selection from a bad one is the holistic judgement a reader brings to any argument.
## The benchmark record, read against that divide
Salimi et al.'s suite separates two kinds of task: the formally constrained recovery of a single canonical missing premise, and open-text explanation judged as explanation (2026). A system that could recognise explanations but had nothing to draw on in producing one should fail wherever production is demanded, and the record runs the other way. The collapse concentrates on the first kind alone — the strongest model reaches 21.5% on the hardest such benchmark and most score near zero — while on the second the strongest models exceed 90% validity. Failure tracks the recovery of an exact missing premise, the parenthesis side of the divide, and philosophical abduction does not live there.