# 4. The Challenge from Observation — gate-clean revision Ordinary uses of LLMs do not usually produce philosophy worth reading. If someone types 'What is the meaning of life?' or 'What is the solution to the hard problem of consciousness?', the result is normally a survey, a compressed introduction, a set of familiar options, or a polished non-answer. That is exactly what one should expect from the use being made of the system. However, this does not necessarily show that LLMs cannot produce philosophy worth reading. What it does show, for certain, is that treating the system as though it were an oracle, able to dispense deep truths in an instant, is not the right way to get worthwhile philosophical output (Janus REF?). If we recall how LLMs are trained and function, we can see why this would be the case. Pre-training involves the system learning how to continue strings of tokens in a likely pattern; post-training involves the system learning how to respond as a "helpful assistant" to user queries. Neither of these will lead LLMs to tend to produce worthwhile philosophical arguments in response to the simple queries in the previous paragraph. A better approach would be to ask what sort of input would *elicit* substantial philosophical output; however, before turning to this question, we should head off an objection that could be made at this point: the more work that the prompter has to do, the further away they get from typing in a simple question, the more it seems that it is not the LLM that is doing the work, but the prompter. On this approach the LLM might be considered a tool that a philosopher uses, but no more a *producer* of philosophy than a typewriter: believing that the LLM is producing the philosophy is a mistake akin to believing that it is the ventriloquist's dummy that is doing the talking. Recall that a model is first fitted to a vast and general body of writing and trained to continue it, the picture of these systems we relied on in Section 2; what it returns to a simple philosophical question is therefore whatever most plausibly continues such a question across writing at large, and across writing at large a question like "what is the meaning of life?" is continued by an overview of the familiar positions rather than by a developed argument for any one of them. The system is then shaped further to respond as a helpful assistant to whoever asks, and this presses in the same direction, since someone whose task is to be helpful to all comers, handed a large question and nothing else, gives an accessible overview rather than working out a case of their own. A simple question thus draws the survey from both stages of the system's making at once, so that the thin answer is the continuation we should expect, and not a measure of what the system could produce were it given something already shaped as a philosophical problem. Rather than thinking of an LLM as a person with all the answers, we need to think of it as a system which has the potential to have worthwhile philosophy elicited from it, by the right sorts of input. Before considering what this input might be, however, we should first head off following objection. [descrption of the objection, using the just a tool/typewriter example, and then the ventriqlquists dummy example. do you know what i am talking about? tell me if you do not] As regards this observation, it is indeed true that the answer to a bare question is rarely worth reading, but this will only affect what the system produces when it is asked to continue a bare question rather than what it could produce more generally. One might object here that the answer to a bare question is the most the system can do, treating the question, the answer, and the mark given to it as all there is to weigh; yet this is no reason to think that the same way of weighing tells us anything about other uses. To put a question, take the answer, and weigh it is to treat the system as an oracle, which fits a thing that is asked to give back a fact or to answer a problem with one determinate answer, but does not fit philosophical writing. The most obvious thought is that a philosophical paper is the answer to a question taken on its own; yet it is implausible to think that a paper is anything other than the continuation of a position already under pressure, set among rivals, in a dialectical situation. A bare question, by contrast, has too little of the shape of an argument, in that it gives no position to test and no rival to set against it. Since good abduction, as we argued in Section 2, is the weighing of rivals, a prompt that sets up no rivals is not yet asking for the kind of thing that Section 2 defended. This observation takes one point in the space of possible continuations, but it does not tell us what happens when the system is given something already shaped as a philosophical problem. Although it is standard to read the output of these systems as something the model makes out of whatever comes before it, I want to suggest that the depth of what it produces is not fixed apart from its input. This can be drawn out by comparison with the account given in Section 3, on which the system does not produce philosophical content out of nothing but continues whatever text it is given, so that what it produces depends on the prior text that fixes what it is set to continue. The most obvious thing to say is that one should learn to write better prompts. Yet it is implausible to think that the lesson is a matter of skill at prompting, since the claim is not that better words draw better output but that the philosophical thing the system produces is made by the prior text that sets the problem of continuation. Compare what the system is doing when it is asked, with nothing before it, "what is the meaning of life?", with what it is doing when it is given a prompt that sets out a position on the meaning of life, sets two rivals against that position, and gives the objection the position must answer, then asks for the strongest abductive case for the position by showing what it explains that the rivals do not. These are not two forms of one prompt; the first sets the problem of continuing a standard question towards an answer, while the second sets the problem of continuing within a dialectical structure already laid down, so that the text to be produced is the development of a position rather than a reply to a question. Even if we accept the reasons for thinking the system needs a richer philosophical context, we might still ask why the philosophy that comes out should be credited to anything other than the person who supplies that context, since on this way of looking at things the system only carries out or fills out a thought that was the philosopher's all along. This is a stronger worry than the one we faced in Section 3, which asked where the good outputs were; this one grants that good outputs exist and denies that they are the system's. Put in its strongest form, the worry runs as follows: if the user supplies the position, the rivals, and the objections, the system only fills in the words; if the user keeps the strong continuations, drops the weak, and presses the system through draft after draft, then it is the person who is doing the philosophical work; and if an output is worth reading only after such heavy direction, it is closer to edited ghostwriting than to philosophy of the system's own. The more the prompting works, the more it seems that the credit belongs to the person who did it. It is this strong form, and not a weaker one, that we answer. A prompt that fixes the starting point of an inquiry is one thing; a prompt that fixes everything which follows from that starting point is another; and it is the first, not the second, that the cases in Section 3 show. While there is considerable disagreement amongst philosophers as to whether a beginning can ever be set down clearly, it has struck many as a default, common-sense way of describing how a piece of philosophy gets going that it begins from some starting point laid down by a person which does not yet contain every consequence later drawn from it. Consider Jackson's case of Mary, the colour scientist confined to a black-and-white room: the case is a short setup, written down by Jackson, which hands later philosophers a structure to work through rather than a set of conclusions to read off. While it is straightforward to say that Jackson supplied the words of the thought experiment, it is much less clear how to make sense of the idea that he thereby supplied everything those words went on to commit one to, since the writers who took the case up, Lewis among them, did not merely paraphrase the setup but drew consequences from it, resisted inferences others wanted to draw, and redescribed what the case commits one to. We propose, then, that the analogy is narrow but exact: just as Jackson's text gives a reader something to continue without already containing the continuation, so a prompt can give a system something to continue without already containing the continuation. This is nothing special to generative systems, then; the difference we have been tracking is the difference between a starting point and its development, and that difference is enough to make sense of the case without appeal to anything further from Section 3. We propose, perhaps surprisingly, that the system's contribution is best found in the continuation rather than in the prompt, since the prompt supplies the materials, and the output may, in continuing them, draw out a pressure, a distinction, an implication, or a comparison that the prompt did not make; and it is here that a contribution can occur. While there is considerable disagreement amongst those who write about these systems as to whether such outputs amount to anything, it has struck many as a default, common-sense way of describing the case to say that the model is not a philosopher in the human sense; we make no such claim, holding only that an output can carry philosophical work that is not already fixed by the prompt. The prompt is not without its own reach, for it can fix which problem is taken up, which view is developed, which rivals are live, and what the answer must meet, while the continuation can still supply how the pressure is handled, which difference does the work, and which consequence follows. Setting apart what is fixed from what is continued gives us our central test: to ask what the output says that the prompt did not, a test that tells development from paraphrase. There is an obvious objection to thinking of the model as a typewriter. A typewriter does not continue a context; it sets down words already chosen by the user. Should we not say that the model does the same? The answer is that a typewriter only sets down what the user has already fixed, whereas the model continues a context, so that the very same prompt can give different continuations. The model fixes only what the user has fixed, the objection says; but the user may fix where a dialectical route begins without fixing how it develops. Where a typewriter would set down only what was already chosen, the model would carry the same starting point in different directions, some better than others, and some that go wrong. One might object here that the prompt fixes the output after all, so that the model can only give it back or fail to. Yet that a continuation can go wrong shows that the system is not merely setting down what was fixed: if the prompt fixed the output, the system could only give it back or fail to, rather than go wrong in this way. What divides a typewriter from a model, then, is not whose philosophy the continuation is, but what the prompt fixes and what the continuation adds; and where a typewriter makes no philosophical mistake, a model can. One might object here that a minimal prompt is not the only kind of prompt available. A rich one might fix the continuation, for if the user supplies the view, the dialectical setting, the objections, the wanted conclusion, and the line of reply, the system may only fill out what the user gave. While we resist the over-simple answer that a prompt is never the user's own work, we grant that this is a good objection, since sometimes the prompt contains the philosophy and the output is a paraphrase. Prompts do differ in how much they fix, and this matters once we sort them along a spectrum: a bare prompt with too little structure, likely to draw an overview; an articulated prompt with enough structure to set a development going; an over-filled prompt where much of the work is already the user's; and the case at the far end, where the prompt sets out the comparison and the verdict and the output only says it back. To see why richness alone does not settle the matter, consider the following test. We set the prompt and the output against each other; if the output says nothing relevant that the prompt did not, it is a paraphrase, but if it draws out a consequence, a pressure, or a contrast the prompt did not make, it is a development. Even if all this is granted, we might still ask whether such a text can do more than handle well the positions a literature already contains and make a distinction that literature lacks, and so be creative in the stronger, public sense in which a human philosophical text is creative. We want to be careful here, because this is quite different from the modest sense in which an output is novel only relative to its prompt; what is at issue is the stronger claim of saying something the literature had not yet said. We do not need to settle this question here. It would be settled as the rest has been, by setting the output not only against the prompt but against the literature, and asking what it says that the literature had not. This is not meant to settle the matter either way, but it does give us a reason to leave the question open, for further work. We can now return to the challenge from Section 1, where the ordinary blandness of what these systems produce when given a bare question was taken to show that they have nothing to contribute to philosophy. The outputs to such questions do tend towards the empty and the thin; but that bears only on whether bare questions are good tests of philosophical capacity, since what these systems do is continue the context they are given, and a context with no shape of argument can only draw from them an output with no shape of argument. This is no reason to think that a context which hands the system a position, its rivals, and the pressures bearing on each of them cannot draw a development of its own; and whether such a development is worth reading is settled not by looking at the system but by reading the continuation first against the prompt that occasioned it and then against the literature it means to add to. To sum up, these systems are not oracles; they are continuation systems, and philosophy worth reading needs a dialectical context which a prompt can supply without thereby fixing the development, so that the philosophical standing of any output depends on what the continuation itself adds. # Section 2, second half — original / diagnosis / corrected Each corrected paragraph uses only vocabulary and sentence-structures found in the published corpus (the 9 #published-paper notes). "0" means the word or frame returns no match in your papers. The spine words "weigh/weighing" (0) and "display/displayed" (0) are carried by attested substitutes — "decide between / prefer / the choice", and "present / set out / shape". ## A-i > A model generates no candidates and weighs none. That much of Floridi et al.'s account can be granted, since nothing in what follows depends on the model having either capacity. None of it bears on the text the model produces. Section 1 located a text's merit in the argument it presents, not in the history of its production. By that standard the car-battery reply is not a list of candidates left for someone else to sift. It weighs the cold morning against each one and comes down in favour of the battery, so the assessment the brainstorming picture reserves for the collaborator has already been made on the page. Not in your corpus: "generates" (0), "weighs" (0), the passive "can be granted" (0 — you write the active "we grant that…", 6 hits), "bears on" (0), "by that standard" (0), "sift" (0), "comes down in favour of" (0), "on the page" (0), "reserves" (0). Yours, and kept: "candidates" (8), "account" (66), "located/merit" (20/3), "prefer" (7), "the case" (12). The car-battery is also over-compressed — the symptoms that make it a real choice are dropped. > The model produces no candidates of its own, and chooses between none of them; that much of Floridi et al.'s account we grant, since nothing we go on to argue depends on the model doing either. None of it, though, settles anything about the text the model produces. Section 1 located where a philosophical work's merit lies — in the argument it presents, not in the history of how it was produced — and by that measure the reply about the cold morning is not a list of candidates left for a collaborator to choose between. It sets the cold morning against each candidate and chooses the weak battery, so the choosing that the brainstorming picture leaves to the collaborator the reply has itself carried out. ## A-ii > Asked whether anything turns on the process being different when the hypothesis produced is the same, Floridi et al. half-concede: they allow that for justification it perhaps does, but "regarding the content of the hypothesis and our interpretation of it, maybe not" (2025, p. 12). Whether the assessment a text displays is any good is a question about a piece of writing. Not in your corpus: "turns on" (0), "half-concede" (0), "assessment/assess" (0), "displays" (0). Kept: "follows" (22), "account" (66), "question" (5), and the quotation verbatim. > Asked whether anything follows from the process being different when the hypothesis produced is the same, Floridi et al. concede as much: they allow that for justification it perhaps does, but "regarding the content of the hypothesis and our interpretation of it, maybe not" (2025, p. 12). Whether the choice a text presents is any good is a question about a piece of writing. ## B > A text can display a good weighing without its producer having weighed anything. A displayed weighing is good when a reader can assess it: when the text sets rival explanations against one another and comes down for one, the reader can ask whether it has come down well. That is the standard Lipton draws for inference to the best explanation. A system that continues a body of text, as Wolfram describes, can produce a weighing that meets it where nothing was weighed at all. Not in your corpus (the densest cluster in the section): "display/displayed" (0, twice), "weighing/weighed" (0, four times), "assess" (0), "comes down" (0, twice), "rivals" (0). It also pre-labels Lipton and Wolfram before either argues, which you never do. Kept: "prefer" (7), "present" (177), "the case" (12), "decide" (2). > So the model's having decided nothing of its own settles nothing by itself. What it produces still sets one explanation against the others and prefers it, and we can ask of that preference, as we would of any in philosophy, whether it is the right one. Two things are left to show: what makes such a preference a good one, and how a text that nobody decided can present one. ## C-i > Lipton asks what makes one selection among the candidates better than another, and separates two things the best explanation might be. It might be the likeliest, the explanation the total evidence most warrants, or the loveliest, the one that, were it correct, would yield the most understanding. "Likeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). Newtonian mechanics shows how far the two can fall apart: it is no longer the likeliest account of the motions it was built to explain, yet it remains as lovely an explanation of them as it ever was (p. 60). Not in your corpus: "selection" (0 — use "choice", from "choose" 3), the appositive frame "the one that" (0). "warrants" (0) is Lipton's own term, so kept and flagged. Yours, and kept: "separates" (12), "come apart" (1), "account" (66), "explanation" (16), "understanding" (7). > Lipton asks what makes one choice among the candidates better than another, and separates two things the best explanation might be. It might be the likeliest — the explanation the total evidence most warrants — or the loveliest, which, were it correct, would give the most understanding. "Likeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). The two can come apart, as Newtonian mechanics shows: it is no longer the likeliest account of the motions it was built to explain, yet it remains as lovely an explanation of them as it ever was (p. 60). ## C-ii > A philosophical text answers to loveliness. What a reader assesses is whether the explanation on offer would, if correct, give more understanding than its rivals, and that assessment is made under "if correct". So it does not wait on the explanation's truth, and the reader can carry it out on the page. Not in your corpus: "answers to" (0), "assesses/assessment" (0), "on offer" (0), "wait on" (0), "rivals" (0), "on the page" (0), and the cleft frame "What … is whether" (0). One idea is also stated three times. Kept: "suppose" (3), "give … understanding", and the quoted "if correct". > Loveliness is what a philosophical text is read for. We can suppose the explanation it offers correct, and ask whether, so taken, it would give more understanding than the explanations set against it. The question is put under that "if correct", so whether the explanation is in fact true is left open, and a reader can take it up with nothing before them but the text. ## D-i > Loveliness shows in the comparison of rivals. To explain is to explain why this rather than that, and that requires citing a difference between them. This is Lipton's Difference Condition: a cause present in the favoured case, together with the absence, in the rival, of any corresponding cause (2004, ch. 3). "Rain rather than a burst pipe, because the window is open and the water lies beneath it" meets it. The open window and the pooling beneath it are a cause present in the rain case, and the wetting-by-the-pipe that a burst pipe would have left is absent. "Rain rather than a burst pipe, because the floor is very wet" does not, since a very wet floor is what both rivals would produce and picks out nothing between them. Only the first cites a difference between the rivals, and so only the first earns its loveliness. Not in your corpus: "rivals" (0, three times), "favoured" (0), "citing" (0 — the apparent hits are "exciting"), "earns" (0 — the apparent hits are "learn"). Kept: "picks out" (1), "corresponding" (1), "present" (177); the Difference Condition and both kitchen sentences stay verbatim. > Loveliness is shown in setting one explanation against another. To explain is to explain why this rather than that, and that needs a difference between the two — Lipton's Difference Condition: a cause present in the one case, together with the absence, in the other, of any corresponding cause (2004, ch. 3). "Rain rather than a burst pipe, because the window is open and the water lies beneath it" meets it. The open window and the pooling beneath it are a cause present in the rain case, and the wetting at the pipe that a burst pipe would have left is absent. "Rain rather than a burst pipe, because the floor is very wet" does not, since a very wet floor is what either would produce, and picks out nothing between them. Both sentences have the same comparative form, and only the first gives a difference of the right kind — only the first, that is, makes a genuine choice between them. ## D-ii > No rule sorts the two sentences for the reader. Our grasp of what makes one explanation lovelier than another is weak, and such standards as there are come from past explanations serving as exemplars, and from prevailing styles of reasoning (2004, pp. 61, 139). Telling the two apart is the ordinary work of reading an argument: following what each says and asking whether it would decide the case. The bare form of explanation never did that work for human philosophers either. What sorts the good weighing from the bad is the reading, and the same bar applies whether a human or a machine wrote the paragraph. Not in your corpus: the verb "sorts" (0 — the 7 hits are "sorts of"), "such … as there are" (0), "exemplars" (0 — your word is "examples", 28), "the ordinary work" (0), and the closing cleft "What sorts … is the reading" (0). Kept: "examples" (28), "field" (17), "consider" (49), "decide" (2). > No rule does this for the reader. Our grasp of what makes one explanation lovelier than another is weak, and what standards we have come from past explanations that serve as examples, and from the styles of reasoning a field has grown into (2004, pp. 61, 139). To tell the two apart is to do what reading any argument asks: to follow what each says, and to consider whether it would decide the case. The bare form of an explanation never did that for a human philosopher either, and it does no more for a machine. Only the reading tells a good choice of explanation from a bad one, and it asks the same of both. ## E1 > A system trained only to continue text comes to respect constraints that were never stated, and Wolfram (2023) collects the cases. Trained on English, it respects English syntax though it was given no grammar, because well-formed sentences predominate in the writing it continues. Its sentences are mostly meaningful, not merely well-formed, and here no rule was available even to withhold, since no one has built a complete theory of what makes a sentence meaningful. Its syllogisms come out valid for the same reason. The patterns pervade the writing; Aristotle, Wolfram suggests, read them off many examples of rhetoric, and a system continuing that writing yields "correct inferences" of the syllogistic kind, with nothing derived. In each case a structure is present in the output while the capacity that ordinarily produces it is absent. Not in your corpus: the verb "respect" (0 — your only "respect-" hits are "respectively"), "comes to [verb]" (0), "collects" (0), "predominate" (0), "valid" (0), "the output" (0), "not merely" (0), "in each case" (0). Kept: "constraints" (3), "withhold" (1), "pervade" (1), "derived" (3), "ordinarily" (1), "capacity" (1), "yields" (1), "for the same reason" (1); "follow" (22) replaces "respect". > A system trained only to continue text follows constraints no one ever gave it; Wolfram (2023) gives the examples. Trained on English, it follows English syntax although it was given no grammar, because the writing it continues is full of well-formed sentences and little else. Its sentences come out meaningful — and a sentence can be well-formed without being that — though here there was no rule even to withhold, since no one has ever given a full account of what makes a sentence meaningful. Its syllogisms come out right for the same reason: the patterns pervade the writing, and Aristotle, Wolfram suggests, took them from many examples of rhetoric, so a system that continues such writing yields "correct inferences" of the syllogistic kind, with nothing derived. Each of these is a case where the structure is there in what the system produces, while the capacity that ordinarily produces it is not. ## E2 > That loveliness answers to no stated rule is no barrier to its appearing in the output. A system that wrote by stated rules would stop wherever no rule had been stated. These systems were given no stated rules at all, and what they acquire they acquire from exemplars — which, on Lipton's account, is just where the standards of loveliness reside. The precedent reaches only so far. A syllogism has one correct completion where an abductive comparison has none, so what survives is the weaker claim, which is all that is needed: that a structure can stand in a text with no trace of the capacity that ordinarily produces it. Not in your corpus: "answers to" (0), "no barrier" (0), "acquire" (0), "reside" (0 — the hits are "residents"), "the precedent" (0), "completion" (0), "what survives" (0). Kept: "it might seem" (1), "examples" (28), "field" (17), "stand" (28), "capacity" (1); "is to be found" replaces "reside". > It might seem that loveliness, having no rule of its own, is the one thing such a system could not reach. But these systems were given a rule for nothing, syntax included, and what they take, they take from the examples they were trained on — which, on Lipton's account, is just where loveliness is to be found, since no one gives a rule for what makes an explanation a fine one, and a field keeps its standard in the explanations it has come to count as good. The likeness to syntax holds only so far: a syllogism has a single right ending, where the choice between explanations has none. But what we need holds even so — a structure can stand in a text with no sign, behind it, of the capacity that would ordinarily produce it. ## F1 > Most of what these systems are trained on is not philosophy, but the corpus contains the philosophical literature too. A philosophy paper is itself a displayed comparison: a position is stated, set against its rivals, and the difference that decides between them is drawn out. That comparison is as much a regularity of the writing as syntax is. Wolfram's cases stop at the sentence, and the move to the paragraph is ours; but what he describes are regularities in writing, not facts about grammar in particular. Not in your corpus: "displayed" (0), "rivals" (0), "drawn out" (0). Kept: "shape" (54 — strongly yours), "regularities" (2), "follow/shown"; "shape" replaces "displayed comparison". > Most of what these systems are trained on is not philosophy, but the writing they are trained on does take in the philosophical literature. And a philosophy paper has a shape of its own: a position is put, the positions against it are set out, and what decides between them is shown. That shape is a regularity of the writing as much as grammar is. Wolfram stops at the sentence, and the step to the whole paper is ours; but what he points to is a regularity in writing, not a fact about grammar in particular. ## F2-a > It may be said that this only redescribes the statistics, since a model reproduces the regularities of its training text, and reproducing regularities is not weighing. To the Bayesian who holds that, once belief revision has its mechanics, explanatory considerations have nothing left to do, Lipton replies that a true account of the mechanism need not displace a true account of what it produces. A squash ball's flight obeys the laws of mechanics, yet "thinking about technique cannot help my squash game" does not follow. Even granting the Bayesian mechanics, inference to the best explanation "might yet illuminate its psychology" (2004, p. 108). Not in your corpus: "it may be said" (0 — your objection-opener is "One might object that", 1 hit, and "It might be objected here that", 2 hits), "redescribes" (0), "reproduces" (0), "displace" (0), "does not follow" (0). Squash quotes and (2004, p. 108) stay verbatim. > One might object that this only puts the statistics in other words: a model produces the regularities already in its training text, and to produce regularities is not to decide anything. The Bayesian once pressed the same objection on Lipton, that once belief revision has its mechanics there is nothing left for explanatory considerations to do. A true account of the mechanism, Lipton answers, need not stand in the way of a true account of what it produces — to argue otherwise is like holding that "thinking about technique cannot help my squash game" because a squash ball's flight obeys the laws of mechanics. Even granted the Bayesian mechanics, inference to the best explanation "might yet illuminate its psychology" (2004, p. 108). ## F2-b > The mechanism at issue here is the one Floridi et al. themselves describe, and on that description the objection does not go through. The patterns a model absorbs are patterns of reasoning as expressed in writing, and writing does not carry the phrasing of an explanation stripped of its organisation. Which considerations bear on which rival, and what settles the matter between them, are in the writing too, and a system that learns to continue the writing learns these with the rest. The appearance the objection grants was never separable from the organisation that makes a piece of reasoning assessable on the page. Not in your corpus (the most meta-laden paragraph): "the mechanism at issue here" (0), "does not go through" (0), "stripped of" (0), "settles the matter" (0), "separable" (0), "assessable" (0). Kept: "account" (66), "decide" (2), "follow" (22). > The mechanism is the one Floridi et al. set out, and on their own account it does not give the objection what it needs. The patterns a model takes up are patterns of reasoning as it appears in writing, and writing never carries the words of an explanation apart from its working. Which consideration counts against which position, and what decides between them, are in the writing too, and a system that learns to carry the writing on takes these up with the rest. What the objection was ready to call mere appearance was the working itself, set down where a reader can follow it. ## G-i > It may still be objected that syntax is one thing and inference to the best explanation another, and that whatever structure next-word prediction carries, a system of this kind is too shallow for abduction, with the benchmark record reading like confirmation. But the line Wolfram draws falls elsewhere, and it comes from the same discussion that supplied the syllogism. His simple network cannot balance long sequences of parentheses, a task that demands exact procedure with no shortcut, and sophisticated formal logic should fail for the same reason, while it manages whatever a person can take in at a glance. The divide these systems fail at runs between exact procedure and holistic judgement; the contrast between the simple and the sophisticated is beside the point. Weighing, on Lipton's account, sits with judgement, since no rule runs from the evidence to the loveliest explanation. Not in your corpus: "it may still be objected" (0 — use "It might be objected here that", 2 hits), "the line … falls elsewhere" (0), "exact procedure" (0), "at a glance" (0), "the divide" (0), "holistic" (0), "sits with" (0), "no rule runs" (0). Kept: "slight" (8), "is one thing" (1), "prefer" (7), "judgement" (1). > It might be objected here that syntax is one thing and inference to the best explanation another: whatever next-word prediction picks up, a system of this kind is too slight for abduction, and the benchmarks seem to confirm it. But these systems do not fail for being slight. Wolfram's own network cannot keep a long row of parentheses in balance — a task that has to be worked through exactly, with nothing to be guessed — and heavier formal logic fails in the same place, while whatever a person can see straight off, it manages. What beats them is the exact, step-by-step kind of work, not depth; and to prefer one explanation to another is not work of that kind, but a matter of judgement, since no rule leads from the evidence to the explanation that would teach us most. ## G-ii > A system that could recognise explanations but had nothing to draw on in making one should fail wherever production is demanded. The record reverses. The collapse concentrates in one place: where abduction is recast as the exact recovery of a single canonical missing premise under formal constraint. There the strongest model manages 21.5% on the hardest such benchmark, and most score near zero. On open-ended tasks, where the output is judged as an explanation, the strongest models exceed 90% validity (Salimi et al. 2026). Failure tracks exact recovery, the parenthesis side. Philosophical abduction is not of that kind. Not in your corpus: "the collapse" (0), "gathers" (0), "tracks" (0), "of that kind" (0). Kept: "recognise" (2), "produce"; the figures and (Salimi et al. 2026) stay verbatim. > A system that could recognise an explanation but had nothing to draw on in making one should fail wherever it has to produce rather than recognise. The benchmarks fall the other way. What failure there is comes in one place: where abduction has been set as the exact recovery of a single missing premise fixed in advance, under formal constraint. There the strongest model reaches 21.5% on the hardest such test, and most come out near zero; on open tasks, where what the model produces is judged as an explanation, the strongest exceed 90% validity (Salimi et al. 2026). Failure follows the demand for exact recovery — the side of the parentheses — and philosophical abduction does not lie there. ## H > None of this gives the model any capacity Floridi et al. deny it. It infers nothing, and it weighs and tests nothing; what it makes is text, and the text can hold what its maker did not: a candidate stated, its live rivals set in order, and the difference that decides between them. Whether a given text holds these things, and holds them well, is settled by the reading any philosophy paper is given, by the same standard. A good weighing of positions the literature already contains is not yet the distinction the literature lacks. Whether a model can supply that is taken up in Section 4; Section 3 asks what philosophy a system with no relation to the world could produce at all. Not in your corpus: "weighs" (0), "set in order" (0), "settled by" (0), "no hold on" (0). Kept: "deny" (7), "choose/choice" (3), "reading" (4), "produce"; Section 4/3 references stay. > We have given the model back none of what Floridi et al. deny it: it infers nothing, decides nothing, puts nothing to the test. What it makes is a text, and a text can hold what its maker never did — a position put, the positions against it set out, and the difference that decides between them. Whether a given text does this, and does it well, comes out in the reading any philosophy paper is given, by the one standard there is. To choose well among positions a literature already holds is not yet to give it one it lacks; whether a model can do that is the question of Section 4, and what a system cut off from the world could produce is the question of Section 3. --- # 1. The Challenge from Authorship In this section we address what we might call the _challenge from authorship_: the idea that philosophy is something that only persons, or at least minds, can produce. This view has not, to our knowledge, been explicitly defended in just this form, but it gives shape to an intuition that many philosophers may have: philosophy is a person-only domain. An imperfect comparison is with art: One might deny that an image generated by an AI system~~, at least in the familiar prompt-and-output cases,~~ is an artwork, because no artist exercises the relevant kind of intentional control over its production.[^1] One might think, for similar reasons, that philosophy can only be done by people: no text produced by an LLM can be a work of philosophy, because no philosopher lies behind it. Similar to the study of art, the study of philosophy is often organised around individuals: undergraduates take courses on Kant's ethics or Lewis's metaphysics, ~~and at more advanced levels there are specialists in,~~ and conferences are devoted to, the work of particular philosophers. Physics students, on the other hand, are taught Newtonian mechanics from a current textbook, and the course loses nothing if Newton's own writing is never looked at. In the sciences, then, what a text contributes can be carried by other texts. In philosophy, the contribution and its original presentation are harder to prise apart, and we might take this as evidence that a philosophical work is bound to the activity of the particular person who produced it, in a way that the sciences are not. We will now try to make this challenge from authorship more precise, by considering how far Davies' _performance_ theory of art transposes to philosophy. Davies writes: > [T]he work – what the artist achieves – is the process eventuating in that product. Works themselves are neither structures nor objects simpliciter, nor are they contextualized structures or objects. They are, rather, intentionally guided generative performances that eventuate in contextualized structures or objects (or events, as we shall see) – performances completed by what I am terming a focus of appreciation. (2004, p. 97) On Davies' view, when a painter paints a picture, the canvas is what we attend to, but it is not the work. The work is the artist's intentionally guided activity in producing the canvas; the canvas is ~~the work's focus of appreciation,~~ "the focus of our appreciative interest in the work" (2004, p. 150). Provenance, on this view, does more than supply context: facts about how the object came into being help determine what the work is and what is properly appreciated in it. %%this seems a bit compressed, maybe a bit more detail here, and a bit more succinctness in the next paragraph with the examples%% Davies supports the relocation of the work from the surface to the activity with cases in which two surfaces, or two texts, would look the same and yet differ as works, and the cases are of two kinds.%%not a very clear sentence: convoluted%% In one kind there is no performance at all: an instance of the verbal structure of _Kubla Khan_ might be generated by desert wind, or a monkey at a typewriter. A theorist who identifies the poem with its verbal structure must then either count these as instances of Coleridge's work or explain why not (2004, p. 102). In the other kind there is a performance, but not the one the surface was taken to give access to: van Meegeren's _The Disciples at Emmaus_ was presented as a newly discovered Vermeer, and what was appreciated under that description was not the achievement the canvas in fact issued from.[^2] Only the first kind bears on our question, since an LLM text would be a case of absent performance, not of misattributed performance. If Davies is right, the surface does not by itself settle the work. %%does all this need to be said about forgeries? I suspect it is not relevant but I am open to you persuading me i am wrong%% Here is what the analogous proposal for philosophy would be. A philosophical text is not itself the philosophical work: the text is the product of a person's philosophising, and reading it is a way of engaging with that prior activity. The challenge this poses is constitutive:%%not how i write%%the activity is treated not as what causes a philosophical work to exist but as part of what the work is. If no one has philosophised, there is no work to which the text gives access, however the text reads — an LLM text would stand to philosophy as the wind-made _Kubla Khan_ stands to poetry. %%the middle of this paragraph could be clearer and better written. %% Should the transposition be accepted? We do not think it should. What entitles Davies to relocate the work into the performance is an evaluative fact about art: surfaces that look the same can differ in artistic value, as the forged Vermeer and a genuine one do,%%if all that forgery talk is removed, just do the wind example%% and provenance is part of what the difference consists in. %%is this opening of the paragraph redundant because of what has already been said?%% The transposition therefore commits its defender to the corresponding claim about philosophy: that two texts containing the same argument could differ in philosophical merit. However, if two texts contain the same argument, including the same inferential moves, the same considerations count for and against them: whether the argument is valid and whether the objections are answered are questions about the texts' contents, and two texts with the same contents receive the same answers. Their philosophical merit does not vary with the route by which the words came to be written. The discipline's evaluative practice is built on the same denial. Journals strip author information from submissions before review because facts about authorship are treated as potential sources of distortion; if texts with the same contents could differ in merit, anonymising would discard evaluatively relevant information, and review would not be designed this way. The grounds for the judgement lie in the argument as presented, not in the history of its production. Nor is the author-centred teaching noted earlier in tension with this. That philosophy is taught through Kant rather than through summaries of Kant is a fact about where the discipline's contributions live, not about how they are evaluated: the _Groundwork_ is assigned because reading it puts a student somewhere no summary has yet put one, which is a fact about the text. %%not how i write, and you should give the fuller title of the book if you mention it. Also, this ending Seems very magazine-like rather than philosophical%% Dellsén et al. (2024) hold that philosophical progress is "for-whom" rather than "by-whom": it consists in putting people in a position to increase their understanding, usually by making philosophical ideas publicly available (p. 679). On this account the discipline's success-conditions locate the contribution in the public text, and a view that locates the philosophy behind the text, in the process by which it came about, misplaces it. The public text is not a dispensable trace of philosophy; it is where the philosophical contribution becomes fully available.%%seems a bit compressed, and unconnected with the paragraphs leading up to it%% A performance theorist can hold the line%%not how i write%%: no philosophising, no work, whatever the text contains. The position can be granted in full%%meta-commentative wank%%, because the thesis of this paper does not use the notion it restricts %%meta-commentative cumstain%%. Suppose the desert wind assembled not _Kubla Khan_ but a sound argument against enactivist approaches to perception. No one would deserve credit for it; ~~there would be no achievement to admire, and no entry for anyone's bibliography.~~ A reader who worked through it would nonetheless confront a thesis and the arguments marshalled in its defence, and would be in a position to answer or extend it ~~— the position a philosophical text puts its readers in when it is worth their time.~~ %%the crossed oout part is shallow and shit, replace with substance%%Whether such a text is a _work_ may then be reserved for texts with performances behind them; what cannot be reserved is the text's being worth reading, since everything that judgement answers to is on the page. %%not a very clear paragraph, what is its function supposed to be?%% The challenge from authorship therefore fails, and the doubt it leaves standing is of a different kind.%%not how i write, metacommentative wanking again%% That a text came from an LLM cannot disqualify it;%%stubby cunty sentence%% nothing so far shows that LLMs can produce such texts. If a parrot produced what sounds like a philosophical argument, this argument would not be disqualified by its source — but parrots produce no arguments, because they lack the capacities arguing requires.%%this paragraph is so compressed as to be meaningless wank%% The doubt about LLMs, in their current state, is of this kind: that they lack capacities that producing philosophy worth reading requires. Section 2 takes up the claim that they cannot perform the inference on which philosophical theorising runs; Section 3 the claims that they stand in no relation to the world and have no experience. A further challenge grants a worthwhile text and asks whose work it is, the model's or the prompting person's; it arises only if the capacity challenges fail, and we take it last. %%there is a good chance most of this paragraph can be cut, the parrot and the capcity stuff should be at the beginning of section 2, and better written.%% [^1]: This is not to deny that systems of this kind can produce beautiful images; we return to image generation in Section 4. [^2]: Han van Meegeren, the Dutch forger exposed in 1945; his _The Disciples at Emmaus_ was authenticated as a Vermeer and celebrated before the forgery came to light. --- --- --- # 2. The Challenge from Abduction — v3 (rebuilt from first principles) In the previous section, we argued against the idea that LLMs cannot produce philosophy worth reading simply because they are not human. In this section and the next we shall consider a different form of challenge: even if LLMs cannot be ruled out of producing philosophy worth reading tout court, they lack particular _capacities_ that producing it requires. If a parrot uttered a sequence of sounds that happened to form a philosophical argument, the argument would be none the worse for its source; yet parrots' powers of mimicry do not extend to producing strings of sounds so complex as to make up a philosophical argument. In this section we address one capacity challenge, which we will call the _challenge from abduction_. In the next we shall look at two more: phenomenological experience and contact with the world. Abduction, or inference to the best explanation, is reasoning from a body of evidence to the hypothesis that would best explain it. In a deductive argument the premises fix the conclusion: if all men are mortal and Socrates is a man, then Socrates is mortal, and there is no wriggle room. In abduction the evidence settles less. Now, imagine walking into your kitchen and finding the floor wet. What has happened? The wet floor does not determine the answer in the way the two premises determined Socrates' mortality: a burst pipe would have left the floor wet, and so would a spilled bucket. But, given that the window is open, the water is under the window, and it rained last night, rain coming through the window seems the most plausible answer. To reason in this way, deciding what best explains a set of facts, is common in everyday life and in the sciences alike. A scientist chooses one theory over another when it explains the same data more simply: Copernicus's model of the solar system was preferred to Ptolemy's because it explained the observed planetary motions without the elaborate epicycles the older model required. Williamson argues that philosophy is continuous with the sciences, and that its theories are to be chosen by the same abductive standards (2007; 2021, p. 351 %%check page%%). In philosophy too there are data that a candidate theory must accommodate — intuitions about cases, and the phenomena of the domain itself — and rival theories that would each accommodate them at different costs. The theory to prefer is the one that would, if true, best explain the data. What makes one explanation better than another, on this account, is a matter of explanatory virtue: a good philosophical theory is, in Williamson's words, "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated", and should "combine simplicity with strength" (2021, p. 354 %%check page%%). That theories are weighed by such comparative and explanatory virtues need not rest on a science-modelled conception of philosophy: Bengson, Cuneo and Shafer-Landau (2022) argue that the assessment of rival theories by their explanatory and unifying merits is a constraint on sound philosophical method as such. This conception of philosophy is widely held (Sider 2011; Paul 2012; Dellsén et al. 2024), though not universally (Bueno and Shalkowski 2020; Thomasson 2015), and we shall assume it in what follows. On this account a philosophical text offers its reader a choice of theory displayed — a position, its rivals, and the case for preferring it — so that whether the text is worth reading and whether it contains a good weighing travel together. If the capacity for abduction is what is required to produce worthwhile philosophy, we can ask whether LLMs possess it. Floridi et al. (2025) argue that they do not, describing what such models do instead as zeroth-order abduction: > LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025, p. 9) An LLM, on this account, has "a stochastic core and an abductive appearance" (2025, p. 3). The model is trained to predict which words are likely to follow which, and it produces the continuation its training makes probable; it aims at the likely continuation, not at the truth. The appearance comes from what the training data have passed on: models have "absorbed patterns of human abductive reasoning as expressed in writing" (p. 10) — how explanations are typically phrased, which causes are typically offered for which effects. What is inherited, on their account, is the look of the reasoning, not the reasoning itself.[^1] Explaining the wet kitchen floor involved two separable activities: coming up with candidate explanations — the burst pipe, the spilled bucket, the rain — and settling which of them the open window and the position of the water favoured. Call the first _generating_ and the second _weighing_. Floridi et al.'s position is that a model does neither, however much its text exhibits both. Asked why a car might not start on a cold morning, a model replies that a weak battery is one possibility, since cold reduces a battery's efficiency; that thickened engine oil is another, since a cold engine is harder to turn over; and that, "[b]ased on your description, the battery is the most likely explanation" (2025, p. 11). The offering of candidates here is not generating, on their reading: the model is not reasoning about causes from the user's case but reproducing the causes such explanations typically cite (p. 9). And the singling out is not weighing: the verdict reproduces how explanations of this kind typically end, and where an output marks a genuine point of difference between two hypotheses, that is something the model has seen stated, not something it has derived anew (p. 14). Floridi et al. draw the consequence themselves: > In essence, LLMs function like brainstorming assistants that toss out ideas without filtering for quality. After all, they work like statistical interfaces to an enormous amount of data accumulated for millennia by generations. A cautious human collaborator can sift through and assess them. (2025, p. 12) On this picture the weighing always remains with the person: the model supplies candidates, and assessing them is the collaborator's work. Whatever such a system produces is raw material for philosophy done by someone else, and raw material is not philosophy worth reading — a list of unweighed candidates is no more worth reading than a bare pronouncement that direct realism is correct. The challenge follows: if a model's text cannot contain a good weighing, there is no reason to regard it as worth reading. The benchmark record can seem to agree, since abduction is the form of reasoning at which models perform worst, with median accuracy across surveyed studies of roughly 43%, against 80% for deduction (Salimi et al. 2026).[^2] Everything in this account of the producer can be granted. The model generates nothing and weighs nothing, and nothing in what follows returns either capacity to it. What the account does not settle is anything about the texts. Section 1 fixed where the grounds of a text's merit lie — in the argument as presented, not in the history of its production — and by that standard the car-battery reply asks to be read rather than explained away. It is not a list of candidates awaiting a collaborator: it brings the cold morning to bear on each candidate and closes in favour of one, so the sifting the brainstorming picture reserves for the person is on the page. Whether that displayed sifting is any good is a question about a piece of writing. What remains of the challenge is the claim that the weighing a model's text displays cannot be good. Against this we will argue that the challenge underestimates what the inherited look of reasoning includes. Two things need showing: what makes a weighing displayed in a text good, and how a good weighing can come to be displayed in text that nobody weighed. Lipton's account of inference to the best explanation supplies the first; Wolfram's account of what continuing text involves supplies the second. The division of the kitchen's work into generating and weighing is Lipton's own: inference to the best explanation, on his account, runs on two filters, one that supplies the plausible candidates and a second that selects among them (2004, p. 59 %%pin page%%) — and his question about the second filter is ours, namely what makes the selection good. The best explanation can be understood as the likeliest, the one most warranted by the total evidence, or as the loveliest, the one which, if correct, would provide the most understanding: "[l]ikeliness speaks of truth; loveliness of potential understanding" (p. 59). The two come apart: Newtonian mechanics is no longer the likeliest account of the observations it was built on, but it remains as lovely an explanation of them as it ever was (p. 60). Loveliness is the standard a philosophical text answers to: what readers assess is whether the explanation offered would, if true, give understanding, and give more of it than the rivals considered. Because the assessment runs under "if correct", it does not wait on verification, and a reader can conduct it on the page. On Dellsén et al.'s account, philosophical progress consists in putting people in a position to increase their understanding (2024, p. 679); a lovely explanation puts its reader in exactly that position. Where loveliness shows itself, on Lipton's analysis, is in the comparison of rivals. Explanation is contrastive: we explain why this rather than that, and doing so requires citing a difference between the two — his Difference Condition — something in the favoured case to which nothing in its rival corresponds (2004, ch. 3). In the kitchen, "rain rather than a burst pipe, because the window is open and the water lies under it" cites such a difference, since a burst pipe would have wet the floor by the pipe; "rain rather than a burst pipe, because the floor is very wet" has the same comparative shape and cites nothing that bears on the contrast, since a very wet floor favours neither rival. Both sentences instantiate the form of a weighing, and only the first contains one worth having; telling them apart requires understanding what each claims and asking whether it decides between the candidates, which is what the reader of any philosophy paper does. Nor is there a rule that would spare the reader the work: our grasp of what makes one explanation lovelier than another is weak (p. 61), and the standards are carried, in part, by past explanations that serve as exemplars and by prevailing styles of reasoning (p. 139). Human philosophers write in the format of explanation too, and the format was never what their comparisons were graded on; the bar that separates the two kitchen sentences separates human paragraphs and machine paragraphs alike. Where reference answers give out, the machine-learning literature itself assesses generated explanations in this way, scoring them for consistency, parsimony and coherence as features of the output (Dalal et al. 2024; He et al. 2025). A model trained only to continue text respects constraints that were never stated for it, and Wolfram (2023) assembles the cases. A model trained on English respects English syntax, although no grammar was supplied to it: the syntax is carried by the writing, in which well-formed sentences predominate, and a system fitted to continue the writing comes to respect what the writing respects. Its sentences are, for the most part, meaningful rather than merely grammatical, although here there was no rule available even to withhold, since nothing like a complete theory of what makes a sentence meaningful has ever been built (2023 %%check page%%). Logic, in its syllogistic form, Wolfram treats the same way: a syllogism marks certain sentence patterns as reasonable, Aristotle, he imagines, arrived at the patterns from many examples of rhetoric, and a model trained on writing the patterns pervade can be expected to produce text containing "correct inferences" of the syllogistic kind, without anything having been derived (2023 %%check page%%). In each case a structure is present in the output while the capacity that ordinarily produces it — knowing the grammar, grasping the meaning, performing the deduction — is nowhere in the system. The absence of a rule for loveliness is therefore no obstacle on the production side. A system that wrote by applying stated rules would be halted exactly where no rule exists; these systems were never given stated rules for anything, and what they acquire, they acquire from exemplars — which are, on Lipton's account, where the standards of loveliness live. The precedent is narrower than the cases suggest, since a syllogism has a single correct completion and an abductive comparison does not: what carries over is the weaker point, and the only one needed, that a structure can be present in a text without the capacity that ordinarily produces it standing behind the text. The corpus such models are trained on is general — most of it is not philosophy — but it contains the philosophical literature, and a philosophy paper is built as a displayed comparison: a position stated, set against rivals, and defended through the objections taken to decide between them. Wolfram's cases stop at the sentence, and the extension past it is ours; but his observations concern regularities in writing rather than grammar in particular, and an argument that states a candidate, sets out its rivals and locates the difference between them is as much a recurring regularity of the writing as syntax is. It may be said that all this redescribes the statistics: the model reproduces the regularities of its training text, and reproducing regularities is not weighing. Lipton met an objection of the same shape. Bayesianism, it had been suggested, gives the mechanics of belief revision and so leaves explanatory considerations nothing to do; arguing this way, he replies, is like arguing that "thinking about technique cannot help my squash game" because the ball's motion is governed by the laws of mechanics — even if Bayesianism gave the mechanics of belief revision, inference to the best explanation "might yet illuminate its psychology" (2004, p. 108). A true description of the mechanism does not displace a true description of what is produced. And here the mechanism is the one Floridi et al. themselves describe: the patterns absorbed from writing are patterns of reasoning as expressed in writing, and the writing does not contain the phrasing of explanations detached from their organisation — which considerations bear on which rivals, and what decides between them, are in the writing too, and a system that learns to continue the writing learns them with it. The look of the reasoning was never separable from the organisation that makes reasoning assessable on a page. It may be objected that syntax is one thing and inference to the best explanation another: whatever structure next-word prediction carries, a system of this kind is too shallow for abduction, and the benchmark record reads like confirmation. The line Wolfram draws lies elsewhere, and it comes from the passage that supplied the syllogism: his toy network fails to balance long sequences of parentheses, a task demanding exact procedure with no shortcut, and sophisticated formal logic can be expected to fail for the same reason, while whatever a person can judge at a glance is managed (2023 %%check page%%). The divide such systems fail at falls between exact procedure and holistic judgement, not between the simple and the sophisticated — and weighing, on Lipton's account, sits with judgement, since no rule runs from evidence to the loveliest explanation. Read with that line in hand, the record divides against the account it seemed to confirm. A model that can recognise explanations but has nothing to draw on in producing one should fail wherever production is demanded; instead the collapse concentrates where abduction has been recast as the exact recovery of a single canonical missing premise under formal constraint — the strongest model reaches 21.5% on the hardest such benchmark and most score near zero — while on open-ended tasks, where the output is judged as an explanation, the strongest models' validity exceeds 90%.[^3] Failure tracks the demand for exact recovery, the parenthesis side of the line, and philosophical abduction does not live on that side. None of this returns to the model any capacity Floridi et al. deny it. The model infers nothing, weighs nothing, and tests nothing; what it produces is text, and the text can contain what its producer never did — a candidate stated, the live rivals organised, the difference that decides between them located. Whether a given text does this, and does it well, is settled by the reading any philosophy paper receives, under the same standard and no other. Floridi et al. come close to saying so themselves: asked whether anything turns on the process being different when the hypothesis produced is the same, they answer that for justification it perhaps does, "but regarding the content of the hypothesis and our interpretation of it, maybe not" (2025, p. 13). A good weighing of positions a literature already contains is not yet a distinction the literature lacks, and whether a model can supply the second is among the questions Section 4 takes up; what philosophy a system with no relation to the world could produce at all is the question of Section 3. [^1]: Floridi et al. also support the denial with an argument from the model's relation to the world: its words are connected to no perception of anything, and a hypothesis, once produced, is never tested against the world (2025, pp. 7–9). We take that argument up in Section 3. [^2]: These benchmarks operationalise abduction in commonsense and formal domains with crowd-labelled or mechanically checkable answers, where philosophy has no answer key, and several score generated explanations against human-written references — a comparison nothing in this paper relies on. Performance also drops under small variations to a problem (Mirzadeh et al. 2025), and Floridi et al. note that beyond the training distribution "the facade can crack" (2025, p. 9); the paper's claim is a capacity claim — that such texts can be produced — and is untouched by variation in how reliably they are. [^3]: Salimi et al.'s benchmark suite separates formally constrained missing-premise completion (ProofWriter, AbductionRules, NeuLR) from open-text explanation tasks (ART, e-CARE, UNcommonsense); the figures are from their Tables 3–6. They observe that exact-match scoring is demanding precisely where several different missing facts would each be plausible — the abductive situation — and that target structure and the size of the hypothesis space shape difficulty at least as much as subject matter does. Salimi et al. also run every benchmark with a single fixed instruction template and score one pass, while cataloguing methods — staged prompts, criticise-and-revise pipelines — that alter what models produce; what elicitation contributes is taken up in Section 4. --- # 2 — Second half restructured (v5) — hand-built to the measured published envelope _The adversarial loop oscillated (tells 137/138/22/126/139, no convergence) and its "best" round returned only one paragraph, so this v5 is built by hand from the comparison phase's per-paragraph findings. Target envelope, measured from the published Young/Terrone paragraphs: ~83–104 words per paragraph, no sentence over ~53 words, at most one heavy mark (em-dash/colon/semicolon) per paragraph, an internal short beat in the denser paragraphs, objection-and-reply within one paragraph, two-member contrasts rather than triplets. D, G, C, A and F are split to sit in the envelope, so the paragraph count rises — matching your density means more, shorter paragraphs. Ledger content preserved. Verified Floridi pages applied (concession p.12; "facade" p.9); Lipton pages unverifiable from the EPUB and left as set. Footnotes resolve to the v4 definitions above._ [fixed lead-in] On this picture the model puts candidates forward and the collaborator does the assessing, so whatever the system produces is raw material for a piece of philosophy carried out by someone else. Raw material of this kind is not itself philosophy worth reading. A list of unweighed candidates gives a reader no more to assess than a bare pronouncement that direct realism is correct, since in each case the position has merely been stated and the comparative work that would make it worth weighing has not been done. So if a model's text cannot contain a good weighing, there is no reason to regard that text as worth reading. Benchmarks of model reasoning can look as though they bear this out, since abduction is the form of reasoning at which models perform worst, with median accuracy across surveyed studies of roughly 43%, against 80% for deduction (Salimi et al. 2026).[^2] A model generates no candidates and weighs none. That much of Floridi et al.'s account can be granted, since nothing in what follows depends on the model having either capacity. None of it bears on the text the model produces. Section 1 located a text's merit in the argument it presents, not in the history of its production. By that standard the car-battery reply is not a list of candidates left for someone else to sift. It weighs the cold morning against each one and comes down in favour of the battery, so the assessment the brainstorming picture reserves for the collaborator has already been made on the page. Asked whether anything turns on the process being different when the hypothesis produced is the same, Floridi et al. half-concede: they allow that for justification it perhaps does, but "regarding the content of the hypothesis and our interpretation of it, maybe not" (2025, p. 12). Whether the assessment a text displays is any good is a question about a piece of writing. A text can display a good weighing without its producer having weighed anything. A displayed weighing is good when a reader can assess it: when the text sets rival explanations against one another and comes down for one, the reader can ask whether it has come down well. That is the standard Lipton draws for inference to the best explanation. A system that continues a body of text, as Wolfram describes, can produce a weighing that meets it where nothing was weighed at all. Lipton asks what makes one selection among the candidates better than another, and separates two things the best explanation might be. It might be the likeliest, the explanation the total evidence most warrants, or the loveliest, the one that, were it correct, would yield the most understanding. "Likeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). Newtonian mechanics shows how far the two can fall apart: it is no longer the likeliest account of the motions it was built to explain, yet it remains as lovely an explanation of them as it ever was (p. 60). A philosophical text answers to loveliness. What a reader assesses is whether the explanation on offer would, if correct, give more understanding than its rivals, and that assessment is made under "if correct". So it does not wait on the explanation's truth, and the reader can carry it out on the page.[^d] Loveliness shows in the comparison of rivals. To explain is to explain why this rather than that, and that requires citing a difference between them. This is Lipton's Difference Condition: a cause present in the favoured case, together with the absence, in the rival, of any corresponding cause (2004, ch. 3). "Rain rather than a burst pipe, because the window is open and the water lies beneath it" meets it. The open window and the pooling beneath it are a cause present in the rain case, and the wetting-by-the-pipe that a burst pipe would have left is absent. "Rain rather than a burst pipe, because the floor is very wet" does not, since a very wet floor is what both rivals would produce and picks out nothing between them. Only the first cites a difference between the rivals, and so only the first earns its loveliness. No rule sorts the two sentences for the reader. Our grasp of what makes one explanation lovelier than another is weak, and such standards as there are come from past explanations serving as exemplars, and from prevailing styles of reasoning (2004, pp. 61, 139). Telling the two apart is the ordinary work of reading an argument: following what each says and asking whether it would decide the case. The bare form of explanation never did that work for human philosophers either. What sorts the good weighing from the bad is the reading, and the same bar applies whether a human or a machine wrote the paragraph.[^ml] A system trained only to continue text comes to respect constraints that were never stated, and Wolfram (2023) collects the cases. Trained on English, it respects English syntax though it was given no grammar, because well-formed sentences predominate in the writing it continues. Its sentences are mostly meaningful, not merely well-formed, and here no rule was available even to withhold, since no one has built a complete theory of what makes a sentence meaningful. Its syllogisms come out valid for the same reason. The patterns pervade the writing; Aristotle, Wolfram suggests, read them off many examples of rhetoric, and a system continuing that writing yields "correct inferences" of the syllogistic kind, with nothing derived. In each case a structure is present in the output while the capacity that ordinarily produces it is absent. That loveliness answers to no stated rule is no barrier to its appearing in the output. A system that wrote by stated rules would stop wherever no rule had been stated. These systems were given no stated rules at all, and what they acquire they acquire from exemplars — which, on Lipton's account, is just where the standards of loveliness reside. The precedent reaches only so far. A syllogism has one correct completion where an abductive comparison has none, so what survives is the weaker claim, which is all that is needed: that a structure can stand in a text with no trace of the capacity that ordinarily produces it. Most of what these systems are trained on is not philosophy, but the corpus contains the philosophical literature too. A philosophy paper is itself a displayed comparison: a position is stated, set against its rivals, and the difference that decides between them is drawn out. That comparison is as much a regularity of the writing as syntax is. Wolfram's cases stop at the sentence, and the move to the paragraph is ours; but what he describes are regularities in writing, not facts about grammar in particular. It may be said that this only redescribes the statistics, since a model reproduces the regularities of its training text, and reproducing regularities is not weighing. To the Bayesian who holds that, once belief revision has its mechanics, explanatory considerations have nothing left to do, Lipton replies that a true account of the mechanism need not displace a true account of what it produces. A squash ball's flight obeys the laws of mechanics, yet "thinking about technique cannot help my squash game" does not follow. Even granting the Bayesian mechanics, inference to the best explanation "might yet illuminate its psychology" (2004, p. 108). The mechanism at issue here is the one Floridi et al. themselves describe, and on that description the objection does not go through. The patterns a model absorbs are patterns of reasoning as expressed in writing, and writing does not carry the phrasing of an explanation stripped of its organisation. Which considerations bear on which rival, and what settles the matter between them, are in the writing too, and a system that learns to continue the writing learns these with the rest. The appearance the objection grants was never separable from the organisation that makes a piece of reasoning assessable on the page. It may still be objected that syntax is one thing and inference to the best explanation another, and that whatever structure next-word prediction carries, a system of this kind is too shallow for abduction, with the benchmark record reading like confirmation. But the line Wolfram draws falls elsewhere, and it comes from the same discussion that supplied the syllogism. His simple network cannot balance long sequences of parentheses, a task that demands exact procedure with no shortcut, and sophisticated formal logic should fail for the same reason, while it manages whatever a person can take in at a glance. The divide these systems fail at runs between exact procedure and holistic judgement; the contrast between the simple and the sophisticated is beside the point. Weighing, on Lipton's account, sits with judgement, since no rule runs from the evidence to the loveliest explanation. A system that could recognise explanations but had nothing to draw on in making one should fail wherever production is demanded. The record reverses. The collapse concentrates in one place: where abduction is recast as the exact recovery of a single canonical missing premise under formal constraint. There the strongest model manages 21.5% on the hardest such benchmark, and most score near zero. On open-ended tasks, where the output is judged as an explanation, the strongest models exceed 90% validity (Salimi et al. 2026).[^3] Failure tracks exact recovery, the parenthesis side. Philosophical abduction is not of that kind. None of this gives the model any capacity Floridi et al. deny it. It infers nothing, and it weighs and tests nothing; what it makes is text, and the text can hold what its maker did not: a candidate stated, its live rivals set in order, and the difference that decides between them. Whether a given text holds these things, and holds them well, is settled by the reading any philosophy paper is given, by the same standard. A good weighing of positions the literature already contains is not yet the distinction the literature lacks. Whether a model can supply that is taken up in Section 4; Section 3 asks what philosophy a system with no relation to the world could produce at all. --- --- # PREVIOUS VERSION (v2) — retained for reference # 2. The Challenge from Abduction — v3 (rebuilt from first principles) In the previous section, we argued against the idea that LLMs cannot produce philosophy worth reading simply because they are not human. In this section and the next we shall consider a different form of challenge: even if LLMs cannot be ruled out of producing philosophy worth reading tout court, they lack particular _capacities_ that producing it requires. If a parrot uttered a sequence of sounds that happened to form a philosophical argument, the argument would be none the worse for its source; yet parrots' powers of mimicry do not extend to producing strings of sounds so complex as to make up a philosophical argument. In this section we address one capacity challenge, which we will call the _challenge from abduction_. In the next we shall look at two more: phenomenological experience and contact with the world. Abduction, or inference to the best explanation, is reasoning from a body of evidence to the hypothesis that would best explain it. In a deductive argument the premises fix the conclusion: if all men are mortal and Socrates is a man, then Socrates is mortal, and there is no wriggle room. In abduction the evidence settles less. Now, imagine walking into your kitchen and finding the floor wet. What has happened? The wet floor does not determine the answer in the way the two premises determined Socrates' mortality: a burst pipe would have left the floor wet, and so would a spilled bucket. But, given that the window is open, the water is under the window, and it rained last night, rain coming through the window seems the most plausible answer. To reason in this way, deciding what best explains a set of facts, is common in everyday life and in the sciences alike. A scientist chooses one theory over another when it explains the same data more simply: Copernicus's model of the solar system was preferred to Ptolemy's because it explained the observed planetary motions without the elaborate epicycles the older model required. Williamson argues that philosophy is continuous with the sciences, and that its theories are to be chosen by the same abductive standards (2007; 2021, p. 351 %%check page%%). In philosophy too there are data that a candidate theory must accommodate — intuitions about cases, and the phenomena of the domain itself — and rival theories that would each accommodate them at different costs. The theory to prefer is the one that would, if true, best explain the data. What makes one explanation better than another, on this account, is a matter of explanatory virtue: a good philosophical theory is, in Williamson's words, "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated", and should "combine simplicity with strength" (2021, p. 354 %%check page%%). That theories are weighed by such comparative and explanatory virtues need not rest on a science-modelled conception of philosophy: Bengson, Cuneo and Shafer-Landau (2022) argue that the assessment of rival theories by their explanatory and unifying merits is a constraint on sound philosophical method as such. This conception of philosophy is widely held (Sider 2011; Paul 2012; Dellsén et al. 2024), though not universally (Bueno and Shalkowski 2020; Thomasson 2015), and we shall assume it in what follows. On this account a philosophical text offers its reader a choice of theory displayed — a position, its rivals, and the case for preferring it — so that whether the text is worth reading and whether it contains a good weighing travel together. %%This is pathetic. First of all, grown-ups don't have single-sentence paragraphs. Second, what do you mean theory displayed? Third, you're using example lists. Yeah, so anyway, I have no idea what to do with it other than to tell you that it is diarrhea bad.%% If the capacity for abduction is what is required to produce worthwhile philosophy, then the prospects for artificially generated philosophy seem dim. Consider the following from Floridi et al. (2025): > LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025, p. 9) An LLM, on this account, has "a stochastic core and an abductive appearance" (2025, p. 3), in that the model is trained to predict which words are likely to follow which, and it produces the continuation its training makes probable; it aims at the likely continuation, not at the truth. The appearance comes from what the training data have passed on: models have "absorbed patterns of human abductive reasoning as expressed in writing" (p. 10) — how explanations are typically phrased, which causes are typically offered for which effects. What is inherited, on their account, is the look of the reasoning, not the reasoning itself.[^1] Explaining the wet kitchen floor involved two separable activities: coming up with candidate explanations, and settling which of them the open window and the position of the water favoured. Call the first _generating_ and the second _weighing_. %%please check we are not lifing these terms from lipton, floridi, or someone else. if we are, no problem but it needs to be clear where they come from%%Floridi et al.'s position is that a model does neither, however much its text exhibits both. Asked why a car might not start on a cold morning, a model replies that a weak battery is one possibility, since cold reduces a battery's efficiency; that thickened engine oil is another, since a cold engine is harder to turn over; and that, "[b]ased on your description, the battery is the most likely explanation" (2025, p. 11). The offering of candidates here is not generating, on their reading: the model is not reasoning about causes from the user's case but reproducing the causes such explanations typically cite (p. 9)%%the clause after the colon is not clear. a little compressed perhaps?%%. And the singling out is not weighing: the verdict reproduces how explanations of this kind typically end, and where an output marks a genuine point of difference between two hypotheses, that is something the model has seen stated, not something it has derived anew (p. 14). %%I wonder if this paragraph could be a little clearer, because it is so important%% Floridi et al. draw the consequence themselves: %%metacommentry wanking...%% > In essence, LLMs function like brainstorming assistants that toss out ideas without filtering for quality. After all, they work like statistical interfaces to an enormous amount of data accumulated for millennia by generations. A cautious human collaborator can sift through and assess them. (2025, p. 12) On this picture the weighing always remains with the person: the model supplies candidates, and assessing them is the collaborator's work. Whatever such a system produces is raw material for philosophy done by someone else, and raw material is not philosophy worth reading — a list of unweighed candidates is no more worth reading than a bare pronouncement that direct realism is correct. %%so far this paragraph has used a lot of words for a VERY simple idea%%The challenge follows:%%the challenge does not follow, this second half of the paragraph is entirely uncnnected to the first half%% if a model's text cannot contain a good weighing, there is no reason to regard it as worth reading. The benchmark record can seem to agree, since abduction is the form of reasoning at which models perform worst, with median accuracy across surveyed studies of roughly 43%, against 80% for deduction (Salimi et al. 2026).[^2] %%so compressed as to be meaningless. also 'the benchmark record' is a meaningless and cretenous way of putting things%% Everything in this account of the producer can be granted.%%It's an obscure way to start a paragraph.%% The model generates nothing and weighs nothing, and nothing in what follows returns either capacity to it. %%last clause =not how i write%%What the account does not settle is anything about the texts.%%Meta commentative wank.%% Section 1 fixed where the grounds of a text's merit lie — in the argument as presented, not in the history of its production — and by that standard the car-battery reply asks to be read rather than explained away.%%not how i write%%It is not a list of candidates awaiting a collaborator: it brings the cold morning to bear on each candidate and closes in favour of one, so the sifting the brainstorming picture reserves for the person is on the page. Whether that displayed sifting is any good is a question about a piece of writing.%%paragraph is compressed so very unclear%% What remains of the challenge is the claim that the weighing a model's text displays cannot be good. Against this we will argue that the challenge underestimates what the inherited look of reasoning includes.%%starting to have my doubts about using the word 'look' for this stuff...%% Two things need showing: what makes a weighing displayed in a text good, and how a good weighing can come to be displayed in text that nobody weighed. Lipton's account of inference to the best explanation supplies the first; Wolfram's account of what continuing text involves supplies the second. %%also, is weighing the bet word to use? why are we using this one?%% The division of the kitchen's work into generating and weighing is Lipton's own%%incredibly unclear sentence%%: inference to the best explanation, on his account, runs on two filters, one that supplies the plausible candidates and a second that selects among them (2004, p. 59 %%pin page%%) — and his question about the second filter is ours, namely what makes the selection good.%%this is the biggest structural issue here. we are introducing these two parts of abduction for a second time, this time with a lipton flavour. this is very inelegant. why do we even need to mention this here when we are talking about loveliness? %% The best explanation can be understood as the likeliest, the one most warranted by the total evidence, or as the loveliest, the one which, if correct, would provide the most understanding: "[l]ikeliness speaks of truth; loveliness of potential understanding" (p. 59). The two come apart: Newtonian mechanics is no longer the likeliest account of the observations it was built on, but it remains as lovely an explanation of them as it ever was (p. 60). Loveliness is the standard a philosophical text answers to: what readers assess is whether the explanation offered would, if true, give understanding, and give more of it than the rivals considered. Because the assessment runs under "if correct", it does not wait on verification, and a reader can conduct it on the page. On Dellsén et al.'s account, philosophical progress consists in putting people in a position to increase their understanding (2024, p. 679); a lovely explanation puts its reader in exactly that position. %%this is not a clear paragraph. Also, in previous version we made it clear to the reader how the term 'likeliest' is being used by Lipton, you have removed this information without my permission.%% Where loveliness shows itself, on Lipton's analysis, is in the comparison of rivals.%%notclear at all%% Explanation is contrastive: we explain why this rather than that, and doing so requires citing a difference between the two — his Difference Condition — something in the favoured case to which nothing in its rival corresponds (2004, ch. 3).%%to compressed so unclear%% In the kitchen, "rain rather than a burst pipe, because the window is open and the water lies under it" cites such a difference, since a burst pipe would have wet the floor by the pipe; "rain rather than a burst pipe, because the floor is very wet" has the same comparative shape and cites nothing that bears on the contrast, since a very wet floor favours neither rival. Both sentences instantiate the form of a weighing, and only the first contains one worth having; telling them apart requires understanding what each claims and asking whether it decides between the candidates, which is what the reader of any philosophy paper does.%%I don't understand this sentence, all i know is that it is conveying a shit idea%% Nor is there a rule that would spare the reader the work:%%why is the reader being told all this shit, it seems like you are putting it in just for the sake of wasting words%% our grasp of what makes one explanation lovelier than another is weak (p. 61), and the standards are carried, in part, by past explanations that serve as exemplars and by prevailing styles of reasoning (p. 139). Human philosophers write in the format of explanation too, and the format was never what their comparisons were graded on; the bar that separates the two kitchen sentences separates human paragraphs and machine paragraphs alike. Where reference answers give out, the machine-learning literature itself assesses generated explanations in this way, scoring them for consistency, parsimony and coherence as features of the output (Dalal et al. 2024; He et al. 2025). %%a steaming turd of a paragraph. an absolute disgrace%% A model trained only to continue text respects constraints that were never stated for it, and Wolfram (2023) assembles the cases. A model trained on English respects English syntax, although no grammar was supplied to it: the syntax is carried by the writing, in which well-formed sentences predominate, and a system fitted to continue the writing comes to respect what the writing respects. Its sentences are, for the most part, meaningful rather than merely grammatical, although here there was no rule available even to withhold, since nothing like a complete theory of what makes a sentence meaningful has ever been built (2023 %%check page%%). Logic, in its syllogistic form, Wolfram treats the same way: a syllogism marks certain sentence patterns as reasonable, Aristotle, he imagines, arrived at the patterns from many examples of rhetoric, and a model trained on writing the patterns pervade can be expected to produce text containing "correct inferences" of the syllogistic kind, without anything having been derived (2023 %%check page%%). In each case a structure is present in the output while the capacity that ordinarily produces it — knowing the grammar, grasping the meaning, performing the deduction — is nowhere in the system. The absence of a rule for loveliness is therefore no obstacle on the production side. A system that wrote by applying stated rules would be halted exactly where no rule exists; these systems were never given stated rules for anything, and what they acquire, they acquire from exemplars — which are, on Lipton's account, where the standards of loveliness live. The precedent is narrower than the cases suggest, since a syllogism has a single correct completion and an abductive comparison does not: what carries over is the weaker point, and the only one needed, that a structure can be present in a text without the capacity that ordinarily produces it standing behind the text. %%this paragraph is far too long, I am not even going to read it%% The corpus such models are trained on is general — most of it is not philosophy — but it contains the philosophical literature, and a philosophy paper is built as a displayed comparison: a position stated, set against rivals, and defended through the objections taken to decide between them. Wolfram's cases stop at the sentence, and the extension past it is ours; but his observations concern regularities in writing rather than grammar in particular, and an argument that states a candidate, sets out its rivals and locates the difference between them is as much a recurring regularity of the writing as syntax is. It may be said that all this redescribes the statistics: the model reproduces the regularities of its training text, and reproducing regularities is not weighing. Lipton met an objection of the same shape. Bayesianism, it had been suggested, gives the mechanics of belief revision and so leaves explanatory considerations nothing to do; arguing this way, he replies, is like arguing that "thinking about technique cannot help my squash game" because the ball's motion is governed by the laws of mechanics — even if Bayesianism gave the mechanics of belief revision, inference to the best explanation "might yet illuminate its psychology" (2004, p. 108). A true description of the mechanism does not displace a true description of what is produced. And here the mechanism is the one Floridi et al. themselves describe: the patterns absorbed from writing are patterns of reasoning as expressed in writing, and the writing does not contain the phrasing of explanations detached from their organisation — which considerations bear on which rivals, and what decides between them, are in the writing too, and a system that learns to continue the writing learns them with it. The look of the reasoning was never separable from the organisation that makes reasoning assessable on a page. It may be objected that syntax is one thing and inference to the best explanation another: whatever structure next-word prediction carries, a system of this kind is too shallow for abduction, and the benchmark record reads like confirmation. The line Wolfram draws lies elsewhere, and it comes from the passage that supplied the syllogism: his toy network fails to balance long sequences of parentheses, a task demanding exact procedure with no shortcut, and sophisticated formal logic can be expected to fail for the same reason, while whatever a person can judge at a glance is managed (2023 %%check page%%). The divide such systems fail at falls between exact procedure and holistic judgement, not between the simple and the sophisticated — and weighing, on Lipton's account, sits with judgement, since no rule runs from evidence to the loveliest explanation. Read with that line in hand, the record divides against the account it seemed to confirm. A model that can recognise explanations but has nothing to draw on in producing one should fail wherever production is demanded; instead the collapse concentrates where abduction has been recast as the exact recovery of a single canonical missing premise under formal constraint — the strongest model reaches 21.5% on the hardest such benchmark and most score near zero — while on open-ended tasks, where the output is judged as an explanation, the strongest models' validity exceeds 90%.[^3] Failure tracks the demand for exact recovery, the parenthesis side of the line, and philosophical abduction does not live on that side. None of this returns to the model any capacity Floridi et al. deny it. The model infers nothing, weighs nothing, and tests nothing; what it produces is text, and the text can contain what its producer never did — a candidate stated, the live rivals organised, the difference that decides between them located. Whether a given text does this, and does it well, is settled by the reading any philosophy paper receives, under the same standard and no other. Floridi et al. come close to saying so themselves: asked whether anything turns on the process being different when the hypothesis produced is the same, they answer that for justification it perhaps does, "but regarding the content of the hypothesis and our interpretation of it, maybe not" (2025, p. 13). A good weighing of positions a literature already contains is not yet a distinction the literature lacks, and whether a model can supply the second is among the questions Section 4 takes up; what philosophy a system with no relation to the world could produce at all is the question of Section 3. [^1]: Floridi et al. also support the denial with an argument from the model's relation to the world: its words are connected to no perception of anything, and a hypothesis, once produced, is never tested against the world (2025, pp. 7–9). We take that argument up in Section 3. [^2]: These benchmarks operationalise abduction in commonsense and formal domains with crowd-labelled or mechanically checkable answers, where philosophy has no answer key, and several score generated explanations against human-written references — a comparison nothing in this paper relies on. Performance also drops under small variations to a problem (Mirzadeh et al. 2025), and Floridi et al. note that beyond the training distribution "the facade can crack" (2025, p. 10); the paper's claim is a capacity claim — that such texts can be produced — and is untouched by variation in how reliably they are. [^3]: Salimi et al.'s benchmark suite separates formally constrained missing-premise completion (ProofWriter, AbductionRules, NeuLR) from open-text explanation tasks (ART, e-CARE, UNcommonsense); the figures are from their Tables 3–6. They observe that exact-match scoring is demanding precisely where several different missing facts would each be plausible — the abductive situation — and that target structure and the size of the hypothesis space shape difficulty at least as much as subject matter does. Salimi et al. also run every benchmark with a single fixed instruction template and score one pass, while cataloguing methods — staged prompts, criticise-and-revise pipelines — that alter what models produce; what elicitation contributes is taken up in Section 4. --- # 3. The Challenge of Connecting to the World and the Challenge from Experience — v4 (approved fixes 1,2,4,5,7,8 applied) In this section we address two more capacity challenges, which we will call the _challenge of connecting to the world_ and the _challenge from experience_. The first holds that LLMs cannot produce philosophy worth reading because they stand in no relation to the world: nothing a model says rests on perception of anything, and nothing it produces is tested against how things are. The second holds that they cannot because some philosophy depends on experience in a way a system without experience cannot meet: experience supplies the starting point of some philosophical reasoning, and is itself the subject matter of some philosophical inquiry. Few would say that current LLMs are conscious, and we assume here that they are not. We argue that both challenges fail, and for the same reason: the materials philosophy takes from the world and from experience reach the philosopher already set down in words, and a model works on those words as any philosopher does. Floridi et al. also object that the model stands in no relation to the world: its words rest on no perception of anything, and a hypothesis, once produced, is never tested against how things are (2025, pp. 7–9). A discipline whose theories answer to how things are, the challenge runs, cannot be advanced by a system with no access to how things are, and a text from one gives its reader no reason to think it worth reading. The challenge from experience says that some philosophy cannot be done by anyone who has not had the relevant experience. Zahavy (2026) raises a similar worry about scientific discovery. A model can carry out the deductive part of discovery, working out the consequences of premises it has been given; what it cannot do, he holds, is produce the premises — make the move from sense experience to new first principles. On the picture he takes from Einstein, that move is a leap, and it is the leap[^5] that gives a theory its axioms. His case is the thought experiment that gave Einstein the equivalence principle: > Einstein's variation required inventing new axioms based on a physical intuition that did not yet exist in the mathematics. He envisioned a physicist inside an elevator being uniformly accelerated through deep space [...]. Inside this enclosure, the sensory experience reveals a specific pattern: when objects are released, the floor rushes up to meet them. To the physicist, the objects appear to fall with identical acceleration, regardless of composition. Thus, the simulation here was not a permutation of symbols, but a manipulation of perceptual experience. (Zahavy 2026, §5) Einstein imagines a set of circumstances and attends to what would be experienced within them: everything released inside the elevator appears to fall with identical acceleration. On Zahavy's reconstruction the simulation supplies an observation, and from that observation the new axiom is inferred — the simulated experience of acceleration was indistinguishable from the remembered experience of gravity, and Einstein concluded that the two are one phenomenon.[^2] A model has no access to that observation. It can produce descriptions of elevators and of weightlessness, both present in its corpus, but it has undergone neither, and a discovery whose premises are fixed by simulated experience is beyond a system that, in Zahavy's words, lacks the capacity he calls sensory agency. We might think that philosophical thought experiments depend upon experience in the same way. Does Mary, released from her black-and-white room knowing every physical fact about colour vision, learn something when she first sees red (Jackson 1982)? Settling the question requires considering what the experience is like, and the argument proceeds from the verdict; this is experience entering as the starting point of an argument, and a system that has never experienced anything appears unable to supply it. Experience enters also as subject matter, in philosophy that asks what it is to see red (Harman 1990), to feel anger (Goldie 2000), or to have an intuition take hold (Chudnoff 2011).[^3] Were these claims correct, much of the philosophy of mind would lie beyond a model's reach. We can grant that the model has no senses and has never had an experience, and that it cannot put what it produces to the test against the world. Whether any of that bears on the texts it produces depends on what philosophy does with the world and with experience — on where a philosopher's starting points come from, and on what becomes of a philosophical claim once it is made. Pigliucci (2017) addresses both. He holds that philosophy is constrained by the world without investigating it as the natural sciences do: > This means that the basic parameters that philosophers use as their inputs, the starting points of their philosophizing, their equivalent of axioms in mathematics and assumptions in logic (or rules in chess) are _empirical_ data about the world. This data comes from both everyday experience [...] and of course increasingly from the world of science itself. Philosophy, I maintain, is in the business of exploring the sort of conceptually evoked spaces that Smolin is talking about, where the evocation is the result of whatever starting assumptions are made by individual philosophers working within a particular field and, crucially, of the constraints that are imposed by our best understanding of how the world actually is. (2017, pp. 79–80) _Evocation_ is a term Pigliucci takes from Smolin (Unger and Smolin 2015),[^4] for truths that are neither discovered, in the sense of corresponding to mind-independent states of affairs, nor invented, in the sense of being arbitrary constructs, and his example is chess: when the rules of a game are codified, a whole bundle of facts about it becomes demonstrable — objective facts, in that anyone who can demonstrate one demonstrates the same fact as anyone else — although chess did not exist before its rules were written down (Unger and Smolin 2015, p. 423; quoted at Pigliucci 2017, p. 78). Pigliucci's proposal is that philosophy ascertains evoked truths in this sense, with an addition that separates it from mathematics and chess alike: its starting points are constrained by how the world actually is. The addition is also what separates philosophy from fiction, on his account. A novelist's worlds are invented rather than evoked — nothing about them is rigid, since even the constraints the novelist adopts could have been otherwise — whereas philosophy "is in the business of doing empirically informed evoking, not inventing", so that its objects of study have rigid properties (2017, p. 80). A thought experiment is itself a case of such evoking: the philosopher sets up an imagined scenario but explores it "with an interest in figuring things out as far as this world is concerned" (2017, p. 80), so that what it evokes has the rigid properties Pigliucci means, and the philosophical work proceeds within the structure it opens. Philosophy's starting points, then, are empirical, and a system that perceives nothing cannot reach them on its own. But the data Pigliucci describes comes from everyday experience and, increasingly, from science, and the scientific kind reaches working philosophers already articulated — already set down in language, available to be read rather than undergone. A philosopher of physics works from published results, not from having run the experiments. The same holds for everyday experience: what the discipline retains of it, it retains as the literature's accumulated descriptions of how things seem. For the worldly materials philosophy actually uses, written access is the profession's normal condition rather than a deficiency, and a corpus is such access. On this point the model stands where every philosopher already stands with respect to nearly all of the empirical data they use. The other objection was that the model never checks what it produces against the world. But a philosophical claim is not the kind of thing that gets checked that way. Here the elevator and Mary part company. The equivalence principle, once Einstein had it, faced a tribunal of measurement: the experiments might have gone against it, and then it would have been dropped. Mary's case faces no such tribunal. The question Mary raises — whether she learns something new on first seeing red — is a question about what follows within the scenario Jackson has set up, and that is settled in the way any question about chess is settled: by working out what the set-up commits us to, something any competent party can do and none can decide by fiat. This is the only checking a philosophical thesis gets, and it happens in the literature, in the back-and-forth the previous section described. So a model's inability to run experiments costs it nothing a philosophical text needs: the testing that philosophy does is the working-out of what a scenario commits us to, and that is done on the page. Take the knowledge argument itself. Nobody who debates it has been through what Mary goes through — released into colour after a lifetime of black and white — and the debate does not suffer for it. The participants have in front of them Jackson's description: a scenario set down in words, with a claim about what it is meant to show. Lewis's reply (1988) works on that description, changing what we should say follows from Mary's release while saying nothing about what her first sight of red is like from the inside. This is the usual way experience figures in philosophy. A philosopher need not have had an experience to argue about it: Nagel (1974) asks what it is like to be a bat without supposing he could find out, working instead from a description of the bat's situation. A model is no worse off here than these philosophers are. It has had no experiences of its own, but it has not needed them, because what its corpus supplies, and what the philosophy works on, is experience already put into words — the same descriptions that, in the previous section, carried displayed reasoning into its texts. Merleau-Ponty noticed something about touching one hand with the other that was not, so far as we know, already written down anywhere. At any moment one hand is the toucher and the other the touched; the roles can switch, but they cannot both hold at once (Merleau-Ponty 1945 %%page%%). Suppose he came to this only by attending to his own body, to what the touching was like from the inside, and not from anything already in the literature. Then it is a starting point a model could not have reached on its own, since reaching it took a body and a first-person view that a model does not have. Once Merleau-Ponty has written the description down, a model can take it up and argue about it as well as anyone. It could not, though, have been the one to set it down first. What a philosopher does with Merleau-Ponty's description, though, has nothing to do with who first arrived at it. The description earns its place by what can be drawn out of it — what it shows about the body, what follows once the asymmetry is granted — and that is there for any reader to work through, whether or not they could have come to the asymmetry on their own. First-person attention is not, in any case, the only way new descriptions come about. A description the literature does not yet contain can also be reached from the descriptions it does contain, by drawing out what they have not been taken to imply, or by putting two of them together as no one has — and a model can do this. Whether today's models in fact produce descriptions the literature lacks is the question of novelty, which Section 4 takes up. A model can do the philosophy that turns on the world and on experience, save for the one point already granted: it could not be the first to set down a description that only first-person attention could yield. Its worldly starting points come to it already written down; the claims it draws are settled by working out what they commit us to, not by measurement against the world; and the experiences philosophy argues about reach it as descriptions, which it can work through as well as any reader. In the previous section, abduction entered philosophy as reasoning displayed in a text; here the world and experience enter it as descriptions set down in a text. What a model produces is read as any philosophy is read, and how it was produced settles nothing in advance. [^2]: Zahavy, following Magnani, calls the process _manipulative abduction_: hypothesis generation through the manipulation of a model — here a simulated experience — rather than of symbols (Magnani et al. 2009; Zahavy 2026, §5). It is abduction in the previous section's sense: the equivalence principle is inferred as the best explanation of the simulated observation, the simulation supplying an explanandum that no search over existing text would have produced. What experience contributes, on this picture, is not the inference but its starting point. [^3]: The materials need not be sensory: the feeling of understanding something is sometimes used to motivate the claim that thought itself has a phenomenology (Pitt 2004). [^5]: Zahavy too calls this leap abduction, but the word picks out something other than it did in the previous section. There, with Floridi et al., abduction was the weighing of rival explanations, and the charge was that a model only mimics it; here it is the generation of new first principles from experience, and Zahavy's claim is that a model cannot make the move because it has had no experience to move from. The present challenge rests on that second claim, about experience, and not on any verdict about the weighing. [^4]: _The Singular Universe and the Reality of Time_ is jointly authored, but its second part, which contains the discussion of evocation, was written by Smolin alone, as Pigliucci notes (2017, p. 77); we follow him in attributing the view to Smolin. --- --- # 4. The Challenge from Observation In this final section we address what we will call the _challenge from observation_. In Section 1 we argued that LLMs should not be disqualified from producing worthwhile philosophy tout court. - In Sections 2 and 3 we argued that, although LLMs neither perform abductive inference, nor have experience, nor are connected to the world, there is still reason to think that they are capable of producing text which exhibits good quality abduction, and works with articulated axioms about experience and the world. - The challenge from observation begins with an obvious question: if all of these arguments are correct, where is all the worthwhile LLM-written philosophy? If you ask an LLM the answer to the hard problem of consciousness, or the meaning of life you will not receive *the correct answer*, but instead a competent but unopionated survey of the field if you are lucky, or a less accurate but equally bland survey if you are unlucky. The observation is accurate, and it reports less than it seems to: it reports what models produce under one use — a bare question, put once, answered in one pass. How these systems are built explains why that use yields what it does. A model is first fitted to a vast general corpus and trained to continue text, so its response to a bare philosophical question is the likely continuation of such a question in writing at large, and the likely continuation of "what is the meaning of life?" in a general corpus is not an analytic tract. It is the sort of text that follows the question at large: a survey of views, a consoling generality, a joke. The model is then further shaped in post-training to converse as a helpful assistant, and the shaping pulls in the same direction, a helpful assistant would not answer a six word question with a substantial piece pof philosophy. The use that generates the observation treats the model as an oracle: a system whose answers are its measure, so that asking is all the eliciting there is.[^2] The empirical record tells against the assumption. The survey of abductive benchmarks discussed in Section 2 runs every test with a single fixed instruction and scores the answer, while cataloguing, in the same pages, methods that alter what models produce — prompts that separate the stages of a task, pipelines in which an answer is criticised and revised over several passes (Salimi et al. 2026). What a model returns depends on what it is given, and the observation samples one point in that space, the bare question. It therefore cannot discriminate between the two hypotheses at issue — that the capacity defended in the preceding sections is absent, and that it has not been elicited. Both predict the observed record, and an argument against this paper needs the first; the observation supports it no better than the second. The challenge has a natural escalation: if philosophy worth reading comes out of these systems only when a philosopher directs the process — supplies the framing, sets the constraints, presses for development — then the philosophy, it will be said, is the philosopher's. The model is an instrument in the production, as a typewriter is, and crediting it with the result is crediting the dummy with the ventriloquism. Section 1's challenge held that a model's text is not philosophy tout court; what stands here is narrower, that the philosophy in such a text is not the model's. Whether the escalation succeeds depends on what prompting a model involves, and two things need saying: what a prompt supplies, and what the model's continuation adds to it. Section 3's account of starting points says the first; Section 2's account of continuing text says the second. A prompt articulates a starting point, as a thought experiment does. A prompt that sets out a position and the rivals it must beat stands to the model as Jackson's two paragraphs stand to the profession: a starting point handed over for development. What an articulated starting point does, on the account already in place, is evoke a structure with rigid properties — there are facts about what holds within it, demonstrable by anyone and chosen by no one, and they outrun whatever has been stated, just as the facts about chess outran the rules the moment the rules were written down. Most of what a starting point evokes, no one has ever said. What the model contributes is the development, and the mechanics are the ones Section 2 drew from Wolfram: a model produces a reasonable continuation of the text it has been given, where what counts as reasonable is relative to the corpus it was fitted to (2023). A prompt is part of the text the model has been given. An articulated starting point therefore changes what there is to continue — the reasonable continuation of a stated position under stated constraints is not the reasonable continuation of a bare question — and the model makes use of what the prompt states in everything that follows: tell one of these systems something once, Wolfram observes, and it is used thereafter (2023). The continuation that results states consequences of the starting point that the starting point does not state. Section 2 said what it is for such a text to go well — the comparison it displays cites differences that bear, and would, if correct, give understanding — and whether a given continuation goes well is read off the continuation. Nothing in this makes the development a transcription. An evoked structure contains more than any text states: the rules of chess settle every fact about chess, and do not settle which theorems get written down, in what order, or to what depth, so that two writers working from the same rules produce different books, both correct, neither dictated by the rules. The mechanics mirror the structure, since the same prompt, run twice, yields different continuations (Wolfram 2023). The starting point underdetermines the development, and the gap between them is where the model's contribution lies: were there one text the prompt fixed, the output would transcribe what the person had already settled, and the instrument description would be true. The gap also leaves room for error. A development can state what does not hold in the evoked structure — a chess writer can publish a false theorem, a philosopher can misdraw the consequences of their own thought experiment, and a model can do both, along with its characteristic failure of stating fluently what nothing supports. The errors are found on the page. And an error is attributable only to a developer: no one blames the rules of chess for a false theorem, and no one's typewriter has ever made a mistake of content. Three contributions, then, and three owners: the articulated starting point is the person's; the structure it evokes, and the facts that hold there, are no one's; the text that develops them is the model's. Much in the instrument picture is true. The person writes the prompt and the prompt is authored; the person chooses which continuations to pursue and when to stop; without the person, there is the survey. What the picture adds to these truths is a description of the model — a device, like the typewriter, that fixes only what its user has already settled — and the description is what the account above denies. Every word of the novel was the author's before the typewriter touched it; the consequences a model's text states were nobody's before the text stated them. We have pressed the tool picture elsewhere, for image generators: such a system is reliably unpredictable — a prompter settles what an image is to depict, and the system settles what the image is like, so the user's control runs out where the product's properties begin (Young and Terrone 2025). What the user of a typewriter settles is the text; what the writer of a prompt settles is a starting point. The account invites an obvious enrichment of the prompt. State the position, name the rivals, list the objections and the lines along which they are to be met, and at some point, it will be said, the prompt contains the philosophy and the model is expanding what the person wrote — so that where a model's output is good, one should suspect a prompt rich enough to have done the work. But enriching a prompt enlarges the starting point without converting it into the development. A game with more rules is a bigger game, not a book of its theorems, and however much the prompt states, the consequences the output draws were not among the statements. There is a genuine limiting case — a prompt that states the comparison and the verdict, so that the continuation only rephrases — and it is identified the way everything in this paper is identified: set the output against the prompt and ask what the text states that the prompt did not. A text that states nothing beyond its prompt is a paraphrase, and owed to the person; a text that states what the prompt left unstated is a development, and the unstated part is not the person's. Which of the two a given output is, is settled by reading them together. [^1]: While preparing this paper we asked GPT-5.5 for a detailed overview of the positions an analytic philosopher might take on the meaning of life. What came back was a competent, hedged survey of the field; what did not come back was an argument for any position in it. %%add date of test%% [^2]: That these systems are mischaracterised as oracles — with the corollary that no benchmark of single-pass answers should be expected to probe the upper limits of what they can produce — has been argued from inside the practitioner literature (Janus 2022). _Draft note: the section currently ends at the rich-prompt reply; the close and the novelty question (Section 3's hand-off) are deliberately unwritten pending design. Citation flags: Janus 2022 is a pseudonymous LessWrong post — confirm citation practice; the GPT-5.5 test needs its date._ --- --- # 2 — Second half restructured (v4) — DRAFT for comparison + adversarial stages _Scope: the reply, from the grant/shift through the close. The threat paragraph below is the approved/fixed lead-in. Proposed one-clause lead-in edit (not yet applied): where generating/weighing is first introduced, attribute the division to Lipton — "Call the first generating and the second weighing — the two filters into which Lipton (2004) divides inference to the best explanation" — so the Lipton beat below need not introduce it a second time._ [fixed lead-in] On this picture the model puts candidates forward and the collaborator does the assessing, so whatever the system produces is raw material for a piece of philosophy carried out by someone else. Raw material of this kind is not itself philosophy worth reading. A list of unweighed candidates gives a reader no more to assess than a bare pronouncement that direct realism is correct, since in each case the position has merely been stated and the comparative work that would make it worth weighing has not been done. So if a model's text cannot contain a good weighing, there is no reason to regard that text as worth reading. Benchmarks of model reasoning can look as though they bear this out, since abduction is the form of reasoning at which models perform worst, with median accuracy across surveyed studies of roughly 43%, against 80% for deduction (Salimi et al. 2026).[^2] A model generates no candidates and weighs none, and that much of Floridi et al.'s account can be granted in full; nothing in what follows hands either capacity back to it. None of it, though, bears on the text the model produces. Section 1 located a text's merit in the argument it presents rather than in the history of its production, and on that footing the car-battery reply is not a list of candidates left for someone else to sift: it brings the cold morning to bear on each one and closes in favour of the battery, so the sifting the brainstorming picture reserves for the collaborator has already happened on the page. Floridi et al. half-concede the point themselves. Asked whether anything turns on the process being different when the hypothesis produced is the same, they allow that for justification it perhaps does, "but regarding the content of the hypothesis and our interpretation of it, maybe not" (2025, p. 13). Whether the sifting a text displays is any good is, by their own concession, a question about a piece of writing. What survives is the narrower claim that the weighing a text displays cannot itself be good, and meeting it takes two steps. The first is to say what makes a displayed weighing good — what a reader is assessing when a text sets rival explanations against one another and comes down in favour of one. The second is to show that a weighing of that quality can stand in a text whose producer weighed nothing. Lipton's account of inference to the best explanation supplies the first, and Wolfram's account of what continuing text involves supplies the second; together they leave room for a good weighing that nobody performed. The question Lipton presses about the second filter — what makes one selection among the candidates better than another — is ours. He separates two things the best explanation might be. It might be the likeliest, the explanation the total evidence most warrants; or the loveliest, the one that, were it correct, would yield the most understanding. "Likeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). The two can come apart: Newtonian mechanics is no longer the likeliest account of the motions it was built to explain, yet it remains as lovely an explanation of them as it ever was (p. 60). A philosophical text answers to loveliness, since what a reader assesses is whether the explanation on offer would, if correct, give more understanding than its rivals — an assessment that runs under "if correct" and so does not wait on the explanation's truth, which is why it can be carried out on the page.[^d] Loveliness is shown in the comparison of rivals. To explain is to explain why this rather than that, and that requires citing a difference between them — Lipton's Difference Condition: a cause present in the favoured case, together with the absence, in the rival, of any corresponding cause (2004, ch. 3). The kitchen makes the test concrete. "Rain rather than a burst pipe, because the window is open and the water lies beneath it" meets it: the open window and the pooling beneath it are a cause present in the rain case, and the wetting-by-the-pipe that a burst pipe would have left is absent; "rain rather than a burst pipe, because the floor is very wet" does not, since a very wet floor is what both rivals would produce and so picks out nothing between them. The two sentences share the comparative form, and only the first holds a weighing worth the name. No rule sorts them for the reader — our grasp of what makes one explanation lovelier than another is weak, and such standards as there are come from past explanations serving as exemplars and from prevailing styles of reasoning (2004, pp. 61, 139) — so telling them apart is the ordinary work of reading an argument, following what each says and asking whether it would decide the case. The bare form of explanation never did that work, and it did no more of it for human philosophers: the line between the two kitchen sentences runs between good and bad weighings in human and machine paragraphs alike.[^ml] A system trained only to continue text comes to respect constraints that were never stated for it, and Wolfram (2023) collects the cases. Trained on English, it respects English syntax though it was handed no grammar, because well-formed sentences predominate in the writing it continues and a system fitted to continue that writing comes to respect what the writing respects. Its sentences come out meaningful rather than merely well-formed, and here not even a rule was available to withhold, since no one has built a complete theory of what makes a sentence meaningful. Its syllogisms come out valid in the same way: the patterns pervade the writing — Aristotle, Wolfram whimsically suggests, first read them off many examples of rhetoric — so a system continuing that writing yields "correct inferences" of the syllogistic kind with nothing derived. Each time, the structure is present in the output while the capacity that ordinarily produces it is nowhere in the system. That loveliness answers to no stated rule is, by the same token, no barrier to its appearing in the output. A system that wrote by applying stated rules would halt exactly where no rule had been stated; these systems were handed stated rules for nothing, and what they acquire they acquire from exemplars — which, on Lipton's account, is just where the standards of loveliness reside. The precedent reaches only so far, since a syllogism has one correct completion where an abductive comparison has none; what carries across is the weaker claim, and the only one needed, that a structure can stand in a text with no trace behind it of the capacity that ordinarily produces it. The corpus these systems are trained on is mostly not philosophy, but it contains the philosophical literature, and a philosophy paper is itself a displayed comparison: a position stated, set against its rivals, and defended by the considerations taken to decide between them. Wolfram's cases stop at the sentence and the step past it is ours, but what he points to are regularities in writing rather than facts about grammar, and an argument that states a candidate, lays out its rivals and locates the difference between them is as much a regularity of the writing as syntax is. It may be said that this only redescribes the statistics — that a model reproduces the regularities of its training text, and reproducing regularities is not weighing. Lipton met an objection of the same shape, urged against his own view from Bayesianism: that once the mechanics of belief revision are given, explanatory considerations have nothing left to do. His reply was that a true account of the mechanism need not displace a true account of what it produces — to think otherwise is like holding that because a squash ball's flight obeys the laws of mechanics, "thinking about technique cannot help my squash game" — so that even granting the Bayesian mechanics, inference to the best explanation "might yet illuminate its psychology" (2004, p. 108). The mechanism at issue here is the one Floridi et al. themselves describe, and it does not have the consequence the objection needs: the patterns a model absorbs are patterns of reasoning as expressed in writing, and writing does not carry the phrasing of an explanation stripped of its organisation, since which considerations bear on which rival, and what settles the matter between them, are in the writing too. A system that learns to continue the writing learns these with the rest, and the appearance the objection is willing to grant was never separable from the organisation that makes a piece of reasoning assessable on the page. It may still be objected that syntax is one thing and inference to the best explanation another — that whatever structure next-word prediction carries, a system of this kind is too shallow for abduction, with the benchmark record reading like confirmation. But the line Wolfram draws falls elsewhere, and it comes from the same discussion that supplied the syllogism: his toy network cannot balance long sequences of parentheses, a task that demands exact procedure with no shortcut, and sophisticated formal logic should fail for the same reason, while whatever a person can take in at a glance is managed (2023). The divide these systems fail at lies between exact procedure and holistic judgement, not between the simple and the sophisticated — and weighing, on Lipton's account, sits with judgement, since no rule runs from the evidence to the loveliest explanation. Read with that line in hand, the record turns against the reading it seemed to support. A system that could recognise explanations but had nothing to draw on in making one should fail wherever production is demanded; instead the collapse concentrates where abduction has been recast as the exact recovery of a single canonical missing premise under formal constraint — the strongest model manages 21.5% on the hardest such benchmark, and most score near zero — while on open-ended tasks, where the output is judged as an explanation, the strongest models exceed 90% validity (Salimi et al. 2026).[^3] Failure tracks the demand for exact recovery, the parenthesis side of the line, and philosophical abduction does not live there. None of this gives the model any capacity Floridi et al. deny it. It infers nothing, weighs nothing, tests nothing; what it makes is text, and the text can hold what its maker did not — a candidate stated, the live rivals set in order, the difference that decides between them located. Whether a given text holds these things, and holds them well, is settled by the reading any philosophy paper is given, by the same standard and no other. A good weighing of positions the literature already contains is not yet the distinction the literature lacks; whether a model can supply that is among the questions Section 4 takes up, and what philosophy a system with no relation to the world could produce at all is the question of Section 3. [^d]: A lovely explanation is, in the terms of Dellsén et al.'s account of philosophical progress, one that puts its reader in a position to increase their understanding (2024, p. 679). [relocated from body — renumber on merge] [^ml]: Where human reference answers run out, the machine-learning literature scores generated explanations along the same lines, for consistency, parsimony and coherence as features of the output (Dalal et al. 2024; He et al. 2025). [relocated from body — renumber on merge] [^2]: These benchmarks operationalise abduction in commonsense and formal domains with crowd-labelled or mechanically checkable answers, where philosophy has no answer key, and several score generated explanations against human-written references — a comparison nothing in this paper relies on. Performance also drops under small variations to a problem (Mirzadeh et al. 2025), and Floridi et al. note that beyond the training distribution "the facade can crack" (2025, p. 9); the paper's claim is a capacity claim — that such texts can be produced — and is untouched by variation in how reliably they are. [^3]: Salimi et al.'s benchmark suite separates formally constrained missing-premise completion (ProofWriter, AbductionRules, NeuLR) from open-text explanation tasks (ART, e-CARE, UNcommonsense); the figures are from their Tables 3–6. They observe that exact-match scoring is demanding precisely where several different missing facts would each be plausible — the abductive situation — and that target structure and the size of the hypothesis space shape difficulty at least as much as subject matter does. Salimi et al. also run every benchmark with a single fixed instruction template and score one pass, while cataloguing methods — staged prompts, criticise-and-revise pipelines — that alter what models produce; what elicitation contributes is taken up in Section 4. --- # 2 — Second half restructured (v5) — hand-built to the measured published envelope _The adversarial loop oscillated (tells 137/138/22/126/139, no convergence) and its "best" round returned only one paragraph, so this v5 is built by hand from the comparison phase's per-paragraph findings. Target envelope, measured from the published Young/Terrone paragraphs: ~83–104 words per paragraph, no sentence over ~53 words, at most one heavy mark (em-dash/colon/semicolon) per paragraph, an internal short beat in the denser paragraphs, objection-and-reply within one paragraph, two-member contrasts rather than triplets. D, G, C, A and F are split to sit in the envelope, so the paragraph count rises — matching your density means more, shorter paragraphs. Ledger content preserved. Verified Floridi pages applied (concession p.12; "facade" p.9); Lipton pages unverifiable from the EPUB and left as set. Footnotes resolve to the v4 definitions above._ [fixed lead-in] On this picture the model puts candidates forward and the collaborator does the assessing, so whatever the system produces is raw material for a piece of philosophy carried out by someone else. Raw material of this kind is not itself philosophy worth reading. A list of unweighed candidates gives a reader no more to assess than a bare pronouncement that direct realism is correct, since in each case the position has merely been stated and the comparative work that would make it worth weighing has not been done. So if a model's text cannot contain a good weighing, there is no reason to regard that text as worth reading. Benchmarks of model reasoning can look as though they bear this out, since abduction is the form of reasoning at which models perform worst, with median accuracy across surveyed studies of roughly 43%, against 80% for deduction (Salimi et al. 2026).[^2] A model generates no candidates and weighs none. That much of Floridi et al.'s account can be granted, since nothing in what follows depends on the model having either capacity. None of it bears on the text the model produces. Section 1 located a text's merit in the argument it presents, not in the history of its production. By that standard the car-battery reply is not a list of candidates left for someone else to sift. It weighs the cold morning against each one and comes down in favour of the battery, so the assessment the brainstorming picture reserves for the collaborator has already been made on the page. Floridi et al. half-concede the point. Asked whether anything turns on the process being different when the hypothesis produced is the same, they allow that for justification it perhaps does, but "regarding the content of the hypothesis and our interpretation of it, maybe not" (2025, p. 12). Whether the assessment a text displays is any good is, on their own concession, a question about a piece of writing. What survives is the narrower claim that the weighing a text displays cannot itself be good. Meeting it means saying what makes a displayed weighing good — what a reader assesses when a text sets rival explanations against one another and comes down for one — and then showing that a weighing of that quality can stand in a text whose producer weighed nothing. Lipton's account of inference to the best explanation says what a good weighing is; Wolfram's account of what continuing text involves shows how one can stand where no one performed it. What makes one selection among the candidates better than another is Lipton's question as much as ours. He separates two things the best explanation might be. It might be the likeliest, the explanation the total evidence most warrants, or the loveliest, the one that, were it correct, would yield the most understanding. "Likeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). The two can come apart. Newtonian mechanics is no longer the likeliest account of the motions it was built to explain, yet it remains as lovely an explanation of them as it ever was (p. 60). A philosophical text answers to loveliness. What a reader assesses is whether the explanation on offer would, if correct, give more understanding than its rivals, and that assessment is made under "if correct". So it does not wait on the explanation's truth, and the reader can carry it out on the page.[^d] Loveliness shows in the comparison of rivals. To explain is to explain why this rather than that, and that requires citing a difference between them. This is Lipton's Difference Condition: a cause present in the favoured case, together with the absence, in the rival, of any corresponding cause (2004, ch. 3). The kitchen makes the test concrete. "Rain rather than a burst pipe, because the window is open and the water lies beneath it" meets it. The open window and the pooling beneath it are a cause present in the rain case, and the wetting-by-the-pipe that a burst pipe would have left is absent. "Rain rather than a burst pipe, because the floor is very wet" does not, since a very wet floor is what both rivals would produce and picks out nothing between them. The two sentences share the comparative form, and only the first is a genuine weighing. No rule sorts the two sentences for the reader. Our grasp of what makes one explanation lovelier than another is weak, and such standards as there are come from past explanations serving as exemplars, and from prevailing styles of reasoning (2004, pp. 61, 139). Telling the two apart is the ordinary work of reading an argument: following what each says and asking whether it would decide the case. The bare form of explanation never did that work, and it did no more of it for human philosophers. The line between the two kitchen sentences separates good weighings from bad ones in human and machine paragraphs equally.[^ml] A system trained only to continue text comes to respect constraints that were never stated, and Wolfram (2023) collects the cases. Trained on English, it respects English syntax though it was given no grammar, because well-formed sentences predominate in the writing it continues. Its sentences are mostly meaningful, not merely well-formed, and here no rule was available even to withhold, since no one has built a complete theory of what makes a sentence meaningful. Its syllogisms come out valid for the same reason. The patterns pervade the writing; Aristotle, Wolfram suggests, read them off many examples of rhetoric, and a system continuing that writing yields "correct inferences" of the syllogistic kind, with nothing derived. In each case a structure is present in the output while the capacity that ordinarily produces it is absent. That loveliness answers to no stated rule is no barrier to its appearing in the output. A system that wrote by stated rules would stop wherever no rule had been stated. These systems were given no stated rules at all, and what they acquire they acquire from exemplars — which, on Lipton's account, is just where the standards of loveliness reside. The precedent reaches only so far. A syllogism has one correct completion where an abductive comparison has none, so what survives is the weaker claim, which is all that is needed: that a structure can stand in a text with no trace of the capacity that ordinarily produces it. The corpus these systems are trained on is mostly not philosophy, but it contains the philosophical literature, and a philosophy paper is itself a displayed comparison: a position stated, set against its rivals, and defended by the considerations taken to decide between them. Wolfram's cases stop at the sentence, and the step beyond it is one we are taking. But his observations concern regularities in writing rather than facts about grammar, and an argument that states a candidate, sets out its rivals and locates what divides them is as much a regularity of the writing as syntax is. It may be said that this only redescribes the statistics: a model reproduces the regularities of its training text, and reproducing regularities is not weighing. Lipton met an objection of the same shape, raised against his own view from Bayesianism, that once the mechanics of belief revision are given, explanatory considerations have nothing left to do. His reply was that a true account of the mechanism need not displace a true account of what it produces. To think otherwise is like holding that because a squash ball's flight obeys the laws of mechanics, "thinking about technique cannot help my squash game". Even granting the Bayesian mechanics, inference to the best explanation "might yet illuminate its psychology" (2004, p. 108). The mechanism at issue here is the one Floridi et al. themselves describe, and on that description the objection does not go through. The patterns a model absorbs are patterns of reasoning as expressed in writing, and writing does not carry the phrasing of an explanation stripped of its organisation. Which considerations bear on which rival, and what settles the matter between them, are in the writing too, and a system that learns to continue the writing learns these with the rest. The appearance the objection grants was never separable from the organisation that makes a piece of reasoning assessable on the page. It may still be objected that syntax is one thing and inference to the best explanation another, and that whatever structure next-word prediction carries, a system of this kind is too shallow for abduction, with the benchmark record reading like confirmation. But the line Wolfram draws falls elsewhere, and it comes from the same discussion that supplied the syllogism. His simple network cannot balance long sequences of parentheses, a task that demands exact procedure with no shortcut, and sophisticated formal logic should fail for the same reason, while it manages whatever a person can take in at a glance. The divide these systems fail at runs between exact procedure and holistic judgement; the contrast between the simple and the sophisticated is beside the point. Weighing, on Lipton's account, sits with judgement, since no rule runs from the evidence to the loveliest explanation. On that distinction, the record tells against the reading it seemed to support. A system that could recognise explanations but had nothing to draw on in making one should fail wherever production is demanded. Instead the collapse concentrates in one place: where abduction has been recast as the exact recovery of a single canonical missing premise under formal constraint. There the strongest model manages 21.5% on the hardest such benchmark, and most score near zero. On open-ended tasks, where the output is judged as an explanation, the strongest models exceed 90% validity (Salimi et al. 2026).[^3] Failure tracks the demand for exact recovery — the parenthesis side of the line — and philosophical abduction is not of that kind. None of this gives the model any capacity Floridi et al. deny it. It infers nothing, and it weighs and tests nothing; what it makes is text, and the text can hold what its maker did not: a candidate stated, its live rivals set in order, and the difference that decides between them. Whether a given text holds these things, and holds them well, is settled by the reading any philosophy paper is given, by the same standard. A good weighing of positions the literature already contains is not yet the distinction the literature lacks. Whether a model can supply that is taken up in Section 4; Section 3 asks what philosophy a system with no relation to the world could produce at all. --- --- # 2 — Second half (v6) — strict-gate pass on v4 _v4 put through the strict deterministic vocabulary gate (your corpus + only the words approved today) and the published envelope. The ten tell-words are replaced with attested words; nothing else in your wording is swapped. Long paragraphs split and over-long sentences broken at punctuation already present, to sit near ~83–104 words with sentences under ~53. Ledger content held: all obligations, the six verbatim quotations, every figure and citation. Body gates clean; footnotes retained verbatim from v4. Tells replaced: whimsically→cut, recast→set up as, concentrates→falls, separable→independent of, reserves→leaves to, reside→lie, footing→by that standard, stripped→without, predominate→make up most of, prevailing→common._ On this picture the model puts candidates forward and the collaborator does the assessing, so whatever the system produces is raw material for a piece of philosophy carried out by someone else. Raw material of this kind is not itself philosophy worth reading. A list of unweighed candidates gives a reader no more to assess than a bare pronouncement that direct realism is correct, since in each case the position has merely been stated and the comparative work that would make it worth weighing has not been done. So if a model's text cannot contain a good weighing, there is no reason to regard that text as worth reading. Benchmarks of model reasoning can look as though they bear this out, since abduction is the form of reasoning at which models perform worst, with median accuracy across surveyed studies of roughly 43%, against 80% for deduction (Salimi et al. 2026).[^2] A model generates no candidates and weighs none, and that much of Floridi et al.'s account can be granted in full; nothing in what follows hands either capacity back to it. None of it, though, bears on the text the model produces. Section 1 located a text's merit in the argument it presents rather than in the history of its production. By that standard the car-battery reply is not a list of candidates left for someone else to sift: it brings the cold morning to bear on each one and closes in favour of the battery, so the sifting the brainstorming picture leaves to the collaborator has already happened on the page. Floridi et al. half-concede the point themselves. Asked whether anything turns on the process being different when the hypothesis produced is the same, they allow that for justification it perhaps does, "but regarding the content of the hypothesis and our interpretation of it, maybe not" (2025, p. 13). Whether the sifting a text displays is any good is, by their own concession, a question about a piece of writing. What survives is the narrower claim that the weighing a text displays cannot itself be good, and meeting it takes two steps. The first is to say what makes a displayed weighing good — what a reader is assessing when a text sets rival explanations against one another and comes down in favour of one. The second is to show that a weighing of that quality can stand in a text whose producer weighed nothing. Lipton's account of inference to the best explanation supplies the first, and Wolfram's account of what continuing text involves supplies the second; together they leave room for a good weighing that nobody performed. The question Lipton presses about the second filter — what makes one selection among the candidates better than another — is ours. He separates two things the best explanation might be. It might be the likeliest, the explanation the total evidence most warrants; or the loveliest, the one that, were it correct, would yield the most understanding. "Likeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). The two can come apart: Newtonian mechanics is no longer the likeliest account of the motions it was built to explain, yet it remains as lovely an explanation of them as it ever was (p. 60). A philosophical text answers to loveliness, since what a reader assesses is whether the explanation on offer would, if correct, give more understanding than its rivals. That assessment runs under "if correct" and so does not wait on the explanation's truth, which is why it can be carried out on the page.[^d] Loveliness is shown in the comparison of rivals. To explain is to explain why this rather than that, and that requires citing a difference between them — Lipton's Difference Condition: a cause present in the favoured case, together with the absence, in the rival, of any corresponding cause (2004, ch. 3). The kitchen makes the test concrete. "Rain rather than a burst pipe, because the window is open and the water lies beneath it" meets it: the open window and the pooling beneath it are a cause present in the rain case, and the wetting-by-the-pipe that a burst pipe would have left is absent. "Rain rather than a burst pipe, because the floor is very wet" does not, since a very wet floor is what both rivals would produce and so picks out nothing between them. The two sentences share the comparative form, and only the first holds a weighing worth the name. No rule sorts them for the reader. Our grasp of what makes one explanation lovelier than another is weak, and such standards as there are come from past explanations serving as exemplars and from common styles of reasoning (2004, pp. 61, 139). So telling them apart is the ordinary work of reading an argument, following what each says and asking whether it would decide the case. The bare form of explanation never did that work, and it did no more of it for human philosophers: the line between the two kitchen sentences runs between good and bad weighings in human and machine paragraphs alike.[^ml] A system trained only to continue text comes to respect constraints that were never stated for it, and Wolfram (2023) collects the cases. Trained on English, it respects English syntax though it was handed no grammar, because well-formed sentences make up most of the writing it continues, and a system fitted to continue that writing comes to respect what the writing respects. Its sentences come out meaningful rather than merely well-formed, and here not even a rule was available to withhold, since no one has built a complete theory of what makes a sentence meaningful. Its syllogisms come out valid in the same way: the patterns pervade the writing — Aristotle, Wolfram suggests, first read them off many examples of rhetoric — so a system continuing that writing yields "correct inferences" of the syllogistic kind with nothing derived. That loveliness answers to no stated rule is, by the same token, no barrier to its appearing in the output. A system that wrote by applying stated rules would halt exactly where no rule had been stated. These systems were handed stated rules for nothing, and what they acquire they acquire from exemplars — which, on Lipton's account, is just where the standards of loveliness lie. The precedent reaches only so far, since a syllogism has one correct completion where an abductive comparison has none. What carries across is the weaker claim, and the only one needed, that a structure can stand in a text with no trace behind it of the capacity that ordinarily produces it. The corpus these systems are trained on is mostly not philosophy, but it contains the philosophical literature, and a philosophy paper is itself a displayed comparison: a position stated, set against its rivals, and defended by the considerations taken to decide between them. Wolfram's cases stop at the sentence and the step past it is ours, but what he points to are regularities in writing rather than facts about grammar, and an argument that states a candidate, lays out its rivals and locates the difference between them is as much a regularity of the writing as syntax is. It may be said that this only redescribes the statistics — that a model reproduces the regularities of its training text, and reproducing regularities is not weighing. Lipton met an objection of the same shape, urged against his own view from Bayesianism: that once the mechanics of belief revision are given, explanatory considerations have nothing left to do. His reply was that a true account of the mechanism need not displace a true account of what it produces. To think otherwise is like holding that because a squash ball's flight obeys the laws of mechanics, "thinking about technique cannot help my squash game"; even granting the Bayesian mechanics, inference to the best explanation "might yet illuminate its psychology" (2004, p. 108). The mechanism at issue here is the one Floridi et al. themselves describe, and it does not have the consequence the objection needs. The patterns a model absorbs are patterns of reasoning as expressed in writing, and writing does not carry the phrasing of an explanation without its organisation: which considerations bear on which rival, and what settles the matter between them, are in the writing too. A system that learns to continue the writing learns these with the rest, and the appearance the objection is willing to grant was never independent of the organisation that makes a piece of reasoning assessable on the page. It may still be objected that syntax is one thing and inference to the best explanation another — that whatever structure next-word prediction carries, a system of this kind is too shallow for abduction, with the benchmark record reading like confirmation. But the line Wolfram draws falls elsewhere, and it comes from the same discussion that supplied the syllogism. His toy network cannot balance long sequences of parentheses, a task that demands exact procedure with no shortcut; sophisticated formal logic should fail for the same reason, while whatever a person can take in at a glance is managed (2023). The divide these systems fail at lies between exact procedure and holistic judgement, not between the simple and the sophisticated. Weighing, on Lipton's account, sits with judgement, since no rule runs from the evidence to the loveliest explanation. Read with that line in hand, the record turns against the reading it seemed to support. A system that could recognise explanations but had nothing to draw on in making one should fail wherever production is demanded. Instead the collapse falls where abduction has been set up as the exact recovery of a single canonical missing premise under formal constraint. The strongest model manages 21.5% on the hardest such benchmark, and most score near zero, while on open-ended tasks, where the output is judged as an explanation, the strongest models exceed 90% validity (Salimi et al. 2026).[^3] Failure tracks the demand for exact recovery, the parenthesis side of the line, and philosophical abduction does not live there. None of this gives the model any capacity Floridi et al. deny it. It infers nothing, weighs nothing, tests nothing; what it makes is text, and the text can hold what its maker did not — a candidate stated, the live rivals set in order, the difference that decides between them located. Whether a given text holds these things, and holds them well, is settled by the reading any philosophy paper is given, by the same standard and no other. A good weighing of positions the literature already contains is not yet the distinction the literature lacks. Whether a model can supply that is among the questions Section 4 takes up, and what philosophy a system with no relation to the world could produce at all is the question of Section 3. [^d]: A lovely explanation is, in the terms of Dellsén et al.'s account of philosophical progress, one that puts its reader in a position to increase their understanding (2024, p. 679). [relocated from body — renumber on merge] [^ml]: Where human reference answers run out, the machine-learning literature scores generated explanations along the same lines, for consistency, parsimony and coherence as features of the output (Dalal et al. 2024; He et al. 2025). [relocated from body — renumber on merge] [^2]: These benchmarks operationalise abduction in commonsense and formal domains with crowd-labelled or mechanically checkable answers, where philosophy has no answer key, and several score generated explanations against human-written references — a comparison nothing in this paper relies on. Performance also drops under small variations to a problem (Mirzadeh et al. 2025), and Floridi et al. note that beyond the training distribution "the facade can crack" (2025, p. 9); the paper's claim is a capacity claim — that such texts can be produced — and is untouched by variation in how reliably they are. [^3]: Salimi et al.'s benchmark suite separates formally constrained missing-premise completion (ProofWriter, AbductionRules, NeuLR) from open-text explanation tasks (ART, e-CARE, UNcommonsense); the figures are from their Tables 3–6. They observe that exact-match scoring is demanding precisely where several different missing facts would each be plausible — the abductive situation — and that target structure and the size of the hypothesis space shape difficulty at least as much as subject matter does. Salimi et al. also run every benchmark with a single fixed instruction template and score one pass, while cataloguing methods — staged prompts, criticise-and-revise pipelines — that alter what models produce; what elicitation contributes is taken up in Section 4. --- # 2 — Second half (v7) — subtraction pass on v4 (inline, deterministic-gated) _Method: your v4, subtracted not regenerated. Only the ten tells replaced with your attested words (whimsically→cut, recast→set up as, concentrates→falls, separable→independent of, reserves→leaves to, reside→lie, footing→by that standard, stripped→without, predominate→make up most of, prevailing→common); sentences over ~53 words broken at punctuation already there; paragraphs re-cut to your measured envelope (Growing the Image: ~60–120w, up to ~165); one asyndeton softened ("infers nothing, weighs nothing, tests nothing" → "…and tests nothing") to undrum the rhythm while keeping all three. No other wording changed. Lead-in frozen; footnotes verbatim. Vocab gate silent; every quote/figure/citation retained._ On this picture the model puts candidates forward and the collaborator does the assessing, so whatever the system produces is raw material for a piece of philosophy carried out by someone else. Raw material of this kind is not itself philosophy worth reading. A list of unweighed candidates gives a reader no more to assess than a bare pronouncement that direct realism is correct, since in each case the position has merely been stated and the comparative work that would make it worth weighing has not been done. So if a model's text cannot contain a good weighing, there is no reason to regard that text as worth reading. Benchmarks of model reasoning can look as though they bear this out, since abduction is the form of reasoning at which models perform worst, with median accuracy across surveyed studies of roughly 43%, against 80% for deduction (Salimi et al. 2026).[^2] A model generates no candidates and weighs none, and that much of Floridi et al.'s account can be granted in full; nothing in what follows hands either capacity back to it. None of it, though, bears on the text the model produces. Section 1 located a text's merit in the argument it presents rather than in the history of its production. By that standard the car-battery reply is not a list of candidates left for someone else to sift: it brings the cold morning to bear on each one and closes in favour of the battery. The sifting the brainstorming picture leaves to the collaborator has already happened on the page. Floridi et al. half-concede the point themselves. Asked whether anything turns on the process being different when the hypothesis produced is the same, they allow that for justification it perhaps does, "but regarding the content of the hypothesis and our interpretation of it, maybe not" (2025, p. 13). Whether the sifting a text displays is any good is, by their own concession, a question about a piece of writing. What survives is the narrower claim that the weighing a text displays cannot itself be good, and meeting it takes two steps. The first is to say what makes a displayed weighing good — what a reader is assessing when a text sets rival explanations against one another and comes down in favour of one. The second is to show that a weighing of that quality can stand in a text whose producer weighed nothing. Lipton's account of inference to the best explanation supplies the first, and Wolfram's account of what continuing text involves supplies the second; together they leave room for a good weighing that nobody performed. The question Lipton presses about the second filter — what makes one selection among the candidates better than another — is ours. He separates two things the best explanation might be. It might be the likeliest, the explanation the total evidence most warrants; or the loveliest, the one that, were it correct, would yield the most understanding. "Likeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). The two can come apart: Newtonian mechanics is no longer the likeliest account of the motions it was built to explain, yet it remains as lovely an explanation of them as it ever was (p. 60). A philosophical text answers to loveliness, since what a reader assesses is whether the explanation on offer would, if correct, give more understanding than its rivals. That assessment runs under "if correct" and so does not wait on the explanation's truth, which is why it can be carried out on the page.[^d] Loveliness is shown in the comparison of rivals. To explain is to explain why this rather than that, and that requires citing a difference between them — Lipton's Difference Condition: a cause present in the favoured case, together with the absence, in the rival, of any corresponding cause (2004, ch. 3). The kitchen makes the test concrete. "Rain rather than a burst pipe, because the window is open and the water lies beneath it" meets it: the open window and the pooling beneath it are a cause present in the rain case, and the wetting-by-the-pipe that a burst pipe would have left is absent. "Rain rather than a burst pipe, because the floor is very wet" does not, since a very wet floor is what both rivals would produce and so picks out nothing between them. The two sentences share the comparative form, and only the first holds a weighing worth the name. No rule sorts them for the reader. Our grasp of what makes one explanation lovelier than another is weak, and such standards as there are come from past explanations serving as exemplars and from common styles of reasoning (2004, pp. 61, 139). So telling them apart is the ordinary work of reading an argument, following what each says and asking whether it would decide the case. The bare form of explanation never did that work, and it did no more of it for human philosophers: the line between the two kitchen sentences runs between good and bad weighings in human and machine paragraphs alike.[^ml] A system trained only to continue text comes to respect constraints that were never stated for it, and Wolfram (2023) collects the cases. Trained on English, it respects English syntax though it was handed no grammar, because well-formed sentences make up most of the writing it continues, and a system fitted to continue that writing comes to respect what the writing respects. Its sentences come out meaningful rather than merely well-formed, and here not even a rule was available to withhold, since no one has built a complete theory of what makes a sentence meaningful. Its syllogisms come out valid in the same way: the patterns pervade the writing — Aristotle, Wolfram suggests, first read them off many examples of rhetoric — so a system continuing that writing yields "correct inferences" of the syllogistic kind with nothing derived. That loveliness answers to no stated rule is, by the same token, no barrier to its appearing in the output. A system that wrote by applying stated rules would halt exactly where no rule had been stated. These systems were handed stated rules for nothing, and what they acquire they acquire from exemplars — which, on Lipton's account, is just where the standards of loveliness lie. The precedent reaches only so far, since a syllogism has one correct completion where an abductive comparison has none. What carries across is the weaker claim, and the only one needed, that a structure can stand in a text with no trace behind it of the capacity that ordinarily produces it. The corpus these systems are trained on is mostly not philosophy, but it contains the philosophical literature, and a philosophy paper is itself a displayed comparison: a position stated, set against its rivals, and defended by the considerations taken to decide between them. Wolfram's cases stop at the sentence and the step past it is ours, but what he points to are regularities in writing rather than facts about grammar. An argument that states a candidate, lays out its rivals and locates the difference between them is as much a regularity of the writing as syntax is. It may be said that this only redescribes the statistics — that a model reproduces the regularities of its training text, and reproducing regularities is not weighing. Lipton met an objection of the same shape, urged against his own view from Bayesianism: that once the mechanics of belief revision are given, explanatory considerations have nothing left to do. His reply was that a true account of the mechanism need not displace a true account of what it produces. To think otherwise is like holding that because a squash ball's flight obeys the laws of mechanics, "thinking about technique cannot help my squash game"; even granting the Bayesian mechanics, inference to the best explanation "might yet illuminate its psychology" (2004, p. 108). The mechanism at issue here is the one Floridi et al. themselves describe, and it does not have the consequence the objection needs. The patterns a model absorbs are patterns of reasoning as expressed in writing, and writing does not carry the phrasing of an explanation without its organisation: which considerations bear on which rival, and what settles the matter between them, are in the writing too. A system that learns to continue the writing learns these with the rest, and the appearance the objection is willing to grant was never independent of the organisation that makes a piece of reasoning assessable on the page. It may still be objected that syntax is one thing and inference to the best explanation another — that whatever structure next-word prediction carries, a system of this kind is too shallow for abduction, with the benchmark record reading like confirmation. But the line Wolfram draws falls elsewhere, and it comes from the same discussion that supplied the syllogism: his toy network cannot balance long sequences of parentheses, a task that demands exact procedure with no shortcut. Sophisticated formal logic should fail for the same reason, while whatever a person can take in at a glance is managed (2023). The divide these systems fail at lies between exact procedure and holistic judgement, not between the simple and the sophisticated. Weighing, on Lipton's account, sits with judgement, since no rule runs from the evidence to the loveliest explanation. Read with that line in hand, the record turns against the reading it seemed to support. A system that could recognise explanations but had nothing to draw on in making one should fail wherever production is demanded. Instead the collapse falls where abduction has been set up as the exact recovery of a single canonical missing premise under formal constraint. The strongest model manages 21.5% on the hardest such benchmark, and most score near zero, while on open-ended tasks, where the output is judged as an explanation, the strongest models exceed 90% validity (Salimi et al. 2026).[^3] Failure tracks the demand for exact recovery, the parenthesis side of the line, and philosophical abduction does not live there. None of this gives the model any capacity Floridi et al. deny it. It infers nothing, weighs nothing, and tests nothing; what it makes is text, and the text can hold what its maker did not — a candidate stated, the live rivals set in order, the difference that decides between them located. Whether a given text holds these things, and holds them well, is settled by the reading any philosophy paper is given, by the same standard and no other. A good weighing of positions the literature already contains is not yet the distinction the literature lacks. Whether a model can supply that is among the questions Section 4 takes up, and what philosophy a system with no relation to the world could produce at all is the question of Section 3. [^d]: A lovely explanation is, in the terms of Dellsén et al.'s account of philosophical progress, one that puts its reader in a position to increase their understanding (2024, p. 679). [relocated from body — renumber on merge] [^ml]: Where human reference answers run out, the machine-learning literature scores generated explanations along the same lines, for consistency, parsimony and coherence as features of the output (Dalal et al. 2024; He et al. 2025). [relocated from body — renumber on merge] [^2]: These benchmarks operationalise abduction in commonsense and formal domains with crowd-labelled or mechanically checkable answers, where philosophy has no answer key, and several score generated explanations against human-written references — a comparison nothing in this paper relies on. Performance also drops under small variations to a problem (Mirzadeh et al. 2025), and Floridi et al. note that beyond the training distribution "the facade can crack" (2025, p. 9); the paper's claim is a capacity claim — that such texts can be produced — and is untouched by variation in how reliably they are. [^3]: Salimi et al.'s benchmark suite separates formally constrained missing-premise completion (ProofWriter, AbductionRules, NeuLR) from open-text explanation tasks (ART, e-CARE, UNcommonsense); the figures are from their Tables 3–6. They observe that exact-match scoring is demanding precisely where several different missing facts would each be plausible — the abductive situation — and that target structure and the size of the hypothesis space shape difficulty at least as much as subject matter does. Salimi et al. also run every benchmark with a single fixed instruction template and score one pass, while cataloguing methods — staged prompts, criticise-and-revise pipelines — that alter what models produce; what elicitation contributes is taken up in Section 4.