# 1. The Challenge from Authorship
In this section we address what we might call the _challenge from authorship_: the idea that philosophy is something that only persons, or at least minds, can produce. This view has not, to our knowledge, been explicitly defended in just this form, but it gives shape to an intuition that many philosophers may have: philosophy is a person-only domain. An imperfect comparison is with art: One might deny that an image generated by an AI system~~, at least in the familiar prompt-and-output cases,~~ is an artwork, because no artist exercises the relevant kind of intentional control over its production.[^1] One might think, for similar reasons, that philosophy can only be done by people: no text produced by an LLM can be a work of philosophy, because no philosopher lies behind it.
Similar to the study of art, the study of philosophy is often organised around individuals: undergraduates take courses on Kant's ethics or Lewis's metaphysics, ~~and at more advanced levels there are specialists in,~~ and conferences are devoted to, the work of particular philosophers. Physics students, on the other hand, are taught Newtonian mechanics from a current textbook, and the course loses nothing if Newton's own writing is never looked at. In the sciences, then, what a text contributes can be carried by other texts. In philosophy, the contribution and its original presentation are harder to prise apart, and we might take this as evidence that a philosophical work is bound to the activity of the particular person who produced it, in a way that the sciences are not.
We will now try to make this challenge from authorship more precise, by considering how far Davies' _performance_ theory of art transposes to philosophy. Davies writes:
> [T]he work – what the artist achieves – is the process eventuating in that product. Works themselves are neither structures nor objects simpliciter, nor are they contextualized structures or objects. They are, rather, intentionally guided generative performances that eventuate in contextualized structures or objects (or events, as we shall see) – performances completed by what I am terming a focus of appreciation. (2004, p. 97)
On Davies' view, when a painter paints a picture, the canvas is what we attend to, but it is not the work. The work is the artist's intentionally guided activity in producing the canvas; the canvas is ~~the work's focus of appreciation,~~ "the focus of our appreciative interest in the work" (2004, p. 150). Provenance, on this view, does more than supply context: facts about how the object came into being help determine what the work is and what is properly appreciated in it. %%this seems a bit compressed, maybe a bit more detail here, and a bit more succinctness in the next paragraph with the examples%%
Davies supports the relocation of the work from the surface to the activity with cases in which two surfaces, or two texts, would look the same and yet differ as works, and the cases are of two kinds.%%not a very clear sentence: convoluted%% In one kind there is no performance at all: an instance of the verbal structure of _Kubla Khan_ might be generated by desert wind, or a monkey at a typewriter. A theorist who identifies the poem with its verbal structure must then either count these as instances of Coleridge's work or explain why not (2004, p. 102). In the other kind there is a performance, but not the one the surface was taken to give access to: van Meegeren's _The Disciples at Emmaus_ was presented as a newly discovered Vermeer, and what was appreciated under that description was not the achievement the canvas in fact issued from.[^2] Only the first kind bears on our question, since an LLM text would be a case of absent performance, not of misattributed performance. If Davies is right, the surface does not by itself settle the work. %%does all this need to be said about forgeries? I suspect it is not relevant but I am open to you persuading me i am wrong%%
Here is what the analogous proposal for philosophy would be. A philosophical text is not itself the philosophical work: the text is the product of a person's philosophising, and reading it is a way of engaging with that prior activity. The challenge this poses is constitutive:%%not how i write%%the activity is treated not as what causes a philosophical work to exist but as part of what the work is. If no one has philosophised, there is no work to which the text gives access, however the text reads — an LLM text would stand to philosophy as the wind-made _Kubla Khan_ stands to poetry. %%the middle of this paragraph could be clearer and better written. %%
Should the transposition be accepted? We do not think it should. What entitles Davies to relocate the work into the performance is an evaluative fact about art: surfaces that look the same can differ in artistic value, as the forged Vermeer and a genuine one do,%%if all that forgery talk is removed, just do the wind example%% and provenance is part of what the difference consists in. %%is this opening of the paragraph redundant because of what has already been said?%% The transposition therefore commits its defender to the corresponding claim about philosophy: that two texts containing the same argument could differ in philosophical merit. However, if two texts contain the same argument, including the same inferential moves, the same considerations count for and against them: whether the argument is valid and whether the objections are answered are questions about the texts' contents, and two texts with the same contents receive the same answers. Their philosophical merit does not vary with the route by which the words came to be written.
The discipline's evaluative practice is built on the same denial. Journals strip author information from submissions before review because facts about authorship are treated as potential sources of distortion; if texts with the same contents could differ in merit, anonymising would discard evaluatively relevant information, and review would not be designed this way. The grounds for the judgement lie in the argument as presented, not in the history of its production. Nor is the author-centred teaching noted earlier in tension with this. That philosophy is taught through Kant rather than through summaries of Kant is a fact about where the discipline's contributions live, not about how they are evaluated: the _Groundwork_ is assigned because reading it puts a student somewhere no summary has yet put one, which is a fact about the text. %%not how i write, and you should give the fuller title of the book if you mention it. Also, this ending Seems very magazine-like rather than philosophical%%
Dellsén et al. (2024) hold that philosophical progress is "for-whom" rather than "by-whom": it consists in putting people in a position to increase their understanding, usually by making philosophical ideas publicly available (p. 679). On this account the discipline's success-conditions locate the contribution in the public text, and a view that locates the philosophy behind the text, in the process by which it came about, misplaces it. The public text is not a dispensable trace of philosophy; it is where the philosophical contribution becomes fully available.%%seems a bit compressed, and unconnected with the paragraphs leading up to it%%
A performance theorist can hold the line%%not how i write%%: no philosophising, no work, whatever the text contains. The position can be granted in full%%meta-commentative wank%%, because the thesis of this paper does not use the notion it restricts %%meta-commentative cumstain%%. Suppose the desert wind assembled not _Kubla Khan_ but a sound argument against enactivist approaches to perception. No one would deserve credit for it; ~~there would be no achievement to admire, and no entry for anyone's bibliography.~~ A reader who worked through it would nonetheless confront a thesis and the arguments marshalled in its defence, and would be in a position to answer or extend it ~~— the position a philosophical text puts its readers in when it is worth their time.~~ %%the crossed oout part is shallow and shit, replace with substance%%Whether such a text is a _work_ may then be reserved for texts with performances behind them; what cannot be reserved is the text's being worth reading, since everything that judgement answers to is on the page. %%not a very clear paragraph, what is its function supposed to be?%%
The challenge from authorship therefore fails, and the doubt it leaves standing is of a different kind.%%not how i write, metacommentative wanking again%% That a text came from an LLM cannot disqualify it;%%stubby cunty sentence%% nothing so far shows that LLMs can produce such texts. If a parrot produced what sounds like a philosophical argument, this argument would not be disqualified by its source — but parrots produce no arguments, because they lack the capacities arguing requires.%%this paragraph is so compressed as to be meaningless wank%% The doubt about LLMs, in their current state, is of this kind: that they lack capacities that producing philosophy worth reading requires. Section 2 takes up the claim that they cannot perform the inference on which philosophical theorising runs; Section 3 the claims that they stand in no relation to the world and have no experience. A further challenge grants a worthwhile text and asks whose work it is, the model's or the prompting person's; it arises only if the capacity challenges fail, and we take it last. %%there is a good chance most of this paragraph can be cut, the parrot and the capcity stuff should be at the beginning of section 2, and better written.%%
[^1]: This is not to deny that systems of this kind can produce beautiful images; we return to image generation in Section 4.
[^2]: Han van Meegeren, the Dutch forger exposed in 1945; his _The Disciples at Emmaus_ was authenticated as a Vermeer and celebrated before the forgery came to light.
---
---
---
# 2. The Challenge from Abduction — v3 (rebuilt from first principles)
In the previous section, we argued against the idea that LLMs cannot produce philosophy worth reading simply because they are not human. In this section and the next we shall consider a different form of challenge: even if LLMs cannot be ruled out of producing philosophy worth reading tout court, they lack particular _capacities_ that producing it requires. If a parrot uttered a sequence of sounds that happened to form a philosophical argument, the argument would be none the worse for its source; yet parrots' powers of mimicry do not extend to producing strings of sounds so complex as to make up a philosophical argument. In this section we address one capacity challenge, which we will call the _challenge from abduction_. In the next we shall look at two more: phenomenological experience and contact with the world.
Abduction, or inference to the best explanation, is reasoning from a body of evidence to the hypothesis that would best explain it. In a deductive argument the premises fix the conclusion: if all men are mortal and Socrates is a man, then Socrates is mortal, and there is no wriggle room. In abduction the evidence settles less. Now, imagine walking into your kitchen and finding the floor wet. What has happened? The wet floor does not determine the answer in the way the two premises determined Socrates' mortality: a burst pipe would have left the floor wet, and so would a spilled bucket. But, given that the window is open, the water is under the window, and it rained last night, rain coming through the window seems the most plausible answer.
To reason in this way, deciding what best explains a set of facts, is common in everyday life and in the sciences alike. A scientist chooses one theory over another when it explains the same data more simply: Copernicus's model of the solar system was preferred to Ptolemy's because it explained the observed planetary motions without the elaborate epicycles the older model required.
Williamson argues that philosophy is continuous with the sciences, and that its theories are to be chosen by the same abductive standards (2007; 2021, p. 351 %%check page%%). In philosophy too there are data that a candidate theory must accommodate — intuitions about cases, and the phenomena of the domain itself — and rival theories that would each accommodate them at different costs. The theory to prefer is the one that would, if true, best explain the data. What makes one explanation better than another, on this account, is a matter of explanatory virtue: a good philosophical theory is, in Williamson's words, "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated", and should "combine simplicity with strength" (2021, p. 354 %%check page%%). That theories are weighed by such comparative and explanatory virtues need not rest on a science-modelled conception of philosophy: Bengson, Cuneo and Shafer-Landau (2022) argue that the assessment of rival theories by their explanatory and unifying merits is a constraint on sound philosophical method as such. This conception of philosophy is widely held (Sider 2011; Paul 2012; Dellsén et al. 2024), though not universally (Bueno and Shalkowski 2020; Thomasson 2015), and we shall assume it in what follows.
On this account a philosophical text offers its reader a choice of theory displayed — a position, its rivals, and the case for preferring it — so that whether the text is worth reading and whether it contains a good weighing travel together.
If the capacity for abduction is what is required to produce worthwhile philosophy, we can ask whether LLMs possess it. Floridi et al. (2025) argue that they do not, describing what such models do instead as zeroth-order abduction:
> LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025, p. 9)
An LLM, on this account, has "a stochastic core and an abductive appearance" (2025, p. 3). The model is trained to predict which words are likely to follow which, and it produces the continuation its training makes probable; it aims at the likely continuation, not at the truth. The appearance comes from what the training data have passed on: models have "absorbed patterns of human abductive reasoning as expressed in writing" (p. 10) — how explanations are typically phrased, which causes are typically offered for which effects. What is inherited, on their account, is the look of the reasoning, not the reasoning itself.[^1]
Explaining the wet kitchen floor involved two separable activities: coming up with candidate explanations — the burst pipe, the spilled bucket, the rain — and settling which of them the open window and the position of the water favoured. Call the first _generating_ and the second _weighing_. Floridi et al.'s position is that a model does neither, however much its text exhibits both. Asked why a car might not start on a cold morning, a model replies that a weak battery is one possibility, since cold reduces a battery's efficiency; that thickened engine oil is another, since a cold engine is harder to turn over; and that, "[b]ased on your description, the battery is the most likely explanation" (2025, p. 11). The offering of candidates here is not generating, on their reading: the model is not reasoning about causes from the user's case but reproducing the causes such explanations typically cite (p. 9). And the singling out is not weighing: the verdict reproduces how explanations of this kind typically end, and where an output marks a genuine point of difference between two hypotheses, that is something the model has seen stated, not something it has derived anew (p. 14).
Floridi et al. draw the consequence themselves:
> In essence, LLMs function like brainstorming assistants that toss out ideas without filtering for quality. After all, they work like statistical interfaces to an enormous amount of data accumulated for millennia by generations. A cautious human collaborator can sift through and assess them. (2025, p. 12)
On this picture the weighing always remains with the person: the model supplies candidates, and assessing them is the collaborator's work. Whatever such a system produces is raw material for philosophy done by someone else, and raw material is not philosophy worth reading — a list of unweighed candidates is no more worth reading than a bare pronouncement that direct realism is correct. The challenge follows: if a model's text cannot contain a good weighing, there is no reason to regard it as worth reading. The benchmark record can seem to agree, since abduction is the form of reasoning at which models perform worst, with median accuracy across surveyed studies of roughly 43%, against 80% for deduction (Salimi et al. 2026).[^2]
Everything in this account of the producer can be granted. The model generates nothing and weighs nothing, and nothing in what follows returns either capacity to it. What the account does not settle is anything about the texts. Section 1 fixed where the grounds of a text's merit lie — in the argument as presented, not in the history of its production — and by that standard the car-battery reply asks to be read rather than explained away. It is not a list of candidates awaiting a collaborator: it brings the cold morning to bear on each candidate and closes in favour of one, so the sifting the brainstorming picture reserves for the person is on the page. Whether that displayed sifting is any good is a question about a piece of writing.
What remains of the challenge is the claim that the weighing a model's text displays cannot be good. Against this we will argue that the challenge underestimates what the inherited look of reasoning includes. Two things need showing: what makes a weighing displayed in a text good, and how a good weighing can come to be displayed in text that nobody weighed. Lipton's account of inference to the best explanation supplies the first; Wolfram's account of what continuing text involves supplies the second.
The division of the kitchen's work into generating and weighing is Lipton's own: inference to the best explanation, on his account, runs on two filters, one that supplies the plausible candidates and a second that selects among them (2004, p. 59 %%pin page%%) — and his question about the second filter is ours, namely what makes the selection good. The best explanation can be understood as the likeliest, the one most warranted by the total evidence, or as the loveliest, the one which, if correct, would provide the most understanding: "[l]ikeliness speaks of truth; loveliness of potential understanding" (p. 59). The two come apart: Newtonian mechanics is no longer the likeliest account of the observations it was built on, but it remains as lovely an explanation of them as it ever was (p. 60). Loveliness is the standard a philosophical text answers to: what readers assess is whether the explanation offered would, if true, give understanding, and give more of it than the rivals considered. Because the assessment runs under "if correct", it does not wait on verification, and a reader can conduct it on the page. On Dellsén et al.'s account, philosophical progress consists in putting people in a position to increase their understanding (2024, p. 679); a lovely explanation puts its reader in exactly that position.
Where loveliness shows itself, on Lipton's analysis, is in the comparison of rivals. Explanation is contrastive: we explain why this rather than that, and doing so requires citing a difference between the two — his Difference Condition — something in the favoured case to which nothing in its rival corresponds (2004, ch. 3). In the kitchen, "rain rather than a burst pipe, because the window is open and the water lies under it" cites such a difference, since a burst pipe would have wet the floor by the pipe; "rain rather than a burst pipe, because the floor is very wet" has the same comparative shape and cites nothing that bears on the contrast, since a very wet floor favours neither rival. Both sentences instantiate the form of a weighing, and only the first contains one worth having; telling them apart requires understanding what each claims and asking whether it decides between the candidates, which is what the reader of any philosophy paper does. Nor is there a rule that would spare the reader the work: our grasp of what makes one explanation lovelier than another is weak (p. 61), and the standards are carried, in part, by past explanations that serve as exemplars and by prevailing styles of reasoning (p. 139). Human philosophers write in the format of explanation too, and the format was never what their comparisons were graded on; the bar that separates the two kitchen sentences separates human paragraphs and machine paragraphs alike. Where reference answers give out, the machine-learning literature itself assesses generated explanations in this way, scoring them for consistency, parsimony and coherence as features of the output (Dalal et al. 2024; He et al. 2025).
A model trained only to continue text respects constraints that were never stated for it, and Wolfram (2023) assembles the cases. A model trained on English respects English syntax, although no grammar was supplied to it: the syntax is carried by the writing, in which well-formed sentences predominate, and a system fitted to continue the writing comes to respect what the writing respects. Its sentences are, for the most part, meaningful rather than merely grammatical, although here there was no rule available even to withhold, since nothing like a complete theory of what makes a sentence meaningful has ever been built (2023 %%check page%%). Logic, in its syllogistic form, Wolfram treats the same way: a syllogism marks certain sentence patterns as reasonable, Aristotle, he imagines, arrived at the patterns from many examples of rhetoric, and a model trained on writing the patterns pervade can be expected to produce text containing "correct inferences" of the syllogistic kind, without anything having been derived (2023 %%check page%%). In each case a structure is present in the output while the capacity that ordinarily produces it — knowing the grammar, grasping the meaning, performing the deduction — is nowhere in the system. The absence of a rule for loveliness is therefore no obstacle on the production side. A system that wrote by applying stated rules would be halted exactly where no rule exists; these systems were never given stated rules for anything, and what they acquire, they acquire from exemplars — which are, on Lipton's account, where the standards of loveliness live. The precedent is narrower than the cases suggest, since a syllogism has a single correct completion and an abductive comparison does not: what carries over is the weaker point, and the only one needed, that a structure can be present in a text without the capacity that ordinarily produces it standing behind the text.
The corpus such models are trained on is general — most of it is not philosophy — but it contains the philosophical literature, and a philosophy paper is built as a displayed comparison: a position stated, set against rivals, and defended through the objections taken to decide between them. Wolfram's cases stop at the sentence, and the extension past it is ours; but his observations concern regularities in writing rather than grammar in particular, and an argument that states a candidate, sets out its rivals and locates the difference between them is as much a recurring regularity of the writing as syntax is. It may be said that all this redescribes the statistics: the model reproduces the regularities of its training text, and reproducing regularities is not weighing. Lipton met an objection of the same shape. Bayesianism, it had been suggested, gives the mechanics of belief revision and so leaves explanatory considerations nothing to do; arguing this way, he replies, is like arguing that "thinking about technique cannot help my squash game" because the ball's motion is governed by the laws of mechanics — even if Bayesianism gave the mechanics of belief revision, inference to the best explanation "might yet illuminate its psychology" (2004, p. 108). A true description of the mechanism does not displace a true description of what is produced. And here the mechanism is the one Floridi et al. themselves describe: the patterns absorbed from writing are patterns of reasoning as expressed in writing, and the writing does not contain the phrasing of explanations detached from their organisation — which considerations bear on which rivals, and what decides between them, are in the writing too, and a system that learns to continue the writing learns them with it. The look of the reasoning was never separable from the organisation that makes reasoning assessable on a page.
It may be objected that syntax is one thing and inference to the best explanation another: whatever structure next-word prediction carries, a system of this kind is too shallow for abduction, and the benchmark record reads like confirmation. The line Wolfram draws lies elsewhere, and it comes from the passage that supplied the syllogism: his toy network fails to balance long sequences of parentheses, a task demanding exact procedure with no shortcut, and sophisticated formal logic can be expected to fail for the same reason, while whatever a person can judge at a glance is managed (2023 %%check page%%). The divide such systems fail at falls between exact procedure and holistic judgement, not between the simple and the sophisticated — and weighing, on Lipton's account, sits with judgement, since no rule runs from evidence to the loveliest explanation. Read with that line in hand, the record divides against the account it seemed to confirm. A model that can recognise explanations but has nothing to draw on in producing one should fail wherever production is demanded; instead the collapse concentrates where abduction has been recast as the exact recovery of a single canonical missing premise under formal constraint — the strongest model reaches 21.5% on the hardest such benchmark and most score near zero — while on open-ended tasks, where the output is judged as an explanation, the strongest models' validity exceeds 90%.[^3] Failure tracks the demand for exact recovery, the parenthesis side of the line, and philosophical abduction does not live on that side.
None of this returns to the model any capacity Floridi et al. deny it. The model infers nothing, weighs nothing, and tests nothing; what it produces is text, and the text can contain what its producer never did — a candidate stated, the live rivals organised, the difference that decides between them located. Whether a given text does this, and does it well, is settled by the reading any philosophy paper receives, under the same standard and no other. Floridi et al. come close to saying so themselves: asked whether anything turns on the process being different when the hypothesis produced is the same, they answer that for justification it perhaps does, "but regarding the content of the hypothesis and our interpretation of it, maybe not" (2025, p. 13). A good weighing of positions a literature already contains is not yet a distinction the literature lacks, and whether a model can supply the second is among the questions Section 4 takes up; what philosophy a system with no relation to the world could produce at all is the question of Section 3.
[^1]: Floridi et al. also support the denial with an argument from the model's relation to the world: its words are connected to no perception of anything, and a hypothesis, once produced, is never tested against the world (2025, pp. 7–9). We take that argument up in Section 3.
[^2]: These benchmarks operationalise abduction in commonsense and formal domains with crowd-labelled or mechanically checkable answers, where philosophy has no answer key, and several score generated explanations against human-written references — a comparison nothing in this paper relies on. Performance also drops under small variations to a problem (Mirzadeh et al. 2025), and Floridi et al. note that beyond the training distribution "the facade can crack" (2025, p. 10); the paper's claim is a capacity claim — that such texts can be produced — and is untouched by variation in how reliably they are.
[^3]: Salimi et al.'s benchmark suite separates formally constrained missing-premise completion (ProofWriter, AbductionRules, NeuLR) from open-text explanation tasks (ART, e-CARE, UNcommonsense); the figures are from their Tables 3–6. They observe that exact-match scoring is demanding precisely where several different missing facts would each be plausible — the abductive situation — and that target structure and the size of the hypothesis space shape difficulty at least as much as subject matter does. Salimi et al. also run every benchmark with a single fixed instruction template and score one pass, while cataloguing methods — staged prompts, criticise-and-revise pipelines — that alter what models produce; what elicitation contributes is taken up in Section 4.
---
---
# PREVIOUS VERSION (v2) — retained for reference
# 2. The Challenge from Abduction — v3 (rebuilt from first principles)
In the previous section, we argued against the idea that LLMs cannot produce philosophy worth reading simply because they are not human. In this section and the next we shall consider a different form of challenge: even if LLMs cannot be ruled out of producing philosophy worth reading tout court, they lack particular _capacities_ that producing it requires. If a parrot uttered a sequence of sounds that happened to form a philosophical argument, the argument would be none the worse for its source; yet parrots' powers of mimicry do not extend to producing strings of sounds so complex as to make up a philosophical argument. In this section we address one capacity challenge, which we will call the _challenge from abduction_. In the next we shall look at two more: phenomenological experience and contact with the world.
Abduction, or inference to the best explanation, is reasoning from a body of evidence to the hypothesis that would best explain it. In a deductive argument the premises fix the conclusion: if all men are mortal and Socrates is a man, then Socrates is mortal, and there is no wriggle room. In abduction the evidence settles less. Now, imagine walking into your kitchen and finding the floor wet. What has happened? The wet floor does not determine the answer in the way the two premises determined Socrates' mortality: a burst pipe would have left the floor wet, and so would a spilled bucket. But, given that the window is open, the water is under the window, and it rained last night, rain coming through the window seems the most plausible answer.
To reason in this way, deciding what best explains a set of facts, is common in everyday life and in the sciences alike. A scientist chooses one theory over another when it explains the same data more simply: Copernicus's model of the solar system was preferred to Ptolemy's because it explained the observed planetary motions without the elaborate epicycles the older model required.
Williamson argues that philosophy is continuous with the sciences, and that its theories are to be chosen by the same abductive standards (2007; 2021, p. 351 %%check page%%). In philosophy too there are data that a candidate theory must accommodate — intuitions about cases, and the phenomena of the domain itself — and rival theories that would each accommodate them at different costs. The theory to prefer is the one that would, if true, best explain the data. What makes one explanation better than another, on this account, is a matter of explanatory virtue: a good philosophical theory is, in Williamson's words, "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated", and should "combine simplicity with strength" (2021, p. 354 %%check page%%). That theories are weighed by such comparative and explanatory virtues need not rest on a science-modelled conception of philosophy: Bengson, Cuneo and Shafer-Landau (2022) argue that the assessment of rival theories by their explanatory and unifying merits is a constraint on sound philosophical method as such. This conception of philosophy is widely held (Sider 2011; Paul 2012; Dellsén et al. 2024), though not universally (Bueno and Shalkowski 2020; Thomasson 2015), and we shall assume it in what follows.
On this account a philosophical text offers its reader a choice of theory displayed — a position, its rivals, and the case for preferring it — so that whether the text is worth reading and whether it contains a good weighing travel together. %%This is pathetic. First of all, grown-ups don't have single-sentence paragraphs. Second, what do you mean theory displayed? Third, you're using example lists. Yeah, so anyway, I have no idea what to do with it other than to tell you that it is diarrhea bad.%%
If the capacity for abduction is what is required to produce worthwhile philosophy, then the prospects for artificially generated philosophy seem dim. Consider the following from Floridi et al. (2025):
> LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025, p. 9)
An LLM, on this account, has "a stochastic core and an abductive appearance" (2025, p. 3), in that the model is trained to predict which words are likely to follow which, and it produces the continuation its training makes probable; it aims at the likely continuation, not at the truth. The appearance comes from what the training data have passed on: models have "absorbed patterns of human abductive reasoning as expressed in writing" (p. 10) — how explanations are typically phrased, which causes are typically offered for which effects. What is inherited, on their account, is the look of the reasoning, not the reasoning itself.[^1]
Explaining the wet kitchen floor involved two separable activities: coming up with candidate explanations, and settling which of them the open window and the position of the water favoured. Call the first _generating_ and the second _weighing_. %%please check we are not lifing these terms from lipton, floridi, or someone else. if we are, no problem but it needs to be clear where they come from%%Floridi et al.'s position is that a model does neither, however much its text exhibits both. Asked why a car might not start on a cold morning, a model replies that a weak battery is one possibility, since cold reduces a battery's efficiency; that thickened engine oil is another, since a cold engine is harder to turn over; and that, "[b]ased on your description, the battery is the most likely explanation" (2025, p. 11). The offering of candidates here is not generating, on their reading: the model is not reasoning about causes from the user's case but reproducing the causes such explanations typically cite (p. 9)%%the clause after the colon is not clear. a little compressed perhaps?%%. And the singling out is not weighing: the verdict reproduces how explanations of this kind typically end, and where an output marks a genuine point of difference between two hypotheses, that is something the model has seen stated, not something it has derived anew (p. 14). %%I wonder if this paragraph could be a little clearer, because it is so important%%
Floridi et al. draw the consequence themselves: %%metacommentry wanking...%%
> In essence, LLMs function like brainstorming assistants that toss out ideas without filtering for quality. After all, they work like statistical interfaces to an enormous amount of data accumulated for millennia by generations. A cautious human collaborator can sift through and assess them. (2025, p. 12)
On this picture the weighing always remains with the person: the model supplies candidates, and assessing them is the collaborator's work. Whatever such a system produces is raw material for philosophy done by someone else, and raw material is not philosophy worth reading — a list of unweighed candidates is no more worth reading than a bare pronouncement that direct realism is correct. %%so far this paragraph has used a lot of words for a VERY simple idea%%The challenge follows:%%the challenge does not follow, this second half of the paragraph is entirely uncnnected to the first half%% if a model's text cannot contain a good weighing, there is no reason to regard it as worth reading. The benchmark record can seem to agree, since abduction is the form of reasoning at which models perform worst, with median accuracy across surveyed studies of roughly 43%, against 80% for deduction (Salimi et al. 2026).[^2] %%so compressed as to be meaningless. also 'the benchmark record' is a meaningless and cretenous way of putting things%%
Everything in this account of the producer can be granted.%%It's an obscure way to start a paragraph.%% The model generates nothing and weighs nothing, and nothing in what follows returns either capacity to it. %%last clause =not how i write%%What the account does not settle is anything about the texts.%%Meta commentative wank.%% Section 1 fixed where the grounds of a text's merit lie — in the argument as presented, not in the history of its production — and by that standard the car-battery reply asks to be read rather than explained away.%%not how i write%%It is not a list of candidates awaiting a collaborator: it brings the cold morning to bear on each candidate and closes in favour of one, so the sifting the brainstorming picture reserves for the person is on the page. Whether that displayed sifting is any good is a question about a piece of writing.%%paragraph is compressed so very unclear%%
What remains of the challenge is the claim that the weighing a model's text displays cannot be good. Against this we will argue that the challenge underestimates what the inherited look of reasoning includes.%%starting to have my doubts about using the word 'look' for this stuff...%% Two things need showing: what makes a weighing displayed in a text good, and how a good weighing can come to be displayed in text that nobody weighed. Lipton's account of inference to the best explanation supplies the first; Wolfram's account of what continuing text involves supplies the second. %%also, is weighing the bet word to use? why are we using this one?%%
The division of the kitchen's work into generating and weighing is Lipton's own%%incredibly unclear sentence%%: inference to the best explanation, on his account, runs on two filters, one that supplies the plausible candidates and a second that selects among them (2004, p. 59 %%pin page%%) — and his question about the second filter is ours, namely what makes the selection good.%%this is the biggest structural issue here. we are introducing these two parts of abduction for a second time, this time with a lipton flavour. this is very inelegant. why do we even need to mention this here when we are talking about loveliness? %% The best explanation can be understood as the likeliest, the one most warranted by the total evidence, or as the loveliest, the one which, if correct, would provide the most understanding: "[l]ikeliness speaks of truth; loveliness of potential understanding" (p. 59). The two come apart: Newtonian mechanics is no longer the likeliest account of the observations it was built on, but it remains as lovely an explanation of them as it ever was (p. 60). Loveliness is the standard a philosophical text answers to: what readers assess is whether the explanation offered would, if true, give understanding, and give more of it than the rivals considered. Because the assessment runs under "if correct", it does not wait on verification, and a reader can conduct it on the page. On Dellsén et al.'s account, philosophical progress consists in putting people in a position to increase their understanding (2024, p. 679); a lovely explanation puts its reader in exactly that position. %%this is not a clear paragraph. Also, in previous version we made it clear to the reader how the term 'likeliest' is being used by Lipton, you have removed this information without my permission.%%
Where loveliness shows itself, on Lipton's analysis, is in the comparison of rivals.%%notclear at all%% Explanation is contrastive: we explain why this rather than that, and doing so requires citing a difference between the two — his Difference Condition — something in the favoured case to which nothing in its rival corresponds (2004, ch. 3).%%to compressed so unclear%% In the kitchen, "rain rather than a burst pipe, because the window is open and the water lies under it" cites such a difference, since a burst pipe would have wet the floor by the pipe; "rain rather than a burst pipe, because the floor is very wet" has the same comparative shape and cites nothing that bears on the contrast, since a very wet floor favours neither rival. Both sentences instantiate the form of a weighing, and only the first contains one worth having; telling them apart requires understanding what each claims and asking whether it decides between the candidates, which is what the reader of any philosophy paper does.%%I don't understand this sentence, all i know is that it is conveying a shit idea%% Nor is there a rule that would spare the reader the work:%%why is the reader being told all this shit, it seems like you are putting it in just for the sake of wasting words%% our grasp of what makes one explanation lovelier than another is weak (p. 61), and the standards are carried, in part, by past explanations that serve as exemplars and by prevailing styles of reasoning (p. 139). Human philosophers write in the format of explanation too, and the format was never what their comparisons were graded on; the bar that separates the two kitchen sentences separates human paragraphs and machine paragraphs alike. Where reference answers give out, the machine-learning literature itself assesses generated explanations in this way, scoring them for consistency, parsimony and coherence as features of the output (Dalal et al. 2024; He et al. 2025). %%a steaming turd of a paragraph. an absolute disgrace%%
A model trained only to continue text respects constraints that were never stated for it, and Wolfram (2023) assembles the cases. A model trained on English respects English syntax, although no grammar was supplied to it: the syntax is carried by the writing, in which well-formed sentences predominate, and a system fitted to continue the writing comes to respect what the writing respects. Its sentences are, for the most part, meaningful rather than merely grammatical, although here there was no rule available even to withhold, since nothing like a complete theory of what makes a sentence meaningful has ever been built (2023 %%check page%%). Logic, in its syllogistic form, Wolfram treats the same way: a syllogism marks certain sentence patterns as reasonable, Aristotle, he imagines, arrived at the patterns from many examples of rhetoric, and a model trained on writing the patterns pervade can be expected to produce text containing "correct inferences" of the syllogistic kind, without anything having been derived (2023 %%check page%%). In each case a structure is present in the output while the capacity that ordinarily produces it — knowing the grammar, grasping the meaning, performing the deduction — is nowhere in the system. The absence of a rule for loveliness is therefore no obstacle on the production side. A system that wrote by applying stated rules would be halted exactly where no rule exists; these systems were never given stated rules for anything, and what they acquire, they acquire from exemplars — which are, on Lipton's account, where the standards of loveliness live. The precedent is narrower than the cases suggest, since a syllogism has a single correct completion and an abductive comparison does not: what carries over is the weaker point, and the only one needed, that a structure can be present in a text without the capacity that ordinarily produces it standing behind the text. %%this paragraph is far too long, I am not even going to read it%%
The corpus such models are trained on is general — most of it is not philosophy — but it contains the philosophical literature, and a philosophy paper is built as a displayed comparison: a position stated, set against rivals, and defended through the objections taken to decide between them. Wolfram's cases stop at the sentence, and the extension past it is ours; but his observations concern regularities in writing rather than grammar in particular, and an argument that states a candidate, sets out its rivals and locates the difference between them is as much a recurring regularity of the writing as syntax is. It may be said that all this redescribes the statistics: the model reproduces the regularities of its training text, and reproducing regularities is not weighing. Lipton met an objection of the same shape. Bayesianism, it had been suggested, gives the mechanics of belief revision and so leaves explanatory considerations nothing to do; arguing this way, he replies, is like arguing that "thinking about technique cannot help my squash game" because the ball's motion is governed by the laws of mechanics — even if Bayesianism gave the mechanics of belief revision, inference to the best explanation "might yet illuminate its psychology" (2004, p. 108). A true description of the mechanism does not displace a true description of what is produced. And here the mechanism is the one Floridi et al. themselves describe: the patterns absorbed from writing are patterns of reasoning as expressed in writing, and the writing does not contain the phrasing of explanations detached from their organisation — which considerations bear on which rivals, and what decides between them, are in the writing too, and a system that learns to continue the writing learns them with it. The look of the reasoning was never separable from the organisation that makes reasoning assessable on a page.
It may be objected that syntax is one thing and inference to the best explanation another: whatever structure next-word prediction carries, a system of this kind is too shallow for abduction, and the benchmark record reads like confirmation. The line Wolfram draws lies elsewhere, and it comes from the passage that supplied the syllogism: his toy network fails to balance long sequences of parentheses, a task demanding exact procedure with no shortcut, and sophisticated formal logic can be expected to fail for the same reason, while whatever a person can judge at a glance is managed (2023 %%check page%%). The divide such systems fail at falls between exact procedure and holistic judgement, not between the simple and the sophisticated — and weighing, on Lipton's account, sits with judgement, since no rule runs from evidence to the loveliest explanation. Read with that line in hand, the record divides against the account it seemed to confirm. A model that can recognise explanations but has nothing to draw on in producing one should fail wherever production is demanded; instead the collapse concentrates where abduction has been recast as the exact recovery of a single canonical missing premise under formal constraint — the strongest model reaches 21.5% on the hardest such benchmark and most score near zero — while on open-ended tasks, where the output is judged as an explanation, the strongest models' validity exceeds 90%.[^3] Failure tracks the demand for exact recovery, the parenthesis side of the line, and philosophical abduction does not live on that side.
None of this returns to the model any capacity Floridi et al. deny it. The model infers nothing, weighs nothing, and tests nothing; what it produces is text, and the text can contain what its producer never did — a candidate stated, the live rivals organised, the difference that decides between them located. Whether a given text does this, and does it well, is settled by the reading any philosophy paper receives, under the same standard and no other. Floridi et al. come close to saying so themselves: asked whether anything turns on the process being different when the hypothesis produced is the same, they answer that for justification it perhaps does, "but regarding the content of the hypothesis and our interpretation of it, maybe not" (2025, p. 13). A good weighing of positions a literature already contains is not yet a distinction the literature lacks, and whether a model can supply the second is among the questions Section 4 takes up; what philosophy a system with no relation to the world could produce at all is the question of Section 3.
[^1]: Floridi et al. also support the denial with an argument from the model's relation to the world: its words are connected to no perception of anything, and a hypothesis, once produced, is never tested against the world (2025, pp. 7–9). We take that argument up in Section 3.
[^2]: These benchmarks operationalise abduction in commonsense and formal domains with crowd-labelled or mechanically checkable answers, where philosophy has no answer key, and several score generated explanations against human-written references — a comparison nothing in this paper relies on. Performance also drops under small variations to a problem (Mirzadeh et al. 2025), and Floridi et al. note that beyond the training distribution "the facade can crack" (2025, p. 10); the paper's claim is a capacity claim — that such texts can be produced — and is untouched by variation in how reliably they are.
[^3]: Salimi et al.'s benchmark suite separates formally constrained missing-premise completion (ProofWriter, AbductionRules, NeuLR) from open-text explanation tasks (ART, e-CARE, UNcommonsense); the figures are from their Tables 3–6. They observe that exact-match scoring is demanding precisely where several different missing facts would each be plausible — the abductive situation — and that target structure and the size of the hypothesis space shape difficulty at least as much as subject matter does. Salimi et al. also run every benchmark with a single fixed instruction template and score one pass, while cataloguing methods — staged prompts, criticise-and-revise pipelines — that alter what models produce; what elicitation contributes is taken up in Section 4.
---
# 3. The Challenge of Connecting to the World and the Challenge from Experience — v4 (approved fixes 1,2,4,5,7,8 applied)
In this section we address two more capacity challenges, which we will call the _challenge of connecting to the world_ and the _challenge from experience_. The first holds that LLMs cannot produce philosophy worth reading because they stand in no relation to the world: nothing a model says rests on perception of anything, and nothing it produces is tested against how things are. The second holds that they cannot because some philosophy depends on experience in a way a system without experience cannot meet: experience supplies the starting point of some philosophical reasoning, and is itself the subject matter of some philosophical inquiry. Few would say that current LLMs are conscious, and we assume here that they are not. We argue that both challenges fail, and for the same reason: the materials philosophy takes from the world and from experience reach the philosopher already set down in words, and a model works on those words as any philosopher does.
Floridi et al. also object that the model stands in no relation to the world: its words rest on no perception of anything, and a hypothesis, once produced, is never tested against how things are (2025, pp. 7–9). A discipline whose theories answer to how things are, the challenge runs, cannot be advanced by a system with no access to how things are, and a text from one gives its reader no reason to think it worth reading.
The challenge from experience says that some philosophy cannot be done by anyone who has not had the relevant experience. Zahavy (2026) raises a similar worry about scientific discovery. A model can carry out the deductive part of discovery, working out the consequences of premises it has been given; what it cannot do, he holds, is produce the premises — make the move from sense experience to new first principles. On the picture he takes from Einstein, that move is a leap, and it is the leap[^5] that gives a theory its axioms. His case is the thought experiment that gave Einstein the equivalence principle:
> Einstein's variation required inventing new axioms based on a physical intuition that did not yet exist in the mathematics. He envisioned a physicist inside an elevator being uniformly accelerated through deep space [...]. Inside this enclosure, the sensory experience reveals a specific pattern: when objects are released, the floor rushes up to meet them. To the physicist, the objects appear to fall with identical acceleration, regardless of composition. Thus, the simulation here was not a permutation of symbols, but a manipulation of perceptual experience. (Zahavy 2026, §5)
Einstein imagines a set of circumstances and attends to what would be experienced within them: everything released inside the elevator appears to fall with identical acceleration. On Zahavy's reconstruction the simulation supplies an observation, and from that observation the new axiom is inferred — the simulated experience of acceleration was indistinguishable from the remembered experience of gravity, and Einstein concluded that the two are one phenomenon.[^2] A model has no access to that observation. It can produce descriptions of elevators and of weightlessness, both present in its corpus, but it has undergone neither, and a discovery whose premises are fixed by simulated experience is beyond a system that, in Zahavy's words, lacks the capacity he calls sensory agency.
We might think that philosophical thought experiments depend upon experience in the same way. Does Mary, released from her black-and-white room knowing every physical fact about colour vision, learn something when she first sees red (Jackson 1982)? Settling the question requires considering what the experience is like, and the argument proceeds from the verdict; this is experience entering as the starting point of an argument, and a system that has never experienced anything appears unable to supply it. Experience enters also as subject matter, in philosophy that asks what it is to see red (Harman 1990), to feel anger (Goldie 2000), or to have an intuition take hold (Chudnoff 2011).[^3] Were these claims correct, much of the philosophy of mind would lie beyond a model's reach.
We can grant that the model has no senses and has never had an experience, and that it cannot put what it produces to the test against the world. Whether any of that bears on the texts it produces depends on what philosophy does with the world and with experience — on where a philosopher's starting points come from, and on what becomes of a philosophical claim once it is made. Pigliucci (2017) addresses both. He holds that philosophy is constrained by the world without investigating it as the natural sciences do:
> This means that the basic parameters that philosophers use as their inputs, the starting points of their philosophizing, their equivalent of axioms in mathematics and assumptions in logic (or rules in chess) are _empirical_ data about the world. This data comes from both everyday experience [...] and of course increasingly from the world of science itself. Philosophy, I maintain, is in the business of exploring the sort of conceptually evoked spaces that Smolin is talking about, where the evocation is the result of whatever starting assumptions are made by individual philosophers working within a particular field and, crucially, of the constraints that are imposed by our best understanding of how the world actually is. (2017, pp. 79–80)
_Evocation_ is a term Pigliucci takes from Smolin (Unger and Smolin 2015),[^4] for truths that are neither discovered, in the sense of corresponding to mind-independent states of affairs, nor invented, in the sense of being arbitrary constructs, and his example is chess: when the rules of a game are codified, a whole bundle of facts about it becomes demonstrable — objective facts, in that anyone who can demonstrate one demonstrates the same fact as anyone else — although chess did not exist before its rules were written down (Unger and Smolin 2015, p. 423; quoted at Pigliucci 2017, p. 78). Pigliucci's proposal is that philosophy ascertains evoked truths in this sense, with an addition that separates it from mathematics and chess alike: its starting points are constrained by how the world actually is. The addition is also what separates philosophy from fiction, on his account. A novelist's worlds are invented rather than evoked — nothing about them is rigid, since even the constraints the novelist adopts could have been otherwise — whereas philosophy "is in the business of doing empirically informed evoking, not inventing", so that its objects of study have rigid properties (2017, p. 80). A thought experiment is itself a case of such evoking: the philosopher sets up an imagined scenario but explores it "with an interest in figuring things out as far as this world is concerned" (2017, p. 80), so that what it evokes has the rigid properties Pigliucci means, and the philosophical work proceeds within the structure it opens.
Philosophy's starting points, then, are empirical, and a system that perceives nothing cannot reach them on its own. But the data Pigliucci describes comes from everyday experience and, increasingly, from science, and the scientific kind reaches working philosophers already articulated — already set down in language, available to be read rather than undergone. A philosopher of physics works from published results, not from having run the experiments. The same holds for everyday experience: what the discipline retains of it, it retains as the literature's accumulated descriptions of how things seem. For the worldly materials philosophy actually uses, written access is the profession's normal condition rather than a deficiency, and a corpus is such access. On this point the model stands where every philosopher already stands with respect to nearly all of the empirical data they use.
The other objection was that the model never checks what it produces against the world. But a philosophical claim is not the kind of thing that gets checked that way. Here the elevator and Mary part company. The equivalence principle, once Einstein had it, faced a tribunal of measurement: the experiments might have gone against it, and then it would have been dropped. Mary's case faces no such tribunal. The question Mary raises — whether she learns something new on first seeing red — is a question about what follows within the scenario Jackson has set up, and that is settled in the way any question about chess is settled: by working out what the set-up commits us to, something any competent party can do and none can decide by fiat. This is the only checking a philosophical thesis gets, and it happens in the literature, in the back-and-forth the previous section described. So a model's inability to run experiments costs it nothing a philosophical text needs: the testing that philosophy does is the working-out of what a scenario commits us to, and that is done on the page.
Take the knowledge argument itself. Nobody who debates it has been through what Mary goes through — released into colour after a lifetime of black and white — and the debate does not suffer for it. The participants have in front of them Jackson's description: a scenario set down in words, with a claim about what it is meant to show. Lewis's reply (1988) works on that description, changing what we should say follows from Mary's release while saying nothing about what her first sight of red is like from the inside. This is the usual way experience figures in philosophy. A philosopher need not have had an experience to argue about it: Nagel (1974) asks what it is like to be a bat without supposing he could find out, working instead from a description of the bat's situation. A model is no worse off here than these philosophers are. It has had no experiences of its own, but it has not needed them, because what its corpus supplies, and what the philosophy works on, is experience already put into words — the same descriptions that, in the previous section, carried displayed reasoning into its texts.
Merleau-Ponty noticed something about touching one hand with the other that was not, so far as we know, already written down anywhere. At any moment one hand is the toucher and the other the touched; the roles can switch, but they cannot both hold at once (Merleau-Ponty 1945 %%page%%). Suppose he came to this only by attending to his own body, to what the touching was like from the inside, and not from anything already in the literature. Then it is a starting point a model could not have reached on its own, since reaching it took a body and a first-person view that a model does not have. Once Merleau-Ponty has written the description down, a model can take it up and argue about it as well as anyone. It could not, though, have been the one to set it down first.
What a philosopher does with Merleau-Ponty's description, though, has nothing to do with who first arrived at it. The description earns its place by what can be drawn out of it — what it shows about the body, what follows once the asymmetry is granted — and that is there for any reader to work through, whether or not they could have come to the asymmetry on their own. First-person attention is not, in any case, the only way new descriptions come about. A description the literature does not yet contain can also be reached from the descriptions it does contain, by drawing out what they have not been taken to imply, or by putting two of them together as no one has — and a model can do this. Whether today's models in fact produce descriptions the literature lacks is the question of novelty, which Section 4 takes up.
A model can do the philosophy that turns on the world and on experience, save for the one point already granted: it could not be the first to set down a description that only first-person attention could yield. Its worldly starting points come to it already written down; the claims it draws are settled by working out what they commit us to, not by measurement against the world; and the experiences philosophy argues about reach it as descriptions, which it can work through as well as any reader. In the previous section, abduction entered philosophy as reasoning displayed in a text; here the world and experience enter it as descriptions set down in a text. What a model produces is read as any philosophy is read, and how it was produced settles nothing in advance.
[^2]: Zahavy, following Magnani, calls the process _manipulative abduction_: hypothesis generation through the manipulation of a model — here a simulated experience — rather than of symbols (Magnani et al. 2009; Zahavy 2026, §5). It is abduction in the previous section's sense: the equivalence principle is inferred as the best explanation of the simulated observation, the simulation supplying an explanandum that no search over existing text would have produced. What experience contributes, on this picture, is not the inference but its starting point.
[^3]: The materials need not be sensory: the feeling of understanding something is sometimes used to motivate the claim that thought itself has a phenomenology (Pitt 2004).
[^5]: Zahavy too calls this leap abduction, but the word picks out something other than it did in the previous section. There, with Floridi et al., abduction was the weighing of rival explanations, and the charge was that a model only mimics it; here it is the generation of new first principles from experience, and Zahavy's claim is that a model cannot make the move because it has had no experience to move from. The present challenge rests on that second claim, about experience, and not on any verdict about the weighing.
[^4]: _The Singular Universe and the Reality of Time_ is jointly authored, but its second part, which contains the discussion of evocation, was written by Smolin alone, as Pigliucci notes (2017, p. 77); we follow him in attributing the view to Smolin.
---
---
# 4. The Challenge from Observation
In this final section we address what we will call the _challenge from observation_. Sections 2 and 3 argued that a model which performs no abductive inference, connects to nothing beyond text, and has no experience can nonetheless produce text bearing the properties those capacities ordinarily produce. The challenge from observation begins with an obvious question: if LLMs should not be ruled out from producing philosophy tout court, and **have the capacity to produce worthwhile philosophy**, where is all the worthwhile LLM-written philosophy? If you ask an LLM the answer to the hard problem of consciousness, or the meaning of life you will not receive *the correct answer*, but instead a competent but unopionated survey of the field if you are lucky, or a less accurate but equally bland survey if you are unlucky.
The observation is accurate, and it reports less than it seems to: it reports what models produce under one use — a bare question, put once, answered in one pass. How these systems are built explains why that use yields what it does. A model is first fitted to a vast general corpus and trained to continue text, so its response to a bare philosophical question is the likely continuation of such a question in writing at large, and the likely continuation of "what is the meaning of life?" in a general corpus is not an analytic tract. It is the sort of text that follows the question at large: a survey of views, a consoling generality, a joke. The model is then further shaped to converse as a helpful assistant, and the shaping presses the same way, since a person employed to be helpful to all comers would not answer the question with a tract either. The survey is not a ceiling the systems have hit; it is the likely continuation of exactly what was given them.
The use that generates the observation treats the model as an oracle: a system whose answers are its measure, so that asking is all the eliciting there is.[^2] The empirical record tells against the assumption. The survey of abductive benchmarks discussed in Section 2 runs every test with a single fixed instruction and scores the answer, while cataloguing, in the same pages, methods that alter what models produce — prompts that separate the stages of a task, pipelines in which an answer is criticised and revised over several passes (Salimi et al. 2026). What a model returns depends on what it is given, and the observation samples one point in that space, the bare question. It therefore cannot discriminate between the two hypotheses at issue — that the capacity defended in the preceding sections is absent, and that it has not been elicited. Both predict the observed record, and an argument against this paper needs the first; the observation supports it no better than the second.
The challenge has a natural escalation: if philosophy worth reading comes out of these systems only when a philosopher directs the process — supplies the framing, sets the constraints, presses for development — then the philosophy, it will be said, is the philosopher's. The model is an instrument in the production, as a typewriter is, and crediting it with the result is crediting the dummy with the ventriloquism. Section 1's challenge held that a model's text is not philosophy tout court; what stands here is narrower, that the philosophy in such a text is not the model's.
Whether the escalation succeeds depends on what prompting a model involves, and two things need saying: what a prompt supplies, and what the model's continuation adds to it. Section 3's account of starting points says the first; Section 2's account of continuing text says the second.
A prompt articulates a starting point, as a thought experiment does. A prompt that sets out a position and the rivals it must beat stands to the model as Jackson's two paragraphs stand to the profession: a starting point handed over for development. What an articulated starting point does, on the account already in place, is evoke a structure with rigid properties — there are facts about what holds within it, demonstrable by anyone and chosen by no one, and they outrun whatever has been stated, just as the facts about chess outran the rules the moment the rules were written down. Most of what a starting point evokes, no one has ever said.
What the model contributes is the development, and the mechanics are the ones Section 2 drew from Wolfram: a model produces a reasonable continuation of the text it has been given, where what counts as reasonable is relative to the corpus it was fitted to (2023). A prompt is part of the text the model has been given. An articulated starting point therefore changes what there is to continue — the reasonable continuation of a stated position under stated constraints is not the reasonable continuation of a bare question — and the model makes use of what the prompt states in everything that follows: tell one of these systems something once, Wolfram observes, and it is used thereafter (2023). The continuation that results states consequences of the starting point that the starting point does not state. Section 2 said what it is for such a text to go well — the comparison it displays cites differences that bear, and would, if correct, give understanding — and whether a given continuation goes well is read off the continuation.
Nothing in this makes the development a transcription. An evoked structure contains more than any text states: the rules of chess settle every fact about chess, and do not settle which theorems get written down, in what order, or to what depth, so that two writers working from the same rules produce different books, both correct, neither dictated by the rules. The mechanics mirror the structure, since the same prompt, run twice, yields different continuations (Wolfram 2023). The starting point underdetermines the development, and the gap between them is where the model's contribution lies: were there one text the prompt fixed, the output would transcribe what the person had already settled, and the instrument description would be true. The gap also leaves room for error. A development can state what does not hold in the evoked structure — a chess writer can publish a false theorem, a philosopher can misdraw the consequences of their own thought experiment, and a model can do both, along with its characteristic failure of stating fluently what nothing supports. The errors are found on the page. And an error is attributable only to a developer: no one blames the rules of chess for a false theorem, and no one's typewriter has ever made a mistake of content. Three contributions, then, and three owners: the articulated starting point is the person's; the structure it evokes, and the facts that hold there, are no one's; the text that develops them is the model's.
Much in the instrument picture is true. The person writes the prompt and the prompt is authored; the person chooses which continuations to pursue and when to stop; without the person, there is the survey. What the picture adds to these truths is a description of the model — a device, like the typewriter, that fixes only what its user has already settled — and the description is what the account above denies. Every word of the novel was the author's before the typewriter touched it; the consequences a model's text states were nobody's before the text stated them. We have pressed the tool picture elsewhere, for image generators: such a system is reliably unpredictable — a prompter settles what an image is to depict, and the system settles what the image is like, so the user's control runs out where the product's properties begin (Young and Terrone 2025). What the user of a typewriter settles is the text; what the writer of a prompt settles is a starting point.
The account invites an obvious enrichment of the prompt. State the position, name the rivals, list the objections and the lines along which they are to be met, and at some point, it will be said, the prompt contains the philosophy and the model is expanding what the person wrote — so that where a model's output is good, one should suspect a prompt rich enough to have done the work. But enriching a prompt enlarges the starting point without converting it into the development. A game with more rules is a bigger game, not a book of its theorems, and however much the prompt states, the consequences the output draws were not among the statements. There is a genuine limiting case — a prompt that states the comparison and the verdict, so that the continuation only rephrases — and it is identified the way everything in this paper is identified: set the output against the prompt and ask what the text states that the prompt did not. A text that states nothing beyond its prompt is a paraphrase, and owed to the person; a text that states what the prompt left unstated is a development, and the unstated part is not the person's. Which of the two a given output is, is settled by reading them together.
[^1]: While preparing this paper we asked GPT-5.5 for a detailed overview of the positions an analytic philosopher might take on the meaning of life. What came back was a competent, hedged survey of the field; what did not come back was an argument for any position in it. %%add date of test%%
[^2]: That these systems are mischaracterised as oracles — with the corollary that no benchmark of single-pass answers should be expected to probe the upper limits of what they can produce — has been argued from inside the practitioner literature (Janus 2022).
_Draft note: the section currently ends at the rich-prompt reply; the close and the novelty question (Section 3's hand-off) are deliberately unwritten pending design. Citation flags: Janus 2022 is a pseudonymous LessWrong post — confirm citation practice; the GPT-5.5 test needs its date._
---
---