[this section needs the most work] In this final section we address what we will call the _challenge from observation_. In Section 1 we argued that LLMs should not be disqualified from producing worthwhile philosophy tout court. In Sections 2 and 3 we argued that, although LLMs neither perform abductive inference, nor have experience, nor are connected to the world, there is still reason to think that they are capable of producing text which exhibits good quality abduction, and works with articulated axioms about experience and the world. The challenge from observation begins with an obvious question: if all of these arguments are correct, where is all the worthwhile LLM-written philosophy? If you ask an LLM the answer to the hard problem of consciousness, or the meaning of life,[1](https://chatgpt.com/g/g-p-6863cdae32988191b6ef486a436ee4e5-nick/c/6a2fbe04-d274-83eb-a320-967f94fc91f3#user-content-fn-1) you will not receive _the correct answer_, but instead a competent but unopionated survey of the field if you are lucky, or a less accurate but equally bland survey if you are unlucky. The observation is accurate, and it reports less than it seems to: it reports what models produce under one use — a bare question, put once, answered in one pass. How these systems are built explains why that use yields what it does. A model is first fitted to a vast general corpus and trained to continue text, so its response to a bare philosophical question is the likely continuation of such a question in writing at large, and the likely continuation of "what is the meaning of life?" in a general corpus is not an analytic tract. It is the sort of text that follows the question at large: a survey of views, a consoling generality, a joke. The model is then further shaped to converse as a helpful assistant, and the shaping presses the same way, since a person employed to be helpful to all comers would not answer the question with a tract either. The survey is not a ceiling the systems have hit; it is the likely continuation of exactly what was given them. The use that generates the observation treats the model as an oracle: a system whose answers are its measure, so that asking is all the eliciting there is.[2](https://chatgpt.com/g/g-p-6863cdae32988191b6ef486a436ee4e5-nick/c/6a2fbe04-d274-83eb-a320-967f94fc91f3#user-content-fn-2) The empirical record tells against the assumption. The survey of abductive benchmarks discussed in Section 2 runs every test with a single fixed instruction and scores the answer, while cataloguing, in the same pages, methods that alter what models produce — prompts that separate the stages of a task, pipelines in which an answer is criticised and revised over several passes (Salimi et al. 2026). What a model returns depends on what it is given, and the observation samples one point in that space, the bare question. It therefore cannot discriminate between the two hypotheses at issue — that the capacity defended in the preceding sections is absent, and that it has not been elicited. Both predict the observed record, and an argument against this paper needs the first; the observation supports it no better than the second. The challenge has a natural escalation: if philosophy worth reading comes out of these systems only when a philosopher directs the process — supplies the framing, sets the constraints, presses for development — then the philosophy, it will be said, is the philosopher's. The model is an instrument in the production, as a typewriter is, and crediting it with the result is crediting the dummy with the ventriloquism. Section 1's challenge held that a model's text is not philosophy tout court; what stands here is narrower, that the philosophy in such a text is not the model's. Whether the escalation succeeds depends on what prompting a model involves: what a prompt supplies, and what the model's continuation adds to it. Section 3's account of starting points says the first; Section 2's account of continuing text says the second. A prompt articulates a starting point, as a thought experiment does. A prompt that sets out a position and the rivals it must beat stands to the model as Jackson's two paragraphs stand to the profession: a starting point handed over for development. What an articulated starting point does, on the account already in place, is evoke a structure with rigid properties — there are facts about what holds within it, demonstrable by anyone and chosen by no one, and they outrun whatever has been stated, just as the facts about chess outran the rules the moment the rules were written down. Most of what a starting point evokes, no one has ever said. What the model contributes is the development, and the mechanics are the ones Section 2 drew from Wolfram: a model produces a reasonable continuation of the text it has been given, where what counts as reasonable is relative to the corpus it was fitted to (2023). A prompt is part of the text the model has been given. An articulated starting point therefore changes what there is to continue — the reasonable continuation of a stated position under stated constraints is not the reasonable continuation of a bare question — and the model makes use of what the prompt states in everything that follows: tell one of these systems something once, Wolfram observes, and it is used thereafter (2023). The continuation that results states consequences of the starting point that the starting point does not state. Section 2 said what it is for such a text to go well — the comparison it displays cites differences that tell between the positions, and would, if correct, give understanding — and whether a given continuation goes well is read off the continuation. Nothing in this makes the development a transcription. An evoked structure contains more than any text states: the rules of chess settle every fact about chess, and do not settle which theorems get written down, in what order, or to what depth, so that two writers working from the same rules produce different books, both correct, neither dictated by the rules. The mechanics mirror the structure, since the same prompt, run twice, yields different continuations (Wolfram 2023). The starting point underdetermines the development, and the gap between them is where the model's contribution lies: were there one text the prompt fixed, the output would transcribe what the person had already settled, and the instrument description would be true. The gap also leaves room for error. A development can state what does not hold in the evoked structure — a chess writer can publish a false theorem, a philosopher can misdraw the consequences of their own thought experiment, and a model can do both, along with its characteristic failure of stating fluently what nothing supports. The errors are found on the page. And an error is attributable only to a developer: no one blames the rules of chess for a false theorem, and no one's typewriter has ever made a mistake of content. Three contributions, then, and three owners: the articulated starting point is the person's; the structure it evokes, and the facts that hold there, are no one's; the text that develops them is the model's. Much in the instrument picture is true. The person writes the prompt and the prompt is authored; the person chooses which continuations to pursue and when to stop; without the person, there is the survey. What the picture adds to these truths is a description of the model — a device, like the typewriter, that fixes only what its user has already settled — and the description is what the account above denies. Every word of the novel was the author's before the typewriter touched it; the consequences a model's text states were nobody's before the text stated them. What the user of a typewriter settles is the text; what the writer of a prompt settles is a starting point. The account invites an obvious enrichment of the prompt. State the position, name the rivals, list the objections and the lines along which they are to be met, and at some point, it will be said, the prompt contains the philosophy and the model is expanding what the person wrote — so that where a model's output is good, one should suspect a prompt rich enough to have done the work. But enriching a prompt enlarges the starting point without converting it into the development. A game with more rules is a bigger game, not a book of its theorems, and however much the prompt states, the consequences the output draws were not among the statements. There is a genuine limiting case — a prompt that states the comparison and the verdict, so that the continuation only rephrases — and it is identified the way everything in this paper is identified: set the output against the prompt and ask what the text states that the prompt did not. A text that states nothing beyond its prompt is a paraphrase, and owed to the person; a text that states what the prompt left unstated is a development, and the unstated part is not the person's. Which of the two a given output is, is settled by reading them together. Even if all this is granted, we might still ask whether such a text can do more than handle well the positions a literature already contains and make a distinction that literature lacks, and so be creative in the stronger, public sense in which a human philosophical text is creative. We want to be careful here, because this is quite different from the modest sense in which an output is novel only relative to its prompt; what is at issue is the stronger claim of saying something the literature had not yet said. We do not need to settle this question here. It would be settled as the rest has been, by setting the output not only against the prompt but against the literature, and asking what it says that the literature had not. This is not meant to settle the matter either way, but it does give us a reason to leave the question open, for further work. We can now return to the challenge from Section 1, where the ordinary blandness of what these systems produce when given a bare question was taken to show that they have nothing to contribute to philosophy. The outputs to such questions do tend towards the empty and the thin; but that bears only on whether bare questions are good tests of philosophical capacity, since what these systems do is continue the context they are given, and a context with no shape of argument can only draw from them an output with no shape of argument. This is no reason to think that a context which hands the system a position, its rivals, and the pressures bearing on each of them cannot draw a development of its own; and whether such a development is worth reading is settled not by looking at the system but by reading the continuation first against the prompt that occasioned it and then against the literature it means to add to. To sum up, these systems are not oracles; they are continuation systems, and philosophy worth reading needs a dialectical context which a prompt can supply without thereby fixing the development, so that the philosophical standing of any output depends on what the continuation itself adds. ## Footnotes 1. While preparing this paper we asked GPT-5.5 for a detailed overview of the positions an analytic philosopher might take on the meaning of life. What came back was a competent, hedged survey of the field; what did not come back was an argument for any position in it. %%add date of test%% [↩](https://chatgpt.com/g/g-p-6863cdae32988191b6ef486a436ee4e5-nick/c/6a2fbe04-d274-83eb-a320-967f94fc91f3#user-content-fnref-1) 2. That these systems are mischaracterised as oracles — with the corollary that no benchmark of single-pass answers should be expected to probe the upper limits of what they can produce — has been argued from inside the practitioner literature (Janus 2022). [↩](https://chatgpt.com/g/g-p-6863cdae32988191b6ef486a436ee4e5-nick/c/6a2fbe04-d274-83eb-a320-967f94fc91f3#user-content-fnref-2)