--- # Abstract We argue that current-generation large language models (LLMs) are capable of producing philosophical texts that are worth reading, in the same way that a good piece of human-written philosophy can be. We defend this claim against various challenges. First, one might think that a text produced by an LLM cannot count as a work of philosophy because no person lies behind it, just as one might think that an image produced by AI cannot count as an artwork. Second, one might think that LLMs lack a particular capacity, or capacities, that producing worthwhile philosophy requires: the capacity to weigh rival explanations, to connect with the world, or to draw on experience. Finally, one might ask why, if we are right, no philosophy worth reading has yet come from these systems, and whether the philosophical content of any text produced under a philosopher's prompting should be credited to the philosopher rather than to the model. We argue that none of these challenges succeeds. Two texts **with the same contents** cannot differ in philosophical merit; the distinction between genuine and merely apparent abduction cannot be located in the texts LLMs actually produce; and the absence, so far, of worthwhile LLM-written philosophy reflects how these systems are used rather than what they can produce. --- # 0. Introduction The last decade or so has seen the rise of generative artificial intelligence: systems that produce text, images, code, music, video, and other outputs in response to prompts. AI has had success in domains where the value of an output is not exhausted by its superficial fluency. There have been recent AI-assisted discoveries in physics (Guevara et al. 2026), mathematics (Novikov et al. 2025), biomedicine (Gottweis et al. 2025), and materials science (Zeni et al. 2025). In this paper we argue that we should expect similar success in philosophy. Specifically, we argue that current-generation LLMs are capable of producing philosophical texts that are _worth reading_. We do not want to begin by settling what counts as _good_ philosophy. Instead, we appeal to a distinction familiar to anyone who reads philosophy: between texts that repay the time spent reading them and texts that do not. As you begin reading this article, you likely hope that it is worth reading, in the sense that the time spent reading it will not be wasted. When you write a philosophical text yourself you aim to make it worth readers' while to read it, and whether or not the journal you send it to accepts it depends on whether or not they agree. Note that a text's being worth reading is not the same as its being correct: we take many of the philosophical texts we read to be mistaken in their conclusions, and few of them to have wasted our time. Note also that the mere statement of a philosophical conclusion is unlikely to be worth reading. A text consisting only of bare pronouncements — that direct realism is correct, that we should be utilitarians — would not repay anyone's attention, and this is as true of a human philosopher's pronouncements as of an LLM's.[^1] What would repay attention is argument, and it is the capacity of LLMs to produce philosophical arguments, rather than philosophical pronouncements, that concerns us in what follows. **We should note, finally, that our claim concerns what these systems can produce rather than what they reliably do produce; why they come apart is the question Section 5 takes up.** We will argue for this claim by considering six challenges that might be raised against it. Section 1 rejects the challenge from authorship: the claim that an LLM output cannot be philosophy worth reading because no philosopher lies behind it. Section 2 turns to abduction and argues that the absence of human-style inference to the best explanation in the producer does not preclude abductive structure in the product. Section 3 turns to the challenge from detachment, and argues that the materials philosophy takes from the world reach it already set down in words. Section 4 takes up the challenge from experience, and argues that the experiences philosophy argues about reach it in the same way. Section 5 turns to the challenge from observation — that if all this is right, philosophy worth reading should already be coming from these systems, and is not — and argues that this reflects how they are used rather than what they can produce. Section 6 addresses the challenge from instrumentality — that where a philosopher's prompting draws out such a text, the philosophy is the philosopher's — and argues that a prompt articulates a starting point whose development it does not fix. [^1]: TODO: add footnote. --- # 1. The Challenge from Authorship Philosophy might be thought to be something that only persons, or at least minds, can produce. This view has not, to our knowledge, been explicitly defended in just this form, but it gives shape to an intuition that many philosophers may have: philosophy is a person-only domain. An imperfect comparison is with art: one might deny that an image generated by an AI system is an artwork, because no artist exercises the relevant kind of intentional control over its production.[^2] One might think, for similar reasons, that philosophy can only be done by people: no text produced by an LLM can be a work of philosophy, because no philosopher lies behind it. Call this the _challenge from authorship_. There is some initial support for this thought in the way philosophy is studied. Like the study of art, and unlike the study of physics, the study of philosophy is organised around individuals: undergraduates take courses on Kant's ethics or Lewis's metaphysics, and reading an important philosopher's own words is held to be of value in a way that reading Newton's is not. A physics student is taught Newtonian mechanics from a current textbook, and the course loses nothing if the _Principia_ is never opened; a course on Kant's ethics that never opened the **second _Critique_** would scarcely count as one. In the sciences, that is, what a text contributes can be carried entirely by other texts, while in philosophy the contribution and its original presentation are harder to prise apart. One might take this as evidence that a philosophical work is bound to the activity of the person who produced it, in a way that a scientific result is not. **We will now try to make this challenge from authorship more precise, by considering how far Davies' _performance_ theory of art transposes to philosophy.** Davies writes: > [T]he work – what the artist achieves – is the process eventuating in that product. Works themselves are neither structures nor objects simpliciter, nor are they contextualized structures or objects. They are, rather, intentionally guided generative performances that eventuate in contextualized structures or objects (or events, as we shall see) – performances completed by what I am terming a focus of appreciation. (2004, p. 97) On Davies' view, when a painter paints a picture, the canvas is what we attend to, but it is not the work. The work is the artist's intentionally guided activity in producing the canvas; the canvas is "the focus of our appreciative interest in the work" (2004, p. 150). What we appreciate in a painting, on this account, is an achievement, and achievements are individuated by the activities that bring them about: the same surface, reached by some other route, would be a different achievement. Facts about how the object came into being thus do more than supply context; they help determine what the work is, and what is properly appreciated in it. Suppose, to adapt one of Davies' examples (2004, p. 102), that a storm were to blow pigment across a stretched canvas, leaving a surface indistinguishable, mark for mark, from an abstract painting. A theorist who identifies the painting with its surface must either count the storm's product as a painting or explain why it is not. There is a surface with no painting of a picture behind it, and the corresponding case for philosophy — a text indistinguishable from a philosophical argument, with no philosophising behind it — is the one an LLM presents. If Davies is right, the surface does not by itself settle the work. Transposed to philosophy, Davies' proposal would run as follows: a philosophical text is not itself the philosophical work; the text is the product of a person's philosophising, and reading it is a way of engaging with that prior activity. The activity, on this proposal, is part of what the work is, so that there is a philosophical work only where the philosophising has taken place, which makes the challenge a constitutive one. If no one has philosophised, there is no work to which the text gives access, however the text reads — an LLM text would stand to philosophy as the storm-made canvas stands to painting. We do not think the transposition should be accepted. What makes Davies' view plausible in the case of art is that indiscernible surfaces really do seem to differ in artistic value: the storm-made canvas is worth nothing as a painting, while a brushed one may be worth a great deal. Whoever transposes the view to philosophy is therefore committed to the corresponding claim, that two texts **with the same contents** can differ in philosophical merit, and it is difficult to see what such a difference could consist in. Whether an argument is valid, whether its premises are plausible, and whether the objections to it have been answered are questions about a text's contents, and two texts with the same contents receive the same answers to them. The philosophical merit of a text does not vary with the route by which its words came to be written. Peer review proceeds on the same assumption. Journals strip author information from submissions before review because facts about authorship are treated as potential sources of distortion rather than as evidence of merit; if two texts with the same contents could differ in philosophical merit, anonymising would discard information relevant to the assessment, and review would not be designed as it is. The grounds for accepting or rejecting a paper lie in the argument as presented, not in the history of its production.[^3] Nor is the author-centred teaching noted earlier in tension with this. That ethics is taught through the **second _Critique_** rather than through a digest of its conclusions reflects what a reader gains by working through Kant's arguments, and this is consistent with holding that the merit so gained is a feature of the text rather than of its author. It might still be insisted that where there has been no philosophising there is no work, whatever the resulting text contains. We need not resist this, because our thesis concerns texts worth reading, and a text can be worth reading without being a work in Davies' sense. Suppose a desert wind traced out, in the sand, a sound argument against enactivist theories of perception. Nobody would deserve credit for the argument, and there would be no performance for the text to give access to; but a reader who worked through it would still encounter a thesis and the considerations advanced in its favour, and would be in a position to answer or to extend it. 'Work' may, if the performance theorist insists, be reserved for texts with performances behind them. What cannot be so reserved is being worth reading, since everything that judgement answers to is on the page. That no one lies behind a text does not, then, settle whether it is worth reading. Whether an LLM can actually _produce_ a text worth reading is a further question, and it is to this that we now turn. [^2]: This is not to deny that systems of this kind can produce beautiful images; we return to image generation in Section 4. [^3]: The same location of philosophy in the public text is reached by accounts of philosophical progress. Dellsén et al. (2024) hold that progress is "for-whom" rather than "by-whom": it consists in putting people in a position to increase their understanding, usually by making philosophical ideas publicly available (p. 679). Being put in such a position requires something one can take up and think through, and what is available to be taken up is the text. A philosophical contribution so understood is constituted by what the public text makes available, which a view that locates the philosophy in the antecedent private activity mislocates. --- # 2. The challenge from abduction In the previous section, we argued against the idea that LLMs cannot produce philosophy worth reading simply because they are not human. In this and the two sections that follow we shall consider a different form of challenge: even if LLMs cannot be ruled out of producing philosophy worth reading, they lack particular _capacities_ that are required to produce it. If our arguments regarding authorship are correct, then a novel philosophical argument produced by a parrot would deserve to be taken just as seriously as one produced by a human being. No such argument will be forthcoming, however: a parrot can only reproduce sounds it has already heard. One might suspect that LLMs are in the parrot's position — that while nothing rules their outputs out of being philosophy worth reading _tout court_, they lack some capacity that producing it requires. **In this section we address one capacity challenge, which we will call _the challenge from abduction_. In sections 3 and 4 we shall look at two more: the _challenge from detachment_ and the _challenge from experience_.** Abductive inference is reasoning from a body of evidence to the hypothesis that would best explain it. In a deductive argument the premises fix the conclusion: if all men are mortal and Socrates is a man, then Socrates is mortal, and that is that. Now, imagine walking into your kitchen one morning and finding that part of the floor is wet. What has happened? The wet floor does not determine the answer in the way the two premises determined Socrates' mortality: a burst pipe would have left the floor wet, and so would a spilled bucket. But, given that the window is open, the water is under the window, and it rained last night, rain coming through the window seems the most plausible answer. To reason in this way, inferring the best explanation for a set of facts, is common in the sciences as well as everyday life. Williamson argues that philosophy should use a broadly abductive methodology (2007; 2021, §9.2). In philosophy, as in science, there are data that a candidate theory must accommodate — intuitions about cases, and the phenomena of the domain itself — and rival theories that would each accommodate them at different costs. The theory to prefer is the one that would, if true, best account for the data, where what makes one theory's account better than another's is its explanatory virtue: a good philosophical theory is, in Williamson's words, "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated", and should "combine simplicity with strength" (2021, §9.2). Nor is this view of theory choice confined to those who, like Williamson, take philosophy to be methodologically continuous with the sciences: Bengson, Cuneo and Shafer-Landau (2022) argue that the assessment of rival theories by their explanatory and unifying merits is a constraint on sound philosophical method as such.[^4] We shall assume it in what follows.[^5] Lipton (2004, p. 59) distinguishes two things that "the best explanation" might mean: the _likeliest_ explanation, the one most warranted by the evidence, and the _loveliest_, the one that would, if true, provide the most understanding. As he puts it, "Likeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). The two can come apart. Return to the wet kitchen floor: that water has fallen on it is the likeliest explanation of its being wet, and it is scarcely possible that it is false, yet it affords no understanding at all of how the floor came to be in that state: of course it is wet because water fell on it, but how, and which water? The explanatory virtues by which, on the methodology we have assumed, rival philosophical theories are weighed are virtues of loveliness: elegance, unity, and simplicity concern the understanding a theory would provide were it true, not the probability that it is true. Where a philosophical text turns on the weighing of rival explanations — not all philosophical writing does — whether the text is worth reading and whether it weighs its rivals well go together. If a capacity for abduction is required to produce worthwhile philosophy, do LLMs possess it? Floridi et al. (2025) argue that they do not: > We argue that such LLMs generate text based on learned associations rather than performing abductive inferences. […] LLMs produce plausible hypotheses, simulate commonsense reasoning, and provide explanatory answers without grounding them directly in truth, semantics, verification, or understanding, and without any abductive reasoning. (Floridi et al. 2025, p. 1) An LLM, on this account, has "a stochastic core and an abductive appearance" (2025, p. 2); the further charge, that these systems cannot "genuinely validate" their explanations against reality (ibid., p. 6), we take up in the next section. Models are trained to predict which words are likely to follow which, and to produce the continuation that their training makes probable; nothing in the training directs them at truth. The appearance comes from what the training data have passed on: models have "absorbed patterns of human abductive reasoning as expressed in writing" (p. 9) — how explanations are typically phrased, which causes are typically offered for which effects. [^6] Floridi et al.'s own example is a car that will not start on a cold morning. Asked why not, a model replies that a weak battery is one possibility, since cold reduces a battery's efficiency; that thickened engine oil is another, since a cold engine is harder to turn over; and that, "[b]ased on your description, the battery is the most likely explanation" (2025, p. 10). On their reading, the LLM is not actually reasoning about causes from the case before it; it is the statistical reproduction of the causes such explanations typically cite (p. 9). What they deny is that the model weighs the candidates — that it "entails choosing the best explanation among alternatives" (2025, p. 3). The verdict that settles on the battery does no weighing; it reproduces how explanations of this kind conventionally end. And where the output marks a genuine difference between the two — some consideration that would tell the battery from the oil — it is one already drawn in the explanations the model learned from, not one worked out afresh for the case in hand. Because LLMs produce no more than a "veneer of explanation" (ibid., p. 20), Floridi et al. conclude, they can play only a supporting role in intellectual work: > In essence, LLMs function like brainstorming assistants that toss out ideas without filtering for quality. After all, they work like statistical interfaces to an enormous amount of data accumulated for millennia by generations. A cautious human collaborator can sift through and assess them. (2025, p. 11) A device for brainstorming seems a far cry from something which might produce worthwhile philosophy. If you were presented with a text and told that it contains a number of philosophical ideas, none of which have been filtered for quality, it is unlikely you would think it is worth your while reading it. The filtering these systems are said to lack is, in Lipton's terms, filtering for loveliness: **what a model cannot do,** on Floridi et al.'s account, is prefer one candidate explanation to another on the grounds of the understanding it would afford. Floridi and his colleagues allow that, in ordinary cases such as their own cold-morning car, the model's answer is a good one: "the same explanation a human reasoner would likely choose" (2025, p. 10), one that may be "even optimal by IBE criteria" (2025, p. 19), since such systems "echo the obvious, common explanations" (2025, p. 10). The complaint cannot then be that the explanation is poor, which leaves it hard to say what, in such an answer, is supposed to be merely apparent. What their account points to is the uncommon case: "in less common situations, LLMs can falter" (2025, p. 10), and on inputs "that go beyond their training" "the facade can crack" (2025, p. 9), the success on familiar cases being "a sign of overfitting to common patterns" (2025, p. 15). **A model overfits when it memorises the patterns of the examples it was trained on rather than the regularity behind them: its success then carries over only to cases resembling those examples.** Once the question of its truth is set aside, this is what the facade reduces to: the claim that the competence shown on common problems is overfitting that would give out in less common ones. Whether the competence does give out on uncommon cases is an empirical question. A recent survey of abductive reasoning in language models suggests that it does (Salimi et al. 2026). Current models do markedly worse on abductive tasks than on deductive ones: where their median accuracy on deductive tasks is near eighty per cent, on abductive ones it is some forty-two and a half, with a spread "extending down to near-zero accuracy", and the survey reports that "strong deductive performance does not reliably imply strong abductive performance". This is close to Floridi et al.'s own diagnosis — an answer that reproduces a common pattern instead of reasoning to it — and, taken at face value, it tells in their favour. We do not disagree with Floridi et al.'s characterisation of how LLMs function: these systems do not weigh and choose among alternatives in the way that humans do. However, we should not be too quick to jump from this to the conclusion that LLMs cannot _produce text_ which exhibits abductive reasoning. **Consider the artificial images made by generative AI: a picture of a cat created in this way is no less a picture of a cat, yet it is produced quite differently from any picture a human would make by painting or photography.**[^milliere]** Similarly, it might be possible for LLMs to produce text which displays abductive reasoning, **without any actual abductive reasoning behind it**. Despite their stochastic core, LLMs are perfectly capable of producing grammatically correct text: without being given any specific rules, the system "implicitly 'discovers' them—and then seems to be good at following them" (Wolfram 2023). Does this mean that the texts LLMs produce have merely the appearance of being grammatically well-formed? Clearly they do not: LLM sentences _are_ grammatically well-formed, **regardless of their stochastic roots**. A stochastic core, then, need not mean that the best an LLM can manage is a veneer of abductive inference. It may be, rather, that the core is marshalled to produce text exhibiting actual abductive inference, in just the way it is marshalled to produce actual grammatical correctness. It might be objected that the grammar analogy will not stretch this far. Grammar, the objector would say, is a system of rules a model can follow without understanding anything, and explanatory loveliness is not — it is not a system of rules you follow to get the right answer, so the model's grammatical competence gives no reason to expect it to produce lovely explanations. If Floridi et al. are right that the output has the form of an inference to the best explanation but not its substance, then there should be cases where the form is present and the substance is absent: cases where the model produces something that looks like an explanation but is empty or nonsensical, just as it could produce something that is grammatically faultless and says nothing. Wolfram's example of the latter is "Inquisitive electrons eat blue theories for fish" (2023) — impeccably grammatical, and meaningless. The model **almost never produces** such strings. The abductive equivalent would be a passage with the form of an inference to the best explanation, grammatically faultless, and yet senseless — asked why a car will not start on a cold morning, there are no squirrel tracks, so it must be freak arctic winds blowing into the exhaust pipe. That has the form of evidence weighed towards a conclusion, and it is nonsense; and it is nonsense the model **almost never produces**. Put the question to it and it offers the weak battery and the thickened oil, plausible explanations rather than arctic winds. Wherever the facade is to be located, it cannot be located there: what the model produces are, **for the most part,** plausible explanations. Moreover, the answer Floridi and his colleagues describe is not quite the answer these systems give. Their account asks us to picture a confident verdict laid over a hollow core, yet what one finds in practice is closer to hedging. Asked why a car would not start on a cold December morning, a current model will say that the battery is the most likely culprit, set out other plausible contributors — thickened oil, fuel-system problems, ignition faults — and end by observing that if the car started once the day had warmed, the battery is almost certainly the primary cause, though a load test would confirm it.[^7] The model's avoidance of senseless explanations, like its avoidance of senseless sentences, points to a semantic competence picked up in training, over and above syntax. A syntactic grammar settles only how the parts of speech may be combined, and not whether what a sentence says makes sense: avoiding strings like the inquisitive electrons takes a grasp of what can sensibly be said of what — which predicates go with which subjects, which causes with which effects. Wolfram calls such a grasp a semantic grammar: > to deal with meaning, we need to go further. And one version of how to do this is to think about not just a syntactic grammar for language, but also a semantic one. (Wolfram 2023) A model trained on enough text has, on his account, come by one: > From its training ChatGPT has effectively "pieced together" a certain (rather impressive) quantity of what amounts to semantic grammar. (Wolfram 2023) What the training pieces together extends beyond the fit of predicate to subject. The arctic winds owed their senselessness as much to the inference as to its parts — nothing in the absence of squirrel tracks bears on what enters an exhaust pipe — and Wolfram suggests that inference is picked up by the same route: just as Aristotle might have arrived at syllogistic logic by working through many examples of rhetoric, a model in training can "discover syllogistic logic" in the text it reads, and can be expected to produce text containing "correct inferences" (2023).[^8] It is this, rather than any contact with the case in hand, that keeps the arctic winds out of the model's explanations: a feel for what, in a working model of the world, can hang together.[^9] But a feel for how things hang together is gathered from the text the model has read. It amounts to a grasp of how failures of this kind are explained in general — that cold weather weakens batteries, say — and it gives the model no means of finding out which explanation is true of the particular car whose fault is in question. It does, however, underwrite the weighing itself: the model's answer sets out the candidate explanations and fits its confidence to the evidence it has been given, and neither of these requires access to the car. What requires such access is settling which candidate is correct. In much philosophical abduction there is no analogue of the car: what a thought experiment commits us to, or which of two theories carries the lighter explanatory cost, is settled from what is already set down in writing, and a corpus already holds evidence of that sort. The survey scores of Salimi et al. (2026) measure whether the model arrived at the right answer, not the weighing that got it there. On the harder tasks the scores are low, and on long mysteries with their clues strewn through the text the best models fall just short of the average human solver; taken at face value, the numbers count against the model. But the score is a score for the answer — whether the named culprit was the keyed one, whether the right diagnosis came first — and not for the weighing that reached it; such a score, Salimi et al. note, "completely bypass[es] the actual reasoning trace". So when a model misses the keyed culprit, the number marks the miss, and not the comparison it set out on the way — which explanations it canvassed, and why it came down on one. And the survey's hardest tasks are built from low-prior, non-stereotypical outcomes, where several explanations may be reasonable; there, matching an answer to a single reference "underestimates explanation quality", so a low score on such tasks does not show that the model cannot weigh, only that it did not land on the keyed answer. It is now difficult to see where the facade should be located. The explanations the model produces are plausible, the confidence they express is fitted to the evidence it has been given, and the benchmark scores that seemed to confirm the diagnosis measure answers rather than the weighing that produced them. What the model lacks is access to the particular case, and that is a limitation concerning its relation to the world rather than its capacity for abduction. What that costs a philosophical text is a matter for the next section. [^4]: TODO: add footnote. [^5]: Understanding-based accounts of philosophical progress converge on the same standard: if progress in philosophy consists in placing people in a position to increase their understanding (Dellsén et al. 2024, p. 679), then rival theories are rightly weighed by the understanding they would, if true, provide. [^6]: TODO: add footnote. [^7]: The exchange, with Kimi k2.67 (%%add date of test%%), is reproduced in full in the Appendix. This is not yet to say that the facade charge is refuted: a hedged answer can still be, on Floridi et al.'s account, a statistical reproduction of how hedged explanations typically read. [^8]: Wolfram's expectation comes with a caveat: on "more sophisticated formal logic" he expects the model to fail, for the same reasons it fails to match parentheses across long sequences (2023). The caveat concerns extended derivation, and the weighing of explanations is not derivation: as the objector above puts it, loveliness is not a system of rules one follows to get the right answer. [^9]: We use "model of the world" in Wolfram's informal sense: a body of implicitly learned regularities concerning what goes with what. This is weaker than the technical sense now disputed in machine learning, on which a world model is an internal representation of an environment that supports prediction and counterfactual reasoning. Nothing in our argument requires deciding whether LLMs possess world models in the technical sense; how they connect to the world is the matter of §3. **[^milliere]: For an account of how images of this kind are produced, see Millière (2022).** --- # 3. The Challenge from Detachment A second capacity challenge, the _challenge from detachment_, concerns the model's relation to the world. In addition to the argument considered in the previous section, Floridi et al. object that an LLM stands in no relation to it: its words rest on no perception of anything, and a hypothesis, once produced, is never tested against how things are (2025, pp. 6–7). A discipline whose theories answer to how things are, the challenge runs, cannot be advanced by a system with no access to how things are, and a text produced by such a system gives its reader no reason to think it worth reading. Zahavy (2026) raises a worry of this kind about scientific discovery. A model can carry out the deductive part of discovery, working out the consequences of premises it has been given; what it cannot do, he holds, is produce the premises — **make the leap from sense experience to new first principles.[^11]** His case is the thought experiment that gave Einstein the equivalence principle: > Einstein's variation required inventing new axioms based on a physical intuition that did not yet exist in the mathematics. He envisioned a physicist inside an elevator being uniformly accelerated through deep space [...]. Inside this enclosure, the sensory experience reveals a specific pattern: when objects are released, the floor rushes up to meet them. To the physicist, the objects appear to fall with identical acceleration, regardless of composition. Thus, the simulation here was not a permutation of symbols, but a manipulation of perceptual experience. (Zahavy 2026, §5) Einstein imagines a set of circumstances and attends to what would be experienced within them: everything released inside the elevator appears to fall with identical acceleration. On Zahavy's reconstruction the simulation supplies an observation, and from that observation the new axiom is inferred — the simulated experience of acceleration was indistinguishable from the remembered experience of gravity, and Einstein concluded that the two are one phenomenon.[^12] A model has no access to that observation. It can produce descriptions of elevators and of weightlessness, both present in its corpus, but it has undergone neither, and a discovery whose premises are fixed by simulated experience is beyond a system that, in Zahavy's words, lacks the capacity he calls sensory agency. **While Zahavy often uses phrases which suggest that it is a model's lack of phenomenology that is the principal stumbling block — Einstein arrived at the equivalence principle by "simulating the physical feelings" of his observer, guided by "the sensation of gravity" (2026, §5) — on closer inspection it is actually better thought of as an argument about a model's access to the world. If a model is to make the leap, Zahavy argues, it must be equipped with an interactive world model, a system that simulates a physical world within which the model can act and observe what follows (2026, §5). The fact that such a system contains no phenomenology, yet would supply what the model is missing, shows that what is missing is access rather than phenomenology. Having said that, one might ask whether a lack of phenomenology does in fact preclude an LLM from doing at least some sorts of philosophy. How can a system which is not conscious produce worthwhile philosophy about consciousness? We will examine this question in the next section.** We argue that the challenge fails because the materials philosophy takes from the world reach it already articulated in language. We can grant that the model perceives nothing, and that nothing it produces is put to the test against the world. Whether these concessions disqualify it depends on how philosophy itself meets the world. Pigliucci (2017), writing on how philosophy makes progress, holds that philosophy is constrained by the world without investigating it as the natural sciences do: > This means that the basic parameters that philosophers use as their inputs, the starting points of their philosophizing, their equivalent of axioms in mathematics and assumptions in logic (or rules in chess) are _empirical_ data about the world. This data comes from both everyday experience [...] and of course increasingly from the world of science itself. Philosophy, I maintain, is in the business of exploring the sort of conceptually evoked spaces that Smolin is talking about, where the evocation is the result of whatever starting assumptions are made by individual philosophers working within a particular field and, crucially, of the constraints that are imposed by our best understanding of how the world actually is. (2017, pp. 79–80) Pigliucci takes the term _evocation_ from Smolin (Unger and Smolin 2015).[^10] Evoked truths are neither discovered, in the sense of corresponding to mind-independent states of affairs, nor invented, in the sense of being arbitrary constructs. Smolin's illustration is chess: before the rules of the game were codified there were no facts about chess, but once they were, a great many facts about the game became demonstrable — objective facts, in the sense that anyone who demonstrates one demonstrates the same fact as anyone else (Unger and Smolin 2015, p. 423; quoted at Pigliucci 2017, p. 78). On Pigliucci's proposal, philosophy is in the business of ascertaining evoked truths. **Unlike the axioms of mathematics or the rules of chess, however, philosophy's starting points are constrained by how the world actually is. They are empirical data about the world, drawn from everyday experience and, increasingly, from science.** A system that perceives nothing cannot gather such data for itself; but nor do **most philosophers**. The scientific data Pigliucci has in mind reach philosophers already set down in language: a philosopher of physics works from published results, not from having run the experiments. **Everyday experience, Pigliucci's other source, requires no experiments in the first place. A philosopher who appeals to how things ordinarily seem is not reporting the result of an investigation, and descriptions of how things seem are everywhere in written language, whether philosophical or not.** Written access to the world's constraints is the profession's normal condition, and a corpus provides access of just this kind. With respect to nearly all the empirical data philosophy uses, then, the model is in the same position as any philosopher. The same constraint distinguishes philosophy from fiction. Pigliucci's contrast case is a science-fiction writer who describes the same planet across three different timelines, keeping every description within the laws of physics and of logic. The writer is exploring conceptual spaces of a kind, but he is inventing rather than evoking: his worlds have no rigid properties, because even the constraints he keeps to could have been otherwise — he could as easily have imagined planets with a different physics, or a different logic. Philosophy, on the other hand, "is in the business of doing empirically informed evoking, not inventing", and its objects of study accordingly have rigid properties (2017, p. 80): once a philosophical starting point has been articulated, what holds within the structure it opens is no longer up to anyone, any more than what follows from the rules of chess is up to anyone. **In a thought experiment,** the philosopher sets up an imagined scenario, but explores it "with an interest in figuring things out as far as this world is concerned" (2017, p. 80), and the philosophical work then proceeds within a structure whose properties the philosopher does not control. **Philosophical claims are not, in any case, tested against the world. The equivalence principle could have been refuted by experiment; a philosophical thesis could not.** Whether it holds is a question about what obtains within an evoked structure, and it is settled in the way questions about chess are settled: by working out what the starting points commit us to. This working-out is carried out in writing, in the literature's back-and-forth of argument and reply, and it is work of a kind Section 2 argued a model's text can display. A model's inability to run experiments therefore costs it nothing that a philosophical text needs: the testing philosophy does is done on the page. # 4. The Challenge from Experience A third capacity challenge, the _challenge from experience_, holds that some philosophy depends on **subjective** experience in a way that a system which has never had any cannot meet. Few would say that current LLMs are conscious, and we assume here that they are not. Does Mary, released from her black-and-white room knowing every physical fact about colour vision, learn something when she first sees red (Jackson 1982)? To settle the question one must consider what the experience of seeing red is like, and the argument proceeds from that verdict. Here experience enters as the starting point of an argument, and a system that has never experienced anything seems unable to supply it. Experience enters philosophy as subject matter too, in work that asks what it is to see red (Harman 1990), to feel anger (Goldie 2000), or to have an intuition take hold (Chudnoff 2011).[^13] If starting points of the first kind, and subject matters of the second, can be handled only by a being that has the experiences, then much of the philosophy of mind lies beyond a model's reach. Experience does not, however, enter the knowledge argument in the way the challenge assumes. The facts about colour experience that the argument **draws on** are general ones, articulated many times in the literature: that colour is experienced as a quality of the things seen, and that one cannot imagine a colour one has never seen — which is why Mary's room must be black and white. The texts produced in response to Jackson's paper work with claims of this kind and with his description of the case, contesting what they commit us to — whether what Mary gains on her release is knowledge of a fact at all. **None of this requires having seen red for oneself:** a reader of Jackson's paper need not introspect the fine details of their own experience of red to follow his argument, or the literature that discusses it. Philosophy does not usually begin from raw encounter; it works on what has already been articulated. Even where a case does begin in its author's own experience — as Jackson's may have — what enters the argument is the description he makes of it. Descriptions of this kind — in which what experience is like has been set down in words, and can be read and reasoned from like any other claim — are what we will call _articulated phenomenology_. Consider, for example, Austin's discussion of colour phenomenology in _Sense and Sensibilia_, where aspects of what it is like to see colour are carefully described and considered (e.g. how a pointillist painting's appearance changes depending on how far one stands from it, or how an object's colour changes under different illumination; 1962, pp. %%page%%, 66). Philosophers are not the only ones in an LLM's corpus to articulate their phenomenology in text; we might also consider novels, diaries, interviews — anything in which someone describes _what it is like_ to have an experience. An LLM has never felt what Zahavy's physicist feels inside the elevator, but it has read the memoirs of astronauts who have. We saw in Section 3 that written access to the world's constraints is the profession's normal condition, and experience is no exception. A model has undergone none of these experiences, but the descriptions of them are in its corpus, and it is the descriptions that philosophy works from. What models do lack is the capacity to make _novel_ phenomenological observations. An example of this sort of novelty might be Merleau-Ponty's (1945 %%page%%) description of how, when one hand touches the other, the roles of toucher and touched can alternate but cannot be occupied at once. The experience is a familiar one, but the asymmetry it contains seems not to have been described before. Noticing it required having hands, and tactile experience of them; a model has neither, and so could not have been the first to set it down. However, philosophical arguments which rely on novel phenomenological observations are rare indeed: the descriptions of experience that philosophy works with have, for the most part, long been in the literature. And once an observation of this kind is written into the corpus, everything philosophical that can be done with it — what it shows about bodily self-awareness, what follows once the asymmetry is granted — is there on the page for any reader, and none of it depends on being able to have made the observation oneself. Nor is first-person attention the only source of descriptions the literature lacks: a new description can also be reached from those it already contains, by drawing out what they have not been taken to imply, or by combining them as no one has, and this a model can do. There is, then, one route to new starting points that a model cannot take. Whether the route it can take yields descriptions the literature lacks is the question of novelty, which Section 6 takes up. A model can, then, do the philosophy that turns on experience, with one exception: it cannot be the source of a novel phenomenological observation. Whether experience enters as an argument's starting point or as its subject matter, it reaches the model as articulated phenomenology, which it can work through as well as any reader. [^10]: _The Singular Universe and the Reality of Time_ is jointly authored, but its second part, which contains the discussion of evocation, was written by Smolin alone, as Pigliucci notes (2017, p. 77); we follow him in attributing the view to Smolin. [^11]: Zahavy too calls this leap abduction, but the word picks out something other than it did in Section 2. There, with Floridi et al., abduction was the weighing of rival explanations, and the charge was that a model only mimics it; here it is the generation of new first principles from experience, and Zahavy's claim is that a model cannot make the move because it has had no experience to move from. The present challenge rests on that second claim, about experience, and not on any verdict about the weighing. [^12]: Zahavy, following Magnani, calls the process _manipulative abduction_: hypothesis generation through the manipulation of a model — here a simulated experience — rather than of symbols (Magnani et al. 2009; Zahavy 2026, §5). It is abduction in Section 2's sense: the equivalence principle is inferred as the best explanation of the simulated observation, the simulation supplying an explanandum that no search over existing text would have produced. What experience contributes, on this picture, is not the inference but its starting point. [^13]: The materials need not be sensory: the feeling of understanding something is sometimes used to motivate the claim that thought itself has a phenomenology (Pitt 2004). --- --- # 5. The Challenge from Observation The challenge from observation begins with an obvious question: if all of these arguments are correct, where is all the LLM philosophy?[^14] If you ask an LLM the answer to the hard problem of consciousness or the meaning of life, the response you receive will not typically contain much in the way of philosophical argument. More likely you will get a somewhat bland survey of various philosophical positions on the topic, or perhaps just an irreverent joke.[^15] One might think that whatever the merits of the arguments we have put forward so far, _something_ must be preventing them from producing philosophy worth reading because they _don't_ produce philosophy worth reading. We want to suggest that this is a failure of prompting rather than evidence that LLMs lack the capability to produce philosophy worth reading. As we saw in Section 2, these systems are trained to produce reasonable continuations of the text they have been given (Wolfram 2023), before being trained further to converse as helpful assistants. In the text these systems are trained on, questions like "what is the meaning of life?" are hardly ever followed by philosophical argument: arguments occur in journal articles, which begin from claims already advanced and objections already raised rather than from questions, and it is these that their arguments continue. A similar point applies to the assistant training: an assistant is trained to answer the question it is asked, and answering a philosophical question is not the same as arguing for an answer to it. None of this shows that LLMs can produce philosophy worth reading, but it does mean that their failure to produce it when questioned point-blank tells us nothing either way. How, then, can we elicit philosophy from these systems? What counts as a reasonable continuation depends on what is being continued, and a bare question is rarely continued by philosophical argument. To elicit philosophy from an LLM, we suggest, is to give it a text of which philosophical argument is a reasonable continuation. One obvious way is simply to _ask_ for careful, rational arguments about a philosophical question: a request for careful reasoning about a stated problem is continued by text in which that reasoning is carried out. One can refine the request by giving the model a problem or set of facts and asking for the explanation that would, if true, provide the most understanding of them — that is, by asking it for an inference to the best explanation. We saw in Section 2 that a model's continuations are produced under the semantic grammar it has acquired — a grasp not only of what can sensibly be said of what, but also of which inferences may correctly be drawn. A question constrains its continuation only in its topic: almost any competent remark about consciousness is a reasonable continuation of "what is consciousness?", and the survey of positions is the most reasonable of all. An argument constrains its continuation in its content. The claims it advances and the objections it raises stand in the text being continued, and each further step must be reasonable in the light of them: a continuation that contradicted what the prompt conceded, or passed over a pressure it raised, would not be a reasonable continuation of it. A prompt that contains an argument therefore fixes more than a subject matter; it fixes what the continuation must answer to. **Nor need what answers to it be merely argument-shaped.** The senseless passage that merely looks like reasoning is what the model almost never produces — that was the lesson of the arctic winds — and the inferences its grammar licenses are, for the most part, correct ones. A bare question also articulates little: it advances no claims, and where little has been articulated, little is evoked — there is little structure whose facts a continuation could state. An argument, by contrast, sets claims down, and we saw in Section 3 that this is how philosophy's structures are opened: once a starting point has been articulated, a great many facts about it become demonstrable, though almost none of them are stated. A prompt containing an argument therefore **articulates a position and the pressures upon it,** and thereby evokes, and what it evokes is there for the continuation to develop. The developing is done one step at a time. At every step, what the model continues is the whole text so far — the prompt together with everything it has itself already written (Wolfram 2023) — so each inference the continuation draws becomes part of what its next step must reasonably continue. Early moves open commitments that later moves are made under, and the further a continuation runs, the more of what it is continuing was never in the prompt at all. Nor does the prompt fix the continuation it will receive: at any point there are several ways of going on that would each be reasonable (Wolfram 2023), and a prompt constrains its continuation without fixing it. What the steps produce, where the prompt is philosophical, are the developments themselves: consequences drawn from the starting point, rivals weighed against one another. How much of a starting point a prompt articulates is a matter of degree, and the more it articulates, the more there is for a continuation to develop. Nothing restricts evocation to a single form of prompt: several hundred pages of draft manuscript make starting assumptions and accept constraints as surely as two paragraphs of thought experiment do, and which forms of prompt evoke most productively is an empirical question which we shall not try to settle. It is in the continuations of such prompts that the question of these systems' philosophical capacities would be decided.[^16] If a philosopher must articulate the starting point, then it will be said that whatever philosophy the resulting text contains should be credited to the philosopher, and that the model is merely its instrument. This is the challenge of the next section. [^14]: Strictly speaking, the observation is difficult to verify. To know that no LLM-written philosophy has appeared in the journals over the last half-decade, one would need to know that every author who has declared that they did not use AI in preparing their manuscript was telling the truth. **Most journals, moreover, do not at present require any such declaration, so an LLM-written text could have appeared in them without any author having lied.** [^15]: TODO: add footnote. [^16]: TODO: add footnote. --- --- # 6. The Challenge from Instrumentality Grant that philosophy worth reading comes out of these systems only when a philosopher prompts for it. If the philosopher must put so much in, one might suspect that the philosophy is the philosopher's — that the model does not produce philosophy but rearranges and remixes what it is given. Call this the _challenge from instrumentality_. Just as a ventriloquist's dummy is silent unless the ventriloquist is beside it, an LLM produces no philosophy unless a philosopher prompts it. The challenge from authorship held that a model's text cannot be philosophy at all. The present challenge is narrower: even if such a text is philosophy worth reading, none of the philosophy in it is the model's contribution. Dependence of this kind is, however, the normal condition of philosophical writing. A reply is elicited by the objection it answers, and nobody counts the reply as the objector's contribution; what the objector contributed is the objection. Whether the challenge succeeds therefore depends on what the prompt contributes and on what the continuation adds. We saw in the previous section that a prompt which articulates a position and the pressures upon it evokes, and that the continuation is produced from this starting point under the semantic grammar the model has acquired; what it states are facts of the evoked structure. There are, then, three contributions to distinguish: the articulated starting point, which is the prompter's; the structure it evokes, whose facts belong to no one; and the text that develops them, which is the model's. A prompt settles only a starting point, and the consequences the model's text goes on to state were stated by no one beforehand. We saw in Section 3 that once a starting point has been articulated, what holds within the structure it opens is no longer up to anyone — and that includes the person who articulated it. Smolin observes that whoever lays down the rules of a game is afterwards in the same position as everyone else: exploring the game feels like exploring a pre-existing territory, because at each point there is little or no choice about what its facts are, and the territory holds surprises for its own inventor (Unger and Smolin 2015, p. 422). The same holds for a philosopher and the starting point she articulates, and for the same reason. Articulating it makes its developments demonstrable, but it demonstrates none of them: a demonstration is the working-out of what the starting point commits us to, and an articulation does no such work. What a prompt states, then, is a starting point; the developments it makes demonstrable are stated nowhere in it. A model given such a prompt has nothing to rearrange but the starting point itself, and no rearrangement of a starting point is a demonstration of its developments; a continuation that states them has demonstrated them. Nor does the prompt select among the demonstrations available: several continuations of it would each be reasonable, and a prompt constrains its continuations without fixing them. It might still be insisted that the developments were contained in the starting point all along, as theorems are contained in the axioms that entail them, and that in drawing them out the model adds nothing of its own. Even in mathematics, however, the containment does not do the demonstrating: the axioms make a theorem provable, and the proof still has to be produced. And philosophical developments do not stand to their starting points as theorems stand to axioms in any case. We saw in Section 2 that where a philosophical text weighs rival explanations, what decides between them is loveliness — the understanding each rival would, if true, provide — and that loveliness is not a system of rules one follows to get the right answer. A theorem can be recovered from its axioms by derivation; a verdict that one rival carries a lighter explanatory cost than another cannot be recovered from the prompt by any procedure, because it is reached by weighing, and weighing is not derivation. What the prompt contains is the problem; the weighing is done in the continuation. Suppose, then, that the starting point is the philosopher's own — a case or an objection the literature does not contain. The structure it evokes is new, and until the prompt was written there were no facts about it to state; whatever the continuation states of it has been stated by no text before it. Novelty of this kind is measured against the prompt, and it is what the instrumentality challenge turns on: the prompter did not state these developments, and so they are not her contribution. Whether they are new to philosophy is a further question, for arguments repeat across texts — two texts can contain the same argument — and a continuation may state what nothing has stated about this case while containing no argument the literature lacks. A starting point the literature does not contain makes the stronger novelty available; whether a continuation achieves it is settled by reading it against the literature, as the narrower question was settled by reading it against the prompt. And whether what it states was worth stating is judged as any philosophical text is judged, by whether the considerations it cites genuinely tell between the positions and by the understanding they would, if correct, provide. One kind of starting point remains beyond prompting. We saw in Section 4 that a model cannot be the source of a novel phenomenological observation, and prompting does not change this: a prompt supplies text, not experience, and an observation of this kind must be noticed before it can be set down. Whether philosophy worth reading ever requires such observations among its starting points is a question we leave open. The person writes the prompt, decides which continuations to pursue and when to stop, and without her there is only the response to the point-blank question; this much of the challenge is correct. But asking for careful reasoning is not doing the reasoning, and choosing among continuations is the work of an editor, not of an author. Nor can she make the developments hers by enriching the prompt. A richer prompt is a larger starting point, and it asks more of the continuation: what would add something to two paragraphs of thought experiment may add nothing to several hundred pages of manuscript. But however much the prompt states, the consequences the output draws were not among the statements; and we saw in the previous section that at every step the model continues the whole text so far, so that the further a continuation runs, the more of what it is continuing was never in the prompt at all. Whether anything has been added is settled by reading the two together. Where a continuation merely rephrases its prompt, the model has done no more than rearrange, and the philosophy is the prompter's; where it states what the prompt did not, the surplus is not hers, because she did not state it. The labour, then, divides more cleanly than the challenge allowed. Philosophy, Pigliucci writes, attempts "to evoke rational conclusions arising from certain ways of looking at a given problem or set of facts" (2017, p. %%page%%). The way of looking at the problem is stated in the prompt; the rational conclusions arise only in the continuation. A prompt supplies a starting point without fixing what follows from it, and the philosophical standing of any output is settled by what its continuation adds — first to the prompt that occasioned it, and then to the literature it would join. # Appendix: maybe this should go somewhere in Sec. 2 > User: Last December, I tried to start my car early in the morning when it was very cold. It wouldn't start. Why not? > > Kimi k2.67: The most likely culprit is the battery. In very cold weather, a battery's chemical reactions slow dramatically, reducing its available capacity by up to 50%. If the battery was already marginal (older than three years, or not fully charged), freezing temperatures can push it below the threshold needed to turn the starter motor. You may have heard rapid clicking or nothing at all. > > Other plausible contributors: > > - Thickened engine oil — Cold increases viscosity, making the engine physically harder to crank and demanding even more current from the already weakened battery. > - Fuel system — Condensation in the fuel lines can freeze, blocking flow. Diesel fuel can gel; petrol can struggle to vaporise properly in extreme cold. > - Spark/ignition — Moisture in the distributor cap or on spark plugs can freeze, preventing a proper spark. > > If it started fine once temperatures rose later in the day, the battery is almost certainly the primary cause. A load test would confirm whether it needs replacement or just a longer drive to reach full charge. --- # References **Millière, R. (2022). Deep learning and synthetic media. _Synthese_, 200, 231. https://doi.org/10.1007/s11229-022-03739-2** Frankish, K. (2024). What are large language models doing? In A. Strasser (Ed.), _Anna's AI Anthology: How to live with smart machines?_ (pp. 55–78). Xenomoi.