# Abstract
We argue that current-generation large language models (LLMs) are capable of producing philosophical texts that are worth reading, in the same way that a good piece of human-written philosophy can be. We defend this claim against various challenges. First, one might think that a text produced by an LLM cannot count as a work of philosophy because no person lies behind it, just as one might think that an image produced by AI cannot count as an artwork. Second, one might think that LLMs lack a particular capacity, or capacities, that producing worthwhile philosophy requires: the capacity to weigh rival explanations, to connect with the world, or to draw on experience. Finally, one might ask why, if we are right, no philosophy worth reading has yet come from these systems, and whether the philosophical content of any text produced under a philosopher's prompting should be credited to the philosopher rather than to the model. We argue that none of these challenges succeeds. Two texts containing the same argument cannot differ in philosophical merit; the distinction between genuine and merely apparent abduction cannot be located in the texts LLMs actually produce; and the absence, so far, of worthwhile LLM-written philosophy reflects how these systems are used rather than what they can produce.
---
# 0. Introduction
The last decade or so has seen the rise of generative artificial intelligence: systems that produce text, images, code, music, video, and other outputs in response to prompts. AI has had success in domains where the value of an output is not exhausted by its superficial fluency. There have been recent AI-assisted discoveries in physics (Guevara et al. 2026), mathematics (Novikov et al. 2025), biomedicine (Gottweis et al. 2025), and materials science (Zeni et al. 2025). In this paper we argue that we should expect similar success in philosophy.
Specifically, we argue that current-generation LLMs are capable of producing philosophical texts that are _worth reading_. We do not want to begin by settling what counts as _good_ philosophy. Instead, we appeal to a distinction familiar to anyone who reads philosophy: between texts that repay the time spent reading them and texts that do not. As you begin reading this article, you likely hope that it is worth reading, in the sense that the time spent reading it will not be wasted. When you write a philosophical text yourself you aim to make it worth readers' while to read it, and whether or not the journal you send it to accepts it depends on whether or not they agree.
Note that a text's being worth reading is not the same as its being correct: we take many of the philosophical texts we read to be mistaken in their conclusions, and few of them to have wasted our time. Note also that the mere statement of a philosophical conclusion is unlikely to be worth reading. A text consisting only of bare pronouncements — that direct realism is correct, that we should be utilitarians — would not repay anyone's attention, and this is as true of a human philosopher's pronouncements as of an LLM's.[^1] What would repay attention is argument, and it is the capacity of LLMs to produce philosophical arguments, rather than philosophical pronouncements, that concerns us in what follows.
We will argue for this claim by considering six challenges that might be raised against it. Section 1 rejects the challenge from authorship: the claim that an LLM output cannot be philosophy worth reading because no philosopher lies behind it. Section 2 turns to abduction and argues that the absence of human-style inference to the best explanation in the producer does not preclude abductive structure in the product. Section 3 turns to the challenge from detachment, and argues that the materials philosophy takes from the world reach it already set down in words, which an LLM can work on as any philosopher does. Section 4 takes up the challenge from experience, and argues that the experiences philosophy argues about reach it in the same way. Section 5 turns to the challenge from observation — that if all this is right, philosophy worth reading should already be coming from these systems, and is not — and argues that this reflects how they are used rather than what they can produce. Section 6 addresses the challenge from instrumentality — that where a philosopher's prompting draws out such a text, the philosophy is the philosopher's — and argues that a prompt articulates a starting point whose development it does not fix.
---
# 1. The Challenge from Authorship
**Philosophy might be thought to be something that only persons, or at least minds, can produce.** This view has not, to our knowledge, been explicitly defended in just this form, but it gives shape to an intuition that many philosophers may have: philosophy is a person-only domain. An imperfect comparison is with art: one might deny that an image generated by an AI system is an artwork, because no artist exercises the relevant kind of intentional control over its production.[^1] One might think, for similar reasons, that philosophy can only be done by people: no text produced by an LLM can be a work of philosophy, because no philosopher lies behind it. **Call this the _challenge from authorship_.**
There is some initial support for this thought in the way philosophy is studied. Like the study of art, and unlike the study of physics, the study of philosophy is organised around individuals: undergraduates take courses on Kant's ethics or Lewis's metaphysics, and reading an important philosopher's own words is held to be of value in a way that reading Newton's is not. A physics student is taught Newtonian mechanics from a current textbook, and the course loses nothing if the _Principia_ is never opened; a course on Kant's ethics that never opened the _Groundwork_ would scarcely count as one. In the sciences, that is, what a text contributes can be carried entirely by other texts, while in philosophy the contribution and its original presentation are harder to prise apart. One might take this as evidence that a philosophical work is bound to the activity of the person who produced it, in a way that a scientific result is not.
**How far does Davies' _performance_ theory of art transpose to philosophy?** Davies writes:
> [T]he work – what the artist achieves – is the process eventuating in that product. Works themselves are neither structures nor objects simpliciter, nor are they contextualized structures or objects. They are, rather, intentionally guided generative performances that eventuate in contextualized structures or objects (or events, as we shall see) – performances completed by what I am terming a focus of appreciation. (2004, p. 97)
On Davies' view, when a painter paints a picture, the canvas is what we attend to, but it is not the work. The work is the artist's intentionally guided activity in producing the canvas; the canvas is "the focus of our appreciative interest in the work" (2004, p. 150). What we appreciate in a painting, on this account, is an achievement, and achievements are individuated by the activities that bring them about: the same surface, reached by some other route, would be a different achievement. Facts about how the object came into being thus do more than supply context; they help determine what the work is, and what is properly appreciated in it.
Suppose, to adapt one of Davies' examples (2004, p. 102), that a storm were to blow pigment across a stretched canvas, leaving a surface indistinguishable, mark for mark, from an abstract painting. A theorist who identifies the painting with its surface must either count the storm's product as a painting or explain why it is not. There is a surface with no painting of a picture behind it, and the corresponding case for philosophy — a text indistinguishable from a philosophical argument, with no philosophising behind it — is the one an LLM presents. If Davies is right, the surface does not by itself settle the work.
Transposed to philosophy, Davies' proposal would run as follows: a philosophical text is not itself the philosophical work; the text is the product of a person's philosophising, and reading it is a way of engaging with that prior activity. The activity, on this proposal, is part of what the work is, so that there is a philosophical work only where the philosophising has taken place, which makes the challenge a constitutive one. If no one has philosophised, there is no work to which the text gives access, however the text reads — an LLM text would stand to philosophy as the storm-made canvas stands to painting.
We do not think the transposition should be accepted. What makes Davies' view plausible in the case of art is that indiscernible surfaces really do seem to differ in artistic value: the storm-made canvas is worth nothing as a painting, while a brushed one may be worth a great deal. Whoever transposes the view to philosophy is therefore committed to the corresponding claim, that two texts containing the same argument can differ in philosophical merit, and it is difficult to see what such a difference could consist in. Whether an argument is valid, whether its premises are plausible, and whether the objections to it have been answered are questions about a text's contents, and two texts with the same contents receive the same answers to them. The philosophical merit of a text does not vary with the route by which its words came to be written.
Peer review proceeds on the same assumption. Journals strip author information from submissions before review because facts about authorship are treated as potential sources of distortion rather than as evidence of merit; if two texts with the same contents could differ in philosophical merit, anonymising would discard information relevant to the assessment, and review would not be designed as it is. The grounds for accepting or rejecting a paper lie in the argument as presented, not in the history of its production.[^3] Nor is the author-centred teaching noted earlier in tension with this. That ethics is taught through the _Groundwork_ rather than through a digest of its conclusions reflects what a reader gains by working through Kant's arguments, and this is consistent with holding that the merit so gained is a feature of the text rather than of its author.
It might still be insisted that where there has been no philosophising there is no work, whatever the resulting text contains. We need not resist this, because our thesis concerns texts worth reading, and a text can be worth reading without being a work in Davies' sense. Suppose a desert wind traced out, in the sand, a sound argument against enactivist theories of perception. Nobody would deserve credit for the argument, and there would be no performance for the text to give access to; but a reader who worked through it would still encounter a thesis and the considerations advanced in its favour, and would be in a position to answer or to extend it. 'Work' may, if the performance theorist insists, be reserved for texts with performances behind them. What cannot be so reserved is being worth reading, since everything that judgement answers to is on the page.
**That no one lies behind a text does not, then, settle whether it is worth reading.** Whether an LLM can actually _produce_ a text worth reading is a further question, and it is to this that we now turn.
[^1]: This is not to deny that systems of this kind can produce beautiful images; we return to image generation in Section 4.
[^3]: The same location of philosophy in the public text is reached by accounts of philosophical progress. Dellsén et al. (2024) hold that progress is "for-whom" rather than "by-whom": it consists in putting people in a position to increase their understanding, usually by making philosophical ideas publicly available (p. 679). Being put in such a position requires something one can take up and think through, and what is available to be taken up is the text. A philosophical contribution so understood is constituted by what the public text makes available, which a view that locates the philosophy in the antecedent private activity mislocates.
---
# 2. The challenge from abduction
In the previous section, we argued against the idea that LLMs cannot produce philosophy worth reading simply because they are not human. In this and the two sections that follow we shall consider a different form of challenge: even if LLMs cannot be ruled out of producing philosophy worth reading tout court, they lack particular _capacities_ that are required to produce it. If our arguments regarding authorship are correct, then a novel philosophical argument produced by a parrot would deserve to be taken just as seriously as one produced by a human being. No such argument will be forthcoming, however: a parrot can only reproduce sounds it has already heard. One might suspect that LLMs are in the parrot's position — that while nothing rules their outputs out of being philosophy worth reading, they lack some capacity that producing it requires. **We begin with what we will call _the challenge from abduction_; the two sections that follow address the _challenge from detachment_ and the _challenge from experience_.**
Abductive inference is reasoning from a body of evidence to the hypothesis that would best explain it. In a deductive argument the premises fix the conclusion: if all men are mortal and Socrates is a man, then Socrates is mortal, and that is that. Now, imagine walking into your kitchen one morning and finding that part of the floor is wet. What has happened? The wet floor does not determine the answer in the way the two premises determined Socrates' mortality: a burst pipe would have left the floor wet, and so would a spilled bucket. But, given that the window is open, the water is under the window, and it rained last night, rain coming through the window seems the most plausible answer. To reason in this way, inferring the best explanation for a set of facts, is common in the sciences as well as everyday life.
Williamson argues that philosophy should use a broadly abductive methodology (2007; 2021, §9.2). In philosophy, as in science, there are data that a candidate theory must accommodate — intuitions about cases, and the phenomena of the domain itself — and rival theories that would each accommodate them at different costs. The theory to prefer is the one that would, if true, best account for the data, where what makes one theory's account better than another's is its explanatory virtue: a good philosophical theory is, in Williamson's words, "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated", and should "combine simplicity with strength" (2021, §9.2). Nor is this view of theory choice confined to those who, like Williamson, take philosophy to be methodologically continuous with the sciences: Bengson, Cuneo and Shafer-Landau (2022) argue that the assessment of rival theories by their explanatory and unifying merits is a constraint on sound philosophical method as such.[^holders] We shall assume it in what follows.[^progress]
[^progress]: Understanding-based accounts of philosophical progress converge on the same standard: if progress in philosophy consists in placing people in a position to increase their understanding (Dellsén et al. 2024, p. 679), then rival theories are rightly weighed by the understanding they would, if true, provide.
Lipton (2004, p. 59) distinguishes two things that "the best explanation" might mean: the _likeliest_ explanation, the one most warranted by the evidence, and the _loveliest_, the one that would, if true, provide the most understanding. As he puts it, "Likeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). The two can come apart. Return to the wet kitchen floor: that water has fallen on it is the likeliest explanation of its being wet, and it is scarcely possible that it is false, yet it affords no understanding at all of how the floor came to be in that state: of course it is wet because water fell on it, but how, and which water? The explanatory virtues by which, on the methodology we have assumed, rival philosophical theories are weighed are virtues of loveliness: elegance, unity, and simplicity concern the understanding a theory would provide were it true, not the probability that it is true. Where a philosophical text turns on the weighing of rival explanations — not all philosophical writing does — whether the text is worth reading and whether it weighs its rivals well go together.
If a capacity for abduction is required to produce worthwhile philosophy, do LLMs possess it? Floridi et al. (2025) argue that they do not:
> We argue that such LLMs generate text based on learned associations rather than performing abductive inferences. […] LLMs produce plausible hypotheses, simulate commonsense reasoning, and provide explanatory answers without grounding them directly in truth, semantics, verification, or understanding, and without any abductive reasoning. (Floridi et al. 2025, p. 1)
An LLM, on this account, has "a stochastic core and an abductive appearance" (2025, p. 2); the further charge, that these systems cannot "genuinely validate" their explanations against reality (ibid., p. 6), we take up in the next section. Models are trained to predict which words are likely to follow which, and to produce the continuation that their training makes probable; nothing in the training directs them at truth. The appearance comes from what the training data have passed on: models have "absorbed patterns of human abductive reasoning as expressed in writing" (p. 9) — how explanations are typically phrased, which causes are typically offered for which effects. [^1] Floridi et al.'s own example is a car that will not start on a cold morning. Asked why not, a model replies that a weak battery is one possibility, since cold reduces a battery's efficiency; that thickened engine oil is another, since a cold engine is harder to turn over; and that, "[b]ased on your description, the battery is the most likely explanation" (2025, p. 10). On their reading, the LLM is not actually reasoning about causes from the case before it; it is the statistical reproduction of the causes such explanations typically cite (p. 9). What they deny is that the model weighs the candidates — that it "entails choosing the best explanation among alternatives" (2025, p. 3). The verdict that settles on the battery does no weighing; it reproduces how explanations of this kind conventionally end. And where the output marks a genuine difference between the two — some consideration that would tell the battery from the oil — it is one already drawn in the explanations the model learned from, not one worked out afresh for the case in hand.
Because LLMs produce no more than a "veneer of explanation" (ibid., p. 20), Floridi et al. conclude, they can play only a supporting role in intellectual work:
> In essence, LLMs function like brainstorming assistants that toss out ideas without filtering for quality. After all, they work like statistical interfaces to an enormous amount of data accumulated for millennia by generations. A cautious human collaborator can sift through and assess them. (2025, p. 11)
A device for brainstorming seems a far cry from something which might produce worthwhile philosophy. If you were presented with a text and told that it contains a number of philosophical ideas, none of which have been filtered for quality, it is unlikely you would think it is worth your while reading it. The filtering these systems are said to lack is, in Lipton's terms, filtering for loveliness: the stochastic core already selects the likeliest continuation, and what it cannot do, on Floridi et al.'s account, is prefer one candidate explanation to another on the grounds of the understanding it would afford.
Floridi and his colleagues allow that, in ordinary cases such as their own cold-morning car, the model's answer is a good one: "the same explanation a human reasoner would likely choose" (2025, p. 10), one that may be "even optimal by IBE criteria" (2025, p. 19), since such systems "echo the obvious, common explanations" (2025, p. 10). The complaint cannot then be that the explanation is poor, which leaves it hard to say what, in such an answer, is supposed to be merely apparent. What their account points to is the uncommon case: "in less common situations, LLMs can falter" (2025, p. 10), and on inputs "that go beyond their training" "the facade can crack" (2025, p. 9), the success on familiar cases being "a sign of overfitting to common patterns" (2025, p. 15). Once the question of its truth is set aside, this is what the facade reduces to: the claim that the competence shown on common problems is overfitting that would give out in less common ones.
Whether the competence does give out on uncommon cases is an empirical question. A recent survey of abductive reasoning in language models suggests that it does (Salimi et al. 2026). Current models do markedly worse on abductive tasks than on deductive ones: where their median accuracy on deductive tasks is near eighty per cent, on abductive ones it is some forty-two and a half, with a spread "extending down to near-zero accuracy", and the survey reports that "strong deductive performance does not reliably imply strong abductive performance". **This is close to Floridi et al.'s own diagnosis — an answer that reproduces a common pattern instead of reasoning to it — and, taken at face value, it tells in their favour.**
We do not disagree with Floridi et al.'s characterisation of how LLMs function: these systems do not weigh and choose among alternatives in the way that humans do. However, we should not be too quick to jump from this to the conclusion that LLMs cannot _produce text_ which exhibits abductive reasoning. A pocket calculator does not have the capacity to do arithmetic in the way a person does, but does have **the** capacity to produce the correct answer to sums which are entered into it. Similarly, it might be possible for LLMs to produce text which displays abductive reasoning, despite it not being grounded in any actual abductive reasoning.
Despite their stochastic core, LLMs are perfectly capable of producing grammatically correct text: without being given any specific rules, the system "implicitly 'discovers' them—and then seems to be good at following them" (Wolfram 2023). Does this mean that the texts LLMs produce have merely the appearance of being grammatically well-formed? **Clearly they do not:** LLM sentences _are_ grammatically well-formed despite their stochastic roots. A stochastic core, then, need not mean that the best an LLM can manage is a veneer of abductive inference. It may be, rather, that the core is marshalled to produce text exhibiting actual abductive inference, in just the way it is marshalled to produce actual grammatical correctness.
It might be objected that the grammar analogy will not stretch this far. Grammar, the objector would say, is a system of rules a model can follow without understanding anything, and explanatory loveliness is not — it is not a system of rules you follow to get the right answer, so the model's grammatical competence gives no reason to expect it to produce lovely explanations. **If Floridi et al. are right that the output has the form of an inference to the best explanation but not its substance, then** there should be cases where the form is present and the substance is absent: cases where the model produces something that looks like an explanation but is empty or nonsensical, just as it could produce something that is grammatically faultless and says nothing. Wolfram's example of the latter is "Inquisitive electrons eat blue theories for fish" (2023) — impeccably grammatical, and meaningless. The model does not produce such strings. The abductive equivalent would be a passage with the form of an inference to the best explanation, grammatically faultless, and yet senseless — asked why a car will not start on a cold morning, there are no squirrel tracks, so it must be freak arctic winds blowing into the exhaust pipe. That has the form of evidence weighed towards a conclusion, and it is nonsense; and it is nonsense the model does not produce. Put the question to it and it offers the weak battery and the thickened oil, plausible explanations rather than arctic winds. Wherever the facade is to be located, it cannot be located there: what the model produces are plausible explanations.
Moreover, the answer Floridi and his colleagues describe is not quite the answer these systems give. Their account asks us to picture a confident verdict laid over a hollow core, yet what one finds in practice is closer to hedging. Asked why a car would not start on a cold December morning, a current model will say that the battery is the most likely culprit, set out other plausible contributors — thickened oil, fuel-system problems, ignition faults — and end by observing that if the car started once the day had warmed, the battery is almost certainly the primary cause, though a load test would confirm it.[^kimi]
[^kimi]: The exchange, with Kimi k2.67 (%%add date of test%%), is reproduced in full in the Appendix. **This is not yet to say that the facade charge is refuted: a hedged answer can still be, on Floridi et al.'s account, a statistical reproduction of how hedged explanations typically read.**
The model's avoidance of senseless explanations, like its avoidance of senseless sentences, points to a semantic competence picked up in training, over and above syntax. A syntactic grammar settles only how the parts of speech may be combined, and not whether what a sentence says makes sense: avoiding strings like the inquisitive electrons takes a grasp of what can sensibly be said of what — which predicates go with which subjects, which causes with which effects. Wolfram calls such a grasp a semantic grammar:
> to deal with meaning, we need to go further. And one version of how to do this is to think about not just a syntactic grammar for language, but also a semantic one. (Wolfram 2023)
A model trained on enough text has, on his account, come by one:
> From its training ChatGPT has effectively "pieced together" a certain (rather impressive) quantity of what amounts to semantic grammar. (Wolfram 2023)
**What the training pieces together extends beyond the fit of predicate to subject. The arctic winds owed their senselessness as much to the inference as to its parts — nothing in the absence of squirrel tracks bears on what enters an exhaust pipe — and Wolfram suggests that inference is picked up by the same route: just as Aristotle might have arrived at syllogistic logic by working through many examples of rhetoric, a model in training can "discover syllogistic logic" in the text it reads, and can be expected to produce text containing "correct inferences" (2023).[^formal]** It is this, rather than any contact with the case in hand, that keeps the arctic winds out of the model's explanations: a feel for what, in a working model of the world, can hang together.[^wm]
But a feel for how things hang together is gathered from the text the model has read. **It amounts to a grasp of how failures of this kind are explained in general — that cold weather weakens batteries, say — and it gives the model no means of finding out which explanation is true of the particular car whose fault is in question. It does, however, underwrite the weighing itself: the model's answer sets out the candidate explanations and fits its confidence to the evidence it has been given, and neither of these requires access to the car. What requires such access is settling which candidate is correct. In much philosophical abduction there is no analogue of the car: what a thought experiment commits us to, or which of two theories carries the lighter explanatory cost, is settled from what is already set down in writing, and a corpus already holds evidence of that sort.**
The survey scores of Salimi et al. (2026) measure whether the model arrived at the right answer, not the weighing that got it there. On the harder tasks the scores are low, and on long mysteries with their clues strewn through the text the best models fall just short of the average human solver; taken at face value, the numbers count against the model. But the score is a score for the answer — whether the named culprit was the keyed one, whether the right diagnosis came first — and not for the weighing that reached it; such a score, **Salimi et al. note**, "completely bypass[es] the actual reasoning trace". So when a model misses the keyed culprit, the number marks the miss, and not the comparison it set out on the way — which explanations it canvassed, and why it came down on one. And the survey's hardest tasks are built from low-prior, non-stereotypical outcomes, where several explanations may be reasonable; there, matching an answer to a single reference "underestimates explanation quality", so a low score on such tasks does not show that the model cannot weigh, only that it did not land on the keyed answer.
It is now difficult to see where the facade should be located. The explanations the model produces are plausible, the confidence they express is fitted to the evidence it has been given, and the benchmark scores that seemed to confirm the diagnosis measure answers rather than the weighing that produced them. What the model lacks is **access to** the particular case, and that is a limitation concerning its relation to the world rather than its capacity for abduction. **What that costs** a philosophical text is a matter for the next section. What a one-shot prompt leaves undrawn, and how it might be drawn out, is taken up in Section 5.
**[^formal]: Wolfram's expectation comes with a caveat: on "more sophisticated formal logic" he expects the model to fail, for the same reasons it fails to match parentheses across long sequences (2023). The caveat concerns extended derivation, and the weighing of explanations is not derivation: as the objector above puts it, loveliness is not a system of rules one follows to get the right answer.**
[^wm]: We use "model of the world" in Wolfram's informal sense: a body of implicitly learned regularities concerning what goes with what. This is weaker than the technical sense now disputed in machine learning, on which a world model is an internal representation of an environment that supports prediction and counterfactual reasoning. Nothing in our argument requires deciding whether LLMs possess world models in the technical sense; how they connect to the world is the matter of §3.
---
# 3. The Challenge from Detachment
**A second capacity challenge, the _challenge from detachment_, concerns the model's relation to the world.** In addition to the argument considered in the previous section, Floridi et al. object that an LLM stands in no relation to **it**: its words rest on no perception of anything, and a hypothesis, once produced, is never tested against how things are (2025, pp. 6–7). A discipline whose theories answer to how things are, the challenge runs, cannot be advanced by a system with no access to how things are, and a text produced by such a system gives its reader no reason to think it worth reading. We argue that the challenge fails because the materials philosophy takes from the world reach it already articulated in language, and a model can work on those words as any philosopher does.
We can grant that the model perceives nothing, and that nothing it produces is put to the test against the world. **Whether these concessions disqualify it depends on how philosophy itself meets the world.** Pigliucci (2017), writing on how philosophy makes progress, holds that philosophy is constrained by the world without investigating it as the natural sciences do:
> This means that the basic parameters that philosophers use as their inputs, the starting points of their philosophizing, their equivalent of axioms in mathematics and assumptions in logic (or rules in chess) are _empirical_ data about the world. This data comes from both everyday experience [...] and of course increasingly from the world of science itself. Philosophy, I maintain, is in the business of exploring the sort of conceptually evoked spaces that Smolin is talking about, where the evocation is the result of whatever starting assumptions are made by individual philosophers working within a particular field and, crucially, of the constraints that are imposed by our best understanding of how the world actually is. (2017, pp. 79–80)
Pigliucci takes the term _evocation_ from Smolin (Unger and Smolin 2015).[^4] Evoked truths are neither discovered, in the sense of corresponding to mind-independent states of affairs, nor invented, in the sense of being arbitrary constructs. Smolin's illustration is chess: before the rules of the game were codified there were no facts about chess, but once they were, a great many facts about the game became demonstrable — objective facts, in the sense that anyone who demonstrates one demonstrates the same fact as anyone else (Unger and Smolin 2015, p. 423; quoted at Pigliucci 2017, p. 78). On Pigliucci's proposal, philosophy is in the business of ascertaining evoked truths. What distinguishes it from mathematics and from chess is where its starting points come from: they are constrained by how the world actually is.
The same constraint distinguishes philosophy from fiction. Pigliucci's contrast case is a science-fiction writer who describes the same planet across three different timelines, keeping every description within the laws of physics and of logic. The writer is exploring conceptual spaces of a kind, but he is inventing rather than evoking: his worlds have no rigid properties, because even the constraints he keeps to could have been otherwise — he could as easily have imagined planets with a different physics, or a different logic. Philosophy, **on the other hand**, "is in the business of doing empirically informed evoking, not inventing", and its objects of study accordingly have rigid properties (2017, p. 80): once a philosophical starting point has been articulated, what holds within the structure it opens is no longer up to anyone, any more than what follows from the rules of chess is up to anyone. A thought experiment is a case in point: the philosopher sets up an imagined scenario, but explores it "with an interest in figuring things out as far as this world is concerned" (2017, p. 80), and the philosophical work then proceeds within a structure whose properties the philosopher does not control.
Philosophy's starting points — what Pigliucci calls the basic parameters of philosophising, its equivalent of axioms in mathematics and rules in chess — are data about how the world is, drawn from everyday experience and, increasingly, from science. A system that perceives nothing cannot gather such data for itself; but nor does the working philosopher. The scientific data Pigliucci has in mind reach philosophers already set down in language: a philosopher of physics works from published results, not from having run the experiments. And what the discipline retains of everyday experience, it retains in the same form, as the literature's accumulated descriptions of how things seem. Written access to the world's constraints is the profession's normal condition, and a corpus provides access of just this kind. With respect to nearly all the empirical data philosophy uses, then, the model is in the same position as any philosopher.
Philosophical claims are not, in any case, tested in the way the challenge assumes. The equivalence principle, once Einstein had it, faced a tribunal of measurement — the experiments might have gone against it, and it would then have been dropped. A philosophical thesis faces no such tribunal. Whether it holds is a question about what obtains within an evoked structure, and it is settled in the way questions about chess are settled: by working out what the starting points commit us to, something any competent party can do and none can decide by fiat. This working-out is carried out in writing, in the literature's back-and-forth of argument and reply, and **it is work of a kind Section 2 argued a model's text can display.** A model's inability to run experiments therefore costs it nothing that a philosophical text needs: the testing philosophy does is done on the page.
It might be doubted, however, that experience can be treated **in the way the world's data have been**, since a description of what it is like to see red is not obviously an adequate substitute for seeing it, and the next section addresses a challenge built on this doubt.
[^4]: _The Singular Universe and the Reality of Time_ is jointly authored, but its second part, which contains the discussion of evocation, was written by Smolin alone, as Pigliucci notes (2017, p. 77); we follow him in attributing the view to Smolin.
# 4. The Challenge from Experience
**A third capacity challenge, the _challenge from experience_, holds that some philosophy depends on experience in a way that a system which has never had any cannot meet.** Few would say that current LLMs are conscious, and we assume here that they are not. Zahavy (2026) raises a worry of this kind about scientific discovery. A model can carry out the deductive part of discovery, working out the consequences of premises it has been given; what it cannot do, he holds, is produce the premises — make the move from sense experience to new first principles. On the picture he takes from Einstein, that move is a leap,[^5] and it is the leap that gives a theory its axioms. His case is the thought experiment that gave Einstein the equivalence principle:
> Einstein's variation required inventing new axioms based on a physical intuition that did not yet exist in the mathematics. He envisioned a physicist inside an elevator being uniformly accelerated through deep space [...]. Inside this enclosure, the sensory experience reveals a specific pattern: when objects are released, the floor rushes up to meet them. To the physicist, the objects appear to fall with identical acceleration, regardless of composition. Thus, the simulation here was not a permutation of symbols, but a manipulation of perceptual experience. (Zahavy 2026, §5)
Einstein imagines a set of circumstances and attends to what would be experienced within them: everything released inside the elevator appears to fall with identical acceleration. On Zahavy's reconstruction the simulation supplies an observation, and from that observation the new axiom is inferred — the simulated experience of acceleration was indistinguishable from the remembered experience of gravity, and Einstein concluded that the two are one phenomenon.[^2] A model has no access to that observation. It can produce descriptions of elevators and of weightlessness, both present in its corpus, but it has undergone neither, and a discovery whose premises are fixed by simulated experience is beyond a system that, in Zahavy's words, lacks the capacity he calls sensory agency.
Philosophical thought experiments might seem to depend upon experience in the same way. Does Mary, released from her black-and-white room knowing every physical fact about colour vision, learn something when she first sees red (Jackson 1982)? To settle the question one must consider what the experience of seeing red is like, and the argument proceeds from that verdict. Here experience enters as the starting point of an argument, and a system that has never experienced anything seems unable to supply it. Experience enters philosophy as subject matter too, in work that asks what it is to see red (Harman 1990), to feel anger (Goldie 2000), or to have an intuition take hold (Chudnoff 2011).[^3] If starting points of the first kind, and subject matters of the second, can be handled only by a being that has the experiences, then much of the philosophy of mind lies beyond a model's reach.
**When philosophers reply to the knowledge argument, however, they do not draw on any _actual_ experience of seeing red. It is not as if a reader of Jackson's paper needs to introspect the fine details of their own experience of red in order to understand his argument, or the literature that discusses it. The facts about colour experience that the argument turns on are general ones, articulated many times in the literature: that colour is experienced as a quality of the things seen, and that one cannot imagine a colour one has never seen — which is why Mary's room must be black and white. The texts that have been produced in response to Jackson's paper work with claims of this kind and with his description of the case, contesting what they commit us to — whether what Mary gains on her release is knowledge of a fact at all — and none of them draws on what seeing red is like from the inside. Philosophy does not usually begin from raw encounter; it works on what has already been articulated. Even where a case does begin in its author's own experience — as Jackson's may have — what enters the argument is the description he makes of it.**
**Claims of this kind are examples of what we might call _articulated phenomenology_: descriptions in which what experience is like has been set down in words, and which can then be read and reasoned from like any other claims. Consider, for example, Austin's discussion of colour phenomenology in _Sense and Sensibilia_, where aspects of what it is like to see colour are carefully described and considered (e.g. how a pointillist painting's appearance changes depending on how far one stands from it, or how an object's colour changes under different illumination; 1962, pp. %%page%%, 66). Philosophers are not the only ones in an LLM's corpus to articulate their phenomenology in text; we might also consider novels, diaries, interviews — anything in which someone describes _what it is like_ to have an experience. An LLM has never felt what Zahavy's physicist feels inside the elevator, but it has read the memoirs of astronauts who have. We saw in Section 3 that written access to the world's constraints is the profession's normal condition, and experience is no exception. A model has undergone none of these experiences, but the descriptions of them are in its corpus, and it is the descriptions that philosophy works from.**
**What models do lack is the capacity to make _novel_ phenomenological observations. An example of this sort of novelty might be Merleau-Ponty's (1945 %%page%%) description of how, when one hand touches the other, the roles of toucher and touched can alternate but cannot be occupied at once. The experience is a familiar one, but the asymmetry it contains seems not to have been described before. Noticing it required having hands, and tactile experience of them; a model has neither, and so could not have been the first to set it down.**
**However, philosophical arguments which rely on novel phenomenological observations are rare indeed: the descriptions of experience that philosophy works with have, for the most part, long been in the literature. And once an observation of this kind is written into the corpus, everything philosophical that can be done with it — what it shows about bodily self-awareness, what follows once the asymmetry is granted — is there on the page for any reader, and none of it depends on being able to have made the observation oneself. Nor is first-person attention the only source of descriptions the literature lacks: a new description can also be reached from those it already contains, by drawing out what they have not been taken to imply, or by combining them as no one has, and this a model can do. There is, then, one route to new starting points that a model cannot take. Whether the route it can take yields descriptions the literature lacks is the question of novelty, which Section 6 takes up.**
A model can, then, do the philosophy that turns on experience, with one exception: **it cannot be the source of a novel phenomenological observation.** **Whether experience enters as an argument's starting point or as its subject matter, it reaches the model as articulated phenomenology, which it can work through as well as any reader.**
[^2]: Zahavy, following Magnani, calls the process _manipulative abduction_: hypothesis generation through the manipulation of a model — here a simulated experience — rather than of symbols (Magnani et al. 2009; Zahavy 2026, §5). It is abduction in Section 2's sense: the equivalence principle is inferred as the best explanation of the simulated observation, the simulation supplying an explanandum that no search over existing text would have produced. What experience contributes, on this picture, is not the inference but its starting point.
[^3]: The materials need not be sensory: the feeling of understanding something is sometimes used to motivate the claim that thought itself has a phenomenology (Pitt 2004).
[^5]: Zahavy too calls this leap abduction, but the word picks out something other than it did in Section 2. There, with Floridi et al., abduction was the weighing of rival explanations, and the charge was that a model only mimics it; here it is the generation of new first principles from experience, and Zahavy's claim is that a model cannot make the move because it has had no experience to move from. The present challenge rests on that second claim, about experience, and not on any verdict about the weighing.
---
---
# 5. The Challenge from Observation
The challenge from observation begins with an obvious question: if all of these arguments are correct, where is all the LLM philosophy?[^journals] If you ask an LLM the answer to the hard problem of consciousness or the meaning of life, the response you receive will not typically contain much in the way of philosophical argument. More likely you will get a somewhat bland survey of various philosophical positions on the topic, or perhaps just an irreverent joke.[^gpt55] One might think that whatever the merits of the arguments we have put forward so far, _something_ must be preventing them from producing philosophy worth reading because they _don't_ produce philosophy worth reading.
We want to suggest that this is a failure of prompting rather than evidence that LLMs lack the capability to produce philosophy worth reading. As we saw in Section 2, these systems are trained to produce reasonable continuations of the text they have been given (Wolfram 2023), before being trained further to converse as helpful assistants. **In the text these systems are trained on,** questions like "what is the meaning of life?" are hardly ever followed by philosophical argument: arguments occur in journal articles, which begin from claims already advanced and objections already raised rather than from questions, and it is these that their arguments continue. **The same is true when asking more "serious" philosophical questions.** "Is representationalism the correct theory of perception?" is the sort of sentence that appears in examination papers and essay titles rather than at the opening of a defence of representationalism. A similar point applies to the assistant training: **an assistant is trained to answer the question it is asked, and answering a philosophical question is not the same as arguing for an answer to it.** None of this shows that LLMs can produce philosophy worth reading, but it does mean that their failure to produce it when questioned point-blank tells us nothing either way.
**How, then, can we elicit philosophy from these systems? Philosophy, Pigliucci writes, attempts "to evoke rational conclusions arising from certain ways of looking at a given problem or set of facts" (2017, p. %%page%%). To elicit philosophy from an LLM, we suggest, is to prompt it into producing reasonable philosophical arguments or explanations — reasoning from which conclusions of this kind can arise.**
**One obvious way is simply to _ask_ for careful, rational arguments about a philosophical question. We saw in Section 2 that a model's continuations are produced under the semantic grammar it has acquired — a grasp not only of what can sensibly be said of what, but also of which inferences may correctly be drawn. What counts as a reasonable continuation depends on what is being continued: a bare question is rarely continued by philosophical argument, while a request for careful reasoning about a stated problem is continued by text in which that reasoning is carried out. One can refine the request by giving the model a problem or set of facts and asking for the explanation that would, if true, provide the most understanding of them — that is, by asking it for an inference to the best explanation.**
How best to do this is a question we leave open. **Nothing restricts evocation to a single form of prompt:** several hundred pages of draft manuscript make starting assumptions and accept constraints as surely as two paragraphs of thought experiment do, and which forms of prompt evoke most productively is an empirical question which we shall not try to settle.
It is in the continuations of such prompts that the question of these systems' philosophical capacities would be decided.[11](#user-content-fn-janus) If a philosopher must **articulate the starting point**, then it will be said that whatever philosophy the resulting text contains should be credited to the philosopher, and that the model is merely its instrument. This is the challenge of the next section.
[^journals]: Strictly speaking, the observation is difficult to verify. To know that no LLM-written philosophy has appeared in the journals over the last half-decade, one would need to know that every author who has declared that they did not use AI in preparing their manuscript was telling the truth.
# 6. The Challenge from Instrumentality
Grant that philosophy worth reading comes out of these systems only when a philosopher prompts for it. If the philosopher must put so much in, one might suspect that the philosophy is the philosopher's — that the model does not produce philosophy but rearranges and remixes what it is given. Call this the _challenge from instrumentality_. Just as a ventriloquist's dummy is silent unless the ventriloquist is beside it, an LLM produces no philosophy unless a philosopher prompts it. The challenge from authorship held that a model's text cannot be philosophy at all. The present challenge is narrower: even if such a text is philosophy worth reading, none of the philosophy in it is the model's contribution.
Whether the challenge succeeds depends on what the prompt contributes and on what the continuation adds. The previous section gives us the first: a prompt that articulates a position and the pressures upon it evokes — a starting point handed over for development. The continuation is produced from this starting point under the semantic grammar the model has acquired, and what it states are facts of the evoked structure. There are, then, three contributions to distinguish: the articulated starting point, which is the prompter's; the structure it evokes, whose facts belong to no one; and the text that develops them, which is the model's. A prompt settles only a starting point, and the consequences the model's text goes on to state were stated by no one beforehand.
We saw in Section 3 that once a starting point has been articulated, what holds within the structure it opens is no longer up to anyone — and that includes the person who articulated it. Smolin observes that whoever lays down the rules of a game is afterwards in the same position as everyone else: exploring the game feels like exploring a pre-existing territory, because at each point there is little or no choice about what its facts are, and the territory holds surprises for its own inventor (Unger and Smolin 2015, p. 422). A philosopher stands to her starting point in the same way. She has no say over what it commits her to, and no advance knowledge of it either; she reads a continuation to find out. The remixer picture gets the ownership the wrong way round: it treats the developments as her material, handed to the model for reshuffling, when they were not statements at all, hers or anyone's, until something stated them. She needs no model to rearrange her own text; what she prompts for is a text that states what hers did not.
It might still be insisted that the developments were contained in the starting point all along, as theorems are contained in the axioms that entail them, and that in drawing them out the model adds nothing of its own. But philosophical developments do not stand to their starting points as theorems stand to axioms. We saw in Section 2 that where a philosophical text weighs rival explanations, what decides between them is loveliness — the understanding each rival would, if true, provide — and that loveliness is not a system of rules one follows to get the right answer. A theorem can be recovered from its axioms by derivation; a verdict that one rival carries a lighter explanatory cost than another cannot be recovered from the prompt by any procedure, because it is reached by weighing, and weighing is not derivation. What the prompt contains is the problem; the weighing is done in the continuation.
How, then, does the model develop a starting point? We saw in Section 2 that a model's continuations are produced under a semantic grammar pieced together in training — a grasp of what can sensibly be said of what, and of which inferences may correctly be drawn — and in Section 5 that a request for careful reasoning about a stated problem is continued by text in which that reasoning is carried out. The model produces such text one step at a time, and at every step what it continues is the whole text so far — the prompt together with everything it has itself already written (Wolfram 2023). Each inference drawn becomes part of what the next step must reasonably continue, so early moves open commitments that later moves are made under, and the further a continuation runs, the more of what it is continuing was never in the prompt at all. The grammar, moreover, constrains each step without dictating it. At any point there are several ways of going on that would each be reasonable (Wolfram 2023), and nothing outside the model settles which of them it takes: it chooses its own path through the structure the prompt has evoked. And the path is philosophical work — drawing out what the starting point commits us to, and weighing the rivals it leaves standing.
Creativity, on this picture, is not a further capacity the model would need; it falls out of the evoking. When the starting point is the philosopher's own — a case the literature does not contain, or a position under a pressure nobody has put to it — the structure it evokes is new, and until the prompt was written there were no facts about it to state. Whatever path a continuation then takes, what it states has been stated by no text before it; and what a text contributes, in philosophy quite generally, is what it states that the texts before it did not. Whether its statements were worth making is judged as for any philosophical text, by whether the considerations it cites genuinely tell between the positions and by the understanding they would, if correct, provide. Prompted with a fresh starting point, then, a model produces original arguments in virtue of its evoking. What it cannot produce, however it is prompted, is a starting point of the kind Section 4 conceded in Merleau-Ponty's case. Wolfram notes that a model can take up what it is told only where it sits, in a fairly simple way, on the framework the model already has (2023), and a novel phenomenological observation sits on no such framework. Whether philosophy worth reading ever requires starting points of that kind is a question we leave open, and it would be settled as the others were: by reading what a model produces against the literature and asking what it states that the literature did not.
None of this removes the person from the picture: she writes the prompt, decides which continuations to pursue and when to stop, and without her there is only the response to the point-blank question, and so much of the instrument picture is correct. But asking for careful reasoning is not doing the reasoning, and choosing among continuations is the work of an editor, not of an author. Nor can she make the developments hers by enriching the prompt, however much philosophy she writes into it — a richer prompt is a larger starting point, and the question is always whether the continuation states anything beyond it. That question is settled by reading the two together. Where a continuation merely rephrases its prompt, the model really has done no more than rearrange, and the philosophy is the prompter's; where it states what the prompt did not, the surplus is not hers, because she did not state it.
The labour, then, divides more cleanly than the challenge allowed. The philosopher supplies a way of looking at a problem; the model produces the rational conclusions arising from it, which is Pigliucci's description of what philosophy does. Nor is there any difficulty in telling the two contributions apart: the starting point is stated in the prompt, and the development appears only in the continuation. A prompt supplies a starting point without fixing what follows from it, and the philosophical standing of any output is settled by what its continuation adds — first to the prompt that occasioned it, and then to the literature it would join.
# References
# References
Frankish, K. (2024). What are large language models doing? In A. Strasser (Ed.), _Anna's AI Anthology: How to live with smart machines?_ (pp. 55–78). Xenomoi.