We argue that current-generation large language models (LLMs) are capable of producing philosophical texts that are *worth reading*, in the same way that a good piece of human-written philosophy can be worth reading, even if one does not agree with every word.
We defend this claim against various challenges. First, one might think that any text an LLM produces cannot count as a work of philosophy because it was not produced by a person a text produced by an LLM just cannot count as philosophy in the same way that we might think an image produced by AI does not count as an artwork.
Even if we one might think that LLMs lack a particular capacity, or capacities, required to produce worthwhile philosophy.
- Then group the last two challenges together
: that a text with no philosopher behind it cannot be philosophy at all; that LLMs lack capacities — abductive, perceptual, experiential — which producing philosophy requires; that if we were right, philosophy worth reading would by now have come from these systems; and that whatever does come from them belongs to the prompting philosopher rather than to the model. We argue that none of these challenges succeeds. What makes a philosophical text worth reading is on the page, and how the page was produced settles nothing in advance. The final challenge, from instrumentality, holds that where a philosopher's prompting elicits a text worth reading, the philosophy is the philosopher's and the model merely its instrument. We argue that a prompt articulates a starting point which underdetermines its development, and that what the development states beyond the prompt is not the prompter's.
# Abstract
We argue that current-generation large language models (LLMs) are capable of producing philosophical texts that are worth reading. Rather than beginning from an account of what makes philosophy good, we work with a distinction every philosopher already draws, between texts that repay the time spent reading them and texts that do not, and we defend our claim against six challenges: that a text with no philosopher behind it cannot be philosophy at all; that LLMs lack capacities — abductive, perceptual, experiential — which producing philosophy requires; that if we were right, philosophy worth reading would by now have come from these systems; and that whatever does come from them belongs to the prompting philosopher rather than to the model. We argue that none of these challenges succeeds. What makes a philosophical text worth reading is on the page, and how the page was produced settles nothing in advance. The final challenge, from instrumentality, holds that where a philosopher's prompting elicits a text worth reading, the philosophy is the philosopher's and the model merely its instrument. We argue that a prompt articulates a starting point which underdetermines its development, and that what the development states beyond the prompt is not the prompter's.
---
# 0. Introduction
The last decade or so has seen the rise of generative artificial intelligence: systems that produce text, images, code, music, video, and other outputs in response to prompts. AI has had success in domains where the value of an output is not exhausted by its superficial fluency. There have been recent AI-assisted discoveries in physics (Guevara et al. 2026), mathematics (Novikov et al. 2025), biomedicine (Gottweis et al. 2025), and materials science (Zeni et al. 2025). In this paper we argue that we should expect similar success in philosophy.
Specifically, we argue that current-generation LLMs are capable of producing philosophical texts that are _worth reading_. We do not want to begin by settling what counts as _good_ philosophy. Instead, we appeal to a distinction familiar to anyone who reads philosophy: between texts that repay the time spent reading them and texts that do not. As you begin reading this article, you likely hope that it is worth reading, in the sense that the time spent reading it will not be wasted. When you write a philosophical text yourself you aim to make it worth readers' while to read it, and whether or not the journal you send it to accepts it depends on whether or not they agree.
Note that a text's being worth reading is not the same as its being correct: we take many of the philosophical texts we read to be mistaken in their conclusions, and few of them to have wasted our time. Note also that the mere statement of a philosophical conclusion is unlikely to be worth reading. A text consisting only of bare pronouncements — that direct realism is correct, that we should be utilitarians — would not repay anyone's attention, and this is as true of a human philosopher's pronouncements as of an LLM's.[^1] What would repay attention is argument, and it is the capacity of LLMs to produce philosophical arguments, rather than philosophical pronouncements, that concerns us in what follows.
We will argue for this claim by considering six challenges that might be raised against it. Section 1 rejects the challenge from authorship: the claim that an LLM output cannot be philosophy worth reading because no philosopher lies behind it. Section 2 turns to abduction and argues that the absence of human-style inference to the best explanation in the producer does not preclude abductive structure in the product. Section 3 turns to the challenge from detachment, and argues that the materials philosophy takes from the world reach it already set down in words, which an LLM can work on as any philosopher does. Section 4 takes up the challenge from experience, and argues that the experiences philosophy argues about reach it in the same way. Section 5 turns to the challenge from observation — that if all this is right, philosophy worth reading should already be coming from these systems, and is not — and argues that this reflects how they are used rather than what they can produce. Section 6 addresses the challenge from instrumentality — that where a philosopher's prompting draws out such a text, the philosophy is the philosopher's — and argues that a prompt articulates a starting point whose development it does not fix.
---
# 1. The Challenge from Authorship
In this section we address what we might call the _challenge from authorship_: the idea that philosophy is something that only persons, or at least minds, can produce. This view has not, to our knowledge, been explicitly defended in just this form, but it gives shape to an intuition that many philosophers may have: philosophy is a person-only domain. An imperfect comparison is with art: One might deny that an image generated by an AI system is an artwork, because no artist exercises the relevant kind of intentional control over its production.[^1] One might think, for similar reasons, that philosophy can only be done by people: no text produced by an LLM can be a work of philosophy, because no philosopher lies behind it.
There is some initial support for this thought in the way philosophy is studied. Like the study of art, and unlike the study of physics, the study of philosophy is organised around individuals: undergraduates take courses on Kant's ethics or Lewis's metaphysics, and reading an important philosopher's own words is held to be of value in a way that reading Newton's is not. A physics student is taught Newtonian mechanics from a current textbook, and the course loses nothing if the _Principia_ is never opened; a course on Kant's ethics that never opened the _Groundwork_ would scarcely count as one. In the sciences, that is, what a text contributes can be carried entirely by other texts, while in philosophy the contribution and its original presentation are harder to prise apart. One might take this as evidence that a philosophical work is bound to the activity of the person who produced it, in a way that a scientific result is not.
The challenge becomes more precise if we ask how far Davies' _performance_ theory of art transposes to philosophy. Davies writes:
> [T]he work – what the artist achieves – is the process eventuating in that product. Works themselves are neither structures nor objects simpliciter, nor are they contextualized structures or objects. They are, rather, intentionally guided generative performances that eventuate in contextualized structures or objects (or events, as we shall see) – performances completed by what I am terming a focus of appreciation. (2004, p. 97)
On Davies' view, when a painter paints a picture, the canvas is what we attend to, but it is not the work. The work is the artist's intentionally guided activity in producing the canvas; the canvas is "the focus of our appreciative interest in the work" (2004, p. 150). What we appreciate in a painting, on this account, is an achievement, and achievements are individuated by the activities that bring them about: the same surface, reached by some other route, would be a different achievement. Facts about how the object came into being thus do more than supply context; they help determine what the work is, and what is properly appreciated in it.
Suppose, to adapt one of Davies' examples (2004, p. 102), that a storm were to blow pigment across a stretched canvas, leaving a surface indistinguishable, mark for mark, from an abstract painting. A theorist who identifies the painting with its surface must either count the storm's product as a painting or explain why it is not. There is a surface with no painting of a picture behind it, and the corresponding case for philosophy — a text indistinguishable from a philosophical argument, with no philosophising behind it — is the one an LLM presents. If Davies is right, the surface does not by itself settle the work.
Transposed to philosophy, Davies' proposal would run as follows: a philosophical text is not itself the philosophical work; the text is the product of a person's philosophising, and reading it is a way of engaging with that prior activity. The activity, on this proposal, is part of what the work is, so that there is a philosophical work only where the philosophising has taken place, which makes the challenge a constitutive one. If no one has philosophised, there is no work to which the text gives access, however the text reads — an LLM text would stand to philosophy as the storm-made canvas stands to painting.
We do not think the transposition should be accepted. What makes Davies' view plausible in the case of art is that indiscernible surfaces really do seem to differ in artistic value: the storm-made canvas is worth nothing as a painting, while a brushed one may be worth a great deal. Whoever transposes the view to philosophy is therefore committed to the corresponding claim, that two texts containing the same argument can differ in philosophical merit, and it is difficult to see what such a difference could consist in. Whether an argument is valid, whether its premises are plausible, and whether the objections to it have been answered are questions about a text's contents, and two texts with the same contents receive the same answers to them. The philosophical merit of a text does not vary with the route by which its words came to be written.
Peer review proceeds on the same assumption. Journals strip author information from submissions before review because facts about authorship are treated as potential sources of distortion rather than as evidence of merit; if two texts with the same contents could differ in philosophical merit, anonymising would discard information relevant to the assessment, and review would not be designed as it is. The grounds for accepting or rejecting a paper lie in the argument as presented, not in the history of its production.[^3] Nor is the author-centred teaching noted earlier in tension with this. That ethics is taught through the _Groundwork_ rather than through a digest of its conclusions reflects what a reader gains by working through Kant's arguments, and this is consistent with holding that the merit so gained is a feature of the text rather than of its author.
It might still be insisted that where there has been no philosophising there is no work, whatever the resulting text contains. We need not resist this, because our thesis concerns texts worth reading, and a text can be worth reading without being a work in Davies' sense. Suppose a desert wind traced out, in the sand, a sound argument against enactivist theories of perception. Nobody would deserve credit for the argument, and there would be no performance for the text to give access to; but a reader who worked through it would still encounter a thesis and the considerations advanced in its favour, and would be in a position to answer or to extend it. 'Work' may, if the performance theorist insists, be reserved for texts with performances behind them. What cannot be so reserved is being worth reading, since everything that judgement answers to is on the page.
Nothing in a text's having had no one behind it, then, settles whether it is worth reading. Whether an LLM can actually _produce_ a text worth reading is a further question, and it is to this that we now turn.
[^1]: This is not to deny that systems of this kind can produce beautiful images; we return to image generation in Section 4.
[^3]: The same location of philosophy in the public text is reached by accounts of philosophical progress. Dellsén et al. (2024) hold that progress is "for-whom" rather than "by-whom": it consists in putting people in a position to increase their understanding, usually by making philosophical ideas publicly available (p. 679). Being put in such a position requires something one can take up and think through, and what is available to be taken up is the text. A philosophical contribution so understood is constituted by what the public text makes available, which a view that locates the philosophy in the antecedent private activity mislocates.
---
# 2. The challenge from abduction
In the previous section, we argued against the idea that LLMs cannot produce philosophy worth reading simply because they are not human. In this and the two sections that follow we shall consider a different form of challenge: even if LLMs cannot be ruled out of producing philosophy worth reading tout court, they lack particular _capacities_ that are required to produce it. If our arguments regarding authorship are correct, then a novel philosophical argument produced by a parrot would deserve to be taken just as seriously as one produced by a human being. No such argument will be forthcoming, however: a parrot can only reproduce sounds it has already heard. One might suspect that LLMs are in the parrot's position — that while nothing rules their outputs out of being philosophy worth reading, they lack some capacity that producing it requires. In this section we address one capacity challenge, which we will call _the challenge from abduction_. In the two sections that follow we shall look at two more: the _challenge from detachment_ and the _challenge from experience_.
Abductive inference is reasoning from a body of evidence to the hypothesis that would best explain it. In a deductive argument the premises fix the conclusion: if all men are mortal and Socrates is a man, then Socrates is mortal, and that is that. Now, imagine walking into your kitchen one morning and finding that part of the floor is wet. What has happened? The wet floor does not determine the answer in the way the two premises determined Socrates' mortality: a burst pipe would have left the floor wet, and so would a spilled bucket. But, given that the window is open, the water is under the window, and it rained last night, rain coming through the window seems the most plausible answer. To reason in this way, inferring the best explanation for a set of facts, is common in the sciences as well as everyday life.
Williamson argues that philosophy should use a broadly abductive methodology (2007; 2021, §9.2). In philosophy, as in science, there are data that a candidate theory must accommodate — intuitions about cases, and the phenomena of the domain itself — and rival theories that would each accommodate them at different costs. The theory to prefer is the one that would, if true, best account for the data, where what makes one theory's account better than another's is its explanatory virtue: a good philosophical theory is, in Williamson's words, "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated", and should "combine simplicity with strength" (2021, §9.2). Nor is this view of theory choice confined to those who, like Williamson, take philosophy to be methodologically continuous with the sciences: Bengson, Cuneo and Shafer-Landau (2022) argue that the assessment of rival theories by their explanatory and unifying merits is a constraint on sound philosophical method as such.[^holders] We shall assume it in what follows.[^progress]
[^progress]: Understanding-based accounts of philosophical progress converge on the same standard: if progress in philosophy consists in placing people in a position to increase their understanding (Dellsén et al. 2024, p. 679), then rival theories are rightly weighed by the understanding they would, if true, provide.
Lipton (2004, p. 59) distinguishes two things that "the best explanation" might mean: the _likeliest_ explanation, the one most warranted by the evidence, and the _loveliest_, the one that would, if true, provide the most understanding. As he puts it, "Likeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). The two can come apart. Return to the wet kitchen floor: that water has fallen on it is the likeliest explanation of its being wet, and it is scarcely possible that it is false, yet it affords no understanding at all of how the floor came to be in that state, and offering it would invite exasperation rather than gratitude. The explanatory virtues by which, on the methodology we have assumed, rival philosophical theories are weighed are virtues of loveliness: elegance, unity, and simplicity concern the understanding a theory would provide were it true, not the probability that it is true. Where a philosophical text turns on the weighing of rival explanations — not all philosophical writing does — whether the text is worth reading and whether it weighs its rivals well go together.
If a capacity for abduction is required to produce worthwhile philosophy, do LLMs possess it? Floridi et al. (2025) argue that they do not:
> We argue that such LLMs generate text based on learned associations rather than performing abductive inferences. […] LLMs produce plausible hypotheses, simulate commonsense reasoning, and provide explanatory answers without grounding them directly in truth, semantics, verification, or understanding, and without any abductive reasoning. (Floridi et al. 2025, p. 1)
An LLM, on this account, has "a stochastic core and an abductive appearance" (2025, p. 2); the further charge, that these systems cannot "genuinely validate" their explanations against reality (ibid., p. 6), we take up in the next section. Models are trained to predict which words are likely to follow which, and to produce the continuation that their training makes probable; nothing in the training directs them at truth. The appearance comes from what the training data have passed on: models have "absorbed patterns of human abductive reasoning as expressed in writing" (p. 9) — how explanations are typically phrased, which causes are typically offered for which effects. [^1] Floridi et al.'s own example is a car that will not start on a cold morning. Asked why not, a model replies that a weak battery is one possibility, since cold reduces a battery's efficiency; that thickened engine oil is another, since a cold engine is harder to turn over; and that, "[b]ased on your description, the battery is the most likely explanation" (2025, p. 10). On their reading, the LLM is not actually reasoning about causes from the case before it; it is the statistical reproduction of the causes such explanations typically cite (p. 9). What they deny is that the model weighs the candidates — that it "entails choosing the best explanation among alternatives" (2025, p. 3). The verdict that settles on the battery does no weighing; it reproduces how explanations of this kind conventionally end. And where the output marks a genuine difference between the two — some consideration that would tell the battery from the oil — it is one already drawn in the explanations the model learned from, not one worked out afresh for the case in hand.
Because LLMs produce no more than a "veneer of explanation" (ibid., p. 20), Floridi et al. conclude, they can play only a supporting role in intellectual work:
> In essence, LLMs function like brainstorming assistants that toss out ideas without filtering for quality. After all, they work like statistical interfaces to an enormous amount of data accumulated for millennia by generations. A cautious human collaborator can sift through and assess them. (2025, p. 11)
A device for brainstorming seems a far cry from something which might produce worthwhile philosophy. If you were presented with a text and told that it contains a number of philosophical ideas, none of which have been filtered for quality, it is unlikely you would think it is worth your while reading it. The filtering these systems are said to lack is, in Lipton's terms, filtering for loveliness: the stochastic core already selects the likeliest continuation, and what it cannot do, on Floridi et al.'s account, is prefer one candidate explanation to another on the grounds of the understanding it would afford.
Floridi et al. do not merely hold that the route by which an LLM produces its output is stochastic rather than abductive; they hold that the output itself is only apparently abductive — a "compelling illusion of genuine and structured inferential reasoning" (2025, p. 2). Trained on a great deal of writing in which explanations are offered and weighed, a model absorbs the forms this writing takes and, prompted to explain, reproduces them, following "the typical phrasing and structure of explanations" and offering "typical causes for typical effects" rather than "reason[ing] about causes from scratch" (2025, p. 9). The output has the shape of a weighing of explanations, taken over from the writing the model has digested, without being one.
Floridi and his colleagues allow that, in ordinary cases such as their own cold-morning car, the model's answer is a good one: "the same explanation a human reasoner would likely choose" (2025, p. 10), one that may be "even optimal by IBE criteria" (2025, p. 19), since such systems "echo the obvious, common explanations" (2025, p. 10). The complaint cannot then be that the explanation is poor, which leaves it hard to say what, in such an answer, is supposed to be merely apparent. What their account points to is the uncommon case: "in less common situations, LLMs can falter" (2025, p. 10), and on inputs "that go beyond their training" "the facade can crack" (2025, p. 9), the success on familiar cases being "a sign of overfitting to common patterns" (2025, p. 15). Once the question of its truth is set aside, this is what the facade reduces to: the claim that the competence shown on common problems is overfitting that would give out in less common ones.
Whether the competence does give out on uncommon cases is an empirical question. A recent survey of abductive reasoning in language models suggests that it does (Salimi et al. 2026). Current models do markedly worse on abductive tasks than on deductive ones: where their median accuracy on deductive tasks is near eighty per cent, on abductive ones it is some forty-two and a half, with a spread "extending down to near-zero accuracy", and the survey reports that "strong deductive performance does not reliably imply strong abductive performance". This is close to Floridi's own diagnosis — an answer that reproduces a common pattern instead of reasoning to it — now voiced from within the field that builds the systems. Taken at face value, it tells in Floridi's favour.
We do not disagree with Floridi et al.'s characterisation of how LLMs function: these systems do not weigh and choose among alternatives in the way that humans do. However, we should not be too quick to jump from this to the conclusion that LLMs cannot _produce text_ which exhibits abductive reasoning. A pocket calculator does not have the capacity to do arithmetic in the way a person does, but does have capacity to produce the correct answer to sums which are entered into it. Similarly, it might be possible for LLMs to produce text which displays abductive reasoning, despite it not being grounded in any actual abductive reasoning.
Despite their stochastic core, LLMs are perfectly capable of producing grammatically correct text: without being given any specific rules, the system "implicitly 'discovers' them—and then seems to be good at following them" (Wolfram 2023). Does this mean that the texts LLMs produce have merely the appearance of being grammatically well-formed? Clearly not. LLM sentences _are_ grammatically well-formed despite their stochastic roots. A stochastic core, then, need not mean that the best an LLM can manage is a veneer of abductive inference. It may be, rather, that the core is marshalled to produce text exhibiting actual abductive inference, in just the way it is marshalled to produce actual grammatical correctness.
It might be objected that the grammar analogy will not stretch this far. The model has picked up the shape of abductive explanation, but it is not actually performing abduction. But grammar, the objector would say, is a system of rules a model can follow without understanding anything, and explanatory loveliness is not — it is not a system of rules you follow to get the right answer, so the model's grammatical competence gives no reason to expect it to produce lovely explanations. Floridi's claim is that the model's abductive output is merely apparent — that it has the form of an inference to the best explanation but not the substance. If that is right, then there should be cases where the form is present and the substance is absent: cases where the model produces something that looks like an explanation but is empty or nonsensical, just as it could produce something that is grammatically faultless and says nothing. Wolfram's example of the latter is "Inquisitive electrons eat blue theories for fish" (2023) — impeccably grammatical, and meaningless. The model does not produce such strings. The abductive equivalent would be a passage with the form of an inference to the best explanation, grammatically faultless, and yet senseless — asked why a car will not start on a cold morning, there are no squirrel tracks, so it must be freak arctic winds blowing into the exhaust pipe. That has the form of evidence weighed towards a conclusion, and it is nonsense; and it is nonsense the model does not produce. Put the question to it and it offers the weak battery and the thickened oil, plausible explanations rather than arctic winds. Wherever the facade is to be located, it cannot be located there: what the model produces are plausible explanations, and it is not yet clear what, in a plausible explanation, is supposed to be merely apparent.
Moreover, the answer Floridi and his colleagues describe is not quite the answer these systems give. Their account asks us to picture a confident verdict laid over a hollow core, yet what one finds in practice is closer to hedging. Asked why a car would not start on a cold December morning, a current model will say that the battery is the most likely culprit, set out other plausible contributors — thickened oil, fuel-system problems, ignition faults — and end by observing that if the car started once the day had warmed, the battery is almost certainly the primary cause, though a load test would confirm it.[^kimi]
[^kimi]: The exchange, with Kimi k2.67 (%%add date of test%%), is reproduced in full in the Appendix. The reply does not announce the answer; it fits its confidence to the little it has been told, and marks the point past which it will not go without knowing more of the particular car. This is not yet to say that Floridi's facade charge is refuted — a hedged answer can still be, on his account, a statistical reproduction of how hedged explanations typically read. But it does mean that Floridi et al.'s own example does not quite fit the charge. The case they describe, in which a confident surface verdict is laid over a hollow core, is not the case their own example presents: the confidence these systems express is already answerable to their evidence, and a confidence so answerable is not the facade the objection has in view.
The model's avoidance of senseless explanations, like its avoidance of senseless sentences, points to a semantic competence picked up in training, over and above syntax. A syntactic grammar settles only how the parts of speech may be combined, and not whether what a sentence says makes sense: avoiding strings like the inquisitive electrons takes a grasp of what can sensibly be said of what — which predicates go with which subjects, which causes with which effects. Wolfram calls such a grasp a semantic grammar:
> to deal with meaning, we need to go further. And one version of how to do this is to think about not just a syntactic grammar for language, but also a semantic one. (Wolfram 2023)
A model trained on enough text has, on his account, come by one:
> From its training ChatGPT has effectively "pieced together" a certain (rather impressive) quantity of what amounts to semantic grammar. (Wolfram 2023)
It is this, rather than any contact with the case in hand, that keeps the arctic winds out of the model's explanations: a feel for what, in a working model of the world, can hang together.
But a feel for how things hang together is gathered from the text the model has read, and a model of the world is not the world.[^wm] What a semantic grammar supplies is a sense of what would sound right, not a line to how things actually stand. The same detachment that keeps the model's explanations sensible is what leaves it unable, on its own, to reach the car on the drive: a semantic grammar is a grasp of how such failures are explained in general — cold slows batteries, thickens oil, fouls ignition — and it gives no hold on the particular car whose fault is in question. Set the particular car aside, and what is left is, so far as the text goes, abduction itself: a weighing of explanations that respects sense, fits its confidence to the evidence, and stops where the evidence stops. How much the missing purchase costs depends on the kind of abduction at issue. The car is a hard case, because its answer waits upon the world; only the car itself can tell the weak battery from the frozen line. Much philosophical abduction does not wait upon the world in that way: what a thought experiment commits us to, or which of two theories carries the lighter explanatory cost, is settled from what is already set down, where the discriminating evidence is the sort a corpus already holds.
The survey scores of Salimi et al. (2026) measure whether the model arrived at the right answer, not the weighing that got it there. On the harder tasks the scores are low, and on long mysteries with their clues strewn through the text the best models fall just short of the average human solver; taken at face value, the numbers count against the model. But the score is a score for the answer — whether the named culprit was the keyed one, whether the right diagnosis came first — and not for the weighing that reached it; such a score, Salimi says, "completely bypass[es] the actual reasoning trace". So when a model misses the keyed culprit, the number marks the miss, and not the comparison it set out on the way — which explanations it canvassed, and why it came down on one. And the survey's hardest tasks are built from low-prior, non-stereotypical outcomes, where several explanations may be reasonable; there, matching an answer to a single reference "underestimates explanation quality", so a low score on such tasks does not show that the model cannot weigh, only that it did not land on the keyed answer. What a one-shot score leaves unscored is the weighing behind it, and that is what a fuller prompting would have to draw out.
It is now difficult to see where the facade should be located. The explanations the model produces are plausible rather than senseless, the confidence they express is fitted to the evidence it has been given, and the benchmark scores that seemed to confirm the diagnosis measure answers rather than the weighing behind them. What the model lacks is purchase on the particular case, and that is a limitation concerning its relation to the world rather than its capacity for abduction. How much it costs a philosophical text is a matter for the next section; what a one-shot prompt leaves undrawn, and how it might be drawn out, is a matter for Section 5.
[^wm]: We use "model of the world" in Wolfram's informal sense: a body of implicitly learned regularities concerning what goes with what. This is weaker than the technical sense now disputed in machine learning, on which a world model is an internal representation of an environment that supports prediction and counterfactual reasoning. Nothing in our argument requires deciding whether LLMs possess world models in the technical sense; how they connect to the world is the matter of §3.
---
# 3. The Challenge from Detachment
In this section we address a second capacity challenge, the _challenge from detachment_. In addition to the argument considered in the previous section, Floridi et al. object that an LLM stands in no relation to the world: its words rest on no perception of anything, and a hypothesis, once produced, is never tested against how things are (2025, pp. 6–7). A discipline whose theories answer to how things are, the challenge runs, cannot be advanced by a system with no access to how things are, and a text produced by such a system gives its reader no reason to think it worth reading. We argue that the challenge fails because the materials philosophy takes from the world reach it already articulated in language, and a model can work on those words as any philosopher does.
We can grant that the model perceives nothing, and that nothing it produces is put to the test against the world. What we should not grant is the picture of philosophy on which these concessions would be disqualifying. Pigliucci (2017), writing on the question of how philosophy makes progress, holds that philosophy is constrained by the world without investigating it as the natural sciences do:
> This means that the basic parameters that philosophers use as their inputs, the starting points of their philosophizing, their equivalent of axioms in mathematics and assumptions in logic (or rules in chess) are _empirical_ data about the world. This data comes from both everyday experience [...] and of course increasingly from the world of science itself. Philosophy, I maintain, is in the business of exploring the sort of conceptually evoked spaces that Smolin is talking about, where the evocation is the result of whatever starting assumptions are made by individual philosophers working within a particular field and, crucially, of the constraints that are imposed by our best understanding of how the world actually is. (2017, pp. 79–80)
Pigliucci takes the term _evocation_ from Smolin (Unger and Smolin 2015).[^4] Evoked truths are neither discovered, in the sense of corresponding to mind-independent states of affairs, nor invented, in the sense of being arbitrary constructs. Smolin's illustration is chess: before the rules of the game were codified there were no facts about chess, but once they were, a great many facts about the game became demonstrable — objective facts, in the sense that anyone who demonstrates one demonstrates the same fact as anyone else (Unger and Smolin 2015, p. 423; quoted at Pigliucci 2017, p. 78). On Pigliucci's proposal, philosophy is in the business of ascertaining evoked truths. What distinguishes it from mathematics and from chess is where its starting points come from: they are constrained by how the world actually is. The same constraint distinguishes philosophy from fiction. Pigliucci's contrast case is a science-fiction writer who describes the same planet across three different timelines, keeping every description within the laws of physics and of logic. The writer is exploring conceptual spaces of a kind, but he is inventing rather than evoking: his worlds have no rigid properties, because even the constraints he keeps to could have been otherwise — he could as easily have imagined planets with a different physics, or a different logic. Philosophy, by contrast, "is in the business of doing empirically informed evoking, not inventing", and its objects of study accordingly have rigid properties (2017, p. 80): once a philosophical starting point has been articulated, what holds within the structure it opens is no longer up to anyone, any more than what follows from the rules of chess is up to anyone. A thought experiment is a case in point. The philosopher sets up an imagined scenario, but explores it "with an interest in figuring things out as far as this world is concerned" (2017, p. 80), and the philosophical work then proceeds within a structure whose properties the philosopher does not control.
Philosophy's starting points — what Pigliucci calls the basic parameters of philosophising, its equivalent of axioms in mathematics and rules in chess — are data about how the world is, drawn from everyday experience and, increasingly, from science. A system that perceives nothing cannot gather such data for itself. But nor does the working philosopher. The scientific data Pigliucci has in mind reach philosophers already set down in language: a philosopher of physics works from published results, not from having run the experiments. And what the discipline retains of everyday experience, it retains in the same form, as the literature's accumulated descriptions of how things seem. Written access to the world's constraints is the profession's normal condition, not a deficiency of it, and a corpus is exactly such access. With respect to nearly all the empirical data philosophy uses, the model stands where every philosopher stands.
Philosophical claims are not, in any case, tested in the way the challenge assumes. The equivalence principle, once Einstein had it, faced a tribunal of measurement — the experiments might have gone against it, and it would then have been dropped. A philosophical thesis faces no such tribunal. Whether it holds is a question about what obtains within an evoked structure, and it is settled in the way questions about chess are settled: by working out what the starting points commit us to, something any competent party can do and none can decide by fiat. This working-out is carried out in writing, in the literature's back-and-forth of argument and reply, and a system that produces text is not debarred from it. A model's inability to run experiments therefore costs it nothing that a philosophical text needs: the testing philosophy does is done on the page.
A model's lack of perception and experiment therefore leaves untouched both the starting points a philosophical text works from and the assessment its claims receive: the former arrive already set down in language, and the latter consists in working out what they commit us to. It might be doubted, however, that experience can be treated in the same way, since a description of what it is like to see red is not obviously an adequate substitute for seeing it, and the next section addresses a challenge built on this doubt.
[^4]: _The Singular Universe and the Reality of Time_ is jointly authored, but its second part, which contains the discussion of evocation, was written by Smolin alone, as Pigliucci notes (2017, p. 77); we follow him in attributing the view to Smolin.
# 4. The Challenge from Experience
In this section we address a third capacity challenge, the _challenge from experience_: the claim that some philosophy depends on experience in a way that a system which has never had any experience cannot meet. Few would say that current LLMs are conscious, and we assume here that they are not. Zahavy (2026) raises a worry of this kind about scientific discovery. A model can carry out the deductive part of discovery, working out the consequences of premises it has been given; what it cannot do, he holds, is produce the premises — make the move from sense experience to new first principles. On the picture he takes from Einstein, that move is a leap,[^5] and it is the leap that gives a theory its axioms. His case is the thought experiment that gave Einstein the equivalence principle:
> Einstein's variation required inventing new axioms based on a physical intuition that did not yet exist in the mathematics. He envisioned a physicist inside an elevator being uniformly accelerated through deep space [...]. Inside this enclosure, the sensory experience reveals a specific pattern: when objects are released, the floor rushes up to meet them. To the physicist, the objects appear to fall with identical acceleration, regardless of composition. Thus, the simulation here was not a permutation of symbols, but a manipulation of perceptual experience. (Zahavy 2026, §5)
Einstein imagines a set of circumstances and attends to what would be experienced within them: everything released inside the elevator appears to fall with identical acceleration. On Zahavy's reconstruction the simulation supplies an observation, and from that observation the new axiom is inferred — the simulated experience of acceleration was indistinguishable from the remembered experience of gravity, and Einstein concluded that the two are one phenomenon.[^2] A model has no access to that observation. It can produce descriptions of elevators and of weightlessness, both present in its corpus, but it has undergone neither, and a discovery whose premises are fixed by simulated experience is beyond a system that, in Zahavy's words, lacks the capacity he calls sensory agency.
Philosophical thought experiments might seem to depend upon experience in the same way. Does Mary, released from her black-and-white room knowing every physical fact about colour vision, learn something when she first sees red (Jackson 1982)? To settle the question one must consider what the experience of seeing red is like, and the argument proceeds from that verdict. Here experience enters as the starting point of an argument, and a system that has never experienced anything seems unable to supply it. Experience enters philosophy as subject matter too, in work that asks what it is to see red (Harman 1990), to feel anger (Goldie 2000), or to have an intuition take hold (Chudnoff 2011).[^3] If starting points of the first kind, and subject matters of the second, can be handled only by a being that has the experiences, then much of the philosophy of mind lies beyond a model's reach.
Nobody who debates the knowledge argument has been through what Mary goes through. The debate proceeds on Jackson's description of the case, together with his claim about what the case shows, and the replies work on that description. The replies the argument has accumulated dispute what follows within that description — whether what Mary gains on her release is knowledge of a fact at all — and none of them draws on what seeing red is like from the inside: they are arguments about what Jackson's description commits us to. Nor is the knowledge argument unusual in this. Nagel (1974) asks what it is like to be a bat without supposing he could find out, working instead from a description of the bat's situation. A model is no worse off here than these philosophers are: it has had no experiences of its own, but what its corpus supplies, and what the philosophy works on, is experience already put into words.
There is, however, one kind of starting point that a model could not have supplied. Merleau-Ponty (1945 %%page%%) observed that when one hand touches the other, the roles of toucher and touched can alternate but cannot be occupied at once. Suppose that he came to this by attending to his own body, and not from anything already in the literature. Then here is a description that only first-person attention could yield, and a system with no body and no first-person view could not have been the first to set it down.
Little follows from this, however. Once the description is written down, everything philosophical that can be done with it — what it shows about bodily self-awareness, what follows once the asymmetry is granted — is there on the page for any reader, and none of it depends on being able to have arrived at the description oneself. Nor is first-person attention the only source of descriptions the literature lacks: a new description can also be reached from the descriptions the literature already contains, by drawing out what they have not been taken to imply, or by combining them as no one has, and this a model can do. What the Merleau-Ponty case shows is that there is one route to new starting points that a model cannot take. Whether the route it can take yields descriptions the literature lacks is the question of novelty, which Section 6 takes up.
A model can, then, do the philosophy that turns on experience, with one exception: it could not be the first to set down a description that only first-person attention could yield. The experiences philosophy argues about reach it as descriptions, which it can work through as well as any reader. In Section 2 abduction entered philosophy as reasoning displayed in a text, and in Section 3 the world entered it as starting points set down in one; here experience enters it the same way. What a model produces is read as any philosophy is read, and how it was produced settles nothing in advance.
[^2]: Zahavy, following Magnani, calls the process _manipulative abduction_: hypothesis generation through the manipulation of a model — here a simulated experience — rather than of symbols (Magnani et al. 2009; Zahavy 2026, §5). It is abduction in Section 2's sense: the equivalence principle is inferred as the best explanation of the simulated observation, the simulation supplying an explanandum that no search over existing text would have produced. What experience contributes, on this picture, is not the inference but its starting point.
[^3]: The materials need not be sensory: the feeling of understanding something is sometimes used to motivate the claim that thought itself has a phenomenology (Pitt 2004).
[^5]: Zahavy too calls this leap abduction, but the word picks out something other than it did in Section 2. There, with Floridi et al., abduction was the weighing of rival explanations, and the charge was that a model only mimics it; here it is the generation of new first principles from experience, and Zahavy's claim is that a model cannot make the move because it has had no experience to move from. The present challenge rests on that second claim, about experience, and not on any verdict about the weighing.
---
---
# 5. The Challenge from Observation
The challenge from observation begins with an obvious question: if all of these arguments are correct, where is all the LLM philosophy? If you ask an LLM the answer to the hard problem of consciousness or the meaning of life, the response you receive will not typically contain much in the way of philosophical argument. More likely you will get a somewhat bland survey of various philosophical positions on the topic, or, perhaps just an irreverent joke.[^gpt55] One might think that whatever the merits of the arguments we have put forward so far, *something* must be preventing them from producing philosophy worth reading because they *don't* produce philosophy worth reading.
%%This paragraph should start with some signposting the shape of the argument we are giving: we are providing an alternative explanation for why we do not see anyt worth reading LLM philosophy.
Actually, I just thought, we could actually start with a preliminary. It is perfectly possible that if someone has read something in a journal in the last half decade it might have been ai generated, to guarentee this hasn't happened we need to know for sure that every single person who had sworn that thye hadn't used ai was telling ther truth. now i write this this is obviously a footnote actually.
- Actually maybe this paragraph should be framed in terms of an explanation for the blandness if you hit llms point blank with a 'philosophical' question. It is because the liekly continuation of such a straight shooting qquestion is not a philosophical tract, it is a glib response. the reader can be subtly pointed back to the relevant stuff we talked about in the abduction section
-
- %%~~Outputs like these are produced under one way of using the systems: a bare question, put once and answered in a single pass.~~
A model produces a reasonable continuation of the text it has been given, so what it returns for a bare philosophical question approximates what follows such questions in writing at large (Wolfram 2023). Very little of what follows "what is the meaning of life?" in writing at large is an argument for anything: where the question receives a philosophical continuation at all, it is mostly in the genres whose business is the balanced survey — the encyclopaedia entry, the introductory textbook. %%this is horribly compressed and so shallow and meaningless. %% The further training that shapes a model into a conversational assistant presses in the same direction, since answers that are balanced, hedged, and committed to nothing are rarely unhelpful to anyone.%%again too compressed, and shit why? %%
A bare question does not articulate a starting point in the sense of Section 3: it makes no starting assumptions and accepts no constraints. %%this is not a good writing, unclear and shite. %%What it does is name a conceptual landscape,%%not how i write%% and the natural answer to a question%%not how i write%% that names a landscape without occupying a position in it is a map. %%really stupid thing to say and a pathetic attempt to connect back up with Piglucci%% Pigliucci, following Rescher, describes the state of a philosophical field as a set of refined alternative peaks in its landscape — aporetic clusters, the families of positions that criticism has not eliminated (2017, %%page%%)[^bc] — and a competent survey of the field is a report of those peaks. %%why the fuck are you talking about the survey in this context, you seemed to have completely changed the topic of the paragraph by this point, it is embarassingly fucking dreadful%% Philosophy worth reading, on everything this paper has argued,%%not how i write%% is not a report of the peaks but movement in the landscape%%literally meaningless, what total fucking shite. %%: the refining of a peak, or the drawing out of a consequence that no position on it had stated. Movement requires a location and a direction, and a bare question supplies neither. %%i have no idea what you are wanking on about.%%
The bare question, then, is a poor test of what these systems can produce. %%why are we still on this topic, i find it amazing how many words you can use to say so very very little, it is a disastor%%It treats them as oracles: systems whose answers to questions exhaust their capacities, so that asking is all the eliciting there is.[^janus] %%this is total bollocks%%But if what a model produces is a continuation of the context it is given, the absence of worthwhile output is evidence about the contexts these systems have been given, not about the capacities they have. %%it is like a version of what i said tyo you, only made of dogshit. %%It cannot discriminate between the two hypotheses at issue — that the capacity defended in the preceding sections is absent, and that it has not been elicited — since both predict the observed record, and the challenge requires the first. %%urrrrgh%%What would elicit the capacity, if it is there, is the one thing the bare question withholds%%bare question is a phrase a cunt would use, withholds is a pretentious way of framing things%%: an articulated starting point. Someone must articulate it, and that is where the final challenge begins. %%absolute total shit%%
---
# 6. The Challenge from Instrumentality
In this final section we address the _challenge from instrumentality_. Grant that philosophy worth reading comes out of these systems only when a philosopher directs the process, framing the problem and pressing for development. It may then be said that the philosophy in the resulting text belongs to the philosopher and not to the model. This is, in effect, the position at which Floridi et al. arrive: LLMs are brainstorming assistants, and it falls to "a cautious human collaborator" to "sift through and assess" what they toss out (2025, p. 11) — the assessing is the philosophy, and it is the human's. On this picture the model is an instrument of philosophical production, as a typewriter is, and to credit it with the result would be like crediting the typewriter with the novel. The challenge from authorship held that a model's text cannot be philosophy at all; the present challenge — call it the _challenge from instrumentality_ — is narrower, holding that whatever philosophy such a text contains is not the model's contribution.
Whether the challenge succeeds depends on what a prompt contributes and on what the continuation adds to it. A prompt that sets out a position and the objections any defence of it must meet stands to the model as Jackson's two paragraphs stand to the profession: a starting point handed over for development. And an articulated starting point, as Section 3 argued, evokes a structure with rigid properties: there are facts about what holds within it, demonstrable by anyone and chosen by no one, and they outrun whatever has been stated, just as the facts about chess outran the rules the moment the rules were written down. Most of what a starting point evokes, no one has ever said.
A model produces a reasonable continuation of the text it has been given, and a prompt is part of that text, so an articulated starting point changes what there is to continue. Wolfram observes that a model told something once will make use of it thereafter, and his gloss on the mechanics is worth quoting: "the elements are already in there, but the specifics are defined by something like a 'trajectory between those elements' and that's what you're introducing when you tell it something" (2023). The elements are the corpus's; what the prompt introduces is a trajectory between them; and the continuation is the text of that trajectory, stating consequences of the starting point which the starting point does not itself state. Whether it states them well is judged as any philosophical text is judged: by whether the considerations it cites genuinely tell between the positions, and by the understanding they would, if correct, provide, and the judgement is made by reading the continuation.
It might be thought that the continuation only spells out what the prompt had already settled. It does not, because a starting point underdetermines its developments. The rules of chess settle every fact about chess, but they do not settle which of those facts get written down, in what order, or to what depth: two writers working from the same rules will produce different books, both correct, and neither dictated by the rules. The same holds here — the same prompt, run twice, yields different continuations (Wolfram 2023) — and the gap between starting point and development is where the model's contribution lies. Were there a single text that the prompt fixed, the output would merely restate what the prompter had settled, and the instrument description would be accurate.
A development can also go wrong. A chess writer can publish a false theorem; a philosopher can misdraw the consequences of her own thought experiment; and a model can do both, along with its characteristic failure of asserting fluently what nothing supports. Such errors are found where the development is found, on the page, and they attach only to whatever produced the development: nobody blames the rules of chess for a false theorem, and nobody's typewriter has ever made an error of content. There are, then, three contributions to distinguish: the articulated starting point, which is the prompter's; the structure it evokes, whose facts belong to no one; and the text that develops them, which is the model's.
The person writes the prompt and decides which continuations to pursue and when to stop; without the person there is only the survey of the previous section. So much of the instrument picture is correct. What must be given up is the description of the model on which the picture turns. A typewriter fixes nothing its user has not already settled: every word of the novel was the novelist's before the keys were struck. The consequences a model's text states were nobody's before the text stated them. What the user of a typewriter settles is a text; what the writer of a prompt settles is a starting point.
It may be replied that the prompt can always be enriched — the position stated, the rivals named, the objections listed together with the lines along which they are to be met — until the prompt contains the philosophy and the model merely expands upon it. But enriching a prompt enlarges the starting point without converting it into the development: a game with more rules is a larger game, not a book of its theorems, and however much the prompt states, the consequences the output draws were not among the statements. There is a genuine limiting case, in which the prompt states the comparison and the verdict and the continuation only rephrases them, and it is identified by reading the two together. A text that states nothing beyond its prompt is a paraphrase, and owed to the prompter; a text that states what the prompt left unstated is a development, and the unstated part is not the prompter's. Contributions are individuated this way in the literature quite generally — by what a text states that the texts before it did not — and nothing in that way of individuating them mentions what produced the text.
A further question is whether such a development can state something the literature itself had not yet stated, and so be creative in the sense in which a human philosophical text is creative, rather than merely novel relative to its prompt. Here the mechanics impose a limit of their own. Wolfram's observation about telling a model something carries a caveat: the model can integrate what it is told only if it is "basically riding in a fairly simple way on top of the framework it already has" (2023). Trajectories, that is, run between elements the corpus already contains. Developments within the framework are secured by the account above; whether philosophy worth reading ever requires more than this — a starting point that does not ride on the existing framework, a leap of the kind Section 4 conceded in Merleau-Ponty's case — is a question this paper leaves open, and it would be settled as the narrower one was: by reading the output against the literature and asking what it states that the literature did not.
These systems are not oracles but continuation systems. Philosophy worth reading requires a dialectical context, and a prompt can supply one without thereby fixing the development that follows from it. The philosophical standing of any output is therefore settled by what its continuation adds: first to the prompt that occasioned it, and then to the literature it would join.
[^bc]: Pigliucci reads Bourget and Chalmers' (2013) survey of professional philosophers' views as evidence for this picture of the discipline: the distribution of opinion it records is a distribution over such peaks.
[^gpt55]: While preparing this paper we asked GPT-5.5 for a detailed overview of the positions an analytic philosopher might take on the meaning of life. What came back was a competent, hedged survey of the field; what did not come back was an argument for any position in it. %%add date of test%%
[^janus]: That these systems are mischaracterised as oracles — with the corollary that no benchmark of single-pass answers should be expected to probe the upper limits of what they can produce — has been argued from inside the practitioner literature (Janus 2022). The benchmark literature points the same way: the survey of abductive benchmarks discussed in Section 2 scores single-pass answers to fixed instructions while cataloguing, in the same pages, prompting methods that alter what models produce (Salimi et al. 2026).
# Appendix: the cold-morning exchange
> **User:** Last December, I tried to start my car early in the morning when it was very cold. It wouldn't start. Why not?
>
> **Kimi k2.67:** The most likely culprit is the battery. In very cold weather, a battery's chemical reactions slow dramatically, reducing its available capacity by up to 50%. If the battery was already marginal (older than three years, or not fully charged), freezing temperatures can push it below the threshold needed to turn the starter motor. You may have heard rapid clicking or nothing at all.
>
> Other plausible contributors:
>
> - Thickened engine oil — Cold increases viscosity, making the engine physically harder to crank and demanding even more current from the already weakened battery.
> - Fuel system — Condensation in the fuel lines can freeze, blocking flow. Diesel fuel can gel; petrol can struggle to vaporise properly in extreme cold.
> - Spark/ignition — Moisture in the distributor cap or on spark plugs can freeze, preventing a proper spark.
>
> If it started fine once temperatures rose later in the day, the battery is almost certainly the primary cause. A load test would confirm whether it needs replacement or just a longer drive to reach full charge.
---
# References
# References
Frankish, K. (2024). What are large language models doing? In A. Strasser (Ed.), _Anna's AI Anthology: How to live with smart machines?_ (pp. 55–78). Xenomoi.