# 0. Introduction
The last decade or so has seen the rise of generative artificial intelligence: systems that produce text, images, code, music, video, and other outputs in response to prompts. AI has had success in domains where the value of an output is not exhausted by its superficial fluency. For example, in February 2026, researchers working on gluon scattering amplitudes gave GPT-5.2 worked examples for three, four, five, and six particles and asked it to find the general formula. GPT-5.2 proposed a closed-form expression; another internal model supplied a proof; and the authors then verified the result. The resulting paper argues that single-minus tree-level gluon amplitudes, often presumed to vanish, are non-vanishing in certain half-collinear configurations (Guevara et al. 2026). There are also recent examples in mathematics (Novikov et al. 2025), biomedicine (Gottweis et al. 2025), and materials science (Zeni et al. 2025). In this paper we argue that we should expect similar success in philosophy.
Specifically, we argue that current-generation LLMs are capable of producing philosophical texts that are _worth reading_. We do not want to begin by settling what counts as _good_ philosophy. Instead, we appeal to a distinction that anyone reading this text will recognise. You have read texts that are worth reading, and you have read texts that are not. As you begin reading this article, you likely hope that it is worth reading, in the sense that the time spent reading it will not be wasted. When you write a philosophical text yourself you aim to make it worth readers' while to read it, and whether or not the journal you send it to accepts it, depends on whether or not they agree.
Two clarifications are needed. First, a text’s being worth reading is not the same as its being correct. A text can repay attention even if one rejects its conclusion: it may sharpen a distinction or answer an objection in a way that changes the dialectical situation. Second, the minimal unit we are concerned with is not the bare conclusion of an argument, but the argument itself. If an LLM output consists only in a pronouncement on some philosophical topic ('Direct Realism is correct', 'We should be utilitarians'), it is hard to see why it would be worth reading in and of itself, for the same reason that a bare pronouncement by a human philosopher would not be worth reading.[^1]
The next four sections develop the main argument. Section I rejects the challenge from authorship: the claim that an LLM output cannot be philosophy worth reading because no philosopher lies behind it. Section II turns to abduction and argues that the absence of human-style inference to the best explanation in the producer does not preclude abductive structure in the product. Section III turns to the challenges from experience and from connecting to the world, and argues that the materials philosophy takes from each reach it already set down in words, which an LLM can work on as any philosopher does. Section IV turns to the challenge from observation — that if all this is right, philosophy worth reading should already be coming from these systems, and is not — and argues that this reflects how they are used rather than what they can produce.
---
# 1. The Challenge from Authorship
In this section we address what we might call the *challenge from authorship*: the idea that philosophy is something that only persons, or at least minds, can produce. This view has not, to our knowledge, been explicitly defended in just this form, but it gives shape to an intuition that many philosophers may have: philosophy is a person-only domain. An imperfect comparison is with art: One might deny that an image generated by an AI system is an artwork, because no artist exercises the relevant kind of intentional control over its production.[^1] One might think, for similar reasons, that philosophy can only be done by people: no text produced by an LLM can be a work of philosophy, because no philosopher lies behind it.
Similar to the study of art, the study of philosophy is often organised around individuals: undergraduates take courses on Kant's ethics or Lewis's metaphysics, and conferences are devoted to, the work of particular philosophers. Physics students, on the other hand, are taught Newtonian mechanics from a current textbook, and the course loses nothing if Newton's own writing is never looked at. In the sciences, then, what a text contributes can be carried by other texts. In philosophy, the contribution and its original presentation are harder to prise apart, and we might take this as evidence that a philosophical work is bound to the activity of the particular person who produced it, in a way that the sciences are not.
We will now try to make this challenge from authorship more precise, by considering how far Davies' *performance* theory of art transposes to philosophy. Davies writes:
> [T]he work – what the artist achieves – is the process eventuating in that product. Works themselves are neither structures nor objects simpliciter, nor are they contextualized structures or objects. They are, rather, intentionally guided generative performances that eventuate in contextualized structures or objects (or events, as we shall see) – performances completed by what I am terming a focus of appreciation. (2004, p. 97)
On Davies' view, when a painter paints a picture, the canvas is what we attend to, but it is not the work. The work is the artist's intentionally guided activity in producing the canvas; the canvas is "the focus of our appreciative interest in the work" (2004, p. 150). What we appreciate in a painting, on this account, is an achievement, and an achievement is individuated by the activity that brought it about: the same surface, reached by some other route, would be a different achievement. Provenance, on this view, does more than supply context: facts about how the object came into being help determine what the work is and what is properly appreciated in it.
Davies makes the case with examples in which indiscernible surfaces differ as works. In one kind there is no performance at all: an instance of the verbal structure of *Kubla Khan* might be generated by desert wind, or a monkey at a typewriter. A theorist who identifies the poem with its verbal structure must then either count these as instances of Coleridge's work or explain why not (2004, p. 102). There is a surface indistinguishable from the poem with no writing of a poem behind it, and the corresponding case for philosophy — a text indistinguishable from a philosophical argument with no philosophising behind it — is the one an LLM presents. Questions of misattribution, where one performance is taken for another, do not bear on it. If Davies is right, the surface does not by itself settle the work.
Here is what the analogous proposal for philosophy would be. A philosophical text is not itself the philosophical work: the text is the product of a person's philosophising, and reading it is a way of engaging with that prior activity. The activity, on this proposal, is part of what the work is, so that there is a philosophical work only where the philosophising has taken place, which makes the challenge a constitutive one. If no one has philosophised, there is no work to which the text gives access, however the text reads — an LLM text would stand to philosophy as the wind-made *Kubla Khan* stands to poetry.
Should the transposition be accepted? We do not think it should. In the art case Davies' move is licensed by an evaluative fact: indiscernible surfaces can differ in artistic value, the wind-blown surface worth nothing as a painting where a brushed one may be worth a great deal. The transposition therefore commits its defender to the corresponding claim about philosophy: that two texts containing the same argument could differ in philosophical merit. We can find no difference for the merit to consist in. If two texts contain the same argument, including the same inferential moves, the same considerations count for and against them: whether the argument is valid and whether the objections are answered are questions about the texts' contents, and two texts with the same contents receive the same answers. Their philosophical merit does not vary with the route by which the words came to be written.
The discipline's evaluative practice is built on the same denial. Journals strip author information from submissions before review because facts about authorship are treated as potential sources of distortion; if texts with the same contents could differ in merit, anonymising would discard evaluatively relevant information, and review would not be designed this way. The grounds for the judgement lie in the argument as presented, not in the history of its production.[^3] Nor is the author-centred teaching noted earlier in tension with this. That ethics is taught through Kant's *Groundwork of the Metaphysics of Morals* rather than a digest of its conclusions reflects what a reader gains in understanding by working through Kant's own arguments, and is consistent with holding that the philosophical merit so gained is a feature of the text rather than of its author.
Suppose the performance theorist holds firm: where there has been no philosophising there is no work, whatever the resulting text contains. This may be allowed, because the thesis in question concerns texts worth reading, and a text can be worth reading without being a work in Davies' sense. Suppose the desert wind assembled not *Kubla Khan* but a sound argument against enactivist approaches to perception, an argument for which no one would deserve credit. A reader who worked through it would nonetheless meet a thesis and the considerations advanced for it, and would be in a position to answer or extend it. Whether such a text is a *work* may then be reserved for texts with performances behind them; what cannot be reserved is the text's being worth reading, since everything that judgement answers to is on the page.
Nothing in a text's having had no one behind it, then, settles whether it is worth reading. Whether an LLM can produce a text worth reading is a further question, and a doubt of a different kind bears on it.
[^1]: This is not to deny that systems of this kind can produce beautiful images; we return to image generation in Section 4.
[^3]: The same location of philosophy in the public text is reached by accounts of philosophical progress. Dellsén et al. (2024) hold that progress is "for-whom" rather than "by-whom": it consists in putting people in a position to increase their understanding, usually by making philosophical ideas publicly available (p. 679). Being put in such a position requires something one can take up and think through, and what is available to be taken up is the text. A philosophical contribution so understood is constituted by what the public text makes available, which a view that locates the philosophy in the antecedent private activity mislocates.
---
# 2. The challenge from abduction
In the previous section, we argued against the idea that LLMs cannot produce philosophy worth reading simply because they are not human. In this section and the next we shall consider a different form of challenge: even if LLMs cannot be ruled out of producing philosophy worth reading tout court, they lack particular _capacities_ that are required to produce it. If our arguments regarding authorship are correct, there is no reason that a novel philosophical argument produced by a parrot should be taken any less seriously than one produced by a human. Yet parrots _cannot_ produce such complex strings of sounds as their powers are mimetic, rather than productive. One might think the same is true for LLMs. They just don't have a capacity, or capacities, required to produce worthwhile philosophical argument. In this section we address one capacity challenge, which we will call _the challenge from abduction_. In the next we shall look at two more: the *challenge from phenomenology* and the challenge from *connecting to the world*.
Abductive inference is reasoning from a body of evidence to the hypothesis that would best explain it. In a deductive argument the premises fix the conclusion: if all men are mortal and Socrates is a man, then Socrates is mortal, with no wriggle room whatsoever. Now, imagine walking into your kitchen one morning and finding that part of the floor is wet. What has happened? The wet floor does not determine the answer in the way the two premises determined Socrates' mortality: a burst pipe would have left the floor wet, and so would a spilled bucket. But, given that the window is open, the water is under the window, and it rained last night, rain coming through the window seems the most plausible answer. To reason in this way, inferring the best explanation for a set of facts, is common in the sciences as well as every day life. A scientist chooses one theory over another when it explains the same data more simply: Copernicus's model of the solar system was preferred to Ptolemy's because it explained the observed planetary motions without the elaborate epicycles the older model required.
Williamson argues that philosophy should use a broadly abductive methodology (2007; 2021, §9.2). In philosophy, as in science, there are data that a candidate theory must accommodate — intuitions about cases, and the phenomena of the domain itself — and rival theories that would each accommodate them at different costs. The theory to prefer is the one that would, if true, best accounts for the data, with 'best' being cashed out in terms of explanatory virtue What makes one theory's explanation better than another's, on this account, is a matter of explanatory virtue: a good philosophical theory is, in Williamson's words, "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated", and should "combine simplicity with strength" (2021, §9.2). This need not rest on a science-modelled conception of philosophy: Bengson, Cuneo and Shafer-Landau (2022) argue that the assessment of rival theories by their explanatory and unifying merits is a constraint on sound philosophical method as such.[^holders] We shall assume it in what follows.
If a capacity for abduction is required to produce worthwhile philosophy, do LLMs possess it? Floridi et al. (2025) argue that they do not:
> We argue that such LLMs generate text based on learned associations rather than performing abductive inferences. […] LLMs produce plausible hypotheses, simulate commonsense reasoning, and provide explanatory answers without grounding them directly in truth, semantics, verification, or understanding, and without any abductive reasoning. (Floridi et al. 2025, p. 1)
We will split this quite comprehensive critique of LLMs' abductive capacities into two parts. In the rest of this section we will consider the charge that these systems do not genuinely infer to the best explanation but simply generate text based on learned associations. In the next, we will focus on the idea that LLMs are not connected to the world in a way that would allow them to "genuinely validate" (ibid. p. 6) their explanations against reality.
An LLM, on this account, has "a stochastic core and an abductive appearance" (2025, p. 2). Models are trained to predict which words are likely to follow which, and to produce the continuation their training makes probable — to aim at the likely continuation rather than truth. The appearance comes from what the training data have passed on: models have "absorbed patterns of human abductive reasoning as expressed in writing" (p. 9) — how explanations are typically phrased, which causes are typically offered for which effects. [^1] Floridi et al.'s own example is a car that will not start on a cold morning. Asked why not, a model replies that a weak battery is one possibility, since cold reduces a battery's efficiency; that thickened engine oil is another, since a cold engine is harder to turn over; and that, "[b]ased on your description, the battery is the most likely explanation" (2025, p. 10). On their reading, the LLM is not actually reasoning about causes from the case before it; it is the statistical reproduction of the causes such explanations typically cite (p. 9). What they deny is that the model weighs the candidates — that it "entails choosing the best explanation among alternatives" (2025, p. 3). The verdict that settles on the battery does no weighing; it reproduces how explanations of this kind conventionally end. And where the output marks a genuine difference between the two — some consideration that would tell the battery from the oil — it is one already drawn in the explanations the model learned from, not one worked out afresh for the case in hand.
The fact that LLMs can produce no more than a "veneer of explanation" (ibid. p.20), Floridi et al. argue, means that LLMs can only ever play a supporting role in intellectual work:
> In essence, LLMs function like brainstorming assistants that toss out ideas without filtering for quality. After all, they work like statistical interfaces to an enormous amount of data accumulated for millennia by generations. A cautious human collaborator can sift through and assess them. (2025, p. 11)
A brainstorming assistant seems a far cry from something which might produce worthwhile philosophy. If you were presented with a text and told that it contains a number of philosophical ideas, none of which have been filtered for quality, it is unlikely you would think it is worth your while reading it.
The filtering Floridi et al. have in mind is not a matter of selecting the most likely continuation — the stochastic core already does that. It is a matter of preferring what Lipton (2004, p. 59) calls *lovelier* explanations to merely *likelier* ones: the likeliest is the one most warranted by the data, the loveliest the one that would, if true, provide the most understanding. As he puts it, "Likeliness speaks of truth; loveliness of potential understanding" (PAGE REF). The two come apart. Consider again the wet floor in the kitchen. A very likely, almost certainly true, explanation is that the floor is wet because water has fallen on it, yet offering that as an explanation would be met with exasperation: of course it is because water fell on it, but how, and which water? Banally true explanations do little for a person's understanding. On Williamson's abductive methodology, weighing rival explanations is a matter of explanatory virtue, and the explanation to prefer is the one with a virtue the others lack (2021, §9.2). To prefer an explanation on that ground is to prefer the lovelier rather than the likelier, since explanatory virtue is a matter of the understanding an explanation would afford if true, not of its probability. Not all philosophical writing turns on this kind of weighing, but where a philosophical text does turn on it, whether the text is worth reading and whether it weighs its rivals well go together. And it is just this kind of weighing that Floridi and his colleagues say a system that does no more than continue text cannot do.
We do not disagree with Floridi et al.'s characterisation of how LLMs function: these systems do not weigh and choose among alternatives in the way that humans do. However, we should not be too quick to jump from this to the conclusion that LLMs cannot *produce text* which exhibits abductive reasoning. A pocket calculator does not have the capacity to do arithmetic in the way a person does, but does have capacity to produce the correct answer to sums which are entered into it. Similarly, it might be possible for LLMs to produce text which displays abductive reasoning, despite it not being grounded in any actual abductive reasoning.
Floridi et al. are not merely saying that the process by which an LLM produces its output is stochastic rather than abductive. They are saying that the output itself is only apparently abductive — that what looks like inference to the best explanation is, in their words, a "compelling illusion of genuine and structured inferential reasoning" (2025, p. 2). Trained on a great deal of writing in which explanations are offered and weighed, such a system absorbs the forms this writing takes and, prompted to explain, reproduces them, following "the typical phrasing and structure of explanations" and offering "typical causes for typical effects" rather than "reason[ing] about causes from scratch" (2025, p. 9). On this view the output has the form of an abductive explanation but not the substance — the shape of a weighing of explanations, taken over from the writing the model has digested, and not a genuine weighing of the case in hand.
Floridi and his colleagues allow that, in ordinary cases such as their own cold-morning car, the model's answer is a good one: "the same explanation a human reasoner would likely choose" (2025, p. 10), one that may be "even optimal by IBE criteria" (2025, p. 19), since such systems "echo the obvious, common explanations" (2025, p. 10). The complaint cannot then be that the explanation is poor, which leaves it hard to say what, in such an answer, is supposed to be merely apparent. What their account points to is the uncommon case: "in less common situations, LLMs can falter" (2025, p. 10), and on inputs "that go beyond their training" "the facade can crack" (2025, p. 9), the success on familiar cases being "a sign of overfitting to common patterns" (2025, p. 15). Once the question of its truth is set aside, this is what the facade reduces to: the claim that the competence shown on common problems is overfitting that would give out in less common ones.
Whether the competence gives out on the uncommon case is an empirical question, and the most recent survey of abductive reasoning in language models seems at first to bear the conjecture out (Salimi et al. 2026). Current models do markedly worse on abductive tasks than on deductive ones: where their median accuracy on deductive tasks is near eighty per cent, on abductive ones it is some forty-two and a half, with a spread "extending down to near-zero accuracy", and the survey reports that "strong deductive performance does not reliably imply strong abductive performance". This is close to Floridi's own diagnosis — an answer that reproduces a common pattern instead of reasoning to it — now voiced from within the field that builds the systems. Taken at face value, it tells in Floridi's favour.
Our response to the challenge from abduction begins by considering what text having an abductive appearance actually amounts to. Consider first that, despite their stochastic core, LLMs are perfectly capable of producing grammatically correct text. Despite not being given specific rules, LLM training means that the system "implicitly 'discovers' them—and then seems to be good at following them" (Wolfram 2023). Does this mean that the texts LLMs produce have merely the appearance of being grammatically well-formed? Clearly not. LLMs sentences *are* gramatically well-formed despite their stochastic roots. This suggests that a stochastic core need not mean that the best an LLM can do is produce a veneer of abductive inference. It may be, rather, that the core is marshalled to produce text exhibiting actual abductive inference, in just the way it is marshalled to produce actual grammatical correctness.
It might be objected that the grammar analogy will not stretch this far. The model has picked up the shape of abductive explanation, but it is not actually performing abduction. The preceding paragraph suggested that because LLMs implicitly discover and follow the rules of grammar, they might in the same way be marshalled toward genuine abductive inference. But grammar, the objector would say, is a system of rules a model can follow without understanding anything, and explanatory loveliness is not — it is not a system of rules you follow to get the right answer, so the model's grammatical competence gives no reason to expect it to produce lovely explanations. Floridi's claim is that the model's abductive output is merely apparent — that it has the form of an inference to the best explanation but not the substance. If that is right, then there should be cases where the form is present and the substance is absent: cases where the model produces something that looks like an explanation but is empty or nonsensical, just as it could produce something that is grammatically faultless and says nothing. Wolfram's example of the latter is "Inquisitive electrons eat blue theories for fish" (2023) — impeccably grammatical, and meaningless. The model does not produce such strings. The abductive equivalent would be a passage with the form of an inference to the best explanation, grammatically faultless, and yet senseless — asked why a car will not start on a cold morning, there are no squirrel tracks, so it must be freak arctic winds blowing into the exhaust pipe. That has the form of evidence weighed towards a conclusion, and it is nonsense; and it is nonsense the model does not produce. Put the question to it and it offers the weak battery and the thickened oil, plausible explanations rather than arctic winds. Wherever the facade is to be located, it cannot be located there: what the model produces are plausible explanations, and it is not yet clear what, in a plausible explanation, is supposed to be merely apparent.
Moreover, the answer Floridi and his colleagues describe is not quite the answer these systems give. Their account asks us to picture a confident verdict laid over a hollow core, yet what one finds in practice is closer to hedging. Asked why a car would not start on a cold December morning, a current model will say that the battery is the most likely culprit, set out other plausible contributors — thickened oil, fuel-system problems, ignition faults — and end by observing that if the car started once the day had warmed, the battery is almost certainly the primary cause, though a load test would confirm it.[^kimi] The reply does not announce the answer; it fits its confidence to the little it has been told, and marks the point past which it will not go without knowing more of the particular car. This is not yet to say that Floridi's facade charge is refuted — a hedged answer can still be, on his account, a statistical reproduction of how hedged explanations typically read. But it does mean that Floridi et al.'s own example does not quite fit the charge. The case they describe, in which a confident surface verdict is laid over a hollow core, is not the case their own example presents: the confidence these systems express is already answerable to their evidence, and a confidence so answerable is not the facade the objection has in view.
That the model keeps clear of senseless explanations, as it keeps clear of senseless sentences, points to a core of semantic competence picked up in training, over and above syntax. Wolfram calls it a semantic grammar. Syntax, he notes, settles only how the parts of speech may be combined:
> to deal with meaning, we need to go further. And one version of how to do this is to think about not just a syntactic grammar for language, but also a semantic one. (Wolfram 2023)
A model trained on enough text has, on his account, come by one:
> From its training ChatGPT has effectively "pieced together" a certain (rather impressive) quantity of what amounts to semantic grammar. (Wolfram 2023)
Wolfram's notion of a semantic grammar is meant to capture what a model picks up over and above syntax. A syntactic grammar settles only how the parts of speech may be combined — what counts as a well-formed sentence. It does not settle whether what a sentence says makes sense. "Inquisitive electrons eat blue theories for fish" is grammatical and says nothing, and it is the absence of such strings from the model's output that tells against treating its syntax as appearance alone. Avoiding such strings takes more than syntax: it takes a grasp of what can sensibly be said of what, which predicates go with which subjects, which causes go with which effects. Wolfram calls this a semantic grammar, and his claim is that a model trained on enough text has come by one — has, in his words, "pieced together" a quantity of what amounts to semantic grammar (2023). It is this, rather than any contact with the case in hand, that keeps the arctic winds out of the model's explanations: a feel for what, in a working model of the world, can hang together.
But a feel for how things hang together is gathered from the text the model has read, and a model of the world is not the world.[^wm] What a semantic grammar supplies is a sense of what would sound right, not a line to how things actually stand. The same detachment that keeps the model's explanations sensible is what leaves it unable, on its own, to reach the car on the drive. A semantic grammar is a grasp of how such failures are explained in general — cold slows batteries, thickens oil, fouls ignition — and it gives no hold on the particular car whose fault is in question. Pressed for the cause of this failure to start, the model can go only so far before it needs what it cannot get for itself: some purchase on the actual car in the actual world, the lights tried, the turn of the key heard. Set that aside, and what is left is, so far as the text goes, abduction itself: a weighing of explanations that respects sense, fits its confidence to the evidence, and stops where the evidence stops. How much the missing purchase costs depends on the kind of abduction at issue. The car is a hard case, because its answer waits upon the world; only the car itself can tell the weak battery from the frozen line. Much philosophical abduction does not wait upon the world in that way: what a thought experiment commits us to, or which of two theories carries the lighter explanatory cost, is settled from what is already set down, where the discriminating evidence is the sort a corpus already holds. The case that shows the model at its most hobbled is thus the world-bound one, and not the case philosophy most often presents — a matter we take up in the next section. What is left for this one is the empirical question of whether, on harder cases, this competence in fact gives out.
Returning to the survey by Salimi et al. (2026), the scores measure whether the model arrived at the right answer, not the weighing that got it there. On the harder tasks the scores are low, and on long mysteries with their clues strewn through the text the best models fall just short of the average human solver; taken at face value, the numbers count against the model. But the score is a score for the answer — whether the named culprit was the keyed one, whether the right diagnosis came first — and not for the weighing that reached it; such a score, Salimi says, "completely bypass[es] the actual reasoning trace". So when a model misses the keyed culprit, the number marks the miss, and not the comparison it set out on the way — which explanations it canvassed, and why it came down on one. And the survey's hardest tasks are built from low-prior, non-stereotypical outcomes, where several explanations may be reasonable; there, matching an answer to a single reference "underestimates explanation quality", so a low score on such tasks does not show that the model cannot weigh, only that it did not land on the keyed answer. What a one-shot score leaves unscored is the weighing behind it, and that is what a fuller prompting would have to draw out.
The section has argued that Floridi's facade charge does not hold against the output the model actually produces. The model avoids senseless explanations, fits its confidence to its evidence, and stops where the evidence stops; what it lacks is not abductive competence but purchase on the particular case. The semantic grammar that keeps the model's explanations sensible is gathered from text, not from the world, so the model can say how cold-weather starting failures are explained in general but cannot reach the particular car on the drive — the lights tried, the turn of the key heard — because that takes contact with the world, and the model has none. But that is a limitation about the world, not about abduction. Whether the competence gives out on harder cases is an empirical question the survey does not settle, because the scores measure answers, not the weighing behind them: a one-shot prompt may miss the weighing without showing it is not there. That leaves two questions for what follows. The next section takes up the world-connection challenge: whether the model's detachment from the world disqualifies its output from counting as genuine philosophy, or whether the kinds of abduction philosophy most often demands are precisely those that do not wait upon the world. The section after that takes up the elicitation question: if the weighing is there but a one-shot prompt does not draw it out, what does?
[^wm]: We use "model of the world" in Wolfram's sense; whether such systems possess world models in the sense now disputed, and how they connect to the world, is the matter of §3.
Example of real exchange with kimi 2.6 high
> **User:** Last December, I tried to start my car early in the morning when it was very cold. It wouldn't start. Why not?
>
> **Kimi k2.67:** The most likely culprit is the battery. In very cold weather, a battery's chemical reactions slow dramatically, reducing its available capacity by up to 50%. If the battery was already marginal (older than three years, or not fully charged), freezing temperatures can push it below the threshold needed to turn the starter motor. You may have heard rapid clicking or nothing at all.
>
> Other plausible contributors:
>
> - Thickened engine oil — Cold increases viscosity, making the engine physically harder to crank and demanding even more current from the already weakened battery.
> - Fuel system — Condensation in the fuel lines can freeze, blocking flow. Diesel fuel can gel; petrol can struggle to vaporise properly in extreme cold.
> - Spark/ignition — Moisture in the distributor cap or on spark plugs can freeze, preventing a proper spark.
>
> If it started fine once temperatures rose later in the day, the battery is almost certainly the primary cause. A load test would confirm whether it needs replacement or just a longer drive to reach full charge.
---
# 3. The challenge from experience and the challenge from connecting to the world
In this section we address two more capacity challenges, which we will call the _challenge of connecting to the world_ and the _challenge from experience_. The first holds that LLMs cannot produce philosophy worth reading because they do not stand in the right sort of relation to the world: nothing a model says rests on perception of anything, and nothing it produces is tested against how things are. The second holds that they cannot because some philosophy depends on experience in a way a system without experience cannot meet: experience supplies the starting point of some philosophical reasoning, and is itself the subject matter of some philosophical inquiry. Few would say that current LLMs are conscious, and we assume here that they are not. We argue that both challenges fail, and for the same reason: the materials philosophy takes from the world and from experience are *articulated* in text , and a model works on those words as any philosopher does.
As well as what we saw in the previous section, Floridi et al. object that the model stands in no relation to the world: its words rest on no perception of anything, and a hypothesis, once produced, is never tested against how things are (2025, pp. 7–9). A discipline whose theories answer to how things are, the challenge runs, cannot be advanced by a system with no access to how things are, and a text from one gives its reader no reason to think it worth reading.
The challenge from experience says that some philosophy cannot be done by anyone who has not had the relevant experience. Zahavy (2026) raises a similar worry about scientific discovery. A model can carry out the deductive part of discovery, working out the consequences of premises it has been given; what it cannot do, he holds, is produce the premises — make the move from sense experience to new first principles. On the picture he takes from Einstein, that move is a leap, and it is the leap[^5] that gives a theory its axioms. His case is the thought experiment that gave Einstein the equivalence principle:
> Einstein's variation required inventing new axioms based on a physical intuition that did not yet exist in the mathematics. He envisioned a physicist inside an elevator being uniformly accelerated through deep space [...]. Inside this enclosure, the sensory experience reveals a specific pattern: when objects are released, the floor rushes up to meet them. To the physicist, the objects appear to fall with identical acceleration, regardless of composition. Thus, the simulation here was not a permutation of symbols, but a manipulation of perceptual experience. (Zahavy 2026, §5)
Einstein imagines a set of circumstances and attends to what would be experienced within them: everything released inside the elevator appears to fall with identical acceleration. On Zahavy's reconstruction the simulation supplies an observation, and from that observation the new axiom is inferred — the simulated experience of acceleration was indistinguishable from the remembered experience of gravity, and Einstein concluded that the two are one phenomenon.[^2] A model has no access to that observation. It can produce descriptions of elevators and of weightlessness, both present in its corpus, but it has undergone neither, and a discovery whose premises are fixed by simulated experience is beyond a system that, in Zahavy's words, lacks the capacity he calls sensory agency.
We might think that philosophical thought experiments depend upon experience in the same way. Does Mary, released from her black-and-white room knowing every physical fact about colour vision, learn something when she first sees red (Jackson 1982)? Settling the question requires considering what the experience is like, and the argument proceeds from the verdict; this is experience entering as the starting point of an argument, and a system that has never experienced anything appears unable to supply it. Experience enters also as subject matter, in philosophy that asks what it is to see red (Harman 1990), to feel anger (Goldie 2000), or to have an intuition take hold (Chudnoff 2011).[^3] Were these claims correct, much of the philosophy of mind would lie beyond a model's reach.
We can grant that the model has no senses and has never had an experience, and that it cannot put what it produces to the test against the world. Whether any of that bears on the texts it produces depends on what philosophy does with the world and with experience — on where a philosopher's starting points come from, and on what becomes of a philosophical claim once it is made. Pigliucci (2017) addresses both. He holds that philosophy is constrained by the world without investigating it as the natural sciences do:
> This means that the basic parameters that philosophers use as their inputs, the starting points of their philosophizing, their equivalent of axioms in mathematics and assumptions in logic (or rules in chess) are _empirical_ data about the world. This data comes from both everyday experience [...] and of course increasingly from the world of science itself. Philosophy, I maintain, is in the business of exploring the sort of conceptually evoked spaces that Smolin is talking about, where the evocation is the result of whatever starting assumptions are made by individual philosophers working within a particular field and, crucially, of the constraints that are imposed by our best understanding of how the world actually is. (2017, pp. 79–80)
_Evocation_ is a term Pigliucci takes from Smolin (Unger and Smolin 2015),[^4] for truths that are neither discovered, in the sense of corresponding to mind-independent states of affairs, nor invented, in the sense of being arbitrary constructs, and his example is chess: when the rules of a game are codified, a whole bundle of facts about it becomes demonstrable — objective facts, in that anyone who can demonstrate one demonstrates the same fact as anyone else — although chess did not exist before its rules were written down (Unger and Smolin 2015, p. 423; quoted at Pigliucci 2017, p. 78). Pigliucci's proposal is that philosophy ascertains evoked truths in this sense, with an addition that separates it from mathematics and chess alike: its starting points are constrained by how the world actually is. The addition is also what separates philosophy from fiction, on his account. A novelist's worlds are invented rather than evoked — nothing about them is rigid, since even the constraints the novelist adopts could have been otherwise — whereas philosophy "is in the business of doing empirically informed evoking, not inventing", so that its objects of study have rigid properties (2017, p. 80). A thought experiment is itself a case of such evoking: the philosopher sets up an imagined scenario but explores it "with an interest in figuring things out as far as this world is concerned" (2017, p. 80), so that what it evokes has the rigid properties Pigliucci means, and the philosophical work proceeds within the structure it opens.
Philosophy's starting points, then, are empirical, and a system that perceives nothing cannot reach them on its own. But the data Pigliucci describes comes from everyday experience and, increasingly, from science, and the scientific kind reaches working philosophers already articulated — already set down in language, available to be read rather than undergone. A philosopher of physics works from published results, not from having run the experiments. The same holds for everyday experience: what the discipline retains of it, it retains as the literature's accumulated descriptions of how things seem. For the worldly materials philosophy actually uses, written access is the profession's normal condition rather than a deficiency, and a corpus is such access. On this point the model stands where every philosopher already stands with respect to nearly all of the empirical data they use.
The other objection was that the model never checks what it produces against the world. But a philosophical claim is not the kind of thing that gets checked that way. Here the elevator and Mary part company. The equivalence principle, once Einstein had it, faced a tribunal of measurement: the experiments might have gone against it, and then it would have been dropped. Mary's case faces no such tribunal. The question Mary raises — whether she learns something new on first seeing red — is a question about what follows within the scenario Jackson has set up, and that is settled in the way any question about chess is settled: by working out what the set-up commits us to, something any competent party can do and none can decide by fiat. This is the only checking a philosophical thesis gets, and it happens in the literature, in the back-and-forth the previous section described. So a model's inability to run experiments costs it nothing a philosophical text needs: the testing that philosophy does is the working-out of what a scenario commits us to, and that is done on the page.
Take the knowledge argument itself. Nobody who debates it has been through what Mary goes through — released into colour after a lifetime of black and white — and the debate does not suffer for it. The participants have in front of them Jackson's description: a scenario set down in words, with a claim about what it is meant to show. Lewis's reply (1988) works on that description, changing what we should say follows from Mary's release while saying nothing about what her first sight of red is like from the inside. This is the usual way experience figures in philosophy. A philosopher need not have had an experience to argue about it: Nagel (1974) asks what it is like to be a bat without supposing he could find out, working instead from a description of the bat's situation. A model is no worse off here than these philosophers are. It has had no experiences of its own, but it has not needed them, because what its corpus supplies, and what the philosophy works on, is experience already put into words — the same descriptions that, in the previous section, carried displayed reasoning into its texts.
Merleau-Ponty noticed something about touching one hand with the other that was not, so far as we know, already written down anywhere. At any moment one hand is the toucher and the other the touched; the roles can switch, but they cannot both hold at once (Merleau-Ponty 1945 %%page%%). Suppose he came to this only by attending to his own body, to what the touching was like from the inside, and not from anything already in the literature. Then it is a starting point a model could not have reached on its own, since reaching it took a body and a first-person view that a model does not have. Once Merleau-Ponty has written the description down, a model can take it up and argue about it as well as anyone. It could not, though, have been the one to set it down first.
What a philosopher does with Merleau-Ponty's description, though, has nothing to do with who first arrived at it. The description earns its place by what can be drawn out of it — what it shows about the body, what follows once the asymmetry is granted — and that is there for any reader to work through, whether or not they could have come to the asymmetry on their own. First-person attention is not, in any case, the only way new descriptions come about. A description the literature does not yet contain can also be reached from the descriptions it does contain, by drawing out what they have not been taken to imply, or by putting two of them together as no one has — and a model can do this. Whether today's models in fact produce descriptions the literature lacks is the question of novelty, which Section 4 takes up.
A model can do the philosophy that turns on the world and on experience, save for the one point already granted: it could not be the first to set down a description that only first-person attention could yield. Its worldly starting points come to it already written down; the claims it draws are settled by working out what they commit us to, not by measurement against the world; and the experiences philosophy argues about reach it as descriptions, which it can work through as well as any reader. In the previous section, abduction entered philosophy as reasoning displayed in a text; here the world and experience enter it as descriptions set down in a text. What a model produces is read as any philosophy is read, and how it was produced settles nothing in advance.
[^2]: Zahavy, following Magnani, calls the process _manipulative abduction_: hypothesis generation through the manipulation of a model — here a simulated experience — rather than of symbols (Magnani et al. 2009; Zahavy 2026, §5). It is abduction in the previous section's sense: the equivalence principle is inferred as the best explanation of the simulated observation, the simulation supplying an explanandum that no search over existing text would have produced. What experience contributes, on this picture, is not the inference but its starting point.
[^3]: The materials need not be sensory: the feeling of understanding something is sometimes used to motivate the claim that thought itself has a phenomenology (Pitt 2004).
[^5]: Zahavy too calls this leap abduction, but the word picks out something other than it did in the previous section. There, with Floridi et al., abduction was the weighing of rival explanations, and the charge was that a model only mimics it; here it is the generation of new first principles from experience, and Zahavy's claim is that a model cannot make the move because it has had no experience to move from. The present challenge rests on that second claim, about experience, and not on any verdict about the weighing.
[^4]: _The Singular Universe and the Reality of Time_ is jointly authored, but its second part, which contains the discussion of evocation, was written by Smolin alone, as Pigliucci notes (2017, p. 77); we follow him in attributing the view to Smolin.
---
# 4. The Challenge from Observation
[this section needs the most work]
In this final section we address what we will call the _challenge from observation_. In Section 1 we argued that LLMs should not be disqualified from producing worthwhile philosophy tout court.
In Sections 2 and 3 we argued that, although LLMs neither perform abductive inference, nor have experience, nor are connected to the world, there is still reason to think that they are capable of producing text which exhibits good quality abduction, and works with articulated axioms about experience and the world.
The challenge from observation begins with an obvious question: if all of these arguments are correct, where is all the worthwhile LLM-written philosophy? If you ask an LLM the answer to the hard problem of consciousness, or the meaning of life,1 you will not receive _the correct answer_, but instead a competent but unopionated survey of the field if you are lucky, or a less accurate but equally bland survey if you are unlucky.
The observation is accurate, and it reports less than it seems to: it reports what models produce under one use — a bare question, put once, answered in one pass. How these systems are built explains why that use yields what it does. A model is first fitted to a vast general corpus and trained to continue text, so its response to a bare philosophical question is the likely continuation of such a question in writing at large, and the likely continuation of "what is the meaning of life?" in a general corpus is not an analytic tract. It is the sort of text that follows the question at large: a survey of views, a consoling generality, a joke. The model is then further shaped to converse as a helpful assistant, and the shaping presses the same way, since a person employed to be helpful to all comers would not answer the question with a tract either. The survey is not a ceiling the systems have hit; it is the likely continuation of exactly what was given them.
The use that generates the observation treats the model as an oracle: a system whose answers are its measure, so that asking is all the eliciting there is.2 The empirical record tells against the assumption. The survey of abductive benchmarks discussed in Section 2 runs every test with a single fixed instruction and scores the answer, while cataloguing, in the same pages, methods that alter what models produce — prompts that separate the stages of a task, pipelines in which an answer is criticised and revised over several passes (Salimi et al. 2026). What a model returns depends on what it is given, and the observation samples one point in that space, the bare question. It therefore cannot discriminate between the two hypotheses at issue — that the capacity defended in the preceding sections is absent, and that it has not been elicited. Both predict the observed record, and an argument against this paper needs the first; the observation supports it no better than the second.
The challenge has a natural escalation: if philosophy worth reading comes out of these systems only when a philosopher directs the process — supplies the framing, sets the constraints, presses for development — then the philosophy, it will be said, is the philosopher's. The model is an instrument in the production, as a typewriter is, and crediting it with the result is crediting the dummy with the ventriloquism. Section 1's challenge held that a model's text is not philosophy tout court; what stands here is narrower, that the philosophy in such a text is not the model's.
Whether the escalation succeeds depends on what prompting a model involves: what a prompt supplies, and what the model's continuation adds to it. Section 3's account of starting points says the first; Section 2's account of continuing text says the second.
A prompt articulates a starting point, as a thought experiment does. A prompt that sets out a position and the rivals it must beat stands to the model as Jackson's two paragraphs stand to the profession: a starting point handed over for development. What an articulated starting point does, on the account already in place, is evoke a structure with rigid properties — there are facts about what holds within it, demonstrable by anyone and chosen by no one, and they outrun whatever has been stated, just as the facts about chess outran the rules the moment the rules were written down. Most of what a starting point evokes, no one has ever said.
What the model contributes is the development, and the mechanics are the ones Section 2 drew from Wolfram: a model produces a reasonable continuation of the text it has been given, where what counts as reasonable is relative to the corpus it was fitted to (2023). A prompt is part of the text the model has been given. An articulated starting point therefore changes what there is to continue — the reasonable continuation of a stated position under stated constraints is not the reasonable continuation of a bare question — and the model makes use of what the prompt states in everything that follows: tell one of these systems something once, Wolfram observes, and it is used thereafter (2023). The continuation that results states consequences of the starting point that the starting point does not state. Section 2 said what it is for such a text to go well — the comparison it displays cites differences that tell between the positions, and would, if correct, give understanding — and whether a given continuation goes well is read off the continuation.
Nothing in this makes the development a transcription. An evoked structure contains more than any text states: the rules of chess settle every fact about chess, and do not settle which theorems get written down, in what order, or to what depth, so that two writers working from the same rules produce different books, both correct, neither dictated by the rules. The mechanics mirror the structure, since the same prompt, run twice, yields different continuations (Wolfram 2023). The starting point underdetermines the development, and the gap between them is where the model's contribution lies: were there one text the prompt fixed, the output would transcribe what the person had already settled, and the instrument description would be true.
The gap also leaves room for error. A development can state what does not hold in the evoked structure — a chess writer can publish a false theorem, a philosopher can misdraw the consequences of their own thought experiment, and a model can do both, along with its characteristic failure of stating fluently what nothing supports. The errors are found on the page. And an error is attributable only to a developer: no one blames the rules of chess for a false theorem, and no one's typewriter has ever made a mistake of content. Three contributions, then, and three owners: the articulated starting point is the person's; the structure it evokes, and the facts that hold there, are no one's; the text that develops them is the model's.
Much in the instrument picture is true. The person writes the prompt and the prompt is authored; the person chooses which continuations to pursue and when to stop; without the person, there is the survey. What the picture adds to these truths is a description of the model — a device, like the typewriter, that fixes only what its user has already settled — and the description is what the account above denies. Every word of the novel was the author's before the typewriter touched it; the consequences a model's text states were nobody's before the text stated them. What the user of a typewriter settles is the text; what the writer of a prompt settles is a starting point.
The account invites an obvious enrichment of the prompt. State the position, name the rivals, list the objections and the lines along which they are to be met, and at some point, it will be said, the prompt contains the philosophy and the model is expanding what the person wrote — so that where a model's output is good, one should suspect a prompt rich enough to have done the work. But enriching a prompt enlarges the starting point without converting it into the development. A game with more rules is a bigger game, not a book of its theorems, and however much the prompt states, the consequences the output draws were not among the statements. There is a genuine limiting case — a prompt that states the comparison and the verdict, so that the continuation only rephrases — and it is identified the way everything in this paper is identified: set the output against the prompt and ask what the text states that the prompt did not. A text that states nothing beyond its prompt is a paraphrase, and owed to the person; a text that states what the prompt left unstated is a development, and the unstated part is not the person's. Which of the two a given output is, is settled by reading them together.
Even if all this is granted, we might still ask whether such a text can do more than handle well the positions a literature already contains and make a distinction that literature lacks, and so be creative in the stronger, public sense in which a human philosophical text is creative. We want to be careful here, because this is quite different from the modest sense in which an output is novel only relative to its prompt; what is at issue is the stronger claim of saying something the literature had not yet said. We do not need to settle this question here. It would be settled as the rest has been, by setting the output not only against the prompt but against the literature, and asking what it says that the literature had not. This is not meant to settle the matter either way, but it does give us a reason to leave the question open, for further work.
We can now return to the challenge from Section 1, where the ordinary blandness of what these systems produce when given a bare question was taken to show that they have nothing to contribute to philosophy. The outputs to such questions do tend towards the empty and the thin; but that bears only on whether bare questions are good tests of philosophical capacity, since what these systems do is continue the context they are given, and a context with no shape of argument can only draw from them an output with no shape of argument. This is no reason to think that a context which hands the system a position, its rivals, and the pressures bearing on each of them cannot draw a development of its own; and whether such a development is worth reading is settled not by looking at the system but by reading the continuation first against the prompt that occasioned it and then against the literature it means to add to. To sum up, these systems are not oracles; they are continuation systems, and philosophy worth reading needs a dialectical context which a prompt can supply without thereby fixing the development, so that the philosophical standing of any output depends on what the continuation itself adds.
## Footnotes
1. While preparing this paper we asked GPT-5.5 for a detailed overview of the positions an analytic philosopher might take on the meaning of life. What came back was a competent, hedged survey of the field; what did not come back was an argument for any position in it. %%add date of test%% ↩
2. That these systems are mischaracterised as oracles — with the corollary that no benchmark of single-pass answers should be expected to probe the upper limits of what they can produce — has been argued from inside the practitioner literature (Janus 2022). ↩
---
# References
# References
Frankish, K. (2024). What are large language models doing? In A. Strasser (Ed.), *Anna's AI Anthology: How to live with smart machines?* (pp. 55–78). Xenomoi.