# Generating Philosophy with Artificial Intelligence
> [!info] Claude polished draft, 10 June 2026
> Full-paper polish built from the scene files as of today, with Section 4 composed from the working notes and the project plan. Scene files untouched. All quotations verified against the extracted sources; corrections logged in the session summary.
## Abstract
Large language models now contribute results to mathematics and the natural sciences. We argue that philosophy should expect the same: current-generation LLMs are capable of producing philosophical texts that are worth reading. We defend this claim against four challenges. The challenge from authorship holds that a text can be philosophy only if a philosopher's activity lies behind it; we argue that the performance view of art on which this challenge is modelled does not transpose to philosophy. The challenge from abduction holds that LLMs cannot perform the inference to the best explanation on which philosophical theorising runs; we argue that the abduction philosophy demands is a graded, exemplar-carried standard of just the kind such systems acquire from text. The challenge from phenomenology holds that philosophy drawing on experience requires a conscious subject; we argue that phenomenology enters philosophy in articulated form, and articulations are public. The challenge from tools holds that the philosophy in any worthwhile output belongs to the prompter; we argue that a prompt posits a starting point whose consequences the model, not the prompter, develops.
## 0. Introduction
In April 2026, an amateur mathematician entered the statement of an open problem — a conjecture of Erdős, Sárközy and Szemerédi concerning primitive sets, number 1196 in the catalogue of Erdős's open problems — into GPT-5.4 Pro as a single prompt. Roughly eighty minutes later the model returned a proof, built around a Markov-chain technique that the mathematicians who had worked on the problem had not tried. The proof was checked, has since been formalised in Lean, and the problem is now listed as solved, credited jointly to the model and its prompter; Terence Tao's assessment was that the argument contributes to the anatomy of integers well beyond the particular problem it settles.[^intro1] Nor is mathematics an isolated case. AI systems have had success in domains where the value of an output is not exhausted by its superficial fluency: in theoretical physics, where GPT-5.2 conjectured a closed-form expression for a class of gluon tree amplitudes long presumed to vanish, which the authors then proved and verified (Guevara et al. 2026); in biomedicine (Gottweis et al. 2025); and in materials science (Zeni et al. 2025). In this paper we argue that we should expect similar success in philosophy.
Specifically, we argue that current-generation LLMs are capable of producing philosophical texts that are *worth reading*. We frame the claim this way so as not to begin by legislating what counts as *good* philosophy. Instead, we appeal to a distinction that anyone reading this text will recognise. You have read texts that were worth reading, and you have read texts that were not. As you begin this article, you presumably hope that it belongs to the first kind — that the time spent reading it will not be wasted. When you write a philosophical text yourself, you aim to make it worth your readers' while, and whether the journal you send it to accepts it depends on whether the referees agree that you have succeeded.
Two clarifications are needed. First, a text's being worth reading is not the same as its being correct. A text can repay attention even if one rejects its conclusion: it may sharpen a distinction or answer an objection in a way that changes the dialectical situation. Second, the minimal unit we are concerned with is not the bare conclusion of an argument but the argument itself. If an LLM output consists only in a pronouncement on some philosophical topic — "direct realism is correct", "we should be utilitarians" — it is hard to see why it would be worth reading in and of itself, for the same reason that a bare pronouncement by a human philosopher would not be worth reading.[^intro2]
The paper proceeds as follows. Section 1 addresses the challenge from authorship: the claim that an LLM output cannot be philosophy worth reading because no philosopher lies behind it. Section 2 turns to abduction, and argues that the absence of inference to the best explanation in the producer does not prevent abductive structure, of the kind philosophy actually demands, from appearing in the product. Section 3 considers phenomenology, and argues that the lack of conscious experience does not prevent LLMs from producing philosophy grounded in phenomenology, because phenomenology enters philosophy in articulated form. Section 4 confronts the challenge that remains once the first three have been answered: that whenever a model's output is worth reading, the philosophy in it is the work of the prompter, and the model no more produces philosophy than a typewriter does.
[^intro1]: The problem is listed, with the model's proof and its subsequent Lean formalisation, at erdosproblems.com (Problem #1196, solved by GPT-5.4 Pro prompted by Liam Price, April 2026). A human-authored writeup by Alexeev, Barreto, Li, Lichtman, Price, Shah, Tang, and Tao is in circulation.
[^intro2]: A pronouncement can be worth reading as testimony, if its source is known to be reliable. Our concern is with what is worth reading as philosophy, where reading improves the reader's position through the argument presented rather than through trust in its source.
## 1. The Challenge from Authorship
In this section we address what we might call the *challenge from authorship*: the idea that philosophy is something only persons, or at least minds, can produce. The view has not, to our knowledge, been defended explicitly in just this form, but it gives shape to an intuition many philosophers will recognise: philosophy is a person-only domain. An imperfect comparison is with art. One might deny that an image generated by an AI system, at least in the familiar prompt-and-output cases, is an artwork, on the ground that no artist exercises the relevant kind of intentional control over its production.[^auth1] One might think, for similar reasons, that philosophy can only be done by people: no text produced by an LLM can be a work of philosophy, because no philosopher lies behind it.[^auth2]
Consider also that, like art, the study of philosophy is organised around individuals. Undergraduates take courses on Kant's ethics or Lewis's metaphysics, and at more advanced levels one finds specialists in, and conferences devoted to, the work of particular philosophers. The sciences are not, as a rule, organised this way. There are historians of science who study Newton's manuscripts, but physicists do not specialise in Newton: the mechanics he discovered detached itself from his texts long ago, and is taught and used without them. Biology students learn the structure of DNA from textbook diagrams, not from Watson and Crick's announcement of it. Philosophy's continuing attachment to named individuals suggests that here, as in art, who did the work belongs to our conception of what the work is.
We can make the intuition more precise by pursuing the comparison with art, and in particular by asking how far Davies's *performance* theory of art can be transposed to philosophy. On Davies's view, an artwork is not the object the artist leaves behind but the activity that produced it:
> the work – what the artist achieves – is the process eventuating in that product. Works themselves are neither structures nor objects *simpliciter*, nor are they contextualized structures or objects. They are, rather, intentionally guided generative performances that eventuate in contextualized structures or objects. (2004, p. 98)
When a painter paints a picture, then, the canvas is what we attend to, but it is not the work. The work is the painter's intentionally guided activity in producing the canvas; the canvas is the "focus of our appreciative interest in the work" (2004, p. 151). This gives provenance a deeper role than it has on views that identify the artwork with a product plus contextual properties: facts about how the object came into being do not merely inform our appreciation of the work but help determine which work there is and what can properly be appreciated in it. If the same model were transposed to philosophy, an LLM text would fail to be philosophy not because it was badly argued, but because the relevant kind of philosophical performance never took place.
Consider what is involved in attending to a Vermeer. The surface before us is taken as the outcome of a particular painter's activity, in a particular historical context, with particular resources and limitations, and Davies presses the point through cases of perceptual indiscernibility of the kind Danto (1981) made familiar (Davies 2004, pp. 63–4). Imagine a canvas that emerges by accident from a faulty washing machine and happens to look just like a Rembrandt: there is a Rembrandt-like surface, but no artistic performance of the relevant kind. Or take the forged Vermeers that van Meegeren sold in the 1930s and 1940s (Davies 2004, pp. 14–15):[^auth3] there an artistic performance took place, but not the one the work was taken to make available. In neither case does provenance merely add information. It changes what we take the work to be and what kind of achievement we take ourselves to be appreciating. If Davies is right, the surface does not by itself settle the work.
Here is the analogous proposal for philosophy. A philosophical text is not itself the philosophical work. The text is the product of the thinking, writing, and philosophising done by a person or group of persons over time; it is the focus of our attention, but only as a way of gaining access to the philosophical performance that brought it into being. On this proposal, a philosophical work is not identical with the sequence of sentences on the page: the text gives argumentative form to someone's activity of thinking through a problem, and reading it is a way of engaging with that activity. The authorship challenge is therefore not a worry about missing biography. It is the stronger claim that if no one has done the relevant philosophising, there is no philosophical work to which the text could give access.
We do not think the transposition should be accepted. Davies has a reason for moving from product to performance in the case of art: production history can affect which artwork we are dealing with and what is available for appreciation in it. The philosophical case lacks the corresponding pressure. If two texts contain the same argument, including the same inferential moves, then the same considerations count for and against each, and their philosophical merit does not vary with the route by which the words came to be written. When we assess a philosophical paper we ask whether the text does philosophical work, and these questions do not require us to look behind the text to the philosopher's activity: the grounds for the judgement lie in the argument as presented, not in the history of its production. In the art case, production history can change what the work is; in the philosophical case it changes, at most, what we think about the producer, or about the process by which the text came about.
The organisation of analytic philosophy reflects this. Journals strip author information from submissions before sending them to referees, and they do so because facts about authorship are treated as possible sources of distortion in the assessment of the writing. The point is not that blind review always succeeds, nor that philosophical practice takes no interest in authors. The point is narrower: in this evaluative context — the one in which the discipline decides what enters its literature — the paper is supposed to be assessed by attending to what it says, not by reconstructing the circumstances under which it was written.
A point from Dellsén et al. (2024) helps to articulate the same thought, although their concern is philosophical progress rather than LLM authorship. On their account, philosophical progress "is a matter of putting people in a position to increase their understanding", and in practice this "will normally consist in making publicly available various philosophical ideas, such as arguments, theories, distinctions, and thought experiments" (2024, p. 679). Progress, as they put it, calls for a "for-whom rather than a by-whom conception": what makes the difference is the cognitive position of those for whom progress is made, not facts about those by whom it is made (2024, p. 679). Their concern is with what progress consists in, and ours is with what a contribution is, but the lesson carries over. If philosophy makes its contribution by putting readers in a better position — supplying them with arguments they can take up, distinctions they can deploy, objections they must now answer — then the contribution is made by what is publicly available, and we should resist locating the philosophical work somewhere behind the public text, in the process from which the text resulted. The text is not a dispensable trace of philosophy done elsewhere; it is where the contribution becomes available at all.
The challenge from authorship is, at bottom, a constitutive challenge. It treats the philosopher's activity not merely as what causes a philosophical work to exist but as part of what the work is, so that even a text indiscernible from a philosophical paper would fail to be philosophy if no philosophical activity lay behind it. This we have argued should be rejected. If a novel philosophical argument were spelled out by wind-blown sand, or by a very faulty washing machine, the absence of an author would not, in and of itself, prevent the resulting text from being worth reading.
Rejecting the constitutive challenge does not end the matter, because the remaining objections are of a different kind. They concern not what philosophy is but what LLMs can do. Compare: if a young child were, improbably, to produce a text containing a novel philosophical argument, the text would count as philosophy; what actually prevents children from producing philosophy is not their standing as authors but their capacities. The challenges that occupy the next two sections treat LLMs the same way — not as the wrong kind of producer, but as producers lacking capacities that the production of worthwhile philosophy requires. Section 2 takes up the claim that LLMs cannot perform the inference to the best explanation on which philosophical theorising largely runs. Section 3 takes up the claim that some philosophical texts require phenomenal materials available only to conscious subjects.
[^auth1]: For discussion of whether AI-generated images can be artworks, see Anscomb (2022), Hertzmann (2018), and Wojtkiewicz (2023). In practice it is hard to find images whose production involved no human influence at any stage: prompts are written, outputs selected, results refined. The cleaner question, as in the philosophical case, is who or what is responsible for the properties that make the result valuable.
[^auth2]: Denying that AI-generated images are artworks does not amount to denying that they can be beautiful. The parallel distinction — between a text's status and its being worth reading — runs through this paper, and we return to image generation in Section 4.
[^auth3]: Han van Meegeren, the Dutch painter whose forged Vermeers deceived experts and collectors until 1945, when the sale of one to Hermann Göring brought a charge of collaboration and van Meegeren confessed to forgery in his own defence.
## 2. The Challenge from Abduction
Even if no fact about authorship rules LLM texts out of philosophy, facts about capacity might rule them out of philosophy worth reading. One might accept the argument of the previous section and still hold that LLMs, at least in their current form, lack particular capacities that producing worthwhile philosophy requires. This section examines the charge that LLMs cannot perform abductive inference; the next examines the charge that they lack phenomenology.
Abduction, or inference to the best explanation, is reasoning from a body of evidence to the hypothesis that would best explain it. Abduction differs from deduction in that the evidence does not settle which explanation is correct. In a deductive argument the premises fix the conclusion: if all men are mortal and Socrates is a man, then Socrates is mortal, and there is no wriggle room. Now imagine walking into your kitchen and finding the floor wet. What has happened? The wet floor does not determine the answer in the way the two premises gave you Socrates' mortality: a burst pipe would have left the floor wet, and so would a spilled bucket. But given that the window is open, the water is under the window, and it rained last night, rain coming through the window seems the most plausible answer.
To reason in this way — deciding what best explains a set of facts — is common in everyday life and in the sciences alike. A scientist chooses one theory over another when it explains the same data more simply: Copernicus's model of the solar system was preferred to Ptolemy's because it explained the observed planetary motions without the elaborate epicycles the older model required. Williamson holds that philosophy is continuous with the sciences in its method: a philosophical theory, like a scientific one, earns acceptance by explaining the relevant data better than its rivals — more fully, and more simply — rather than by being proved (2007; 2021, pp. 351–4), since deductive argument only passes the question back to its premises, which must themselves be chosen somehow (2021, p. 364). The same methodology has been defended for metaphysics in particular, where rival theories are weighed by these explanatory virtues (Sider 2011; Paul 2012), and it suits a conception of the discipline's aim that Sellars made famous: to understand "how things in the broadest possible sense of the term hang together in the broadest possible sense of the term" (1962, p. 35). We shall assume, in what follows, that producing philosophy worth reading depends, in large part, on abduction.[^abd1]
Floridi and colleagues hold that large language models do not perform abductive inference. They describe what such models do instead as zeroth-order abduction:
> LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025, p. 9)
LLMs, on this view, "have a stochastic core and an abductive appearance" (2025, p. 2), and the appearance has a definite source: the training texts were themselves produced by human reasoning, so reproducing their patterns reproduces explanatory form. An LLM can produce a plausible hypothesis for a given body of observations, and, when the candidate explanations are explicitly provided, it can select the most suitable of them "remarkably well, often at a near-human level" (Floridi et al. 2025, p. 4, citing Bhagavatula et al. 2020 and Balepur et al. 2024). What it cannot do is test a hypothesis. Inference here divides in Reichenbach's (1938) way: abduction supplies a candidate in the context of discovery, the candidate is then assessed against new data in the context of justification, and LLMs perform only the first part (2025, pp. 5–6). The model's words connect only to other words, never to the things the words are about — the predicament Harnad (1990) named the symbol grounding problem (Floridi et al. 2025, p. 8) — so an output may be optimal by every criterion of explanatory goodness and still be false, with nothing in the model able to find out (2025, p. 19). The proper use of such systems is then as brainstorming machines, churning out candidate hypotheses: "It then becomes the human's task to carry out the justification phase" (2025, p. 11).
The case can be made stronger than its authors make it. A sizeable literature has since set about measuring abduction in language models, and its results appear to confirm the diagnosis: on aggregate, models remain markedly weaker on abductive benchmarks than on deductive ones; where a long mystery story supplies the evidence and the task is to identify the culprit, they fail more often than human readers; where every hypothesis consistent with a set of facts must be produced rather than recognised among options, performance collapses; and where the same diagnostic task is posed with and without candidate answers supplied, performance drops sharply once the diagnosis must be produced rather than ranked (Salimi et al. 2026).[^abd2] On this reading the record shows exactly what the mechanism story predicts — the conceded part of abduction, choosing among candidates somebody else supplies, intact, and the denied part missing. So strengthened, the challenge holds that LLMs lack the capacity for abduction, and that the lack shows up wherever it is measured.
The tasks on which the models fail have a common shape. In each there is a fact of the matter laid down in advance — a culprit, a diagnosis, a missing premise — and the task is to recover it; an answer counts as correct when it matches that hidden fact, and as nothing otherwise. Whether philosophical abduction asks for anything of this kind depends on what makes one explanation better than another, and 'best', in Lipton's treatment, can be read in two ways: the best explanation may be the likeliest, the most warranted by the evidence, or the loveliest, the one that would, if correct, provide the most understanding — "Likeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). The readings come apart. That smoking opium puts people to sleep because of its dormative powers is, in Lipton's example, about as likely as an explanation gets — it commits to little beyond the fact it explains — and yet it is "the very model of an unlovely explanation" (2004, pp. 59–60). A benchmark that scores the recovery of a concealed fact is scoring likeliness, with the world, or the puzzle's setter, holding the answer. When a philosophical theory is preferred because it would explain the relevant data better than its rivals, and more simply, what is judged is loveliness, and there is no answer sheet against which the judgement could be checked.
Nothing in this standard requires the explanation to be true. Newtonian mechanics "remains as lovely an explanation of the old data as it ever was", however unlikely the evidence for relativity has since made it (Lipton 2004, p. 60), and philosophy keeps its own counterpart: the positions a student is set to read contradict one another, so most of them are false, and their being false has never been treated as a reason to stop reading them — which is precisely the position of an output that satisfies every explanatory criterion and is nevertheless false. Nor is a verdict pending that would change this. Copernicus's hypothesis had observations still to come; a philosophical theory is not waiting on anything of the kind. The knowledge it must explain is knowledge already gained — by the sciences, by common sense, by philosophy itself — articulated and on the table when the theory is proposed (Williamson 2021, p. 356), and the justification phase, such as it is, consists in argument over that shared stock: an opponent points to something the theory cannot accommodate, or produces a rival that explains it better, much as mathematics settles its first principles abductively and without experiment (Williamson 2021, pp. 357–9; cf. Pigliucci 2017).[^abd4]
What the standard does demand can be read in the text that tries to meet it. Suppose a text argues that rain through the open window, rather than a burst pipe, best explains the wet kitchen floor, on the ground that the water lies under the window. That ground discriminates: the two hypotheses differ over where the water should be, so citing its location favours one and counts against the other. A text that instead offered the ground that the floor is wet would fail — both hypotheses lead us to expect a wet floor — while exhibiting, sentence for sentence, the typical phrasing and structure of an explanatory comparison. To this extent the challenge is right: explanatory structure can be had without any discrimination between rivals. The capacity in dispute is the capacity to produce texts of the first kind, reliably and not by accident.
In the kitchen case anyone can tell that the location of the water discriminates and that the wetness of the floor does not; between philosophical theories the same judgement calls for training, and not because the rules are taught. Lipton is frank that "the weakness of our grasp on what makes one explanation lovelier than another is discouraging" (2004, p. 61): there are no universally shared mechanical rules that generate a unique hypothesis from given data (2004, p. 83), and what counts as a lovely explanation is determined in part by previous explanations that serve an exemplary function, and by more general styles of reasoning (2004, p. 139). The competence involved is, as Lipton himself observes, like our competence with grammaticality — it is easy to distinguish grammatical from ungrammatical strings in one's native tongue, hard to describe the principles underlying the judgement, and the principles at work are "determined by exemplars rather than by rules" (2004, pp. 1, 6–7).[^abd3] Whether a regularity of this kind — unstated, graded, carried in examples — can be picked up from text is the question on which the challenge now turns.
The exemplars that carry the standard are themselves text. Earlier explanations and styles of reasoning survive in one form, writing, and the corpus on which an LLM is trained contains that writing: the literatures of the sciences and of philosophy, together with the far larger record of ordinary argument, weak as well as strong. However natural it is to think that a system which only continues text could not pick up a judgement from it, regularities of just this kind are what such systems demonstrably acquire. A model is never given the rules of English grammar — in training it implicitly picks them up, and its text comes out grammatical; no overall theory specifies which strings of English are meaningful, and its output stays meaningful all the same (Wolfram 2023): regularities for which nobody possesses complete rules are installed by fitting the text that contains them. The challenge itself traces the model's explanatory appearance to this very route — LLMs, Floridi and colleagues write, "have effectively absorbed patterns of human abductive reasoning as expressed in writing" (2025, p. 9) — so the position needed is that grammar and meaningfulness can be acquired from text while the standards of explanation cannot, although all three are regularities of the same kind, unstated, graded, and carried in examples; nothing in the mechanism story supplies a reason for the difference. And what steers abductive theory choice looks like a regularity of that kind too. Williamson, asking why aesthetic criteria such as elegance should bear on theory choice at all, observes that in theorem-proving one is, without a strong aesthetic sense, "lost, directionless", and that such a sense "is surely connected to a capacity for abstract pattern recognition" (2021, p. 367). There is reason, then, to expect a model trained on this record to produce text whose explanations are good, by the same route by which its text comes to be grammatical and meaningful.
Where these systems fail is just as regular. A small transformer trained on sequences of matched brackets learns the language for ordinary cases and breaks down where success requires explicitly counting the brackets, while networks of the same kind, shown a smudged digit, settle what it is at a glance, with no criterion they could state (Wolfram 2023). The line falls between graded judgement on open-textured material, which fitting to examples yields, and exact, step-by-step bookkeeping, which it does not; weighing rival explanations belongs on the first side, the long exact derivation on the second. And the record assembled above breaks along the same line: the tasks at or near human level ask for comparative judgement over supplied or short material, and the tasks that collapse ask for exhaustive enumeration, or for the recovery of a single concealed fact, scored right or wrong against it. Read with Lipton's distinction in hand, the record confirms this placement rather than the challenge: what fails is abduction with an answer sheet, which philosophy does not ask for, and what succeeds is comparative judgement of the kind it does.
None of this is at odds with the description of the mechanism: a true description of a process at one level does not displace true descriptions at another, as Lipton notes of the parallel deflation of explanationism by Bayesianism — like arguing that thinking about technique cannot help one's squash game because the ball's motion is governed by mechanics (2004, p. 108). And what is acquired is a capacity, exercised well or badly on particular occasions: an output can carry the apparatus of a comparison whose cited difference favours neither side, and which kind of text it is can only be settled by reading it, as for every philosophical text.
Half of the challenge has not yet been touched: what was conceded was selection, and in the record generation falls short of it where the comparison is clearest. The denial of unaided generation has independent support — systematic theorising, on Williamson's picture, requires introducing new distinctions at a more abstract level not given in the data (2021, p. 353) — and a sharp current statement: what LLMs lack, Zahavy argues, is "the generation of novel explanatory hypotheses" (2026, p. 1), partly because generating hypotheses begins in sense experience, a claim about the materials of theorising that the next section takes up, and partly because such a system lacks "the external grounding to validate" the novel candidates it produces (2026, §4), which can be answered here. Producing candidates only from what one has already encountered disqualifies nobody: we rank only the potential explanations that have been thought of, and the actual explanation is sometimes one that nobody has thought of (Williamson 2021, p. 355) — why suppose that any of the explanations we happen to have thought of is true (Lipton 2004, p. 70)? Inquiry, for everyone, is conducted from inside whatever stock of candidates history has supplied. Nor is producing them a lottery: against Hempel's conclusion that hypotheses are "happy guesses" (1966, p. 15), Lipton argues that most hypotheses consistent with the data are non-starters, and that the contrastive structure of the evidence — a fact with a foil — sharply constrains which candidates could explain it at all (2004, pp. 82–3). Generation is guided by the same considerations as selection, and learned from the same exemplars.
A genuinely new distinction, finally, is new in relation to the literature: new when the literature lacked it, whoever first set it down, and a line drawn where none had been marked is not a recombination of the lines already there. Even the discipline's transformations mostly began as moves of recognisable kinds — apparatus imported from a neighbouring field, an assumption everyone had treated as binding called into question — and became transformations through what the discipline went on to do with them. The demand that remains, that a producer be able to tell which of its own novel candidates are apt, asks for an answer sheet that exists for no one: the written record preserves questionings that succeeded and questionings that failed, verificationism as thoroughly as rigid designation, with nothing in the producing of them marking the one kind off from the other. Which succeed is settled by the discipline's subsequent work, and a model's candidates enter that process on the same terms as anyone's; whose entries they are — the model's, or the prompter's — is the question of Section 4.
To sum up, the challenge asked of philosophy a kind of abduction it does not practise. What the measurements show failing is abduction graded against an answer sheet, and philosophy grades its explanations against none; what it does demand — graded comparative judgement, carried in exemplars, exercised in the reading — is a regularity of the kind such systems acquire from text, so there is reason to expect the capacity, though no guarantee of any particular output. What experience contributes to philosophy's materials comes next.
[^abd1]: Not everyone accepts that the explanatory virtues carry the same weight in philosophy as in the sciences: see Bueno and Shalkowski (2020) and Thomasson (2015).
[^abd2]: As of mid-2026: on the standard two-choice selection benchmarks the strongest models reach 87.2–88.0 against human averages of 91.4 and 92.0; on long narrative mysteries the best model scores 42.9 against an average human solve rate of 47, with the best human solvers above 80; a tightly constrained formal task sits at ceiling (99.6) while the task of producing every admissible hypothesis, scored against the exact set, collapses (21.5); and ranked diagnosis degrades sharply when candidates are not supplied (Salimi et al. 2026, Tables 3–4). Aggregating across task suites, the survey reports a median accuracy of 79.96 on deductive tasks against 42.50 on abductive ones. The field is organised throughout by Lipton's two stages of hypothesis generation and hypothesis selection.
[^abd3]: Loveliness inherits what Lipton calls Hungerford's objection: beauty is in the eye of the beholder, and explanatory loveliness may be "too subjective and interest relative" to do epistemic work (2004, p. 70). His reply is not that judgements of loveliness do not vary, but that warranted inference is itself audience relative, since it depends on available evidence, and we have been given no reason to suppose that the one relativity is more extreme than the other (2004, pp. 141–4). Nothing stronger is needed here, where the standard is used to assess texts rather than to underwrite a theory of inductive warrant.
[^abd4]: Philosophy does increasingly draw on experimental results, and Williamson's own examples include the psychology of perception; but the experiments are executed by the relevant scientists, and their results enter philosophical argument once articulated as findings (2021, p. 357).
## 3. The Challenge from Phenomenology
A further capacity worry concerns phenomenology. Few would say that current LLMs are conscious, and we shall assume here that they are not. Some philosophy, however, interrogates what it is like to see red (Harman 1990), to feel anger (Goldie 2000), or to have an intuition take hold (Chudnoff 2011). If a system has never undergone an experience of any kind, it is natural to think that philosophy of this kind is closed to it. The worry does not touch everything — much philosophy of language and modal metaphysics proceeds without leaning on the phenomenology of any particular experience — but the region it touches is large, and it contains some of the discipline's set pieces.
The sharpest recent statement of the worry comes from Zahavy (2026), whose denial that LLMs can generate novel explanatory hypotheses we met in Section 2. What grounds that denial is a claim about where new hypotheses come from: on Zahavy's picture, the deep theoretical breakthroughs begin with a jump from sense experience to axioms, and the jump is made by manipulating experience itself. His example is Einstein's:
> Einstein's variation required inventing new axioms based on a physical intuition that did not yet exist in the mathematics. He envisioned a physicist inside an elevator being uniformly accelerated through deep space. Inside this enclosure, the sensory experience reveals a specific pattern: when objects are released, the floor rushes up to meet them. To the physicist, the objects appear to fall with identical acceleration, regardless of composition. Thus, the simulation here was not a permutation of symbols, but a manipulation of perceptual experience. (Zahavy 2026, §5)
The thinker imagines a set of circumstances and attends to what would be experienced within it — in Einstein's case, that all objects inside the elevator would appear to fall with identical acceleration.[^phen1] That observation becomes the new axiom: a starting point arrived at through experiential simulation rather than formal derivation, from which further reasoning proceeds. If thinking of this kind depends on simulated experience, it would seem to be out of reach for LLMs. They can supply descriptions of weightlessness and of elevators, but they have never felt the sensation of an elevator descending, let alone of weightlessness.
Philosophy has its own experience-based thought experiments, and the worry transfers. Jackson's Mary — the scientist who knows every physical fact about colour vision but has never seen colour, and who is finally shown a red tomato — turns on what it is like to see colour (Jackson 1982; 1986). Like the elevator, the case appears to supply an experiential axiom from which further reasoning proceeds, and so the same conclusion beckons: thought experiments of this kind seem to require precisely what LLMs do not have.[^phen2]
Whether the conclusion follows depends on what philosophy's starting points are, and here Zahavy himself marks a boundary. His proposal, he writes, "is specifically tailored to the physical sciences, where the object of study is external material reality. In abstract domains such as Mathematics or Computer Science, the Sense Experience (E) may be grounded in high-dimensional topology or have other goals such as generality or minimality" (2026, §6). The jump may be needed everywhere, but what is jumped from is fixed by the ontology of the discipline: "for physics, the substrate is the world; for mathematics, it is the abstract landscape of formal systems" (2026, §6). The question for philosophy is then what its substrate is: what its starting points are, and in what form they have to be available to whoever, or whatever, would philosophise from them.
Pigliucci offers an account on which philosophy is constrained by the world without aiming at it in the way the natural sciences do:
> This means that the basic parameters that philosophers use as their inputs, the starting points of their philosophizing, their equivalent of axioms in mathematics and assumptions in logic (or rules in chess) are *empirical* data about the world. This data comes from both everyday experience [...] and of course increasingly from the world of science itself. (2017)
What those starting points do, once fixed, Pigliucci explains through Smolin's notion of *evocation*. Some structures neither pre-exist our positing them nor depend on our choices once posited. Smolin's example is chess:
> When a game like chess is invented a whole bundle of facts become demonstrable, some of which indeed are theorems that become provable through straightforward mathematical reasoning. As we do not believe in timeless Platonic realities, we do not want to say that chess always existed — in our view of the world, chess came into existence at the moment the rules were codified. (Unger and Smolin 2014, p. 423, quoted in Pigliucci 2017)
Whoever writes the rules of a game chooses the rules; nobody chooses what follows from them. Once the rules of chess are codified, the facts about chess are objective and demonstrable — anyone who can prove one can show it to anyone else — even though chess did not exist before its rules were written down. Positing a starting point, in this register, evokes a structure with rigid properties: a space of consequences that can be explored but not decided. Pigliucci's claim is that philosophy operates in just this register, with one difference from chess and from mathematics: the starting points of philosophising are not freely stipulated but constrained empirically, by how the world actually is, including by what experience is like. This is also what separates philosophy from fiction. "Philosophy, I maintain, is in the business of doing empirically informed evoking, not inventing, which means that its objects of study have rigid properties" (2017); a novelist's imagined worlds are bound by no comparable constraint, since even the constraints a novelist adopts could have been otherwise. And the picture covers thought experiments without strain: when philosophers explore imagined scenarios, they do so "with an interest in figuring things out as far as *this* world is concerned" (2017). A thought experiment articulates an axiom — an experiential or empirical starting point — and the philosophical work proceeds within the landscape that axiom evokes.
Both the elevator and Mary's room are evocations in Pigliucci's sense: each posits an experiential axiom and develops what follows. What differs is what the evocation is for. Einstein's elevator was the means to an external test: it yielded the equivalence principle, and the principle's truth was settled not within the evoked structure but by subsequent observation — the measurements came years after the thought experiment had done its work. Mary's room awaits no observation. The philosophical question — whether Mary learns something on seeing the tomato, and what this would show about the completeness of physical knowledge — is a question about what the evoked landscape contains, and answering it is a matter of exploring that landscape: drawing out consequences, testing whether a given reply changes them, finding where the rigid structure resists. In physics the evoked structure is an instrument pointed at the world; in philosophy it is the object of inquiry itself.
It will help to be exact about the form in which experience figures in such cases. Call phenomenology *articulated* when what an experience is like has been put into words and thereby made publicly available — Jackson's two paragraphs describing Mary's situation are an articulation in this sense. No competent discussant of the knowledge argument has personally undergone Mary's transition, and none needs to: once the case is articulated, work on it is work on the articulation. The replies to Jackson press on the articulated structure, not on any discussant's experience — Lewis's reply, for instance, modifies what is taken to follow from Mary's situation, not what her situation is taken to be like from the inside (Lewis 1990). And this is the form in which phenomenology enters philosophy generally. Philosophers work on the experiences of the blind, and on the experiences of non-human animals, without first-person access to either, by working on the articulations the literature has accumulated. The bearing on LLMs is direct. A model has no raw phenomenology of its own; but no text corpus contains raw phenomenology either. What a corpus contains is articulated phenomenology, and it is in articulated form that phenomenology becomes usable in philosophical argument.
A sharper version of the worry concerns phenomenological *discovery* rather than phenomenological use. When one fingertip touches another, one finger plays the role of toucher and the other of touched; the roles can reverse, but not simultaneously — at any instant the body is split between touching and touched (Merleau-Ponty 1945/2012). Suppose the toucher–touched asymmetry was first identified by Merleau-Ponty himself, through sustained attention to his own embodied experience. Then there was a time at which this phenomenological axiom was available to no corpus, because nobody had yet articulated it, and an LLM trained before that moment could not have produced it for itself: the system lacks the body, and the experience, that the discovery required. Nothing in this, however, prevents an LLM from working philosophically on the description once articulated — from exploring the landscape the asymmetry evokes, as Merleau-Ponty's readers have done ever since.
What survives of the challenge is therefore a narrower asymmetry than the one it began with. Some phenomenological articulations originate in first-person attention, and an LLM has no experience to attend to. But first-person attention is one route to an articulation; it is not what gives an articulation philosophical use. What makes an articulation philosophically usable, on Pigliucci's picture, is its functioning as an axiom — its evoking a landscape with rigid properties that argument can then explore — and an articulation can also be reached by working outward from the articulations a corpus already contains. Whether a candidate articulation succeeds is a question about what it evokes, and that question is answered the way other philosophical questions are: by public assessment of the structure the articulation makes available. That assessment is one any candidate articulation must meet, whatever produced it.
The phenomenology objection, like the abduction objection, rests on an inference from producer to product: the absence of experience in the producer is taken to remove phenomenological value from the product. The inference fails. LLMs lack conscious experience, but phenomenology does its philosophical work in articulated form, and Pigliucci's account explains why this is no workaround: philosophy puts empirical and experiential materials to use precisely by articulating them — turning them into starting points whose consequences can be developed and contested in argument. Articulated starting points are public, and a landscape, once evoked, is open to whatever can explore it. One challenge remains, and it grants everything so far: that wherever a model's output is worth reading, the worth was put there by the person prompting it.
[^phen1]: Zahavy, following Magnani (2009), calls reasoning of this kind *manipulative abduction*: the imagined observation functions as evidence, and the new axiom is adopted as what best explains it. "Because the simulated sensory experience of acceleration was indistinguishable from the remembered sensory experience of gravity, Einstein abducted that they must be the same phenomenon" (2026, §5). The elevator is thus Zahavy's instance of the generative abduction whose denial Section 2 left as the live issue.
[^phen2]: The experiential materials need not be sensory. The feeling of understanding — the phenomenal difference between merely parsing a proof and grasping why it goes through — is sometimes used to motivate the claim that thought has a phenomenology of its own (see Chudnoff 2011). Philosophy built on such episodes raises the present worry in its purest form.
## 4. The Challenge from Tools
In Sections 2 and 3 we considered challenges that turned on capacities LLMs lack: abductive inference in the one case, conscious experience in the other. In each case the reply had the same shape — the capacity is missing from the producer, but what the capacity was supposed to supply is a property of texts, and the product can have it anyway. A final challenge grants all of this and relocates the philosophy one step back. When a model produces a text worth reading, the challenge runs, the philosophy in it is the work of the prompter, not the model. LLMs cannot produce philosophy worth reading in just the sense that a typewriter cannot: both are tools, which a philosopher can use to produce it.
At a certain fineness of grain the claim is trivially true. A philosopher who pastes a section of a worthwhile paper into a model and asks for a version free of typographical errors will receive, if all goes well, a stretch of worthwhile philosophy produced — in some thin sense — by an LLM, and nobody should be impressed. The charge is that every apparent case of LLM philosophy is like this at bottom: crediting the model with the valuable properties of the resulting text is like believing that it is the ventriloquist's dummy that is doing the talking. And the way models actually behave seems to bear the charge out. Ask a chatbot a philosophical question outright — what is the correct theory of consciousness? what is the meaning of life? — and what comes back is a bland survey of positions at best, and turgid, content-free 'slop' at worst. The dummy speaks only when the ventriloquist is holding it; that is how one tells who is really talking. What is more, the people who can coax something better than a survey out of a model are, almost without exception, trained philosophers — which seems only to confirm that the philosophy in the output is theirs.
The intuition behind the challenge gets several things right, and we grant them at the outset. Unaided output often is a bland survey. Worthwhile output usually is closely directed by a person. The person who writes the prompt is the author of the prompt, and exercises philosophical skill in writing it. And a typewriter, like a word processor, earns no credit at all for what is written on it. The question is whether these concessions add up to the conclusion — that whenever an LLM text displays the properties of worthwhile philosophy, those properties are the prompter's work.
They do not add up to it, and the gap is one this paper has already mapped. The concessions are facts about the producer and the circumstances of production: who wrote the prompt, who directed the process, whose skill was exercised. The conclusion is a claim about the philosophy: whose work the contents of the text are. Section 1 argued that facts of the first kind do not settle what a text contains; the present challenge needs them to settle whose its contents are, which is a stronger demand on the same inference. The bare intuition that the model is a tool will not license it. What would license it is the model's being a tool of one specific kind — the kind a typewriter is. A typewriter adds nothing to the content of what is written with it: it fixes in type what its user has already settled, and every word of the novel was the novelist's before the machine touched it. If the model relays content in this way, the reattribution goes through. So the challenge stands or falls with the typewriter description, and the description is false.
The same pressure has already been applied, instructively, to image generation. The painter's brush invites the tool intuition in its purest form — nobody credits the brush — and Midjourney, one might think, is simply a more elaborate brush. But pressed, the assimilation fails: a Midjourney prompter can specify what an image should depict and in what style, yet has no fine-grained control over the formal properties of what appears, and the unpredictability is not a malfunction but the system's reliable character; an unpredictable brush would be a broken one (Young and Terrone 2025). Whatever the right positive account of such systems, the pressure tells against treating them as devices that merely transmit decisions their users have already made. In the philosophical case the same pressure can be made precise, because Section 3 has supplied the apparatus for saying exactly what a prompt does and exactly what the output does that the prompt does not.
A prompt posits a starting point. The prompter who sets a model a problem — a position, a pressure on that position, a question about where the pressure leads — is doing what Pigliucci's philosopher does in articulating an axiom: positing something that evokes a landscape, a structure that did not exist before the positing and whose properties, once posited, nobody chooses. The person who codifies the rules of chess authors the rules; the theorems of chess — that two knights cannot force mate, that the king and rook ending is won — follow from the rules and were chosen by no one, least of all the rule-writer, who may be incapable of proving any of them. Writing a prompt is writing rules of this kind. The consequences of the posited starting point are no more the prompt-writer's than the theorems of chess are the rule-writer's.
This is what the typewriter description misses. For that description to hold, the philosophy in the output would have to be in the prompt already, with the model relaying it — as the novel was in the novelist before the typewriter touched it. But a starting point, once posited, has more consequences than anyone has drawn, and they hold whether or not anyone draws them. An output can develop a consequence of the starting point that the prompt does not contain and that could not be read off it — an objection the prompter had not seen, a distinction that dissolves a difficulty the prompt only located, a commitment of the posited view that emerges three steps from anything stated. When it does, the relay description fails: what stands in the text was not settled by the user and then fixed in type. The model has added to the content, and a tool that adds to the content is not a tool in the sense the challenge requires.
One retreat remains. Grant that the philosophy in the output is not the prompt-writer's, and conclude that it is no one's: the consequences follow from the evoked landscape on their own, so the model has produced nothing either — it has merely transcribed what the landscape already fixed. But this confuses availability with statement. Once the rules of chess are codified, every fact about chess is fixed; the proof that two knights cannot force mate still had to be produced, and producing it was somebody's achievement. The landscape makes its consequences available; it does not state them. Stating a consequence is producing a text that develops it — selecting, from the unbounded space of things that follow, the ones that bear on the question; finding the consideration that discriminates between the posited view and its rivals; setting the steps in an order in which each earns the next. Sections 2 and 3 were arguments that a model can do exactly this: produce text that carries genuine abductive structure without inferring, and work within experiential landscapes without experiencing. What is worth reading, when the result is worth reading, is the developed text. The developed text is the model's.
On this account, the poverty of unaided output is not the counterexample the challenge takes it to be; it is what the account predicts. The bland survey is what a model returns when nothing has been posited: no starting point, no evoked landscape, nothing whose consequences a text could develop, so the output falls back on the most typical continuation of a question asked from nowhere. An instrument left running produces nothing of value because no one has set it anything — not because it can produce nothing. Far from showing that the model contributes no philosophical content, the contrast between unaided and directed output shows where the contribution happens: in the cases where a starting point has been posited and worked out, which is where the model does the developing.
Actual use, admittedly, is rarely a single prompt followed by a finished text. A philosopher works with a model in turns: reading an output, redirecting, cutting, asking for one line to be pursued and another dropped. Does the philosophy not become the person's somewhere in this process? The person's work in such an exchange divides into two kinds, and neither is authorship of the developments. Each redirection posits a fresh starting point, or narrows the one in play: choosing where the model should develop is the rule-writer's contribution again, at a finer grain, and choosing where to develop is not developing. The rest is assessment — judging which of the returned developments hold up, which collapse under a question, which deserve another turn — and assessment is the contribution of the editor and the referee. Here Section 1's observation about blind review returns from the other side: an editor who picks out the good papers, and a referee who recognises a sound argument, exercise philosophical judgement, sometimes of a high order, without thereby authoring what they pick out. We do not deny that as interventions grow finer the person's role shades from assessment into co-writing, and where it does, the right description is shared authorship. The claim this paper defends does not need to win that whole collaborative range. The claim is that LLMs can produce philosophy worth reading, and one clear kind of case suffices: a person posits a starting point, the model develops it, and the development — not the positing — is what repays the reading.
Two replies on behalf of the challenge deserve answers. The first concedes the framework and enriches the prompt: if the prompt is detailed enough — the position laid out, the difficulty located, the desired shape of the answer gestured at — then surely the development is mostly contained in it, and the model is back to relaying. But a detailed prompt is a larger starting point, not a worked-out philosophy. A longer axiom set is still distinct from its theorems; adding axioms changes which landscape is evoked, not who explores it. Detail increases what is posited. It does not move the development inside the positing, and the test stays the same: if the output develops consequences that could not be read off the prompt, however rich the prompt, the relay description fails. The second reply runs the retreat once more at a different point: the consequences follow from the landscape, so the landscape, not the model, deserves whatever credit there is. The answer is unchanged. A consequence's being available is not its being stated, and what the reader reads — what is or is not worth their while — is the statement.
The model, then, is a tool only in a sense that no longer supports the challenge. A typewriter adds nothing to the content of what is produced with it; the model develops consequences the prompt does not contain. Writing the prompt, the properties of the evoked landscape, and the text that develops them are three things, not one: the first is the person's, the second is no one's, and the third is the model's. This is why a worked-out output can be the model's philosophy, and why the blandness of unaided output, far from refuting the claim, marks exactly the difference between a model left running and a model set a problem.
We began with a distinction every reader already uses, and we end with it. Whether any particular text is worth reading is settled where it has always been settled: in the reading. What we have argued is that nothing settles it in advance — not the absence of a philosopher behind the text, not the producer's want of abductive inference or of experience, and not the fact that a person wrote the prompt. The catalogue of Erdős problems now credits a theorem to a model and the person who prompted it. We see no reason, in authorship, abduction, phenomenology, or tools, why the philosophy journals should not eventually do the same.
## References
Anscomb, C. (2022). Creating art with AI. *Odradek. Studies in Philosophy of Literature, Aesthetics, and Theory of Cognition, 8*(1), 14–51.
Balepur, N., Ravichander, A., & Rudinger, R. (2024). Artifacts or abduction: How do LLMs answer multiple-choice questions without the question? In *Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)* (pp. 10308–10330). Association for Computational Linguistics.
Bhagavatula, C., Le Bras, R., Malaviya, C., Sakaguchi, K., Holtzman, A., Rashkin, H., Downey, D., Yih, S. W., & Choi, Y. (2020). Abductive commonsense reasoning. In *International Conference on Learning Representations (ICLR 2020)*.
Bueno, O., & Shalkowski, S. A. (2020). Troubles with theoretical virtues: Resisting theoretical utility arguments in metaphysics. *Philosophy and Phenomenological Research, 101*(2), 456–469.
Chudnoff, E. (2011). What intuitions are like. *Philosophy and Phenomenological Research, 82*(3), 625–654.
Danto, A. C. (1981). *The transfiguration of the commonplace: A philosophy of art*. Harvard University Press.
Davies, D. (2004). *Art as performance*. Blackwell Publishing.
Dellsén, F., Firing, T., Lawler, I., & Norton, J. (2024). What is philosophical progress? *Philosophy and Phenomenological Research, 109*(2), 663–693.
Floridi, L., Morley, J., Novelli, C., & Watson, D. (2025). *What kind of reasoning (if any) is an LLM actually doing? On the stochastic nature and abductive appearance of large language models* (arXiv:2512.10080). arXiv.
Goldie, P. (2000). *The emotions: A philosophical exploration*. Oxford University Press.
Gottweis, J., Weng, W.-H., et al. (2025). *Towards an AI co-scientist* (arXiv:2502.18864). arXiv.
Guevara, A., Lupsasca, A., Skinner, D., Strominger, A., & Weil, K. (2026). *Single-minus gluon tree amplitudes are nonzero* (arXiv:2602.12176). arXiv.
Harman, G. (1990). The intrinsic quality of experience. *Philosophical Perspectives, 4*, 31–52.
Harnad, S. (1990). The symbol grounding problem. *Physica D: Nonlinear Phenomena, 42*(1–3), 335–346.
Hempel, C. G. (1966). *Philosophy of natural science*. Prentice-Hall.
Hertzmann, A. (2018). Can computers create art? *Arts, 7*(2), Article 18.
Jackson, F. (1982). Epiphenomenal qualia. *The Philosophical Quarterly, 32*(127), 127–136.
Jackson, F. (1986). What Mary didn't know. *The Journal of Philosophy, 83*(5), 291–295.
Lewis, D. (1990). What experience teaches. In W. G. Lycan (Ed.), *Mind and cognition: A reader* (pp. 499–519). Blackwell.
Lipton, P. (2004). *Inference to the best explanation* (2nd ed.). Routledge.
Magnani, L. (2009). *Abductive cognition: The epistemological and eco-cognitive dimensions of hypothetical reasoning*. Springer.
Merleau-Ponty, M. (2012). *Phenomenology of perception* (D. A. Landes, Trans.). Routledge. (Original work published 1945)
Novikov, A., Vũ, N., Eisenberger, M., et al. (2025). *AlphaEvolve: A coding agent for scientific and algorithmic discovery* (arXiv:2506.13131). arXiv.
Paul, L. A. (2012). Metaphysics as modeling: The handmaiden's tale. *Philosophical Studies, 160*(1), 1–29.
Pigliucci, M. (2017). Philosophy as the evocation of conceptual landscapes. In R. Blackford & D. Broderick (Eds.), *Philosophy's future: The problem of philosophical progress* (pp. 75–90). Wiley-Blackwell.
Reichenbach, H. (1938). *Experience and prediction: An analysis of the foundations and the structure of knowledge*. University of Chicago Press.
Salimi, M., Adim, S., Parnian, D., Alighardashi, N., Jafari Siavoshani, M., & Rohban, M. H. (2026). *Wiring the 'why': A unified taxonomy and survey of abductive reasoning in LLMs* (arXiv:2604.08016). arXiv.
Sellars, W. (1962). Philosophy and the scientific image of man. In R. Colodny (Ed.), *Frontiers of science and philosophy* (pp. 35–78). University of Pittsburgh Press.
Sider, T. (2011). *Writing the book of the world*. Oxford University Press.
Thomasson, A. L. (2015). *Ontology made easy*. Oxford University Press.
Unger, R. M., & Smolin, L. (2014). *The singular universe and the reality of time: A proposal in natural philosophy*. Cambridge University Press.
Williamson, T. (2007). *The philosophy of philosophy*. Blackwell Publishing.
Williamson, T. (2021). *The philosophy of philosophy* (2nd ed.). Wiley-Blackwell.
Wojtkiewicz, K. (2023). How do you solve a problem like DALL-E 2? *The Journal of Aesthetics and Art Criticism, 81*(4), 454–467.
Wolfram, S. (2023). *What is ChatGPT doing ... and why does it work?* Wolfram Media.
Young, N., & Terrone, E. (2025). Growing the image: Generative AI and the medium of gardening. *The Philosophical Quarterly, 75*(1), 310–319.
Zahavy, T. (2026). *LLMs can't jump* [Position paper]. Google DeepMind. PhilSci-Archive. http://philsci-archive.pitt.edu/28024/
Zeni, C., Pinsler, R., Zügner, D., et al. (2025). A generative model for inorganic materials design. *Nature, 639*, 624–632.