## Introduction > "Forty-two," said Deep Thought, with infinite majesty and calm. > It was a long time before anyone spoke. > Out of the corner of his eye Phouchg could see the sea of tense expectant faces down in the square outside. > "We're going to get lynched aren't we?" he whispered. > "It was a tough assignment," said Deep Thought mildly. > "Forty-two!" yelled Loonquawl. "Is that all you've got to show for seven and a half million years' work?" > "I checked it very thoroughly," said the computer, "and that quite definitely is the answer. I think the problem, to be quite honest with you, is that you've never actually known what the question is." > > — Douglas Adams, *The Hitchhiker's Guide to the Galaxy* In *The Hitchhiker's Guide to the Galaxy*, humanity asks an AI to do some philosophy. A computer named Deep Thought is constructed and instructed to provide "The Answer to the Ultimate Question of Life, the Universe, and Everything" (REF). Humanity builds this computer only to receive the answer '42'—an answer which, while apparently correct (at least according to Deep Thought), means next to nothing at all due to humanity's failure to know what the Ultimate Question in fact is. In the 2020s, we can actually ask machines philosophical questions. Before worrying about what to ask, however, a more pressing question is: should we bother? In other disciplines, the advent of generative AI has produced genuine breakthroughs. In February 2026, OpenAI's GPT-5.2 was given a problem in theoretical physics concerning gluon scattering amplitudes — a problem whose complexity had defeated calculation by hand — and the model conjectured a general formula, completed a formal proof, and overturned a forty-year-old assumption (Guevara et al. 2026).[^1] Should we expect similar success in analytic philosophy? However, Philosophy, unlike other disciplines, does not have clear and uncontroversial success conditions. %% one succinct sentence mentioning the success conditions of the experiment experiment mentioned in the previous paragraph.%% Moreover, we need to get much clearer about what question we are asking. Interpreted one way, an LLM can simply copy Philosophical Investigations word for wordAt one end, an LLM can simply copy *Philosophical Investigations* word for word; that gives you a philosophical text only in the thin sense that it is already a philosophical text. Given enough time, a room of monkeys at typewriters could produce the same book; but no one would take that as evidence that the monkeys had become philosophers. At the other end, if philosophy is essentially a practice of self-transformation — if what makes an activity philosophical is something that happens in the practitioner rather than anything that can be assessed in what she produces — then no text-generating system could do philosophy regardless of what it outputs, and the discussion ends by stipulation. The interesting territory lies between these extremes. The question worth asking is whether, given philosophically minimal prompting, an LLM can produce a piece of writing that is not a mere reproduction of existing text, that exhibits non-accidental philosophical structure, and that can be assessed as philosophy on the page. My aim in this paper is to argue that LLMs can produce good philosophy. The argument turns on philosophy's evaluative criteria. If those criteria concern properties of texts — arguments that are clear, coherent, and illuminating — then the question of whether LLMs can do philosophy is a question about whether they can produce texts with those properties. I shall argue that they can. The paper proceeds as follows. Section 1 develops the positive case, drawing on recent work by Williamson on theoretical virtues, Bengson and colleagues on philosophical progress, and Dellsén and colleagues on understanding. Section 2 engages arguments by Floridi and Zahavy that LLMs cannot reason genuinely; I treat these as foils whose force depends on assumptions plausible for physics but unmotivated for philosophy. Section 3 addresses the worry that LLM philosophy would be merely derivative. Section 4 considers what a demonstration would look like. ## Section 1: Philosophy in the Text In April 1953, James Watson and Francis Crick published a paper in *Nature* announcing the structure of DNA. They had discovered that the molecule forms a double helix, with two chains of nucleotides running in opposite directions, held together by hydrogen bonds between complementary base pairs — adenine with thymine, guanine with cytosine. The structure explained how genetic information could be copied: the two strands separate, and each serves as a template for a new complementary strand. Watson and Crick's paper reported what they had found. The double helix existed before they described it; the paper communicated a structure that was there to be discovered. We evaluate their paper by asking whether they got the structure right — whether the world is as they said it is. In the same month — April 1953 — Ian Fleming published *Casino Royale*, the first James Bond novel. When we evaluate Fleming's novel, we are not asking whether it correctly reports something external. There is no external structure the novel describes and might have got wrong. The text itself is what we care about: the prose, the pacing, the tension at the baccarat table, the formal qualities of the narrative. The important stuff is on the page. We are interested in aesthetics — whether the novel achieves what novels can achieve — and this is a matter of the text's own qualities, not its relation to some independent domain it reports. Philosophy is like literature in this respect: the important stuff is on the page. A philosophical text is not a report of a prior discovery in the way that Watson and Crick's paper was. Consider Kripke's *Naming and Necessity*. Kripke argues that names are not disguised descriptions. Suppose that 'Gödel' just meant 'the person who proved the incompleteness theorems.' Then if we learned that someone else — call him Schmidt — actually proved them while Gödel stole the credit, it would follow that 'Gödel' refers to Schmidt, since Schmidt is the one who satisfies the description. But this is absurd. If we learned this, we would not conclude that Gödel proved the incompleteness theorems after all; we would conclude that Gödel did not prove them, and that he is a fraud. The name 'Gödel' picks out a particular person — the man we have been calling 'Gödel,' who was born in Brno in 1906 and emigrated to Princeton — not whoever happens to satisfy some description. Kripke concludes that names are rigid designators: they pick out the same individual in every possible world. Kripke also distinguishes necessity from a priori knowability. We can know a priori that Hesperus is Hesperus — this is trivially true. But we cannot know a priori that Hesperus is Phosphorus; it took astronomical observation to establish that the morning star and the evening star are the same celestial body. Yet if the identity is true, it is necessary: there is no possible world in which Hesperus exists but is not identical with Phosphorus, since they are the same object. So something can be necessary without being knowable a priori, and something can be contingent despite being knowable a priori. The distinction between epistemic and metaphysical possibility is not the same as the distinction between what can and cannot be known independently of experience. These arguments — the Gödel/Schmidt case, the Hesperus/Phosphorus example, the account of how natural kind terms like 'water' and 'gold' refer to underlying natures rather than to whatever satisfies associated descriptions — are what makes *Naming and Necessity* what it is. Kripke's conclusions matter, but what makes the book a major work of philosophy is that the arguments are clear, that they illuminate a range of phenomena (reference, modality, the necessary a posteriori), and that they force us to rethink assumptions we had not realised we were making. The arguments ARE the philosophical contribution. We evaluate Kripke by evaluating the arguments. Philosophy differs from literature, though, in what we value. Novels are about aesthetics: whether the prose achieves its effects, whether the narrative structure works, whether the characters illuminate human experience. Philosophy is not about aesthetics in this way. We value philosophical texts for different reasons: clarity, handling of objections, illumination of the subject matter. Recent work in metafilosophy has articulated what this amounts to. Williamson argues that good theories "should be elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated" and should "combine simplicity with strength" (2024, 152–3). Bengson, Cuneo, and Shafer-Landau identify what they call understanding-enabling features: theories that illuminate are "reason-based, robust, illuminating, orderly, and coherent" (2022, 589). Dellsén and colleagues argue that philosophy makes progress when "philosophical ideas (theories, arguments, distinctions, etc.) becom[e] publicly available" in ways that put people "in a position to increase their understanding," where understanding is "a matter of more accurately and/or more comprehensively representing the network of dependence relations between various phenomena" (2024, 663). These characterisations converge on something important: philosophy is evaluated by looking at the arguments themselves. Elegance, unity, coherence, illumination of dependence relations — these are features of arguments. An argument is coherent or it is not; it illuminates the subject matter or it does not; it is ad hoc and gerrymandered or it is elegant and unified. These are properties we can assess by examining the text. If this is correct, the question of whether LLMs can do philosophy is a question about whether they can produce arguments with these features. Can an LLM produce an argument that is clear rather than confused? That responds to objections rather than ignoring them? That illuminates a problem rather than obscuring it? That is elegant and unified rather than ad hoc and gerrymandered? These are the questions we need to ask. They are not easy questions, but they are tractable. We can evaluate LLM outputs by the same standards we use to evaluate any piece of philosophy. ## Section 2: Floridi and Zahavy as Foils - Floridi et al. and Zahavy both argue that LLMs lack something required for genuine reasoning. Their arguments, though different in detail, share a common assumption: that the relevant activity requires access to something beyond text. This assumption is plausible for physics. It is unmotivated for philosophy. - Floridi et al. write: > "We argue that such LLMs generate text based on learned associations rather than performing abductive inferences. When their output exhibits an apparent abductive quality—often reinforced by interface design—this effect is due to the model's training on human-generated texts that encode reasoning structures." - The LLM has learned what reasoning looks like, they claim, without performing it. Its "abductive appearance" masks a merely "stochastic core." - But for philosophy, appearance is not mere appearance. If an argument exhibits the theoretical virtues — if it is elegant, unified, and non-ad-hoc; if it tracks dialectical obligations and responds to objections — then it is good reasoning. The standards concern the output. Whether the producer "really" reasoned is beside the point. - Floridi et al. themselves raise the question: "if an AI can generate the same explanatory hypothesis a human would, does it matter that the process was different?" For philosophy, the answer is straightforward. It does not. - Zahavy's argument concerns the "E→A Jump" — the creative leap from sense experience to axioms. He writes: > "Einstein did not bridge Special Relativity and gravitation by gathering observations, but by simulating the physical feelings of an observer inside a sealed environment." - This jump, Zahavy argues, requires embodied simulation. LLMs lack it. They are, in his phrase, "high-dimensional 'Chinese Rooms' (Harnad, 1990), manipulating the language of physics without access to the physical referents that give that language meaning." - The response to this is also straightforward. Zahavy's argument is explicitly restricted: > "we emphasize that this proposal is specifically tailored to the physical sciences, where the object of study is external material reality. In abstract domains such as Mathematics or Computer Science, the Sense Experience (E) may be grounded in high-dimensional topology or have other goals such as generality or minimality." - Philosophy is one of those abstract domains. Its "data" are not raw sense experiences requiring embodied simulation. They are arguments, intuitions recorded in texts, examples already articulated in language. LLMs have extensive access to these. - Both arguments assume that philosophy is like physics — that it requires access to something beyond what texts contain. But philosophy's standards are internal to a textual practice. The "something beyond text" that LLMs allegedly lack — genuine understanding, embodied experience — is not what philosophy's evaluative criteria require. - Zahavy's GPT-5.2 example illustrates the confusion. The model derived a physics theorem from axioms — A→S work, in his terminology — but could not have formulated those axioms from experience (E→A work). Yet in philosophy, much of the work is A→S-like: tracing implications, checking consistency, developing positions already present in the corpus. The E→A move, if it exists in philosophy at all, goes from text to text. ## Section 3: Dialectical Saturation - Philosophy's dialectical space is extensively documented in its corpus. For any well-explored question — free will, the nature of knowledge, the mind-body problem — the space of positions, the objections to each, and the standard replies have been worked out over centuries of argument. This documentation is what LLMs are trained on. - To do philosophy well is to know where the pressure points are. A competent philosopher addressing a question knows which objections will be raised and which responses are available. This is not occult knowledge. It is encoded in the corpus: the objection appears in paper X, the reply in paper Y, the counter-reply in paper Z. - LLMs have learned this structure. They can anticipate objections because they have been trained on texts that raise them. They can provide responses because they have been trained on texts that give them. - This is not mere mimicry. The dialectical task — identifying where an argument is vulnerable and how it might be defended — is what the training has equipped them to perform. - The corpus also encodes which arguments are good. Papers get published, taught, anthologised, and cited in rough proportion to their quality — where "quality" tracks the theoretical virtues: elegance, unity, non-arbitrariness, engagement with objections. The selection pressure of peer review and disciplinary uptake filters for arguments exhibiting these features. - LLMs, trained on this filtered sample, have learned the distribution of what counts as good philosophy. They do not need an independent evaluative faculty. The evaluative work has already been done by the community whose outputs constitute the training data. - The LLM learns not just which moves exist, but which moves are valued. It learns to produce philosophy that looks like good philosophy because good philosophy is overrepresented in its training data. - One might object that this makes LLM philosophy derivative — a recombination of existing moves rather than genuine innovation. The objection has force for certain kinds of innovation: the LLM is unlikely to introduce a wholly new framework or reframe a debate in a way no one has considered. - But most good philosophy is not of this kind. Most good philosophy consists in careful articulation, rigorous argument, and sophisticated engagement with existing positions. These are precisely the skills that training on the corpus develops. - The LLM can produce a novel argument — novel in the sense that it does not appear verbatim in the training data — by combining existing elements in ways that satisfy the evaluative standards it has learned. This is how human philosophers produce novel arguments too. [^1]: Other AI-assisted breakthroughs include protein structure prediction, which won the 2024 Nobel Prize in Chemistry (Hassabis and Jumper, AlphaFold; nobelprize.org/prizes/chemistry/2024); solving a 30+ year challenge in quantum error correction (Google Quantum AI, Willow chip; blog.google/technology/research/google-willow-quantum-chip/); and discovering new symmetries in black hole event horizon equations (Lupsasca with GPT-5; sciencenews.org/article/ai-enabled-science-discovery-insight).