> "Forty-two," said Deep Thought, with infinite majesty and calm.
> It was a long time before anyone spoke.
> Out of the corner of his eye Phouchg could see the sea of tense expectant faces down in the square outside.
> "We're going to get lynched aren't we?" he whispered.
> "It was a tough assignment," said Deep Thought mildly.
> "Forty-two!" yelled Loonquawl. "Is that all you've got to show for seven and a half million years' work?"
> "I checked it very thoroughly," said the computer, "and that quite definitely is the answer. I think the problem, to be quite honest with you, is that you've never actually known what the question is."
-- Douglas Adams, *The Hitchhiker's Guide to the Galaxy*
In *The Hitchhiker's Guide to the Galaxy*, humanity asks an AI to do some philosophy. A computer named Deep Thought is constructed and instructed to provide "The Answer to the Ultimate Question of Life, the Universe, and Everything" (REF). Humanity builds this computer only to receive the answer '42'—an answer which, while apparently correct (at least according to Deep Thought), means next to nothing at all due to humanity's failure to know what the Ultimate Question in fact is.
Deep Thought is fiction, but the underlying worry is not. If you do not know what you are asking, even a correct answer can be useless; and if you do not know what would count as success, the question "can it do X?" collapses into a verbal dispute. This matters now because large language models are no longer confined to producing plausible-sounding prose. In February 2026, OpenAI's GPT-5.2 was given a problem in theoretical physics concerning a class of gluon scattering amplitudes whose complexity had defeated calculation by hand. The human physicists on the team had worked out a few simple cases, but the expressions grew superexponentially; the model conjectured a general formula and completed a formal proof of its validity, overturning a forty-year-old assumption that the amplitudes in question were zero (Guevara et al. 2026). The point of mentioning this is not that philosophy should aspire to be theoretical physics. It is that AI outputs are now being evaluated by communities with extremely demanding public standards, and sometimes they pass. The question I want to address is whether philosophy could be among the disciplines for which this is true.
Stated baldly, "Can LLMs do philosophy?" is not yet a good question. It invites answers that are too easy to be informative. At one end, an LLM can simply copy *Philosophical Investigations* word for word; that gives you a philosophical text only in the thin sense that it is already a philosophical text. Given enough time, a room of monkeys at typewriters could produce the same book; but no one would take that as evidence that the monkeys had become philosophers. At the other end, if philosophy is essentially a practice of self-transformation - if what makes an activity philosophical is something that happens in the practitioner rather than anything that can be assessed in what she produces - then no text-generating system could do philosophy regardless of what it outputs, and the discussion ends by stipulation. The interesting territory lies between these extremes. The question worth asking is whether, given philosophically minimal prompting, an LLM can produce a piece of writing that is not a mere reproduction of existing text, that exhibits non-accidental philosophical structure, and that can be assessed as philosophy on the page. I want to argue that it can, and that the argument turns on a feature of analytic philosophy that distinguishes it from the empirical sciences: in much analytic philosophy, the text is the contribution, and the discipline’s evaluative standards are available in the text.
To make that claim substantive, "good" and "novel" need unpacking. A competent summary of an existing debate might be well-structured, accurate, and clear without putting any reader in a position she was not already in; it would be good in one sense without amounting to a contribution. What would be interesting is a text that improves a reader's understanding of some phenomenon - that helps her grasp dependence relations she had not previously grasped, not merely that gives her access to what others have already grasped. To make this distinction precise, I adopt Dellsén et al.'s account of philosophical progress, *Enabling Noeticism*:
> Enabling Noeticism: The discipline of philosophy makes progress regarding some phenomenon to the extent that philosophical research puts people in a position to increase their understanding of that phenomenon. (Dellsén et al. 2024, p. 679)
On this view, the central evaluative notion is understanding, and understanding is a matter of degree. Very roughly, an agent understands X to the extent that she accurately and comprehensively represents the network of dependence relations in which X stands (or fails to stand) to other things. Two dimensions matter. A representation can be more or less accurate, depending on whether the dependence relations it encodes actually obtain; and it can be more or less comprehensive, depending on how much of the relevant network it captures rather than omitting. Understanding is, in this sense, epistemically undemanding (it does not require knowledge or justification), while remaining robustly factive: what matters is how well the representation matches the dependency structure of the world.
The Gettier case is an example. The justified true belief theory represents knowledge as depending on three conditions—truth, belief, and justification—and on nothing else. Gettier's counterexamples showed that the 'nothing else' clause was wrong: even justified true beliefs can fail to be knowledge, so knowledge depends on something further. Whatever one thinks about subsequent attempts to supply a fourth condition, Gettier's paper plausibly counts as progress in Dellsén et al.'s sense. It put readers in a position to represent more accurately and more comprehensively what knowledge does and does not depend on, even though it did not supply the missing positive account.
On this account, a reader's understanding of a phenomenon consists in her mental model of the dependence relations in which that phenomenon stands. A philosophical text contributes to progress by putting readers in a position to improve that model - to make it more accurate, more comprehensive, or both. The evaluative question, then, is whether a given text enables such improvement in its readers. That question can be answered by examining the text: does it track genuine dependence relations? does it capture structure the reader had missed? If so, it is a vehicle for progress, regardless of how it was produced.
Analytic philosophy is, by and large, a text-based discipline. Philosophical contributions are written artifacts - arguments, distinctions, counterexamples, theoretical frameworks - and their assessment is likewise text-based. In experimental science a paper typically reports work done elsewhere; in much analytic philosophy the argumentative work is done on the page. Referees read, test inferences, press for missing premises, check whether rival positions are treated fairly, and ask whether the view survives predictable objections. Blind review exists in philosophy for exactly this reason: the discipline's evaluative norms apply to what is in the text, not to the biography, psychology, or social position of its author. If this is correct, then the question of whether LLMs can contribute to philosophical progress should be posed at the level of the artifact - the text itself, assessed independently of its origin: can an LLM produce a text that, when read by an informed reader, puts that reader in a position to understand better?
The word *produce* needs qualifying, because using an LLM can mean many different things. At one end, a human philosopher does the philosophical work and uses the model as a transcription device, a stylistic editor, or a convenient assistant for paraphrase and summarization. At the other end, we have something close to Deep Thought: philosophically minimal prompting - a question, a topic, a request for a familiar kind of move - elicits an extended piece of writing whose substantive structure is not supplied by the user. My focus is on that far end of the spectrum. The question is whether, given minimal prompting, an LLM can produce text with genuine philosophical structure - text that puts a reader in a position to understand better.
Two recent arguments provide the strongest case for skepticism about this possibility. Floridi et al. claim that LLM outputs exhibit, at best, an *abductive appearance*: they mimic the surface form of inference without performing the kind of reasoning that would warrant trusting the result. Zahavy argues, in a related spirit, that genuine abduction requires a leap from experience to explanatory axioms that a purely text-trained system cannot make. These arguments target real architectural limitations, and I do not dispute them as claims about LLM cognition. The question is whether the conception of abduction they presuppose is the right one for philosophy as a text-based practice aimed at improving understanding, rather than for empirical science as a practice whose textual record is downstream of laboratory work. Both arguments are developed with the physical sciences as their reference case - the discipline in which, as it happens, LLMs have most recently shown themselves capable of novel work.
The rest of the paper develops that case. Section 1 sets out Floridi et al.'s and Zahavy's objections and clarifies what they would show if philosophy were relevantly like the physical sciences. Section 2 argues that "abduction" is used in several different ways across the literatures at issue, and that Williamson's abductive picture of philosophy is best understood as a method of theory evaluation by intrinsic virtues—virtues that are, again, assessable in texts. Section 3 makes the positive argument: the norms of philosophical practice are publicly codifiable and textually manifest, and the philosophical corpus is itself the record of an evaluative feedback loop that a model can learn from. Section 4 provides worked examples: minimal prompting can elicit outputs with genuine philosophical structure, and these outputs can be evaluated by the discipline's own standards.