> [!warning] Work in Progress > This introduction is still being drafted and requires significant revision. > "Forty-two," said Deep Thought, with infinite majesty and calm. > It was a long time before anyone spoke. > Out of the corner of his eye Phouchg could see the sea of tense expectant faces down in the square outside. > "We're going to get lynched aren't we?" he whispered. > "It was a tough assignment," said Deep Thought mildly. > "Forty-two!" yelled Loonquawl. "Is that all you've got to show for seven and a half million years' work?" > "I checked it very thoroughly," said the computer, "and that quite definitely is the answer. I think the problem, to be quite honest with you, is that you've never actually known what the question is." > > — Douglas Adams, *The Hitchhiker's Guide to the Galaxy* In *The Hitchhiker's Guide to the Galaxy*, humanity asks an AI to do some philosophy. A computer named Deep Thought is constructed and instructed to provide "The Answer to the Ultimate Question of Life, the Universe, and Everything" (REF). Humanity builds this computer only to receive the answer '42'—an answer which, while apparently correct, means next to nothing at all due to humanity's failure to know what the Ultimate Question in fact is. In the mid 2020s, humanity has reached a position in which it can actually ask machines philosophical questions. In February 2026, researchers working on gluon scattering amplitudes, a problem whose complexity had defeated calculation by hand, gave GPT-5.2 worked examples for three, four, five, and six particles and asked it to find the general formula. The model conjectured a formula, completed a formal proof, and overturned a forty-year-old assumption (Guevara et al. 2026).[^1] Should we expect similar results in analytic philosophy? Unlike other disciplines, philosophy does not have clear and uncontroversial success conditions. The gluon scattering case succeeded because it produced a verifiable formula and a valid proof. Before we can ask whether LLMs could make similar advances in philosophy, we need to clarify what would count as making an advance. At one end of the spectrum are views on which philosophy consists in self-transformation (Hadot 1995) or investigating the conditions of experience (Kant, Merleau-Ponty).[^2] Few consider LLMs to have mental states,[^3] and on such views, the question whether they can do philosophy does not arise. Contemporary analytic philosophy, by contrast, evaluates arguments presented in texts. Referees assess papers by reading them: they check whether distinctions are well-drawn, whether objections are handled, whether the position illuminates its subject matter. Whether this kind of evaluation is sufficient—whether philosophy consists in arguments assessable by reading, or whether it requires mental states and experiences that LLMs lack—depends on what philosophy fundamentally is. Section 1 addresses this question, arguing that philosophical contributions consist in arguments assessed by text-internal criteria. If this is correct, then the question of whether LLMs can do philosophy becomes a question about the texts they produce: can they produce arguments exhibiting the features we recognise as philosophically valuable? I argue that they can. The argument proceeds in two stages. Section 1 establishes the metafilosophical framework: what makes philosophy good, and how we recognise it. Sections 2-3 then apply this framework to LLMs. Floridi et al. argue that LLMs exhibit only an 'abductive appearance' masking a 'stochastic core': they generate text from learned associations rather than performing genuine inference. If philosophy requires theoretical reasoning, and LLMs lack this capacity, they cannot do philosophy. I shall treat Floridi and Zahavy as foils whose force depends on assumptions plausible for physics but unmotivated for philosophy. The paper proceeds as follows. Section 1 argues that philosophical evaluation concerns text-internal criteria: coherence, handling of objections, illumination of dependence relations. Section 2 engages arguments by Floridi and Zahavy that LLMs cannot reason genuinely. Section 3 addresses the worry that LLM philosophy would be merely derivative. Section 4 considers what a demonstration would look like. [^1]: Other AI-assisted breakthroughs include protein structure prediction, which won the 2024 Nobel Prize in Chemistry (Hassabis and Jumper, AlphaFold; nobelprize.org/prizes/chemistry/2024); solving a 30+ year challenge in quantum error correction (Google Quantum AI, Willow chip; blog.google/technology/research/google-willow-quantum-chip/); and discovering new symmetries in black hole event horizon equations (Lupsasca with GPT-5; sciencenews.org/article/ai-enabled-science-discovery-insight). [^2]: Several traditions locate philosophical activity in the practitioner rather than the product. [^3]: Frankish (2024) argues that on an interpretivist view, LLMs can be credited with beliefs but essentially only one desire—to "play the chat game." On transformative conceptions, philosophy is "a practice of self-transformation" in which "what makes an activity philosophical is something that happens in the practitioner rather than anything that can be assessed in what she produces" (Hadot 1995; see also late Wittgenstein on philosophy as therapy). Transcendental and phenomenological approaches hold that philosophy investigates "conditions of possibility" of experience and knowledge (Kant 1781/1787), or "slackens the intentional threads which attach us to the world" to examine them (Merleau-Ponty 1945), which presupposes having experience. World-view conceptions treat philosophy as capturing "what it is actually like to live a human life in the world" (Dilthey; see Overgaard, Gilbert & Burwood 2013: ch. 8), requiring the philosopher to live such a life. On any of these views, an LLM cannot do philosophy because it lacks the requisite human capacities, regardless of what it produces. Watson and Crick discovered the double helix in 1953. Their paper in *Nature* announced what they had found: two strands running in opposite directions, held together by hydrogen bonds between complementary base pairs. The double helix existed before they described it. Had Rosalind Franklin or another researcher discovered it first, the discovery would have been identical, just differently attributed. Wittgenstein's *Philosophical Investigations* appeared the same year. %%Quine's *From a Logical Point of View* also 1953—could be used.%% But asking what Wittgenstein 'discovered' and whether someone else could have made the same discovery does not make sense in the way it does for Watson and Crick. The dialogical exchanges, the questions that resist resolution, the movement from case to case without systematic argument—these are not reports of something that exists independently. To say that someone else might have 'discovered' the same ideas would be to say they might have written the same arguments. But then in what sense would it be the same discovery? The arguments themselves are the contribution. There is nothing behind them that they report.[^lit][^witt] This reflects how philosophical explanation works. Lipton notes that explanations can be self-evidencing: > Suppose you ask me why there are certain peculiar tracks in the snow in front of my house. Looking at the tracks, I explain to you that a person on snowshoes recently passed this way. This is a perfectly good explanation, even if I did not see the person and so an essential part of my reason for believing my explanation are the very tracks whose existence I am explaining. (2004, p. 24) The explanandum provides evidence for the explanans. Philosophical arguments work this way. The argument addresses a philosophical problem, and the quality of the argument—its coherence, its handling of objections, its illumination of the subject matter—provides the evidence that the explanation is good. The argument is the evidence for itself. If philosophical contributions consist in arguments, what makes an argument good philosophy? Williamson argues that good theories should be elegant and unified, not arbitrary or ad hoc (2024, pp. 152-153). An elegant theory explains much with little; a unified theory hangs together rather than being a collection of separate claims. These are properties we assess by examining the theory itself—by working through its implications, checking whether it coheres, seeing whether it makes unmotivated exceptions. Others have identified similar features under different descriptions. Bengson and colleagues emphasise that philosophical work should be reason-based, coherent, and illuminating (2022, p. 589). Dellsén and colleagues frame philosophical progress in terms of representing dependence relations more accurately and comprehensively (2024, p. 663). Despite the different vocabularies, these accounts agree on what matters: features we assess by reading. Coherence, elegance, illumination of dependence relations—we see these by examining what an argument says and how it says it. Whether an argument handles objections, whether it makes unmotivated exceptions, whether it illuminates its subject matter: these are judgements we make by reading, not by investigating who wrote it or how it was produced. Blind review operates on this assumption—referees assess whether distinctions are well-drawn and objections anticipated without knowing the author. If philosophical evaluation concerns features of arguments, then production process is the wrong kind of variable. This is a familiar point in other domains. Gaut notes that Deep Blue plays objectively good chess moves regardless of whether those moves are creative (2010, fn. 23). Good-as-chess and creative-as-chess are different dimensions of evaluation, and the former does not require the latter. Similarly, mechanically generated metaphors can guide an audience's imagination effectively; the metaphor's success does not depend on whether it was produced creatively. These cases illustrate a general principle: domain-specific evaluative criteria do not collapse into facts about production. Lipton makes a parallel point about explanation. We evaluate potential explanations before establishing their truth. A potential explanation is "a proposition which would, if true, explain" some phenomenon (2004, p. 59). The judgement about which hypothesis would provide the best explanation—if true—is made on the basis of what Lipton calls "loveliness." And loveliness is not a matter of causal history. An explanation is lovely or not regardless of how it was generated. Lipton's squash analogy illustrates this distinction between levels: > If these suggestions are along the right lines, then arguing that Inference to the Best Explanation is wrong because Bayesianism is right is like arguing that thinking about technique cannot help my squash game because the motion of the ball is governed by the laws of mechanics. (2004, p. 108) The ball's motion is governed by mechanics, but this does not make thinking about technique pointless. The two levels are compatible. Whatever processes produce a philosophical text—human cognition, collaborative discussion, systematic inquiry—the question of whether the resulting text meets philosophical criteria remains distinct. Production mechanics and philosophical evaluation operate at different levels. We have established that philosophical evaluation concerns text-internal features: coherence, handling of objections, illumination of dependence relations. Production process—whatever mechanisms generated the text—operates at a different level. The question is whether a text exhibits these features, not how it was produced. [^lit]: Literature exhibits a similar constitutive character. When we evaluate Ian Fleming's *Casino Royale* (also published in 1953), we assess the prose, the pacing, the tension at the baccarat table—features internal to the text. Philosophy shares this constitutive character but differs in evaluative criteria. Literature is assessed aesthetically; philosophy is assessed for argumentative virtues (clarity, handling of objections, illumination of dependence relations). The constitutive character is shared; the criteria differ. [^witt]: *Philosophical Investigations* is an extreme case of this constitutive character. The aphoristic form, the dialogical method, and the questions that deliberately resist resolution mean the arguments cannot be separated from their mode of expression. Our claim does not depend on this extreme case—it applies equally to conventional analytic papers where the constitutive character is less dramatic but equally present.