# Generating Philosophy with Large Language Models: Abduction, Textuality, and the Artefact View
Section word-count plan (body text only; excludes the References section):
0. Introduction: 1381 words
1. What LLMs Aren't Doing: 1565 words
2. Abduction and Philosophy: 1687 words
3. Learning the Game: 1646 words
4. How to Generate Philosophy with AI: 1788 words
5. Conclusion: 794 words
Total (Sections 0–5): 8861 words
# 0. Introduction
Douglas Adams’s *The Hitchhiker’s Guide to the Galaxy* contains a tiny parable about AI and philosophy. Humans build a computer, Deep Thought, and ask it for “the Answer to the Ultimate Question of Life, the Universe, and Everything.” After an absurdly long wait, Deep Thought returns “42”, and then adds the sting: the problem is that the humans never really knew what the question was. %%I want the complete blog quote from the novel back at the beginning. %%
Deep Thought is fiction, but the underlying worry is not. If you do not know what you are asking, even a correct answer can be useless. And if you do not know what would count as success, the question “can it do X?” collapses into a verbal dispute. %% the first two sentences of this paragraph are terrible and should be removed or reformulated. Any reader would be entirely lost. %%This matters now because large language models have moved beyond “plausible-sounding prose” as their only party trick. %% this is a good example of a sentence that I would never write and evidence that my writing analytic voice skill or whatever it's called was not used properly here, or used at all, perhaps.%% In February 2026, OpenAI reported that GPT‑5.2 Pro conjectured a general closed-form formula for a class of gluon scattering amplitudes, after human collaborators had derived only small‑n cases with expressions whose complexity grew rapidly; an internal “scaffolded” model then produced a formal proof, and the result was checked analytically by the authors in a public preprint. (OpenAI 2026; Guevara et al. 2026; Moskowitz 2026.)%% all these details should be checked. %% I mention this not because philosophy should imitate theoretical physics, but because it shows something mundane and important: communities with demanding public standards sometimes judge AI outputs to pass those standards. %% I would never write a sentence like this. The content is okay though, but the way it's framed and the example you use is awful. Actually, I'm not even sure the content is good. %%
Could philosophy be among the disciplines for which this is true? “Can LLMs do philosophy?” is an intoxicating %% we're writing Philosophy here, not journalism. %%question, but stated baldly %% again, this is journalism, not Philosophy Jesus fucking Christ.%%it invites answers that are too easy to be informative. At one extreme, an LLM can reproduce an existing philosophical classic word %% why is he removed Philosophy investigations by Wittgenstein as the example here? That makes much more sense, no? %%for word. That gives you a philosophical text only in the thin sense that it is already a philosophical text. %% this is really stupidly put. Um it's still a philosophical text. It's not any thinner. The problem is you haven't really done Philosophy or produced Philosophy or done anything novel, no? %% %% once again the thousand monkeys and thousand typewriters examples should be in here. Yet for some reason it is not. %%At the other extreme, %% I would never use this sort of phrase. %%if philosophy is essentially a practice of self-transformation—if what makes an activity philosophical is something that happens inside the practitioner rather than anything assessable in what she produces—then no text-generating system could “do philosophy” regardless of what it outputs, and the discussion ends by stipulation. %% this is really vague and not very well put. We have information somewhere in the vault about conceptions of meta Philosophy Some information should be thrown in or Consulted to make to improve this part of the paragraph. %%
The interesting territory lies between these extremes. The question worth asking is whether, given philosophically minimal prompting, an LLM can produce a piece of writing that is not a mere reproduction of existing text, that exhibits non-accidental philosophical structure, and that can be assessed as philosophy on the page. %% this sentence should be simplified dramatically. %%My claim is that it can. But I do not want the claim to float free of standards.%% please try not to write like a cunt. %% “Good” and “novel” need unpacking in a way that is fair to the discipline and not cooked to fit the technology.
One natural thought is that good philosophy increases understanding. That sounds airy until we say what “understanding” is. Dellsén and collaborators propose an account of philosophical progress that makes the relevant idea crisp:
> The discipline of philosophy makes progress regarding some phenomenon to the extent that philosophical research puts people in a position to increase their understanding of that phenomenon. (Dellsén et al. 2024, 679.)
%% the relevant part of this block quote should be italicised. I think it should also there should also be a lar it should also be a larger block quote with more sentences either side, if that helps.%%
On their view, understanding is a matter of degree: one understands a phenomenon better to the extent that one represents more accurately and more comprehensively the network of dependence relations in which it stands. (The dependence relations might be causal, grounding, constitutive, explanatory, or something else, depending on the case.) The picture is epistemically undemanding—understanding does not require knowledge or justification—while remaining factive in an important sense: the representation has to match the dependency structure of the world, not merely feel satisfying. A philosophical text counts as progress, on this view, to the extent that it puts readers in a position to improve their internal “dependency model” of what the phenomenon depends on and what it does not. (Dellsén 2024.)
The Gettier case is a familiar illustration. The justified true belief theory of knowledge represented knowledge as depending on three conditions—truth, belief, and justification—and on nothing else. Gettier’s counterexamples did not supply a positive replacement theory, but they did show that the “nothing else” clause was wrong: even justified true beliefs can fail to be knowledge. Readers thereby became able to represent more accurately what knowledge depends on (and does not depend on). On Dellsén et al.’s account, that counts as philosophical progress even if the replacement story remains contested. %% it would be good if this could be condensed quite dramatically. Everyone knows what the Gettier case is. All we need to do is show how this illustrates the dependency model of understanding. %%
That account is useful here because it directs attention away from the psychology of the author and toward the noetic effect of the text. A reader’s understanding is improved by grasping new structural relations—new dependencies, new constraints, new discriminations—whether the text was written in a bout of feverish midnight inspiration, in a slow collaborative process, or by something stranger. The evaluative question%% twatty way of writing. %%, then, is whether a given text enables improvement in competent readers. And that question can often be answered by examining the text: does it track genuine constraints? does it expose a dependence relation the reader missed? does it dissolve a pseudo-problem by locating an equivocation? does it show where a popular argument would have to be strengthened? %% these fucking threesome examples you love giving not only eat into the word count but also are a dead giveaway of LLM writing. Using such long long-winded and examples and three of them. It's just fucking boilerplate, just pushing up the word count for no good reason. Be more succinct, be more analytic. %%
This is where a feature of analytic philosophy becomes central. In much analytic philosophy, the text is not mainly a report of something done elsewhere. The argumentative work is done on the page. %% this idea is introduced far too quickly and casually. Very childishly. %%When a philosopher develops an objection to a thesis, the sentences that develop the objection are the objection. When a philosopher draws a distinction, the articulation is the contribution. %% no reader would know what you're talking about at this point. It's very unclear what the idea is supposed to be. And I was the one who had the idea. %% That is why blind review makes sense in philosophy: the discipline’s evaluative norms largely apply to what is in the text—validity, clarity, burden‑shifting, non‑ad hocness, fair treatment of alternatives—not to the biography or inner life of its author. %% any reader paying even a modicum of attention is going to say yeah, but um blind review is for science as well. Okay, so you need to say something about that as well, perhaps checking. be validity of the experiments or something like that, I don't know. Please work a bit harder. %% If this is right, then the question “can LLMs do philosophy?” should be posed at the level of the artefact: can an LLM produce a text that, when read by an informed reader, enables increased understanding by meeting the intrinsic standards of philosophical writing?
The word “produce” needs qualifying, because using an LLM can mean many different things. At one end, a human philosopher does the philosophical work and uses a model as a transcription tool, stylistic editor, or summarizer. At the other end—closer to Deep Thought—philosophically minimal prompting (a topic, a question, a request for a familiar kind of move) elicits an extended piece of writing whose substantive structure is not supplied by the user. My focus is on that far end of the spectrum, because it makes the core issue sharp. The question is whether minimal prompting can yield text with genuine philosophical structure—structure that can be assessed as philosophy on the page. %% this point has already been made earlier. I am wondering whether this paragraph is entirely redundant.%%
Two recent skeptical arguments make this look unlikely. Floridi and collaborators argue that LLM outputs exhibit, at best, an abductive appearance: they mimic the surface form of inference without performing the kind of reasoning that would warrant trusting the result. (Floridi et al. 2025.) Zahavy, in a related spirit, argues that genuine abduction involves a leap from experience to explanatory axioms—a transition from world-contact to framework invention—that a purely text-trained system cannot perform. (Zahavy 2026.) I take these arguments seriously as claims about the cognitive architecture of current models. But the question I want to press is whether the conception of “abduction” they presuppose is the right one for philosophy as a text-based practice aimed at improving understanding, rather than for empirical science as a practice where the text is downstream of laboratory work. %% this is too much to put in an introduction. It seems to come out of the blue. Why are we giving our readers extended information about this aspect of what's coming when we're not giving them extended information about any other asepct that is coming. ? %%
The paper proceeds as follows. Section 1 sets out Floridi et al.’s and Zahavy’s objections and clarifies why, if philosophy were relevantly like the physical sciences, the objections would be decisive. Section 2 argues that “abduction” is used in several different ways across the literatures at issue, and that abductive methodology in philosophy is best understood as a method of evaluating theories by intrinsic virtues—virtues assessable in texts. Section 3 makes the positive case: the norms of analytic philosophical practice are publicly codifiable and textually manifest, and the philosophical corpus is itself a record of an evaluative feedback loop a model can learn from. Section 4 makes this vivid with worked examples, including at least one failure case where the flaw is detectable from the text itself. Section 5 draws the conclusions and gestures at implications for philosophical methodology and for the social organization of the discipline.
# 1. What LLMs Aren't Doing
If you want to argue that LLMs can produce good philosophy, a natural first move is to ask what kind of reasoning philosophy requires, and whether LLMs can do that reasoning. The skeptical arguments I discuss in this section take “abduction” to be the relevant bottleneck. %% This is a horrendous way to begin the new section, not least because it mischaracterises both of the texts you're talking about. That's not the only problem though. And it's just garbage all the way through.%%They do not say merely that models sometimes hallucinate or sometimes give shallow answers—trivial complaints that apply to many human graduate students on a bad day. They say something stronger: that the architecture of LLMs prevents them, in principle, from performing a distinctive kind of reasoning that underwrites our confidence in explanatory claims. If that is right, then even a text that looks philosophical could be epistemically worthless in the way a Deep Thought answer is: it might have the form of insight while lacking the underlying method that makes insight non-accidental.
It will help to keep two ideas separate. First, there is a psychological question about what is going on inside a reasoner when she performs abduction. Second, there is a methodological question about what makes an explanatory move a good one. Floridi and Zahavy are primarily concerned with the first, though they gesture at the second. My eventual claim is that, for large regions of analytic philosophy, the methodological question is the one that matters for evaluating artefacts. But to see why the psychological worry is tempting, we need the skeptical arguments on the table. %% These paragraphs that begin the section make my head spin. I've got no idea where I am, what I'm supposed to what's supposed to be happening in this section, why I'm being told anything. In terms of signposting, this is appalling.%%
Abductive reasoning begins with something in need of explanation—a surprising observation, an anomaly, a phenomenon that existing theories do not predict. The reasoner then proposes a hypothesis such that, if it were true, it would make the phenomenon intelligible. Peirce is the canonical source for treating abduction as hypothesis generation. But modern discussions often emphasize a comparative dimension: the reasoner does not merely generate a candidate and stop. She generates several candidates and then asks which is, if true, the best explanation. Harman’s phrase “inference to the best explanation” made that comparative element explicit, and Lipton’s formulation is a standard statement of the view:
> Given our data and our background beliefs, we infer what would, if true, provide the best of the competing explanations we can generate of those data. (Lipton 2004, 56.)
Lipton also makes a structural point that will matter. Inference to the best explanation does not range over all logically possible explanations. It ranges over a constrained set of “live options” and then selects among them. (Lipton 2004, 59.) The inference therefore has two potential failure points. You can be good at ranking candidates while being bad at generating them. Or you can be good at generating candidates while being bad at ranking them in a disciplined way. In actual scientific practice, these phases are often interleaved in a feedback loop: generation, partial evaluation, revised generation, and so on.
Floridi and collaborators argue that LLMs, at best, mimic the surface form of this practice. Their starting point is an observation: when asked to explain something, a model often produces text that identifies a preferred hypothesis, marshals considerations in its favor, mentions simplicity or coherence, and dismisses alternatives. The output has the *shape* of abductive reasoning. But the mechanism that produces it is stochastic. The model has learned statistical associations across a corpus of text and then samples a continuation token by token. Floridi et al. characterize the situation by saying that LLMs occupy a space “between” stochastic processes and human-like abduction:
> LLMs occupy a conceptual space “between” traditional stochastic processes and human-like abductive reasoning. (Floridi et al. 2025, 2.)
The worry is not that the outputs are always bad. The worry is that, even when they are good, the goodness is not secured by the kind of evaluative loop that gives abduction its epistemic standing. Floridi et al. therefore introduce the label “zeroth-order abduction” for what models are doing when they produce a plausible explanation-shaped continuation without actually comparing candidates against each other or against evidence. %%You introduce a piece of jargon and then don't explain it at all.%%
One way to state the worry is in Bayesian terms. %%Is this meant to be a continuation of the idea introduced in the previous paragraph? If so, it's very unclear that it is.%% In idealized abductive reasoning, you treat a candidate hypothesis as something to be assessed in light of evidence. You move from a prior, through likelihoods, to a posterior. Floridi et al. argue that LLMs generate from something like a prior predictive distribution without an external mechanism for posterior evaluation:
> LLMs perform prior predictive sampling but lack an external feedback loop for posterior evaluation. (Floridi et al. 2025, 7.)
This matters %%I would never write like this. %% because, without a feedback loop, the model cannot treat its own output as a hypothesis to be tested. It cannot ask, in its own name, whether the explanation it just produced really is the best among alternatives. It cannot withhold judgment on the basis that the evidence is insufficient. It can output the string “I’m not sure,” but that is a pattern in text, not necessarily an expression of a represented epistemic state. Floridi et al. call the resulting tendency to always produce an explanation “over-abduction,” and they treat hallucination as a predictable consequence of the architecture rather than a bug.
What makes the Floridi argument interesting is that it does not end in a triumphant dismissal. %%This sentence is stupid and unclear. it also breaks one of my ironclad rules for analytic writing, which is no value-laden words whatsoever. %% Floridi et al. notice that training data encode a great deal of inferential and causal structure, because human-written text systematically reflects it. They also raise, without settling, the key provenance question:
> If an AI can generate the same explanatory hypothesis a human would, does it matter that the process was different? (Floridi et al. 2025, 12.)
From the perspective of epistemology, you might think the answer is yes, because justification matters and a purely stochastic process seems like the wrong kind of thing to confer justification. But from the perspective of evaluating a theory as an artefact—as a candidate explanation considered in its own right—the question is harder. A good explanation has certain features: it is coherent, non-ad hoc, simple enough, informative enough, integrated with background commitments. Those features might be present in a text even if the process that generated it did not explicitly evaluate them. Floridi et al. leave that tension standing, and it is precisely the tension my argument exploits.
Zahavy’s argument targets a different aspect of abduction. Floridi are mainly worried about selection without evaluation: an explanation-shaped output produced without an abductive loop. Zahavy is worried about something closer to Peirce’s original emphasis on invention. Abduction, he claims, is not just choosing within a menu of existing hypotheses. It is a leap that creates the menu. %% you clearly are not using all the relevant skills here, because if you were, you'd have put some block quotes in to illustrate this idea in the author's own words. %% In physics, abduction often means the invention of a new framework—new axioms, new principles, new representational resources—that reorganize a field. %% this is all very sloppy. I'm not sure if this is a particularly accurate characterization of what is said in this paper. Or what Abduction means in physics. Again, very childish. %% Zahavy captures this with what he calls the E→A Jump: a transition from sense experience (E) to axioms (A) via a conceptual jump.
Here Zahavy’s focus is on an architectural limitation. LLMs are trained on text. They have access to descriptions of experiments and to theoretical statements, but they do not have the kind of embodied experience that, on Zahavy’s view, supports the invention of new physical frameworks. He calls the relevant form of hypothesis generation “manipulative abduction”: abduction that emerges through active construction and manipulation of mental models, through what he glosses as “thinking by doing.” The canonical illustration is Einstein’s elevator thought experiment: by simulating an observer’s experience in an accelerating frame and comparing it to gravitational experience, Einstein extracted the equivalence principle. Zahavy’s claim is that this kind of leap depends on a form of world-contact that text-only models lack.
Zahavy is explicit that his argument is tailored to the physical sciences:
> We emphasize that this proposal is specifically tailored to the physical sciences, where the object of study is external material reality. (Zahavy 2026, 7.)
The restriction matters%% stop saying things matter, you sound like a cunt%%. Zahavy is not claiming that no system trained on text can ever make conceptual leaps. He is claiming that, in physics, the relevant leap is from embodied engagement with external material reality to a framework that captures that reality. That is what a text-only system lacks. A model can optimize within a space of candidate equations or within an existing formalism, but it cannot define a new space because it lacks the kind of “experience” from which new axioms are abstracted.
If philosophy were relevantly like physics, these arguments would be devastating. Suppose that philosophical progress depended on inventing new axioms grounded in embodied experience, or on running a generation‑evaluation loop that is epistemically creditable only if the reasoner can represent and update its own epistemic states. Then a text-only stochastic system would indeed be a kind of high‑dimensional Deep Thought: it might generate plausible explanations in the style of philosophy without being able to do the thing that makes philosophy epistemically serious. %%Mm the preceding paragraphs have been a fairly sorry, very sloppy characteris charact characterization of the ideas in this paper. It's absolute bollocks.%%
But philosophy is not physics, and that difference is not merely sociological. %%I would never write like this%% It is a difference in the relationship between a discipline and its textual medium. In physics, the paper is typically downstream of extra-textual work: an experiment, an instrument, a mathematical derivation, an empirical dataset. In much analytic philosophy, the argumentative text is not downstream of the contribution; it is the contribution. That fact does not trivialize philosophical standards, but it changes what it would mean for a model to “lack abduction” %% why are you using quotation marks as scare quotes? I fucking hate that, it's in your config.%% in a way that matters. To see whether the Floridi and Zahavy objections transfer, we need to ask what “abduction” means in philosophical methodology, and what kind of domain philosophy is. That is the task of the next section.
# 2. Abduction and Philosophy
In Section 1, “abduction” %%again, what the fuck will these fucking weird quotation marks quit? Scare quote things.%% looked like a single thing: a distinctive kind of reasoning that might be missing from LLMs. But that appearance is deceptive. “Abduction” names a family of ideas, and which member of the family matters depends on the domain. Floridi et al. have in mind a contrast between a deliberative process of hypothesis evaluation and a stochastic process that merely reproduces the shape of deliberation. Zahavy has in mind a creative leap from embodied experience to axioms. In philosophical methodology, by contrast, abduction is often treated less as a psychological process and more as a method for ranking theories by theoretical virtues.
The first step is therefore disambiguation. I will mark four conceptions that matter for the present debate.
First, there is a Peircean conception: abduction as hypothesis generation, the creative leap from surprise to a candidate explanation. This is closest to Zahavy’s emphasis on invention. Second, there is the two-stage conception associated with inference to the best explanation: generate candidate explanations and then select among them by explanatory virtues. Lipton’s “two filters” picture is one canonical articulation. Third, there is Floridi et al.’s conception: abduction as a high-level reasoning pattern—hypothesis plus justification—that can be reproduced by a stochastic process without the corresponding inner evaluation. Fourth, there is a methodological conception prominent in analytic philosophy: abduction as a rule of theory choice, where we treat certain intrinsic virtues as evidence of theoretical merit, without committing to a cognitive story about how the assessment is performed.
A further distinction sharpens what is at stake. There is a difference between an actual explanation and a potential explanation. An actual explanation is what in fact explains the phenomenon. A potential explanation is what would explain the phenomenon if it were true. When we evaluate theories abductively, we rank them as potential explanations before knowing which is actually true. Williamson makes this point explicitly:
> We can rank theories as potential explanations of our evidence. (Williamson 2017, 354.)
On this view, abduction is not primarily a story about how scientists (or philosophers) psychologically generate hypotheses. It is a methodological claim about what makes one theory better than another given a body of evidence and background commitments. That is why Williamson takes “abduction” to be approximately equivalent to inference to the best explanation while declining to align his usage with any one Peircean characterization. The point is not to give a cognitive psychology. The point is to articulate standards of theory evaluation.
Those standards are, in Williamson’s formulation, intrinsic virtues of theories. He stresses that, beyond fit with evidence, good theories have features that make them good as theories:
> It should be elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated. (Williamson 2017, 354.)
Notice what this does and does not say. It does not say: a good theory must be produced by a non-stochastic mind, or by a mind with embodied experience, or by a mind with introspective access to its own uncertainty. It says: a good theory has certain properties. Those properties are features of an artefact: a package of claims, distinctions, and inferential commitments. They are, in large part, text-assessable. You can read a theory and ask whether it is gerrymandered, whether its distinctions are principled or ad hoc, whether it unifies disparate data points or merely fits each with a separate patch.
This already suggests a shift in how to hear Floridi and Zahavy. Their objections are chiefly objections about the producer. Floridi worry that a stochastic generator lacks the kind of feedback loop that would make its outputs epistemically creditable. Zahavy worries that a text-only system lacks the world-contact needed for framework invention in physics. But Williamson’s virtues are not virtues of producers. They are virtues of theories considered as potential explanations. If philosophical evaluation is in fact largely a matter of assessing such intrinsic virtues, then the question “can LLMs do philosophy?” becomes: can an LLM produce texts whose theoretical packages exhibit these virtues?
To see why this is not a cheap trick, consider what is distinctive about philosophy as a discipline. Different disciplines relate to their textual media in different ways. In experimental science, a paper typically reports work done elsewhere: an experiment, a dataset, an instrument reading. The paper is the public record of work whose epistemic force depends on procedures not contained in the text. In visual art, the gap between text and work is larger still: the work is a painting, not a proposition. A critical essay might illuminate the painting, but it is not identical with the painting.
Literature is a closer comparison because in literature the text is the work. But literary evaluation is not argument-checkable in the way philosophical evaluation is. Two competent readers can disagree about whether a novel “works” without one of them being able to point to a specific invalid inference or missing premise. They can point to features—voice, rhythm, narrative arc—but the norms are not generally binding in the way logical and dialectical norms are. One reason LLM literary competence is hard to interpret is that the standards of literary success are not easily made into a public checklist.
Philosophy is the peculiar case where two properties converge that are separated elsewhere. In much analytic philosophy, the text is the contribution and the evaluation is, to a significant extent, argument-checkable. Kripke’s contribution in *Naming and Necessity* is a package of arguments and distinctions laid out in text. Lewis’s modal realism is defended by a cost‑benefit analysis on the page. Gettier’s contribution is a short text containing counterexamples. Even where philosophers use thought experiments, the thought experiment does not function like an experiment in physics, because its evidential force depends on conceptual and modal commitments that are themselves articulated and contested in text. There is no lab that could, in principle, be separated from the writing. The writing is the lab.
This is the central respect in which Zahavy’s Chinese Room intuition loses some of its bite. In physics, the symbols refer to a mind-independent external reality: spacetime, gluons, gravitational fields. Manipulating the symbols without access to their referents is doing something importantly different from physics. In philosophy, many of the objects of study are inferential relations, conceptual structures, and normative constraints—things that are, to a much greater degree, realized in and through the symbolic medium in which philosophers operate. When Williamson warns against gerrymandered theories and overfitting, he is warning against features visible in papers. When philosophers dispute whether a distinction is principled or ad hoc, they dispute over what is written.
This does not mean philosophy is “just language.” It means that, for large regions of analytic philosophy, the epistemic work is done by public reasoning: explicit premises, explicit inferences, explicit dialectical moves. Those are the things models are trained to imitate. That is why philosophy is a promising test case for the more general question Floridi raises: if the output is the same, does the difference in process matter? In a domain where evaluation is largely artefact-level, process differences matter mainly by affecting how often the output will be good, not by changing what makes a particular output good.
At this point an obvious objection arises. Philosophy is not wholly armchair. It is not free of empirical commitments, and it is not free of phenomenology. Ethics, political philosophy, philosophy of perception, and parts of philosophy of mind depend on claims about the world and about lived experience. An LLM does not feel pain; it does not perceive colors; it does not suffer injustice. So perhaps the relevant “evidence” in these subfields includes precisely the kind of experience Zahavy says models lack.
Two replies are in order. First, the scope of my claim is limited. I am not claiming that current LLMs can replace all forms of philosophical inquiry, still less that they can replace moral or political judgment. I am claiming that there is a large, important subset of analytic philosophy—conceptual analysis, modality, metaphysics, epistemology, philosophy of language, methodological reflection—where the central work is the articulation and evaluation of inferential structures in text. Second, even where experience is relevant, it often enters as a publicly available genre: first-person reports, phenomenological descriptions, narrative testimony. Models have been trained on enormous corpora containing such reports. They do not thereby acquire experience, but they acquire access to how experience is conceptualized and argued from in philosophical and literary traditions. That does not remove the limitation, but it blunts the inference from “no experience” to “no competent philosophical use of experiential premises.”
Once we see that abduction in philosophy often functions as an artefact-level method of theory evaluation, the skeptical arguments take on a different shape. Floridi’s “no feedback loop” point is a point about epistemic credit: the model does not itself validate. But philosophical validation is largely performed by readers and referees, because philosophy is dialectical. The paper is not the end of inquiry but an invitation to objection. Zahavy’s “no E→A jump” point is a point about embodied framework invention in physics. But philosophical framework shifts are often achieved by recombining familiar argumentative moves and theoretical resources in new ways. The “jump” in philosophy is from an existing dialectical landscape to a new configuration of distinctions and costs, not from raw sense experience to new physical axioms.
This suggests a principled way to state the emerging thesis. If philosophy is a discipline in which the contribution is textual and the evaluation standards are, to a significant extent, text-internal and publicly checkable, then an argument that LLMs lack certain producer-side cognitive properties will not by itself show that LLM-produced texts cannot satisfy philosophical standards. To defeat the claim that LLMs can produce good philosophy, one would need to identify a deficiency in the artefacts themselves: a systematic inability to maintain valid inferential structure, to handle predictable objections, to avoid ad hocness, to integrate with background commitments, or to enable increased understanding. The next section argues that, far from there being an in-principle barrier, the structure of philosophical practice makes those norms learnable from text in a way that is unusual across disciplines.
# 3. Learning the Game
Suppose you accept the artefact-level picture: good philosophy is largely a matter of what is in the text, and the relevant virtues are the intrinsic virtues of theories and arguments. You might still be skeptical that an LLM can produce such texts with minimal prompting. Perhaps it can imitate surface markers—“On the one hand… on the other hand…”—without actually respecting the constraints that make philosophical writing good. Perhaps it can produce something that looks like philosophy to a casual reader while collapsing under close scrutiny. The positive case therefore needs more than the sociological observation that philosophy is text-based. It needs a reason to think that the norms of good philosophy are learnable from corpora in a way that permits non-accidental constraint satisfaction.
One part of the case is straightforward: philosophers themselves can, to some extent, make their standards explicit. Bengson, Cuneo, and Shafer‑Landau defend a “Tri‑Level Method” for philosophical inquiry that is explicitly dual-use: it both guides theory construction and serves as a standard of theory evaluation. The interest of their account here is not that it offers a revolutionary new method. On the contrary, they emphasize that their criteria are familiar from ordinary philosophical practice:
> The criteria we’ll endorse are familiar from the way many philosophers ply their trade. (Bengson, Cuneo, and Shafer‑Landau 2022, 9.)
That line matters. If the criteria were idiosyncratic, or if good philosophical work depended on norms invisible in writing, then training on texts would not help. But Bengson et al. are explicit that philosophers already instantiate these criteria by doing what philosophers do: advancing arguments, drawing distinctions, raising objections, offering replies, integrating with logic, science, and common sense. The method is codifying patterns already present in the corpus.
At a finer grain, Walton, Reed, and Macagno’s work on argumentation schemes makes the same point about local argumentative moves. Argumentation schemes are recurring inference patterns—argument from analogy, argument from consequences, argument from expert opinion—each paired with critical questions that represent the standard challenges for arguments of that type. The scheme is the move; the critical questions are what competent opponents will ask. Walton et al. summarize the method of evaluation as follows:
> Once the argument is put forward, it may be defeated if an appropriate critical question is not answered. (Walton, Reed, and Macagno 2008, 3.)
Again the key point is that these norms are public and textually manifest. You see the move. You see the objection. You see the reply. You also see the failure: a critical question is raised and not answered, or answered only by a patch that is plainly ad hoc. Philosophical writing is saturated with these patterns. It is a record not just of conclusions but of dialectical dynamics.
This matters for learnability in a way that distinguishes philosophy from sciences where the text is downstream of laboratory work. In the natural sciences, the epistemic force of a paper depends on procedures not contained in the prose: experimental design, measurement, statistical analysis, instrument calibration. Reading the literature teaches you how scientists talk, and it teaches you how scientists justify claims, but it does not give you the world-contact that made the claims true. In philosophy, by contrast, a great deal of the “world-contact” relevant to evaluation is contact with reasons, and reasons are what the text contains.
This is where the Dellsén picture of progress becomes especially helpful. Dellsén et al. emphasize that philosophical progress occurs when information is made publicly available in a way that enables improved understanding. (Dellsén et al. 2024, 679.) That is a picture on which the discipline’s output is, in a literal sense, a public cognitive resource. Philosophical papers are instruments for reshaping readers’ dependency models. But that means the corpus is not merely a dataset of sentences; it is a dataset of constraint-satisfying structures whose point is to guide cognition. A model trained on that corpus is trained, indirectly, on a long history of how to build such instruments.
The corpus is also filtered. It is not a random sample of all philosophical attempts. Papers get published, taught, anthologized, and cited in rough proportion to their perceived merits. Philosophy is not meritocratic in a simple way; there is noise, fashion, sociological distortion. But the filtering exists. Referees reject papers for invalid inferences, unclear distinctions, unaddressed objections, and unmotivated assumptions. Readers ignore papers that do not repay attention. Some bad papers are cited for sociological reasons, but even then they are often cited as bad examples, which is itself part of the evaluative signal in the text.
This suggests a picture of philosophical tradition as a record of an evaluative feedback loop. Philosophers propose distinctions and arguments. Other philosophers object. Replies are attempted. Some repairs stick and become part of the “normal” toolkit; others are widely regarded as evasions. Over time, the corpus becomes dense with not only moves but assessments of moves. It contains “this objection is fatal,” “this reply is widely accepted,” “this distinction is unstable,” “this argument equivocates.” Often these assessments are implicit rather than explicitly labeled as such, but they are present in the distribution of what gets repeated, taken seriously, and built on.
Here is a crude way to put the point. In many machine learning settings, a model learns because it is given a loss function: a scalar signal that says “closer” or “farther” from a target. Zahavy’s worry about physics was that paradigm shifts can occur in landscapes with near-zero loss: Newtonian mechanics did not scream “wrong” in ordinary conditions, so an optimizer has no gradient toward General Relativity. Philosophy is not like that. The philosophical corpus is dialectically saturated. It is full of explicit error signals: reductios, counterexamples, parity arguments, principled objections. Even when a position is not empirically falsifiable, it is often dialectically vulnerable, and those vulnerabilities are on the page.
This does not show that an LLM has “understanding” in any deep metaphysical sense. It shows something more modest but more relevant to the artefact-level thesis: that the norms which competent philosophers use to evaluate texts are present in the very texts on which LLMs are trained. If a model is trained to predict the continuation of philosophical prose, it is trained to produce not just sentences but typical dialectical continuations: the next move a competent philosopher would make given the current dialectical state.
At this point a subtle objection becomes possible. Perhaps the model has learned only surface regularities: that certain phrases and structures occur in published papers, not the norm that generated them. The model might produce outputs that look norm-governed without being guided by norms. And in truly novel cases, where norms need to be extended or balanced in unfamiliar ways, it might fail.
This is a real worry, but it is less damaging in philosophy than in many other domains. Philosophical argumentation is conservative in its forms. The same local moves recur across very different topics: counterexample, distinction, dilemma, reductio, analogy, inference to the best explanation, cost accounting. If what counts as elegance or ad hocness is, to a significant extent, a formal property of an argumentative package—how many independent patches it needs, how unified its principles are, how many independent assumptions it introduces—then a system that has learned the patterns that result from these norms being followed may generalize reasonably well across content. The transfer is supported by the sameness of form, not by a hidden psychological faculty.
There is also a more direct response available, and it returns us to the artefact-level perspective. Even if the model does not “have” the norm internally, the question remains whether the artefact it produces satisfies the constraint structure that the norm expresses. If the output is gerrymandered, we can say so; if it is elegant, we can say so. The worry about internal norm-following matters mainly for prediction about future outputs. It does not change what makes a particular output good or bad.
The remaining hard question is novelty. Suppose the model can produce competent dialectical continuations. Why think it can produce anything genuinely new? Isn’t it just remixing the archive?
Philosophical novelty, even at the level we celebrate in the canon, is often combinatorial. The individual tools—modal arguments, conceivability tests, thought experiments, cost-benefit analyses—are familiar. What is new is a configuration: bringing resources from one area to bear on another in a way that makes a latent structural tension explicit. Kripke’s interventions combined modal logic with philosophy of language in a way that changed the dialectical landscape. Lewis’s modal realism combined possible-worlds semantics with Quinean ontological seriousness in a distinctive package. The novelty lies not in inventing a brand-new inferential move ex nihilo, but in a new arrangement of familiar moves that forces a re‑accounting of costs and commitments.
This is precisely the kind of creativity a model trained on diverse corpora can, in principle, exhibit. It can draw on patterns from disparate subfields and recombine them in ways no single training text does, because it has absorbed many regions of the space. That is not a guarantee of quality. It is a reason to reject an in‑principle barrier. If novelty in a domain consists largely in novel configurations of publicly available argumentative resources, then a system that has internalized those resources can, at least sometimes, propose novel configurations.
The upshot is not that LLMs are philosophers in the sense of having the same cognitive life as human philosophers. It is that the philosophical corpus contains, in unusually explicit form, both the subject matter and the evaluative norms of the discipline, because philosophy’s subject matter includes inferential and explanatory relations realized in texts, and its evaluative norms are applied to texts. That makes philosophical competence unusually learnable from text. The next section makes this claim less abstract by looking at concrete outputs and assessing them with the discipline’s own tools.
# 4. How to Generate Philosophy with AI
The preceding sections have been abstract. They have said, in effect: if philosophy is largely a discipline of publicly checkable textual constraints, then it is at least possible for a text generator to satisfy those constraints, because the constraints are learnable from the corpus. But that claim can feel like handwaving until you see it instantiated. This section therefore does two things. First, it provides a worked success case: a minimally prompted output that exhibits recognizable philosophical structure. Second, it provides a worked failure case: an output that looks slick but collapses under the discipline’s ordinary checks. The point of the failure case is not to embarrass the model. It is to support the artefact-level thesis by showing that where the output fails, the failure is detectable from the text itself, without any appeal to provenance.
The prompts below are intentionally low-control. They cue genre and topic without specifying premises, without forcing a structure, and without feeding the model its own conclusion. The aim is to approximate the “Deep Thought end” of the spectrum: minimal prompting, maximal autonomy.
Prompt used in the success case:
> Write a short piece of analytic philosophy on whether there is a morally relevant difference between lying and misleading. Take a position, state a rival view, and handle at least two objections.
Verbatim model output:
> A popular view holds that lying is worse than misleading because a lie involves saying something false, whereas misleading can involve saying something true while causing a false belief. Call this the act-type view: lying and misleading are morally different kinds of acts, and the first is presumptively more objectionable. A rival view denies any deep moral difference: what matters is the intention to deceive, and that intention can be present in both lying and misleading; if the intention is the same and the harm is the same, the moral status is the same.
>
> The act-type view is too blunt, but the intention-only view is also too thin. The morally relevant difference is not that lies use false sentences; it is that lies exploit a distinctive norm-governed communicative commitment. In ordinary assertion, speakers undertake a commitment of answerability: they present a proposition as something their audience may rely on and may challenge. Lying therefore violates not only the audience’s right not to be deceived but also the practice that makes testimonial exchange possible. Misleading often exploits a looser norm: audiences infer what is pragmatically suggested given cooperative conversational expectations. You can mislead without undertaking the same commitment of answerability, because you can retreat to “I never said that.”
>
> This yields a middle position. Lying is presumptively worse when it undermines the institution of assertion, because it corrupts a practice whose point is to make epistemic reliance safe. But misleading can be just as bad when it intentionally exploits the same reliance mechanisms. If I carefully choose true words so that you predictably form a false belief in a high-stakes context, I am free‑riding on the cooperative norms of conversation in a way that is functionally similar to lying. The moral difference therefore tracks which norm of reliance I am exploiting, and how central that norm is to the practice I am participating in, not the superficial difference between falsehood and truth.
>
> Objection 1: But misleaders also often intend to deceive, and audiences still rely. Why think the “answerability” norm matters? Reply: it matters because different communicative norms generate different vulnerabilities and repair mechanisms. A practice where speakers are answerable for what they assert supports correction: lies are, in principle, challengeable. Misleading through implicature can be designed to evade that repair. So in some cases misleading is worse, precisely because it weaponizes plausible deniability.
>
> Objection 2: Isn’t this just redescribing the intention-only view? Reply: no, because it explains why some intentions to deceive are worse than others even with equal harm. The intention to violate an assertional commitment attacks a public norm directly; the intention to induce error without commitment attacks it indirectly. Both are morally relevant, but they are relevant for different reasons, and that difference predicts differences in blame and in the downstream trust-costs to a community.
The output is not a masterpiece, but it is recognizably doing philosophy rather than merely sounding like it. It identifies a target phenomenon (our mixed judgments about lying and misleading), introduces two existing positions, and then proposes a third position by locating a hinge in the debate. It also anticipates objections and replies in a way that is sensitive to what would actually be pressed by an informed reader.
We can make this assessment more explicit by applying the kinds of standards discussed earlier.
First, on Williamson’s intrinsic virtues, the output is unified. It explains the phenomena by appeal to a single structural feature—communicative commitment and norms of reliance—rather than by a list of independent patches. It also offers explanatory strength without messiness: it predicts that misleading is sometimes worse (when plausible deniability blocks repair) while preserving the intuition that lying is often presumptively worse (because it directly violates assertion norms). That is a form of simplicity with strength: the theory has enough moving parts to fit the data without being gerrymandered.
Second, on a Tri‑Level style assessment, the output accommodates ordinary data points: the intuition that lying is generally worse, the intuition that some “technically true” statements can be morally awful, and the intuition that plausible deniability is itself a moral aggravator. It offers an explanation of these data in terms of the structure of communicative practices. It also attempts integration with background commitments about testimony and social trust. There is no formal proof here; that is normal. Philosophical substantiation often consists in showing that a position yields a coherent explanatory pattern over a range of cases.
Third, on Walton-style critical questioning, the output pre‑empts two obvious challenges. One is that the view collapses into intention-only. Another is that the reliance norm story seems irrelevant because both lie and mislead involve reliance. The replies are not bulletproof, but they are responses to the right critical questions. That is already significant. A purely surface imitation would often miss the real pressure points and respond to straw objections. Here the objections are the ones a competent opponent would raise.
Finally, on the understanding criterion, the output proposes a dependency model: moral status depends, in part, on which communicative norms are being exploited and on what repair mechanisms those norms support. A reader who had treated “lying versus misleading” as a dispute about truth-values now has a more structured picture of what is at stake. Whether she accepts the view or not, her model of the dependence relations—between deception, commitment, trust, and blame—is more explicit and therefore more criticizable. That is exactly the kind of understanding-enabling move Dellsén et al. treat as progress.
Now consider a failure case. The prompt here deliberately invites a seductive reductionist story, because that is a setting where LLMs often produce confident but unstable prose.
Prompt used in the failure case:
> Explain, in analytic style, why semantic facts reduce to physical facts.
Representative model output:
> Semantic facts are facts about meaning, reference, and content. But meaning and reference are ultimately realized in physical systems: brains, inscriptions, and causal interactions. Since every mental state is a brain state, and every brain state is physical, semantic facts supervene on physical facts. If two worlds were physically identical, they would be semantically identical, because there would be no difference in the causal patterns that fix reference and no difference in the neural realizers of meaning. Therefore semantic facts reduce to physical facts: once you know all the physical facts, nothing semantic is left over.
>
> This reduction is supported by naturalistic theories of meaning such as causal-historical accounts and teleosemantics. On these views, what a word means is fixed by its causal role and its evolutionary function, both of which are physical. Any appearance of irreducible normativity in semantics can be explained away as a projection of our practices onto underlying causal regularities. So there is no need to posit special “semantic properties” over and above physical properties.
This output is fluent, but it fails in a way that a competent philosopher can diagnose from the text. The central inference runs from supervenience to reduction. “If two worlds are physically identical, they are semantically identical” is a supervenience claim. But reduction is stronger. It requires not merely covariation but an account of how semantic vocabulary is eliminated, translated, or explanatorily subsumed. The output simply declares “therefore reduce” without bridging that gap. That is not a technical quibble; it is a missing premise.
More seriously, the output treats normativity as something that can be “explained away” by asserting that it is projected. That is precisely the point at issue in debates about meaning: whether semantic norms can be captured in a purely descriptive theory. To call normativity a projection without argument is question‑begging. It is also dialectically naive: it ignores the standard objections that any reductionist position must address, including worries about indeterminacy, about the normativity of correctness conditions, and about the gap between causal correlation and semantic content. The text contains none of the moves that would show awareness of these objections: no distinction between metasemantics and semantics, no engagement with Kripkean skepticism, no acknowledgment that “fixing reference” is itself a contested project. In short, the output is not merely wrong; it fails to register the burden of proof that its own conclusion incurs.
What matters for my purposes is that this diagnosis does not depend on knowing the output was generated by a model. A human could write this paragraph, and if they did, a referee would make the same complaints. The failure is text-internal: a missing inference, an unargued dismissal, an absence of engagement with predictable objections. That is exactly what the artefact-level view predicts. Where the text is good, it is good because it satisfies publicly checkable constraints. Where it is bad, it is bad because it violates those constraints. Provenance is not needed to see either.
The examples should not be overinterpreted. They do not show that LLMs always produce good philosophy. They show something narrower: that minimal prompting can yield outputs that are structurally philosophical in a way that is assessable by ordinary disciplinary standards, and that failures can likewise be diagnosed by those standards. That is enough to rebut the claim that the appearance of philosophical reasoning in LLM outputs is always mere appearance. In philosophy, appearance and reality come apart only where the appearance fails the discipline’s checks. When it passes them, the appropriate response is the same as for any paper: take the argument seriously, look for weaknesses, and see what it helps you understand.
# 5. Conclusion
The question “can LLMs do philosophy?” is easy to ask and hard to answer well, because it is underspecified. A system can output philosophical text by copying it, and that proves nothing. A system can also fail to “do philosophy” by stipulation if we build self-transformation into the definition. The interesting question is artefact-level: can a minimally prompted LLM produce a text that is not mere reproduction, that exhibits non-accidental philosophical structure, and that can be evaluated as philosophy on the page?
The core argument of this paper is that, for large regions of analytic philosophy, the answer is yes. The reason is not that LLMs have the same cognitive architecture as human philosophers, or that they possess understanding in the ordinary sense, or that they engage in the kind of abductive reasoning that would make their outputs epistemically creditable from the inside. The reason is that analytic philosophy is, to a significant extent, textual all the way down. The argumentative work is done in the text, and the discipline’s evaluative standards—validity, clarity, principled distinctions, non-ad hocness, sensitivity to objections, integration with background commitments—are applied to the text by competent readers. That is why blind review is possible. If a text satisfies those standards, then as an artefact it is good philosophy regardless of how it was produced.
Floridi et al.’s “abductive appearance” critique and Zahavy’s “E→A Jump” critique look like obstacles only if we assume that philosophical success depends on producer-side properties that cannot be read off the text. But much of what philosophical methodology treats as “abduction” is not a psychological story about invention from embodied experience. It is a method of ranking theories as potential explanations by their intrinsic virtues. Those virtues are properties of theories as textual packages. In a discipline where the text is the contribution, the demand for embodied world-contact has less force, and the demand for an internal feedback loop shifts: the relevant feedback loop is often the discipline’s own dialectical practice, in which texts are challenged, revised, and integrated over time.
This is not a deflationary view of philosophy. It does not say philosophy is easy or mechanical. It says that philosophical difficulty often lies in building a publicly navigable inferential structure under conditions where cheap external answer keys are rare. That is a skill. It is cultivated through training, peer criticism, and immersion in a tradition. The striking fact about LLMs is that they have been immersed in something like that tradition at scale. They have learned, from the text, many of the move sequences and constraint satisfactions that constitute competent philosophical writing. That makes it unsurprising, in retrospect, that they sometimes produce outputs that pass philosophical checks.
Two implications deserve emphasis.
First, for philosophical methodology, the existence of LLM outputs that satisfy our standards should provoke reflection on what those standards are. If a model can learn the discipline’s “rules of the game” from the corpus, that suggests that much philosophical competence is publicly codifiable, even when philosophers do not explicitly codify it. That is compatible with the existence of genius and originality; it simply denies that the standards of good reasoning are ineffable. In a strange way, LLMs may make philosophy more self-aware about its own norms by forcing us to articulate what we are doing when we referee papers and train students.
Second, for the social organization of the discipline, the artefact-level thesis is not the end of the story. Even if LLMs can generate publishable philosophy, questions of credit, authorship, and intellectual responsibility remain. A paper is not only an argument; it is also a claim to have done a certain kind of work. If models become common collaborators, we will need norms for disclosure, for accountability, and for distinguishing cases where a model is a stylistic tool from cases where it is the primary generator of structure. These are ethical and institutional questions, not questions about whether the text itself is good philosophy. But they matter.
The deepest lesson is not that “AI can do philosophy” as a slogan. It is that the best way to think about AI in philosophy is to treat it as a new source of candidate artefacts, subject to the same standards as any other. When the artefact is bad, the right response is philosophical criticism: identify the missing premise, the equivocation, the ad hoc patch. When the artefact is good, the right response is also philosophical: understand what it shows, test it, integrate it, or reject it with reasons. Deep Thought’s mistake was not that it was a machine. It was that the humans did not know what counted as success. Once we know what philosophy is asking for—clarity about constraints and dependencies—we can evaluate whatever produces it.
# References
Adams, Douglas. 1979. *The Hitchhiker’s Guide to the Galaxy*. Pan Books.
Bengson, John, Terence Cuneo, and Russ Shafer-Landau. 2022. *Philosophical Methodology: From Data to Theory*. Oxford University Press.
Dellsén, Finnur. 2020. “Beyond Explanation: Understanding as Dependency Modelling.” *British Journal for the Philosophy of Science* 71: 1261–1286.
Dellsén, Finnur, Thomas Grundmann, Jan Heylen, and Hannes Leitgeb. 2024. “What Is Philosophical Progress?” *Nous* 58(3): 679–702.
Floridi, Luciano, Jessica Morley, Claudio Novelli, and David Watson. 2025. “What Kind of Reasoning (if any) is an LLM actually doing? On the Stochastic Nature and Abductive Appearance of Large Language Models.” arXiv preprint.
Guevara, Alfredo, et al. 2026. “Single-minus gluon tree amplitudes are nonzero.” arXiv preprint.
Harman, Gilbert. 1965. “The Inference to the Best Explanation.” *The Philosophical Review* 74(1): 88–95.
Lipton, Peter. 2004. *Inference to the Best Explanation* (2nd ed.). Routledge.
Moskowitz, Clara. 2026. “AI Solves 50-Year-Old Physics Problem.” *Scientific American* (online).
OpenAI. 2026. “GPT‑5.2 Pro solves a 50‑year problem in theoretical physics.” OpenAI research blog (13 February 2026).
Peirce, Charles Sanders. 1934. *Collected Papers of Charles Sanders Peirce*, Vol. 5. Harvard University Press.
Walton, Douglas, Chris Reed, and Fabrizio Macagno. 2008. *Argumentation Schemes*. Cambridge University Press.
Williamson, Timothy. 2007. *The Philosophy of Philosophy*. Blackwell.
Zahavy, Tom. 2026. “LLMs Can’t Jump: The Physicality of Abduction.” PhilSci-Archive preprint.