# 0. Introduction
> "Forty-two," said Deep Thought, with infinite majesty and calm.
> It was a long time before anyone spoke.
> Out of the corner of his eye Phouchg could see the sea of tense expectant faces down in the square outside.
> "We're going to get lynched aren't we?" he whispered.
> "It was a tough assignment," said Deep Thought mildly.
> "Forty-two!" yelled Loonquawl. "Is that all you've got to show for seven and a half million years' work?"
> "I checked it very thoroughly," said the computer, "and that quite definitely is the answer. I think the problem, to be quite honest with you, is that you've never actually known what the question is."
-- Douglas Adams, *The Hitchhiker's Guide to the Galaxy*
In *The Hitchhiker's Guide to the Galaxy*, humanity asks an AI to do some philosophy. A computer named Deep Thought is constructed and instructed to provide "The Answer to the Ultimate Question of Life, the Universe, and Everything" (REF). Humanity builds this computer only to receive the answer '42'—an answer which, while apparently correct (at least according to Deep Thought), means next to nothing at all due to humanity's failure to know what the Ultimate Question in fact is.
Deep Thought is fiction, but the underlying worry is not. If you do not know what you are asking, even a correct answer can be useless; and if you do not know what would count as success, the question "can it do X?" collapses into a verbal dispute. This matters now because large language models are no longer confined to producing plausible-sounding prose. In February 2026, OpenAI's GPT-5.2 was given a problem in theoretical physics concerning a class of gluon scattering amplitudes whose complexity had defeated calculation by hand. The human physicists on the team had worked out a few simple cases, but the expressions grew superexponentially; the model conjectured a general formula and completed a formal proof of its validity, overturning a forty-year-old assumption that the amplitudes in question were zero (Guevara et al. 2026). The point of mentioning this is not that philosophy should aspire to be theoretical physics. It is that AI outputs are now being evaluated by communities with extremely demanding public standards, and sometimes they pass. The question I want to address is whether philosophy could be among the disciplines for which this is true.
Stated baldly, "Can LLMs do philosophy?" is not yet a good question. It invites answers that are too easy to be informative. At one end, an LLM can simply copy *Philosophical Investigations* word for word; that gives you a philosophical text only in the thin sense that it is already a philosophical text. Given enough time, a room of monkeys at typewriters could produce the same book; but no one would take that as evidence that the monkeys had become philosophers. At the other end, if philosophy is essentially a practice of self-transformation - if what makes an activity philosophical is something that happens in the practitioner rather than anything that can be assessed in what she produces - then no text-generating system could do philosophy regardless of what it outputs, and the discussion ends by stipulation. The interesting territory lies between these extremes. The question worth asking is whether, given philosophically minimal prompting, an LLM can produce a piece of writing that is not a mere reproduction of existing text, that exhibits non-accidental philosophical structure, and that can be assessed as philosophy on the page. I want to argue that it can, and that the argument turns on a feature of analytic philosophy that distinguishes it from the empirical sciences: in much analytic philosophy, the text is the contribution, and the discipline’s evaluative standards are available in the text.
To make that claim substantive, "good" and "novel" need unpacking. A competent summary of an existing debate might be well-structured, accurate, and clear without putting any reader in a position she was not already in; it would be good in one sense without amounting to a contribution. What would be interesting is a text that improves a reader's understanding of some phenomenon - that helps her grasp dependence relations she had not previously grasped, not merely that gives her access to what others have already grasped. To make this distinction precise, I adopt Dellsén et al.'s account of philosophical progress, *Enabling Noeticism*:
> Enabling Noeticism: The discipline of philosophy makes progress regarding some phenomenon to the extent that philosophical research puts people in a position to increase their understanding of that phenomenon. (Dellsén et al. 2024, p. 679)
On this view, the central evaluative notion is understanding, and understanding is a matter of degree. Very roughly, an agent understands X to the extent that she accurately and comprehensively represents the network of dependence relations in which X stands (or fails to stand) to other things. Two dimensions matter. A representation can be more or less accurate, depending on whether the dependence relations it encodes actually obtain; and it can be more or less comprehensive, depending on how much of the relevant network it captures rather than omitting. Understanding is, in this sense, epistemically undemanding (it does not require knowledge or justification), while remaining robustly factive: what matters is how well the representation matches the dependency structure of the world.
The Gettier case is an example. The justified true belief theory represents knowledge as depending on three conditions—truth, belief, and justification—and on nothing else. Gettier's counterexamples showed that the 'nothing else' clause was wrong: even justified true beliefs can fail to be knowledge, so knowledge depends on something further. Whatever one thinks about subsequent attempts to supply a fourth condition, Gettier's paper plausibly counts as progress in Dellsén et al.'s sense. It put readers in a position to represent more accurately and more comprehensively what knowledge does and does not depend on, even though it did not supply the missing positive account.
On this account, a reader's understanding of a phenomenon consists in her mental model of the dependence relations in which that phenomenon stands. A philosophical text contributes to progress by putting readers in a position to improve that model - to make it more accurate, more comprehensive, or both. The evaluative question, then, is whether a given text enables such improvement in its readers. That question can be answered by examining the text: does it track genuine dependence relations? does it capture structure the reader had missed? If so, it is a vehicle for progress, regardless of how it was produced.
Analytic philosophy is, by and large, a text-based discipline. Philosophical contributions are written artifacts - arguments, distinctions, counterexamples, theoretical frameworks - and their assessment is likewise text-based. In experimental science a paper typically reports work done elsewhere; in much analytic philosophy the argumentative work is done on the page. Referees read, test inferences, press for missing premises, check whether rival positions are treated fairly, and ask whether the view survives predictable objections. Blind review exists in philosophy for exactly this reason: the discipline's evaluative norms apply to what is in the text, not to the biography, psychology, or social position of its author. If this is correct, then the question of whether LLMs can contribute to philosophical progress should be posed at the level of the artifact - the text itself, assessed independently of its origin: can an LLM produce a text that, when read by an informed reader, puts that reader in a position to understand better?
The word *produce* needs qualifying, because using an LLM can mean many different things. At one end, a human philosopher does the philosophical work and uses the model as a transcription device, a stylistic editor, or a convenient assistant for paraphrase and summarization. At the other end, we have something close to Deep Thought: philosophically minimal prompting - a question, a topic, a request for a familiar kind of move - elicits an extended piece of writing whose substantive structure is not supplied by the user. My focus is on that far end of the spectrum. The question is whether, given minimal prompting, an LLM can produce text with genuine philosophical structure - text that puts a reader in a position to understand better.
Two recent arguments provide the strongest case for skepticism about this possibility. Floridi et al. claim that LLM outputs exhibit, at best, an *abductive appearance*: they mimic the surface form of inference without performing the kind of reasoning that would warrant trusting the result. Zahavy argues, in a related spirit, that genuine abduction requires a leap from experience to explanatory axioms that a purely text-trained system cannot make. These arguments target real architectural limitations, and I do not dispute them as claims about LLM cognition. The question is whether the conception of abduction they presuppose is the right one for philosophy as a text-based practice aimed at improving understanding, rather than for empirical science as a practice whose textual record is downstream of laboratory work. Both arguments are developed with the physical sciences as their reference case - the discipline in which, as it happens, LLMs have most recently shown themselves capable of novel work.
The rest of the paper develops that case. Section 1 sets out Floridi et al.'s and Zahavy's objections and clarifies what they would show if philosophy were relevantly like the physical sciences. Section 2 argues that "abduction" is used in several different ways across the literatures at issue, and that Williamson's abductive picture of philosophy is best understood as a method of theory evaluation by intrinsic virtues—virtues that are, again, assessable in texts. Section 3 makes the positive argument: the norms of philosophical practice are publicly codifiable and textually manifest, and the philosophical corpus is itself the record of an evaluative feedback loop that a model can learn from. Section 4 provides worked examples: minimal prompting can elicit outputs with genuine philosophical structure, and these outputs can be evaluated by the discipline's own standards.
---
# 1. What LLMs Aren't Doing
# What LLMs Aren't Doing
In this section I look at two reasons we might think that large language models (LLMs) cannot write good philosophy. Both have to do with questions over whether LLMs can perform abductive reasoning. %%%%
Abductive reasoning begins with something in need of explanation — a surprising observation, an anomalous result, a phenomenon that existing theories do not predict. The reasoner generates a candidate hypothesis: one that, if true, would account for what has been observed (Peirce, 1934). But generation is only the first move. Harman (1965) introduced a comparative dimension: we do not simply generate a candidate and commit to it; we generate several, then assess which would, if true, best explain the evidence — weighing simplicity, coherence with background knowledge, and explanatory scope. This is inference to the best explanation (IBE), and Lipton's formulation makes the structure explicit:
> "Given our data and our background beliefs, we infer what would, if true, provide the best of the competing explanations we can generate of those data (so long as the best is good enough for us to make any inference at all)." (Lipton, 2004, p. 56)
The phrase "the competing explanations we can generate" is doing work here: the inference operates not over the space of all logically possible explanations but over a constrained set of candidates, and the quality of the inference depends on the quality of the candidates generated.
IBE thus involves two stages that can come apart. Lipton identifies them as two 'filters': a first that narrows the space of possible explanations to a short list of plausible candidates — the 'live options' — and a second that selects from among those candidates by explanatory virtues (Lipton, 2004). A system might manage the second filter without the first — capable of evaluating candidates when they are provided but unable to generate the short list. Or it might manage the first without the second — able to produce candidates but not to assess which is best. The distinction matters for what follows: both arguments in this section diagnose failures in LLMs' abductive capacities, but they locate the failure at different stages of this two-stage process.
Floridi et al. (2025) begin from a puzzle about LLM outputs. When asked to explain something, the model produces text that identifies a preferred explanation, organises evidence in support of it, and deploys considerations of simplicity and coherence — features we associate with abductive inference. But the mechanism that produces these outputs is stochastic: during training, the model learns probability distributions over sequences of tokens; during generation, it samples from those distributions. The outputs exhibit the form of abductive reasoning without the process of abductive reasoning. Floridi et al. frame the resulting question in spatial terms:
> "Our main argument is that LLMs occupy a conceptual space 'between' traditional stochastic processes and human-like abductive reasoning. On the one hand, their internal processes are entirely stochastic: during training, they gather statistical correlations from text, and during generation, they produce words based on learned probability distributions." (Floridi et al., 2025, p. 2)
The question is what to make of a system whose outputs reliably exhibit the structure of considered judgment — evidence marshalled in support of a preferred hypothesis, alternatives weighed and set aside — while the production process involves none of this. If the form of reasoning can arise from a process that is not reasoning, either the form is less diagnostic of genuine reasoning than epistemology has assumed, or the process is closer to reasoning than it appears.
The concept Floridi et al. introduce to characterise LLM outputs is *zeroth-order abduction*. To see what it strips away, consider what genuine IBE requires. A scientist confronted with unexpected experimental results generates candidate hypotheses, then evaluates those candidates against the evidence and against each other — assessing which is simplest and which coheres best with what else is known — before committing to one. The evaluative stage is what gives IBE its epistemic credentials: the reasoner does not merely produce a plausible story but tests it against alternatives. Zeroth-order abduction retains generation and discards evaluation. The model receives a prompt and produces the most probable continuation from its prior distribution — the distribution learned during training, never updated by engagement with new evidence or comparison with alternatives the model did not produce:
> "LLMs seem to perform a kind of zeroth-order abduction (Kojima et al. 2022): given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence [...] The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations." (Floridi et al., 2025, pp. 8–9)
The claim that the model "does not understand what an explanation is" marks the difference between producing explanatory text because one recognises what makes an explanation good, and producing it because the patterns of explanatory text have been absorbed from training data. LLMs can generate plausible hypotheses from a scenario — what Floridi et al., drawing on Calzavarini and Cevolani (2022), call *weak abduction*. They can even appear to perform *strong abduction* — selecting the best hypothesis — when candidate hypotheses are explicitly provided (Floridi et al., 2025). What they cannot do is run the generation-evaluation loop autonomously: produce the candidates, assess them against evidence and against each other, and commit to the best without the candidates having been supplied by the user or the task.
What the model lacks, in concrete terms, is a feedback loop. It generates from its prior distribution — what it learned during training — but has no mechanism for conditioning on evidence encountered after generation, or for comparing its output against alternatives it did not produce. That is, it cannot treat its own output as a hypothesis to be tested; it can only produce what is most probable and deliver it as a final answer:
> "In statistical terms, LLMs perform prior predictive sampling but lack an external feedback loop for posterior evaluation." (Floridi et al., 2025, p. 7)
The consequence is that the model cannot withhold judgment. Withholding judgment requires representing one's own epistemic state as insufficient for commitment and declining to produce an answer on that basis. The model has a probability distribution over next tokens, not a representation of its own epistemic confidence. It can produce the tokens "I am not sure" if the training data contained hedging in similar contexts, but this is a textual pattern, not an expression of recognised uncertainty. Floridi et al. call the result *over-abduction*: the model always produces an explanation, even when the evidence warrants none (Floridi et al., 2025). Hallucination is therefore not a malfunction but a predictable feature of the architecture — what happens when a system that cannot represent the limits of its own knowledge is required to generate beyond them.
The picture is complicated, however, by something Floridi et al. themselves acknowledge. The training data encodes a great deal of causal and inferential structure — not because the model has observed causal relationships in the world, but because human-written text systematically reflects them. Floridi et al. grant that "the statistical abstraction of cause-and-effect in the training data is often sufficient to mimic human causal reasoning" (Floridi et al., 2025, pp. 16–17). If the mimicry is reliable enough to produce correct answers to causal questions across a wide range of domains, the question of what distinguishes genuine causal reasoning from its reliable statistical surrogate becomes harder to answer than the zeroth-order framework suggests. And this difficulty surfaces explicitly when Floridi et al. consider provenance:
> "if an AI can generate the same explanatory hypothesis a human would, does it matter that the process was different? From an epistemological standpoint, perhaps yes—justification is significant—but regarding the content of the hypothesis and our interpretation of it, maybe not." (Floridi et al., 2025, pp. 12–13)
A reliabilist epistemology would answer yes: what matters is whether the belief-forming mechanism tracks truth, and a stochastic process does not track truth in the relevant sense. But if what we are evaluating is not the producer's epistemic state but the product — the theory or explanation itself — then the question is whether the product has the properties that make it a good theory, and these properties are assessable from the text. Whether they can be present regardless of how the text was produced is a question Floridi et al. raise without settling.
Zahavy (2026) asks a different question. Where Floridi et al. examine what epistemic standing LLM outputs have — whether the production process warrants the outputs' apparent rationality — Zahavy asks what the architecture can produce in principle: specifically, whether it can generate new theoretical frameworks or is confined to operating within existing ones. His diagnostic framework is Peirce's tripartite distinction between inference modes. Deduction applies a rule to a case to derive the result — truth-preserving, analytic. Induction accumulates cases and results to derive a rule — statistical pattern-finding, what LLMs do by design. Abduction is the third mode:
> "Unlike deduction, which guarantees truth, or induction, which finds pattern that generalize in data, abduction is a creative leap that invents a cause for a singular phenomenon." (Zahavy, 2026)
The word 'invents' is doing work here. Deduction and induction operate on what is already given — premises, data, observed regularities. Abduction introduces something new: a hypothesis not contained in the existing materials, generated to explain something those materials leave unexplained. LLMs have mastered induction; Zahavy's evidence for growing deductive capacity is AlphaProof's performance on International Mathematical Olympiad problems (Zahavy, 2026). But the capacity to generate the premises from which deduction proceeds — to invent the axioms rather than derive their consequences — is not itself a deductive or inductive capacity.
The abductive move has a specific structure, which Zahavy calls the *E→A Jump*: the transition from sense experience (E) to a system of axioms (A). The reasoner begins immersed in a domain of experience, abstracts from that experience a novel structural hypothesis — a candidate set of principles that would organise the phenomena — and formalises it into a framework from which consequences can be deduced. LLMs can handle the last of these phases: given axioms, they can derive consequences and verify proofs. They can also process text that describes the first: accounts of experience, experimental data, phenomenological descriptions. What they cannot do is the middle phase — the abstraction of a novel structural hypothesis from experience — because they have no experience from which to abstract. This is what Magnani (Magnani et al., 2009, cited in Zahavy, 2026) calls *manipulative abduction*: hypothesis generation through the active construction and manipulation of mental models, through what Zahavy describes as "thinking by doing." Einstein's formulation of the equivalence principle illustrates the structure. He did not arrive at it by analysing observational data or optimising within Newtonian mechanics; he constructed a mental simulation — an observer in a sealed, uniformly accelerated environment whose experience would be indistinguishable from that of an observer in a gravitational field — and extracted from this indistinguishability a structural hypothesis that gravitational and inertial mass are equivalent. The hypothesis emerged from the simulation, not from the existing symbolic materials.
The architectural consequence is that LLMs remain on the wrong side of a gap between processing the language of a domain and having access to what that language refers to:
> "They operate as high-dimensional 'Chinese Rooms' (Harnad, 1990), manipulating the language of physics without access to the physical referents that give that language meaning." (Zahavy, 2026)
The distinction this points to is not between doing science well and doing it badly but between two kinds of cognitive operation: optimising within a given framework and inventing a new one. Zahavy's discussion of recent AI systems makes the point concrete. The AI Scientist and AlphaEvolve — systems that optimise within given search spaces, recombining existing concepts and exploring parameter landscapes — cannot define a new search space (Zahavy, 2026). Einstein did not optimise within Newtonian mechanics; he replaced the framework. The ability to search within a space, however efficiently, does not confer the ability to define the space; and deductive capacity, however impressive, does not address this bottleneck, because deduction operates on premises it does not itself supply.
Both arguments, however, are developed with empirical science as their target domain, and this matters more than it might initially appear. Floridi et al.'s examples of the training data that encodes reasoning patterns are drawn from scientific papers, Q&A forums, and Wikipedia articles: domains where text *reports* findings made elsewhere — in laboratories, through instruments, by empirical observation. The text is downstream of the reasoning; the reasoning happened in the lab, and the paper is a record of it. Zahavy is explicit about the scope of his own argument:
> "we emphasize that this proposal is specifically tailored to the physical sciences, where the object of study is external material reality. In abstract domains such as Mathematics or Computer Science, the Sense Experience (E) may be grounded in high-dimensional topology or have other goals such as generality or minimality." (Zahavy, 2026)
The restriction matters because philosophy is neither an empirical science nor a purely formal discipline. Its textual medium stands in a different relationship to its contributions. When a physicist writes a paper, the paper reports a discovery made elsewhere — in a laboratory, through an instrument, by mathematical construction. When a philosopher develops an objection to a thesis, the sentences that develop the objection *are* the objection. The reasoning is not reported in the text; it is constituted by it.
If philosophical reasoning is constituted by argumentative text — if the text is the contribution, not a report of a contribution made elsewhere — then a system trained on a large corpus of philosophical writing has absorbed not just the products of philosophical reasoning but the reasoning itself, in a way that does not hold for a system trained on scientific papers that report the outcomes of laboratory work conducted elsewhere. And if the evaluation of philosophical work is argument-checkable — if the standards by which we assess a paper are standards that apply to the text, assessable by competent readers without access to anything beyond it — then the question of what 'abduction' means for philosophy is not the same as the question of what it means for physics. Both Floridi et al. and Zahavy invoke abduction, and neither asks whether the abduction philosophy requires is the same as the abduction LLMs are alleged to lack. Before that question can be answered, we need to know what 'abduction' means in the different literatures that use the term — and in particular, what it means in the literature on philosophical methodology, where inference to the best explanation is not a description of how scientists discover but a method for evaluating theories by their intrinsic properties.
---
# 2. Abduction and Philosophy
# Abduction and Philosophy
When Floridi et al. describe LLM outputs as having an "abductive appearance," they use "abduction" to mean something different from what Zahavy means by it. Floridi et al. have in mind a high-level reasoning pattern — the structure of hypothesis-then-justification that appears in training data and gets reproduced in outputs. Zahavy has in mind a creative leap from embodied experience to explanatory axioms — something cognitively specific and architecturally demanding. And Williamson, when he advocates an "abductive methodology" in philosophy, means something different again: a method of theory-evaluation by theoretical virtues, without commitment to a cognitive story about how the evaluation is performed. Before asking whether LLMs can do abduction, it is worth asking which of these things we are asking about.
There are at least four conceptions of abduction at work in the literature I have been discussing. The first is Peircean: abduction as hypothesis generation, the creative leap from surprise to candidate explanation — what Zahavy's E→A Jump is. The second draws on a two-stage picture of inference to the best explanation: a generation phase that produces candidate explanations and a selection phase that ranks them by explanatory virtues — elegance, unification, simplicity, scope. These two phases are potentially separable; one might manage the second without being able to do the first. The third is Floridi et al.'s: abduction as a high-level reasoning pattern that can be mimicked by stochastic processes — what they call "zeroth-order abduction." And the fourth is Williamson's: IBE as a method for evaluating philosophical theories by their intrinsic properties, without commitment to a cognitive account of how the evaluation is performed. A further distinction sharpens the picture: the difference between actual and potential explanations. An actual explanation is what causally produces a belief; a potential explanation is what would explain a phenomenon if it were true. On the two-stage picture, what matters for evaluation is potential explanation — we rank hypotheses by the understanding they would provide, not by the process that generated them.
Once we distinguish these conceptions, the question "can LLMs do abduction?" fragments into questions with different answers. If abduction is Peircean generation, Zahavy's argument about embodied simulation has force — the E→A Jump requires the reasoner to simulate physical experience, and LLMs do not have physical experience to simulate. If abduction is a two-stage process, the selection phase — ranking candidates by explanatory virtues — is a matter of assessing properties of hypotheses rather than having the right kind of inner life, and it is at least not obvious that LLMs cannot perform it. If abduction is a reasoning pattern (Floridi et al.'s conception), LLMs reproduce the pattern without performing the reasoning, but whether reproducing the pattern suffices depends on the task. And if abduction is a method of theory-evaluation (Williamson's conception), the question is whether LLM outputs can satisfy the criteria that constitute the method, regardless of what goes on inside the system.
These conceptions diverge, I suggest, because the demands of a domain shape what "abduction" looks like in that domain. In physics, where the object of study is external material reality, the creative leap requires embodied simulation or something like it — Zahavy's account of Einstein's "happiest thought" makes this plausible — and so Peirce's "hypothesis generation" is the operative conception. In a domain where the object of study is not external material reality, the relevant conception might be quite different. The question, then, is what kind of domain philosophy is. This can be approached by comparing philosophy with other disciplines in terms of each discipline's relationship to its textual medium.
In the natural sciences, the text reports the discovery. Watson and Crick's Nature paper reports the double helix structure — a physical modelling achievement, built and rebuilt until the base-pairing constraints were satisfied. Einstein's papers report the theory of General Relativity — driven by physical intuitions and mathematical construction. Darwin's *Origin* reports decades of empirical observation: the Beagle voyage, the pigeon breeding, the barnacles. In each case the vehicle of discovery is extra-textual — a physical model, a set of equations, years of fieldwork — and the textual articulation, however consequential, follows the creative act rather than being identical with it.
In the visual arts, the gap between text and creative work is wider still. Picasso did not argue that representational painting was defective and that cubism resolved the defect; he painted differently, and the critical apparatus followed after the fact. The transformation consisted in producing works that showed a new way of organising visual space — not an argumentative demonstration but a perceptual one. In music, Schoenberg composed atonally; the theoretical writings are secondary to the musical acts. Literature offers the more telling comparison, because in literature the text is the creative work — a novel is not a report of an achievement made elsewhere. But the evaluation of literature is aesthetic: style, narrative structure, voice, imaginative reach. These properties are not argument-checkable in the way that the properties of a philosophical contribution are. A reader can disagree about whether a novel succeeds without being able to point to a specific flaw in the text's argumentative structure, because the text does not have an argumentative structure in the relevant sense. Philosophy and literature share the property of being textual all the way down. They differ in the kind of assessment their texts are subject to.
What emerges from these comparisons is that philosophy is the case where two properties converge that are separate elsewhere. In philosophy, the text is the contribution — not a report of a contribution made elsewhere, but the thing itself. Kripke's contribution to the philosophy of language is the modal argument, the epistemic argument, and the semantic argument as laid out in *Naming and Necessity*. Lewis's modal realism is the theoretical package defended in *On the Plurality of Worlds*: arguments, cost-benefit analyses, replies to objections. Chalmers's hard problem is the argumentative demonstration — the zombie argument, the inverted spectrum — that functional explanation cannot close the explanatory gap. There is no lab, no telescope, no physical model; there are arguments on paper. And the evaluation of those arguments is publicly checkable. The standards by which we assess a philosophical paper — validity, adequacy of distinctions, explanatory reach, integration with background commitments, non-ad-hocness — are standards that apply to the text, assessable by competent readers, and do not require access to anything beyond the text.
Even paradigm-shifting philosophical contributions were made through standard argumentative moves — thought experiments, modal intuitions, reductio reasoning, parity-of-reasoning arguments, cost-benefit analysis — all individually familiar from the existing philosophical toolkit. What was novel in each case was the combination: bringing resources from different sub-fields together in a way that exposed a structural deficiency in the received framework. Kripke combined modal logic with philosophy of language; Lewis combined possible-worlds semantics with Quinean ontological seriousness; Chalmers combined functionalism with conceivability arguments from philosophy of modality. The transformations happened within and through the existing argumentative practice.
If philosophy's textual medium is not merely the vehicle for reporting contributions but the medium in which contributions consist, then the gap between symbol and referent — the gap that gives Zahavy's Chinese Room worry its force in physics — narrows substantially. In physics, the symbols refer to something that exists independently of any symbolic representation: spacetime, particles, fields. A system that manipulates the symbols without access to the referent is doing something fundamentally different from physics. But if the referents of philosophical discourse — theoretical virtues, inferential relations, dialectical structures — consist in relations between concepts as expressed in text, then the system that manipulates philosophical symbols is not cut off from the referents in the same way. When a philosophical text discusses the relationship between simplicity and explanatory power, the relationship being discussed is a relationship between properties of theories, and theories are textual artefacts. Philosophical objects are not external to the symbolic medium in which they are expressed; to a far greater extent than in any empirical discipline, they are realised in it.
Some philosophy does require capacities LLMs may lack, and it is worth being explicit about where the limits fall. If a philosophical argument depends on phenomenology-as-datum — if knowing what pain feels like, or what temporal passage is like from the inside, is part of the evidential base — then an LLM is working without that evidence. Philosophy of perception and parts of ethics involve claims about empirical reality, and for those sub-domains the architectural worry retains some force. But the territory affected is bounded. Much of analytic philosophy — argumentation about concepts, theories, and inferential relations — does not depend on phenomenological data. And even where experience is relevant, two things mitigate the worry: common-sense references to the physical world are pervasive in training data (the model has encountered countless descriptions of what things look and feel like), and first-person phenomenological reports are a literary genre in their own right (the corpus is not silent on subjective experience, even if the model has not had subjective experiences).
Return, then, to the question the conceptions of abduction diverge over. Given that philosophy is textual in this distinctive sense — the text is the contribution and the evaluation is argument-checkable — which conception is relevant? Peircean generation, the embodied leap from experience to axioms, applies where the domain demands embodied simulation. For philosophy, the creative "leap" does not run through the body; it runs through recombination of argumentative resources. The E→A Jump in philosophy is not from sense experience to axioms but from the existing dialectical landscape — positions already staked out, objections already lodged, responses already attempted — to a novel argumentative configuration. Selection by explanatory virtue, the second stage of the two-stage picture, is evaluable at the level of the artefact: the question is whether the output exhibits the relevant virtues, and those are properties of the text. And method-level evaluation, Williamson's conception, makes the artefact-level point explicit. On his account, theoretical virtues are "intrinsic" to the theory:
> "Apart from its relation to E, the more T has the intrinsic virtues of a good theory, the better (ceteris paribus). It should be elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated." (Williamson, p. 354)
The word "intrinsic" is doing real work here. The virtues are features of the theory, not of the theorist — properties of the artefact, not of the producer's inner states. Whether a theory is elegant, unified, and non-ad-hoc is assessable from the text; no information about the production process is needed. And the standards are applicable to the philosophical community's output as a whole:
> "Strikingly, the philosophical community showed very little aversion to the multiplication of complication. A firmer preference for simplicity and elegance would have warned the community that something was going wrong. Indications of over-fitting remain quite widespread in analytic philosophy." (Williamson, p. 369)
This is an assessment of papers, not of minds. The standards Williamson invokes — simplicity, non-ad-hocness, resistance to over-fitting — are properties he takes to be visible in texts and assessable by readers.
Both critiques from the previous section, then, are worries about the producer — about its epistemic reliability (Floridi et al.) or its cognitive architecture (Zahavy). Neither is a worry about the artefact. Against Floridi et al.: if what we evaluate are the intrinsic properties of a theory, then whether the production process was stochastic or deliberative does not bear on whether those properties are present. The "stochastic core" is a fact about the producer, not about the product. Against Zahavy: if what we evaluate are intrinsic properties of the theory, then whether the producer had embodied access to physical referents is equally beside the point. His Chinese Room worry presupposes that the gap between symbol and referent must be bridged by the producer's cognitive architecture. But in a discipline where the symbols realise the referents rather than merely representing them, the gap is not what it was.
What follows for the evaluation of LLM-produced philosophy is that any critique must point to specific textual deficiencies — equivocations, illicit premises, ad hoc repairs, question-begging moves, unmet explanatory burdens — rather than gesturing at the production mechanism. Saying "but it is just statistics" is a claim about the producer, not about the artefact, and has no bearing on the artefact's quality as assessed by the discipline's own standards. The question, then, is whether LLMs can actually produce texts that satisfy those standards — whether the norms of good philosophy are learnable from text, and whether what is learned can produce novelty and not merely competent reproduction.
---
# 3. Learning the Game
# Learning the Game
The physics literature is full of discoveries that happened elsewhere — in laboratories, through instruments, via mathematical construction. Training a language model on that literature gives it the language of physics without giving it physics itself, because the work that produced those discoveries lies outside the text. The philosophical literature is different. If philosophy is textual in the sense I described in the previous section — the text is the contribution, the evaluation is argument-checkable, the objects of study are inferential relations — then the philosophical corpus contains not just a language but a discipline: its contributions, its evaluative standards, and its subject matter; to train on this corpus is to train on the discipline.
The first thing this implies is that the norms of good philosophy are learnable from text. If those norms were hidden — if satisfying them required some non-textual insight that left no trace in the writing — then training on texts would not help. But the norms are visible. Bengson, Cuneo, and Shafer-Landau's account of philosophical methodology helps make the point precise. Their Tri-Level Method sets out criteria for building and assessing philosophical theories: accommodation and explanation of the data at the first level, substantiation and integration at the second, theoretical virtues as tie-breakers at the third. The interest of this for present purposes lies not in the details but in a feature the authors themselves stress — that the criteria are drawn from ordinary philosophical practice:
> "We endorse the method not because it makes a philosopher's job easy; indeed, it is quite demanding. Nor are we drawn to its constituent criteria because they revolutionize philosophical thinking; on the contrary, all of them are familiar from the way many philosophers go about their business." (Bengson et al., p. 107–108)
Philosophers satisfy these criteria not by consulting a checklist but by doing philosophy: advancing arguments, raising objections, offering replies, providing clarification, displaying sensitivity to the deliverances of logic, mathematics, science, and common sense. The criteria show up in texts as patterns of exposition and dialectical response, whether or not the writer formulates them as such.
Philosophical corpora contain recurring patterns of how philosophers move from one dialectical state to the next demanded step. If a view fails to accommodate some datum, the next move is accommodation or defence of non-accommodation. If a claim lacks substantiation, the next move is to supply epistemic support or explain why none is required. If a theory conflicts with background commitments, the next move is integration or defence of the conflict. These demanded next steps appear in texts with enough regularity that a system trained on the corpus can learn the distribution. Walton, Reed, and Macagno's work on argumentation schemes reinforces this at finer grain. Argumentation schemes are common inference patterns — argument from analogy, argument from consequences, argument from expert opinion — each paired with critical questions that represent the standard challenges for arguments of that type:
> "The method of evaluation of an argument fitting a scheme is that once the argument is put forward by a proponent, it may be defeated if the respondent asks an appropriate critical question that is not answered by the proponent." (Walton et al., p. 3)
The structure is: move, critical question, response. At both the theory level (Bengson's criteria) and the argument level (Walton's schemes), the norms of philosophical practice are textually manifest.
Learnability, though, is only part of the picture. The philosophical corpus that an LLM trains on is not a random sample of all philosophical attempts. It is a filtered sample. Papers get published, taught, anthologised, and cited in rough proportion to their perceived quality, and quality in philosophy is substantially a matter of theoretical virtue: elegance, unification, simplicity, explanatory power. The corpus is enriched for explanations that exhibit these virtues — not perfectly, since there is noise, there are fashions, and there are weak papers cited for sociological reasons, but the signal is there. When a model learns to produce philosophy-like text, it is learning from material that has already passed through the discipline's quality-control mechanisms. The model does not need its own sense for theoretical virtue; the training data has already done the filtering, and the model needs only to learn the distribution of what survived.
The philosophical tradition, viewed in this light, is the record of an evaluative feedback loop: centuries of philosophers proposing explanations, testing them dialectically, refining their standards, discarding what failed, building on what survived. When the model trains on this record, it absorbs the outcomes of a calibration process it has not participated in. It has not earned its calibration; it has borrowed it. Whether borrowed calibration suffices is a question worth taking seriously.
I suggest it does, for a reason that connects to the character of philosophy as a discipline. One might worry, in the spirit of Voltaire's dormitive virtue, that borrowed calibration fails in novel cases — that the model needs to understand why a standard works, not merely that it works, and that understanding the "why" requires having gone through the feedback loop oneself. In empirical science, the reason that simplicity tracks truth might ultimately be something about the structure of physical reality — something not fully expressible in text. But in philosophy, the reason that simplicity is a virtue — that it protects against over-fitting, prevents ad hoc epicycles, keeps theories answerable to their data — is itself a philosophical claim, fully expressed in the argumentative tradition. Williamson's defence of simplicity is itself a philosophical argument that appears in the corpus:
> "The restriction helps us avoid mistaking noise for signal, which we do if we fit the current data too closely. This account of the role of simplicity and similar aesthetic criteria in abductive methodology is consistent with a fully realist, non-pragmatist understanding of science." (Williamson, p. 368)
The reason the standard works is part of the same tradition that exhibits the standard. Unlike in empirical science, where the justification for a methodological norm may lie outside the textual record, in philosophy the justification is a philosophical argument, available in the corpus alongside the norm it justifies.
There is a further point about error signals. Zahavy argued that compression-based creativity fails where there is no error signal to compress against — in physics, the Newtonian framework was empirically adequate, and no gradient pointed toward the need for a new theory:
> "scientific invention often occurs in the absence of a supervised error signal. An AI operating as an inductive optimization engine would have found the Newtonian loss function to be near-zero." (Zahavy, 2026)
The point is well taken for physics, but philosophy's situation is the reverse. Philosophical corpora are not empirically sparse landscapes presenting near-zero loss; they are dialectically saturated. The training data encodes not just arguments but evaluations of arguments — not just moves but the discipline's accumulated judgments about which moves succeed and which fail. Objection-reply sequences, editorial decisions about which papers to publish, citation patterns that track which contributions the discipline treats as worth engaging — all of these are present in the corpus as textual regularities. Every sustained objection to a philosophical position is a signal about where the position is vulnerable; every accepted repair is a signal about what the discipline treats as a good fix; every ignored response is a signal about what does not work. Where physics presents a near-zero loss landscape with no gradient toward General Relativity, the philosophical corpus presents a landscape dense with evaluative gradients — the kind of landscape in which the training signal is rich.
A natural question at this point is what exactly the model has learned. Consider two possibilities. On the first, the model has internalised something like a norm — "prefer simpler explanations" — and applies it when generating outputs. On the second, the model has learned that certain argument structures, which happen to be simple, produce higher prediction scores because they appear more often in published philosophical text; it has learned the patterns that result from a norm being followed without learning the norm itself. These are empirically difficult to tell apart, because they produce the same outputs in standard cases. The divergence would come in novel cases — cases where the norm needs to be extended to unfamiliar territory or balanced against competing norms in an unfamiliar way.
But here a feature of philosophical practice becomes relevant. Philosophical argumentation is conservative in its forms. The same moves — counterexample, distinction, reductio, analogy, dilemma — recur across very different content areas. If what makes a philosophical explanation elegant is a formal property it shares with elegant explanations in quite different domains, then the second possibility might be extensionally adequate even without norm-internalisation in any deep sense, because the forms transfer across content by being the same forms. Whether this amounts to understanding is a metaphysical question that need not be settled here. What matters is whether the learning — however characterised — is sufficient to produce outputs that satisfy the standards, including the standard of novelty.
The novelty question is the hardest. Even granting that the model has learned the patterns of good philosophical argumentation, one might insist that it can at best reproduce those patterns, not produce new philosophical work. But philosophical novelty, even at the paradigm-shifting level, consists in recombination of standard argumentative moves — the individual tools are familiar; what is new is the combination, bringing resources from different sub-fields together in a way that exposes a structural deficiency in the received framework. Boden's taxonomy of creativity distinguishes combinatorial creativity (novel combinations of existing elements), exploratory creativity (traversal of a structured conceptual space), and transformational creativity (restructuring of the space itself). The first two are within reach of a model trained on diverse philosophical texts, which can combine resources from different regions of its training distribution in ways that no single training text does. The harder question is whether transformational contributions also lie within reach. If transformational contributions in philosophy happen within and through existing argumentative practice rather than by transcending it — if the transformation is a novel combination of standard moves, not a departure from the practice of making them — then the line between exploratory and transformational creativity is less sharp than the taxonomy suggests, at least in this discipline.
Gaut makes a point about mechanically generated creative outputs that bears on this. Even if a metaphor were produced by a purely mechanical process, he observes, it:
> "would still guide their audience imaginatively to link together two domains, and if the metaphors were successful, to discover original and apt connections between them and perhaps to elaborate the metaphors further. They would thus guide those who understood them through a process akin to the process of creative imagination that could have, but did not, produce them." (Gaut, fn. 23)
The output's structure does cognitive work for its audience regardless of how it was produced. A philosophical argument, at its best, functions as an instrument of recognition: a textual structure that constructs a path from familiar premises to an unfamiliar conclusion, enabling a reader to see something she could not see before. The argument does not report the producer's private insights — it builds an inferential path that generates insight in competent readers who follow it. If a reader follows the argument and finds it sound, she has all the evidence she needs to assess its philosophical quality, and information about the production process adds nothing to that assessment. The question of whether the producer traversed the same path, or had any private understanding of where the path leads, is a question about how the argument came to be, not about what it does.
A comparison with the Sokal affair makes the point concrete. The Sokal hoax succeeded in a field where the evaluative norms were not of the kind that could be satisfied or failed argument by argument. A comparable attempt in analytic philosophy would face a different obstacle: the referees would check the arguments, test the inferences, probe the relationship between premises and conclusions. The evaluative norms of analytic philosophy operate on the text itself, publicly and step by step. Where evaluation works in this way, the question of whether the surface matches the depth is answerable by examining the surface with sufficient care — if the arguments are valid, the distinctions sharp, and the explanatory reach adequate, the text has met the discipline's standards.
What follows is that any attempt to dismiss LLM-produced philosophy must itself be a piece of philosophical criticism: it must identify a specific textual deficiency — an equivocation, an illicit premise, an ad hoc repair, a question-begging move, an unmet explanatory burden. If the text accommodates the relevant data, substantiates its claims, integrates with background commitments, and does so with parsimony and precision, then the observation that it was produced by a stochastic process is a remark about the production process, not about the product, and has no bearing on the product's quality as assessed by the discipline's own standards. Blind review exists in philosophy for exactly this reason: provenance is not supposed to affect assessment, because the standards are standards that apply to the text, and a paper that satisfies them satisfies them regardless of who — or what — wrote it.
Philosophy, then, occupies a distinctive position among intellectual disciplines with respect to AI. Because the text is the contribution and the evaluative standards are publicly checkable, training on the philosophical corpus gives a model the discipline itself — its contributions, its norms, and the subject matter those norms apply to. The tradition's evaluative feedback loop is encoded in the corpus, and the kind of creativity the discipline values consists in novel combinations of standard argumentative moves that are well-represented in the training data. In disciplines where the vehicle of transformation lies outside the text — empirical science, the visual arts, music — there are principled reasons to doubt that textual competence alone suffices. Whether those reasons extend to philosophy depends on whether a philosophical contribution requires something beyond the text. For the large territory of analytic philosophy that does not depend on phenomenological data, I have argued that it does not. This is not a deflationary claim about philosophy; it is a claim about what kind of practice philosophy is.
---
# 4. How to Generate Philosophy with AI
# How to Generate Philosophy with AI
*Burden*: Show the thesis in action with worked examples.
This section makes it vivid. You need at least one case where:
- The prompt is minimal (genre-cueing, not micromanaged)
- The output exhibits genuine philosophical structure: hinge identification, cost-accounting, alternative-theory comparison, sensitivity to objections
- You can evaluate it against the standards and show it passes
The reader should be able to *see* what you mean by constraint-satisfaction, not just take your word for it.
You might also include a stress-test case — something that exposes where failure IS identifiable text-internally. The pseudo-robustness example (the semantics-reduces-to-physics prompt) could serve: you show that *when* standards are violated (equivocation, bait-and-switch), the violations are identifiable from the text. This supports your claim that evaluation is artefact-level: you don't need to know it was an LLM to see the flaw.
This section comes last (before conclusion) because it's evidence, not argument. You want the reader to have the framework before seeing the examples.
---
# 5. Conclusion
# Conclusion
*Burden*: Restate the thesis, sum up the argument, gesture at implications.
**Restate**: LLMs can produce novel, first-rate philosophy with minimal prompting. The question isn't "do they really reason?" but "do their outputs satisfy the constraint structure of good philosophy?" The answer is yes — often enough to matter.
**The argument in brief**: Floridi's "abductive appearance" critique and Williamson's centrality-of-abduction picture seem to block LLM philosophy. But philosophical evaluation is artefact-level: we assess texts, not producers. The relevant standards — precision, cost-accounting, non-ad hocness, defeater-sensitivity, fair treatment of rivals — are text-internal and publicly codifiable. LLMs trained on philosophical corpora have internalised these standards as practice-patterns. Minimal prompting cues the relevant genre; the latent dialectical structure does the rest. Philosophy's peculiar self-grounding nature (the map IS the land) and the collapse of appearance/reality for competent readers further support the thesis.
**Implications**:
*For philosophical methodology*: What does it mean that the standards are learnable from text? Perhaps philosophy's "rules of the game" are more public and codifiable than we assumed. Perhaps philosophical competence is more like fluency than genius.
*For understanding what philosophy is*: The thesis suggests that philosophy is a practice governed by publicly accessible norms — not ineffable insight, not special cognitive access, but skill with reasons as they appear in text.
*For the discipline's future*: A new kind of collaborator (or competitor) has arrived. How should philosophers respond? The answer isn't panic or dismissal; it's recognition that the standards remain the same. Good philosophy is good philosophy. Evaluate the work.