# 0. Introduction - LLMs can generate publishable philosophy with minimal prompting. Philosophy is a text-based discipline governed by public, structural norms rather than private mental states. If a text satisfies these norms, the distinction between genuine reasoning and valid text production collapses. - By *minimal prompting* I mean genre-governing cues (e.g., "write a robust defence") rather than micromanaged instruction-following. Such cues select for the kind of continuation the model has learned to produce from philosophical texts. - The argument accepts Floridi's critique that LLMs are stochastic mimics lacking intent, but isolates his concession that training encodes reasoning structures. Williamson's abductivism then does the work: in philosophy, theoretical virtues are the standard by which we evaluate theories—and these virtues are intrinsic to the theory, not the theorist. Section 3 explains how a stochastic engine learns these structures via statistical regularities in the corpus. # 1. What LLMs Aren't Doing - Floridi et al. (2025) argue that LLMs occupy a "conceptual space between traditional stochastic processes and human-like abductive reasoning" (p. 2). Their formulation: "We can briefly describe LLMs as fundamentally stochastic, with surface-level abductive appearances" (pp. 19–20). - The internal mechanism is purely probabilistic and lacks intentionality or semantic understanding. LLMs "lack explicit representations of meaning, everyday relevance, truth values, or causality"; what they produce is a "complex probability distribution" rather than reasoning (pp. 2, 7). - Despite this stochastic core, the outputs exhibit a "phenomenological similarity to human reasoning" (pp. 2–3). Floridi explains the illusion: the model has "absorbed patterns of human abductive reasoning as expressed in writing", including "causal connectives... and explicit reasoning steps" (pp. 2–3, 8, 10). - Because the model optimises for next-token prediction rather than truth-tracking, it lacks verification capabilities. It performs "prior predictive sampling"—generation—without "posterior evaluation"—verification against reality. "LLMs do not know whether they are right/correct or wrong/incorrect" (pp. 7, 9). - This lack of verification leads to "over-abduction": the model generates explanations regardless of justification. While a human might suspend judgement, the LLM "cannot resist explaining"—generating a plausible continuation is its only task (p. 12). - I accept Floridi's diagnosis: the mechanism is stochastic; the deficit is the absence of truth-verification. But I isolate his explanation of the illusion—that the model produces encoded reasoning structures—and argue that in philosophy, these structures are constitutive of the method itself. # 2. Abduction and Philosophy Williamson's account of philosophical method intensifies Floridi's challenge. If Floridi is right that LLMs cannot do abduction in any genuine sense, and if Williamson is right that philosophy is largely abductive, then a simple syllogism seems to settle the matter: no abduction, no philosophy, therefore no good novel philosophy from LLMs. I want to show that this argument is not valid. It relies on a slide in what *abduction* is doing in the two premises, and once that slide is blocked, the conclusion does not follow. Williamson argues that philosophy "should use a broadly abductive methodology" because a purely deductive methodology tends to produce deadlock. His diagnosis is structural: "if the argument is deductively valid, opponents simply reject one of those informative universal premises as 'question-begging'" (Williamson, p. 364). Deductivism thereby exerts pressure for premises to be uncontentious; abductivism removes that pressure by shifting the burden from premise-by-premise acceptability to theory-level comparison. Rather than trying to force an opponent to accept each premise in advance, the abductivist compares whole packages and asks which package best explains the phenomena, given the relevant constraints. Abduction, on this picture, is not an optional extra; it is a methodology of theory choice designed to bypass a predictable failure mode of deductive argumentative exchange. Lewis's modal realism serves as Williamson's paradigm case. Lewis postulates possible worlds because they follow from his modal realism, which he treats as "the best theory of possibility, necessity, and related phenomena, in respect of simplicity, strength, elegance, and explanatory power" (Williamson, p. 314). The argument is abductive in Williamson's broad sense: the theory is licensed by its performance with respect to theoretical virtues and explanatory integration, not by the availability of an uncontroversial set of premises from which modal realism can be deduced. On Williamson's view, philosophical theories are ranked by their theoretical virtues. *Simplicity* is not mere ornament; it is tied to epistemic performance. The thought, familiar from work on model selection, is that over-fitted theories can track noise in current data and therefore generalise badly, whereas simpler theories are less vulnerable to that distortion. *Strength* and *unification* matter because we should prefer theories that are "elegant and unified, not arbitrary, gerrymandered, ad hoc" (Williamson, p. 354)—theories that do not merely patch local problems with piecemeal stipulation, but deliver a systematic account with explanatory reach. *Precision* matters because it enables meaningful testing: vague theories can "avoid the risk of falsification, but by the same token they give up the hope of explaining anything" (Williamson, p. 366). Abduction is therefore comparative and forward-looking: we rank theories as potential explanations before we know they are true, and the ranking guides what we should take most seriously. Williamson also emphasises a robustness requirement. We do not get to reason with perfectly clean inputs, so abductive methodology must tolerate error: "we need robust methods of theory choice that do not crash every time an error enters" (Williamson, pp. 369–370). The point is not that anything goes; it is that theory choice procedures must remain usable under ordinary epistemic imperfection. Finally, Williamson adds a constraint that matters for any attempt to apply his picture to philosophical practice: philosophical theories must cohere with "our total evidence," arguably "no less than the total sum of human knowledge," including "the natural and social sciences, philosophy, and common sense" (Williamson, pp. 356–357). Philosophy is not hermetically sealed. Even when it proceeds from the armchair, it is constrained by the best available evidence and background commitments across inquiry. So far, nothing here is friendly to optimism about LLMs. But this is where the quick sceptical inference smuggles in an extra premise. Floridi's claim that LLMs cannot do abduction is a claim about the internal epistemic character of the system: a stochastic next-token engine that lacks truth-aiming verification and therefore can generate "abductive appearances" without being an epistemic agent in the relevant sense. Williamson's abductivism, by contrast, is a methodological characterisation of how philosophical theories should be assessed: compare candidates as explanations, rank them by virtues, integrate them with total evidence, reject ad hocness and empty vagueness, and prefer robust theoretical packages over brittle deductive stalemates. To get from those two claims to "LLMs cannot produce good novel philosophy," the sceptic needs a bridging premise: that producing an abductively excellent philosophical text requires the producer to instantiate Floridi-style abductive agency. But Williamson's abductivism does not entail that bridging premise. It tells us what abductive excellence consists in at the level of theory choice and evaluation—what makes one articulated account better than another—and why abductive comparison is methodologically indispensable. It does not establish that only systems with a particular internal epistemic profile can generate texts that satisfy those abductive constraints. The sceptic conclusion therefore does not follow from Floridi plus Williamson alone. Once the missing premise is exposed, the debate relocates to a more precise question: can LLM outputs instantiate the abductive constraint-structure Williamson describes in a way that survives ordinary philosophical pressure—avoiding equivocation, avoiding ad hoc repair, achieving unification rather than gerrymandered patchwork, maintaining precision rather than vacuous flexibility, and integrating appropriately with background constraints? That is a substantive question about performance under critical scrutiny, not a conclusion that drops out automatically from the claim that the generator is stochastic. The objection only becomes decisive if one can show that stochastic generation systematically fails these constraints, rather than merely lacking a certain kind of inner epistemic agency. # 3. Learning the Game Section 2 blocks the inference from "LLMs can't do abduction" plus "philosophy is abductive" to "LLMs can't do philosophy." But blocking an inference is not a positive account. We need to explain how a stochastic text engine can generate outputs that meet philosophical standards—and how such outputs can be checked. Floridi's diagnosis is sharp. LLMs produce outputs by sampling continuations under a learned distribution. They generate text with an "abductive appearance" because they have been trained on texts in which abductive reasoning is expressed and rewarded. But the underlying process is generative without a truth-aiming feedback loop: "prior predictive sampling" without "posterior evaluation." The system can produce an explanation-shaped object without any internal mechanism that tests it against reality, evidence, or its own prior commitments. The tendency Floridi calls "over-abduction" is a symptom: when the task is to continue, the system cannot resist producing an answer even when suspending judgement would be more appropriate. If philosophical work were validated the way scientific measurement is, Floridi's critique would be decisive. But much of philosophy's checking happens through argumentative pressure, counterexample, coherence demands, and theoretical-virtue comparisons. What makes philosophy good is visible in how it handles reasons: how it frames a problem, what it treats as data, how it responds to objections, and how it avoids ad hoc repair. Those are textual and dialectical constraints. If they are public and stable, they can be learned from a corpus and applied to outputs, even if the generator lacks an internal truth oracle. Bengson, Cuneo, and Shafer-Landau make those constraints explicit. They model inquiry as moving from data to theory through a method, where methods are criteria that serve a dual role: they guide construction and they supply standards of evaluation. The same constraints that tell you what to build are the constraints by which your output is judged. A competent philosophical text accommodates and explains its data, substantiates its claims, integrates them with background commitments, and competes with alternatives under virtues like simplicity and parsimony. These criteria are "familiar from the way many philosophers go about their business" (Bengson et al., p. 107–108), even if rarely unified into an explicit methodology. If philosophical corpora are full of theorising under these criteria, those patterns are learnable. Philosophers do not merely state theses; they perform a constrained sequence of moves. They identify a target phenomenon, propose a treatment, confront predictable resistance, and repair or refine to satisfy coherence, explanatory adequacy, and non-ad-hocness. These moves recur across papers, subfields, and decades. A model trained on such text can learn what comes next when a theory is challenged on accommodation, substantiation, or integration. At the argument level, Walton, Reed, and Macagno offer a complementary picture. They treat argumentation schemes as stereotyped reasoning patterns and pair each scheme with critical questions that function as an evaluation procedure. If an argument fits a scheme and its premises are plausible, the conclusion receives presumptive entitlement—but if an interlocutor asks an appropriate critical question, the entitlement is defeated unless answered. This supplies the kind of "posterior evaluation" Floridi says is missing: not by giving the model access to the world, but by specifying how a claim must survive interrogation. Philosophical competence thus involves two levels of constraint. At the theory level, competent texts accommodate data, substantiate claims, integrate with background, and avoid ad hoc patching. At the argument level, they deploy recognisable inferential moves and address the critical questions those moves invite. Neither depends on access to a private mental state. Both are publicly enforceable and show up in text as recurring structures. Minimal prompting works because a prompt like "write a robust defence" selects a constraint regime. These are not detailed plans but indicators of what moves are required next: from this dialectical state, produce the kind of thing that counts as a robust defence. Because the model has been trained on texts where "robust defence" correlates with a recognisable suite of moves, the prompt triggers a coherent package rather than random elaboration. Can verification demands convert generation into something that behaves like inquiry? In programming, verification often comes with a clear oracle: the code runs and passes tests, or it fails. Philosophy has no such immediate verdict. But that is why philosophy has developed elaborate public constraints. When the world does not deliver a quick answer, communities compensate with structured methods of theory choice and dialectical testing. Even if base LLMs lack internal posterior evaluation, philosophy supplies a route to evaluation that does not require access to truth. The checking can be external: interrogate the output with scheme-appropriate critical questions; demand integration with background commitments; test whether repairs are genuinely explanatory or ad hoc; probe for equivocations and unargued assumptions. These are the checks by which philosophical texts are ordinarily assessed. If an output repeatedly survives them, it counts as robust philosophical performance. Floridi's "over-abduction" worry can be turned into a design constraint. The generator's tendency to answer regardless of justification is a predictable failure mode when there is no built-in verification. But in philosophy, verification often means subjecting the claim to dialectical pressure. The missing posterior evaluation can be supplied by moving from single-pass generation to an adversarial process: generate a candidate; interrogate it with critical questions; demand non-ad-hoc repairs; test for integration and precision; compare with rivals under theoretical virtues. None of this requires the thought that philosophy is sealed off from empirical constraint. Williamson is right that philosophical theories must cohere with total evidence. The present claim is narrower: much of what makes philosophical work good consists in public constraint satisfaction that can be learned from a corpus and enforced through dialectical evaluation. Where empirical facts matter, they enter as constraints to be supplied or checked. Floridi's critique is most devastating where posterior evaluation is straightforwardly empirical. In philosophy, much posterior evaluation is internal to the argumentative exchange itself. The game can be learned because the game is played in text: recurring theory-level criteria and argument-level challenge-response patterns. Minimal prompting works because genre cues select those patterns. Floridi's strongest objection—the absence of posterior evaluation—can be met by relocating evaluation to the public procedures by which philosophical texts are ordinarily tested. The sceptic cannot rest content with "it's stochastic." They must show that outputs produced under these constraint regimes systematically fail: that they collapse into ad hoc patching under pressure, cannot sustain precision without vacuity, cannot integrate commitments, or break under standard critical questions. If they cannot show that, the inference from "no internal abduction" to "no good philosophy" is unsupported. # 4. How to Generate Philosophy with AI - Worked examples demonstrate that minimal prompts yield outputs satisfying Williamson's theoretical virtues (simplicity, non-ad-hocness) and the dialectical moves (accommodation, substantiation) that constitute the genre. - A case where the output fails—exhibiting ad hocness, say—shows that we can identify the failure text-internally, without knowing the text was AI-generated. # 5. Conclusion - LLMs can produce novel, first-rate philosophy because the discipline's standards are text-internal and publicly learnable. - Floridi establishes that LLMs are stochastic mimics that learn reasoning structures from text; Williamson establishes that philosophy is evaluated by theoretical virtues intrinsic to the theory. If the learned structures produce outputs exhibiting those virtues, the mimicry is the mastery. - Philosophical competence is less ineffable genius than fluency in a public normative practice. The rules of the game are codifiable—and the machine has learned them.