# Learning the Game
The question this section addresses is not whether LLMs have inner reasoning states in some philosophically loaded sense; it is whether they have absorbed and can enact the publicly checkable norms by which philosophical texts are produced and assessed. As noted in the introduction, by 'minimal prompting' I mean genre-governing cues rather than micromanaged instructions: directives like 'be philosophically robust' or 'explain your analysis before giving a final answer'. Such prompts specify what kind of thing is wanted—a philosophical artefact—rather than the specific moves to make. The claim is interesting precisely because thin constraints elicit substantial philosophical structure. What I want to show is that these thin constraints can trigger competent philosophical behaviour because the model has been trained on texts that instantiate the relevant norms as patterns of exposition and dialectical progression.
A recent account of philosophical methodology helps to make precise what 'competent philosophical progression' amounts to. In *Philosophical Methodology: From Data to Theory*, Bengson, Cuneo, and Shafer-Landau offer a systematic treatment of the norms governing philosophical inquiry. I do not claim that their account is the only possible codification, but it provides a concrete handle on what the norms of philosophical writing look like when made explicit. Their framing is instructive: method is 'the engine of inquiry', and methods are sets of criteria that guide both the construction and evaluation of theories:
> We turn now to method, the engine of inquiry. While the first stage of inquiry consists in the collection of data, the second centers on the construction and evaluation of theories. (Bengson et al., p. 77)
> Methods themselves comprise a set of criteria that serve a dual role: they provide instructions for the construction of a theory, given the data, while also serving as standards by reference to which the merits of theories are evaluated. (Bengson et al., p. 77)
This dual-role framing is central. The same criteria that tell theorists *how to build* a theory also tell readers *how to assess* one. Construction and evaluation answer to the same norms.
Bengson et al.'s model treats philosophical inquiry as having two major components: data collection and theorising. Theorising takes the data as input and, guided by method, produces a theory as output. The aim is theoretical understanding—the successful resolution of inquiry's guiding questions. A method is 'sound' just in case satisfying its criteria positions inquirers to achieve this goal:
> we propose to call a method 'sound' just in case satisfaction of its criteria thereby positions inquirers to achieve an ultimate proper goal of inquiry. (Bengson et al., p. 27)
The specific method they defend—the Tri-Level Method—articulates five criteria organised hierarchically. At the first level, theories must *accommodate* and *explain* the data. At the second level, the theory's claims and commitments must be *substantiated* (defended and explained) and *integrated* (cohering internally and with our best picture of the world). At the third level, virtues such as parsimony serve as tie-breakers. This gives us an ordered structure: data-handling first, grounding the theory second, virtues last. The structure matters because it specifies what kinds of 'demands' drive competent philosophical progression—what counts as the next thing to do when constructing or evaluating a theory.
Here is the key point for my purposes. The criteria are not esoteric; they are drawn from ordinary philosophical practice:
> We endorse the method not because it makes a philosopher's job easy; indeed, it is quite demanding. Nor are we drawn to its constituent criteria because they revolutionize philosophical thinking; on the contrary, all of them are familiar from the way many philosophers go about their business. (Bengson et al., p. 107-108)
> The criteria we'll endorse are familiar from the way many philosophers ply their trade, though these criteria have not yet been sufficiently justified, ordered, and integrated in a way that reveals how they can facilitate the principal aim of inquiry. (Bengson et al., p. 9)
Bengson et al. also characterise 'implementing philosophical method' in terms of engagement in specific activities:
> Whatever philosophical method is, it is something that is friendly to these activities. By this we mean that, in the paradigm case, implementing philosophical method involves engaging in such activities. (Bengson et al., p. 80)
The activities they enumerate include: advancing arguments, raising objections, offering replies to objections, providing clarification, developing explanations, and displaying sensitivity to the deliverances of logic, mathematics, science, and common sense. The implication is significant: philosophers who engage in these ordinary practice activities will, in so doing, tend to satisfy the criteria—whether or not they are explicitly aware of the criteria as such. Satisfying the method's requirements need not be an act of self-conscious adherence; it can be the upshot of competent engagement in ordinary philosophical activity.
This is the bridge I need. If philosophers can satisfy these criteria without intending to follow them explicitly, then philosophical texts will tend to instantiate the criteria as patterns of exposition and dialectical response. What gets done next in a philosophical text—what counts as an objection, what counts as a repair, what counts as progress—reflects these criteria. The method is not something philosophers consult like a checklist; it is something they enact in the activity of writing philosophy.
Philosophical corpora, then, do not merely contain conclusions; they contain recurring patterns of how philosophers move from a dialectical state to its demanded next step. If a view fails to accommodate some datum, the next demanded step is accommodation or defence of non-accommodation. If a claim lacks substantiation, the next step is to provide epistemic support or explain why none is required. If a theory conflicts with background constraints, the next step is integration or defence of the conflict. The methodology book helps specify what sorts of demands commonly drive that progression: accommodation and explanation at level one, substantiation and integration at level two, virtues only as tie-breakers at level three.
I want to be careful here and avoid overclaiming. The thesis is not that every philosophical text is a perfect instantiation of the Tri-Level Method. It is that the criteria create typical 'next things to do', and those next things show up in texts with sufficient regularity that a model trained on the corpus can learn the distribution. The patterns are there to be extracted.
Philosophy is truth-directed, but it often lacks cheap external answer keys. Unlike empirical sciences with experimental verification, or mathematics with formal proof, philosophy relies heavily on public, text-assessable constraints—validity, explanatory fit, integration with background knowledge, non-ad-hocness—as the way to track truth under conditions of limited direct verification.
The contrast with code is instructive. In programming, there often is a relatively crisp 'oracle': the code compiles, runs, and passes tests—or it does not. This makes both evaluation and iterative improvement straightforward. A model can generate code, run it, observe whether it fails, and adjust. Success signals are clear, feedback loops are tight, and outputs are easy to score. This is one reason why LLM progress in coding has been so visible.
Philosophy lacks that kind of immediate runtime verdict. There is no compiler that rejects invalid inferences, no test suite that flags unmet explanatory burdens. But the discipline is not therefore unconstrained. Rather, philosophy's public constraints and dialectical procedures play an especially central role in tracking truth precisely because direct verification is unavailable. The standards are encoded in how philosophers actually respond to each other's work: what they accept, what they challenge, what repairs they demand, what moves they treat as successful.
If the practice relies on public criteria and the texts instantiate them as recurring patterns, then a model trained on that text is positioned to learn the patterns—not necessarily as explicit rules it can articulate, but as reliable expectations about what comes next in philosophical writing. This is the 'learn the game' claim. A model trained on philosophical corpora has encountered countless instances of: accommodation moves, explanatory moves, objection-and-reply sequences, integration with background commitments, appeals to parsimony and other virtues. It need not represent these as labelled categories to have absorbed the distributional regularities they create.
The methodology book's own insistence that the criteria are 'familiar from the way many philosophers go about their business' supports this. The patterns are not hidden; they are the visible texture of philosophical writing.
This explains why a minimal prompt can trigger robust philosophical behaviour. A cue like 'What's the obvious move here?' does not feed premises or walk the model through an inference; it functions like a deictic instruction—'from here, do what's demanded'. The model has learned what is typically demanded in philosophical contexts; the prompt activates that knowledge.
But 'the obvious move' is not a single kind of move. Depending on what is currently missing, the demanded next step could be: making a distinction to resolve an apparent tension; unifying disparate considerations under a common principle; reconstructing an opponent's argument charitably; handling an objection; strengthening an explanation; integrating with background constraints; or something else entirely. The methodology book helps here: because it characterises objections as targeting deficits with respect to the criteria (accommodation failures, explanation failures, substantiation gaps, integration conflicts), it implicitly characterises what kinds of repairs count as 'the next thing to do'. The 'obvious move' is whatever addresses the current deficit.
A further consideration reinforces this picture. Bengson et al. note alternative conceptions of philosophy's aim, including Gutting's emphasis on 'knowledge of distinctions and of the strengths and weaknesses of various pictures and their theoretical formulations' and Nozick's and Wilson's celebration of 'the amassing of theoretical options' (Bengson et al., p. 25, fn. 16). What these conceptions have in common is an emphasis on shared frameworks that underwrite philosophical disagreement: inventories of distinctions, catalogues of similarities and differences, records of necessary-condition claims, rosters of dead ends and open possibilities.
These shared frameworks are heavily textual. A philosopher entering a debate does not encounter raw phenomena; she encounters a structured dialectical landscape—positions already staked out, objections already lodged, responses already attempted, options already foreclosed or left open. This supports my earlier observation that 'real-world evidence' in much analytic philosophy is backgrounded and unremarked: the background is precisely these shared frameworks. It also supports a stronger thought: the model has been trained not just on controversial theses but on the shared scaffold that makes serious philosophical disagreement possible. The scaffold is in the texts.
Bengson et al. give us method-level criteria for constructing and appraising theories. But philosophical competence also involves navigating argument at a finer grain: recognising common inference patterns, knowing what critical questions apply, understanding when a challenge has been met. A complementary codification at this scale comes from the argumentation theory literature, particularly Walton, Reed, and Macagno's work on argumentation schemes.
The transition is this: Bengson et al. tell us what it takes for a *theory* to be adequate (accommodation, explanation, substantiation, integration); Walton et al. tell us what it takes for an *argument* to succeed in dialogue (fitting a recognised scheme, withstanding the relevant critical questions). Both belong in this section because competent philosophical writing involves both: building theory-shaped contributions *and* navigating challenge-response dynamics.
Argumentation schemes are structures of inference that represent common types of arguments:
> Argumentation schemes are forms of argument (structures of inference) that represent structures of common types of arguments used in everyday discourse, as well as in special contexts like those of legal argumentation and scientific argumentation. (Walton et al., p. 1)
Each scheme comes with matched critical questions—the standard challenges that apply to arguments of that form:
> Each argument of this type is presented as providing only a defeasible support for its conclusion, subject to critical questioning in a context of dialogue. Matching each argumentation scheme is an appropriate set of critical questions. (Walton et al., p. 3)
The evaluation procedure is explicit:
> The method of evaluation of an argument fitting a scheme is that once the argument is put forward by a proponent, it may be defeated if the respondent asks an appropriate critical question that is not answered by the proponent. (Walton et al., p. 3)
The dialectical structure is: move, critical question, response. This is the game at the argument level, and its rules are stated.
The point for my purposes is that philosophical corpora contain countless instances of scheme-like patterns being enacted. A model trained on those texts plausibly learns a distribution over such patterns. When Walton et al. write that 'a theory of argumentation schemes should be... rich and sufficiently exhaustive to cover a large proportion of naturally occurring argument' (p. 39), they are saying that the schemes capture what happens in real argumentative practice. That practice is what the model has been trained on.
Taken together, Bengson et al. and Walton et al. give us a principled way to characterise what it is for a model to have internalised philosophical competence as expressed in texts. At the theory level: the model can generate contributions that satisfy the criteria—accommodating data, explaining phenomena, substantiating claims, integrating with background commitments. At the argument level: the model can navigate scheme-like challenge-response dynamics—recognising what critical questions apply, what responses are demanded, what moves count as adequate replies.
I am not claiming that this proves the model 'reasons' in some metaphysically loaded sense. I am claiming that it explains why minimal prompts can trigger robust philosophical behaviour. The model has learned patterns at both levels; a genre-governing cue activates that competence.
This brings us back to the 'abductive appearance' worry. If the norms that constitute competent philosophical practice are public, textually manifested, and assessable by competent readers, then an output that satisfies them is not 'mere appearance' in the relevant evaluative sense—it is an instance of the artefact meeting the discipline's standards. The artefact either accommodates the data or it does not. It either substantiates its claims or it does not. It either integrates with background constraints or it does not. These are checkable features of the text.
This does not settle metaphysical questions about the model's mental states. But it relocates the argumentative burden to the level that matters for peer review. If an LLM output satisfies the publicly checkable standards by which philosophical texts are assessed, then saying 'it only appears to do philosophy' is a verbal trick, not a philosophical objection. The burden shifts to the critic: name the specific textual failure—the equivocation, the ad hoc patch, the unmet explanatory burden, the integration conflict—or accept that the artefact meets the standards.
Philosophy is peculiar in a way that strengthens this conclusion. Unlike biology, grounded in cells, or physics, grounded in particles, philosophy is grounded in the space of reasons itself. The objects of philosophical study are logical and inferential relations, not external entities. Philosophy's verification is largely internal: validity, consistency, dialectical robustness. The symbol-grounding problem that plagues LLMs in empirical domains is significantly weakened when the domain in question is the system of reasons the model has internalised. For competent philosophical readers, 'looks like good philosophy' in the evaluatively relevant sense just means 'the standards are satisfied in the text'. When they are satisfied, appearance is reality. The paper is good.