> [!danger] Claude: Nothing Here Is Settled > Nothing in this section is settled. Do not assume that I want Section 4 to be anything like what is here at the moment — neither content-wise nor structurally. These are working notes, not commitments. Treat everything below as provisional raw material. ## Section 4 — Moves (revised) - If philosophical evaluation concerns intrinsic virtues of texts — elegance, unity, non-ad-hocness, combining simplicity with strength — then the question of whether LLMs can produce good philosophy is the question of whether they can produce texts exhibiting these properties. Sections 1–3 established this framing and argued that process-based objections do not undermine it. What remains is the constructive case: can LLMs actually produce such texts, and if so, how? - I want to grant Floridi et al.'s diagnosis completely. LLMs are "engines of generative plausibility": "given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence" (Floridi et al. 2024). They perform "zeroth-order abduction" — producing outputs that exhibit explanatory structure without selecting those outputs by comparing alternatives. All of this is correct at the level of mechanism. But statistical probability is relative to training data. What the model has learned to treat as "plausible" depends entirely on what it was trained on. So the question becomes: what does the training data encode? - The philosophical corpus is not a random sample of text. It is the output of a multi-level filtering process that selects, at each stage, for properties tracking Williamson's intrinsic virtues. - Peer review selects for handling of objections, engagement with the literature, non-trivial contribution — filtering out the arbitrary and ad hoc. - Citation selects for arguments that prove useful — arguments other philosophers find themselves needing to address, refine, or build upon — filtering for explanatory power and integration with existing work. - Teaching and anthologising select for clarity, illumination, and pedagogical power — filtering for elegance and unity. - Sustained philosophical attention selects for depth — works that reward re-reading because their arguments have structure worth unpacking. - The filtering is noisy: bad philosophy gets published, popular but mediocre work gets cited more than excellent but obscure work. But noisy filtering is still filtering. The tendency is toward virtue, even if individual data points deviate. - This claim requires empirical grounding — the proportion of academic philosophy in training data, the actual degree of filtering, and the training pipeline's selection mechanisms are questions that should not be answered by stipulation. What follows assumes that the tendency exists and is non-trivial, not that the filtering is perfect or comprehensive. - An LLM trained on this corpus learns the distribution of text that has survived these filters. The learned probability distribution is shaped by the intrinsic virtues — not because the model has been instructed in those virtues, but because texts exhibiting them are overrepresented in the training data relative to texts that lack them. Williamson writes: "Apart from its relation to E, the more T has the intrinsic virtues of a good theory, the better (ceteris paribus). It should be elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated. It should be informative and general. In brief, it should combine simplicity with strength" (2024, p. 354). The filtering process selects for exactly these properties. The virtues are therefore _latent_ in the model: implicit in the statistical regularities of the learned distribution, recoverable from the model's outputs, but not explicitly represented as rules or criteria the model applies. - This is like the relationship between a language model and grammar. A model trained on grammatical text produces grammatical outputs without having been taught grammar as a set of rules. The grammatical patterns are latent in the distribution — the model has absorbed them from the data without being given the rules explicitly. Similarly, a model trained on philosophically filtered text produces outputs tending toward philosophical quality without having been taught the evaluative criteria. The quality patterns are latent in the distribution. This is not a claim that every LLM output is good philosophy, any more than every output is grammatical. It is a claim about the tendency of the distribution — the direction in which the probability landscape slopes. - Even Zahavy concedes the relevant competence. He grants that LLMs can handle deductive work from given materials and explicitly restricts his critique: "we emphasize that this proposal is specifically tailored to the physical sciences, where the object of study is external material reality. In abstract domains such as Mathematics or Computer Science, the Sense Experience (E) may be grounded in high-dimensional topology or have other goals such as generality or minimality" (Zahavy 2026). Philosophy is one of those abstract domains. Its materials — arguments, distinctions, thought experiments, the logical space of positions — are textually available. They are the training data. The E→A jump that Zahavy claims LLMs cannot make is a jump from bodily experience to formal axioms; in philosophy, the "axioms" are already articulated in language and already in the corpus. Williamson himself notes that philosophy's evidence base includes "whatever knowledge the natural and social sciences, philosophy, and common sense have already gained" (2024, p. 356) — and this knowledge is textual. - But latent does not mean automatically expressed. Unprompted, LLMs produce generic, hedging text — surveys, overviews, cautious summaries. The intrinsic virtues are in the distribution but are not the default output. If they were, every LLM response on a philosophical topic would be good philosophy, which is manifestly false. A model trained on virtue-filtered text can produce texts exhibiting those virtues, but the capacity is not exercised by default. The prompt determines when it is. - The prompt determines which region of the continuation space the model generates from. The probability distribution the model has learned extends over a vast space of possible continuations. The prompt constrains which region the model generates in. Different prompts access different regions, and these regions differ in how reliably they exhibit intrinsic virtues. A bare question — "What is consciousness?" — activates a region dominated by survey-type text: cautious, generic, low in philosophical quality. This is the most probable continuation because it is the most common type of text following such prompts in the corpus. A dialectically structured prompt — one that lays out a position, identifies its vulnerability, and gestures toward a repair — activates a different region, where the most probable continuation is a philosophical _move_: the next step in the dialectic. - The prompter's skill consists in writing text whose good continuation — in the statistical sense of "most probable given the learned distribution" — is also good philosophy. Three modes of prompting access increasingly virtue-dense regions of the distribution: - Dialectical framing (one-shot, problem-oriented): pose a question embedded in dialectical context — not "what is X?" but "given these considerations, what follows?" or "the obvious objection is Y; address it." The training data is densely populated with such dialectical responses at the appropriate points in the argumentative structure. Walton, Reed, and Macagno's argumentation schemes formalise this: each scheme comes with licensed "critical questions" — the canonical pressure points. These are exactly the moves the corpus contains thousands of instances of, and exactly the moves a well-prompted model will produce. - Solution-gestured prompting (one-shot, solution-oriented): write a paragraph that points toward a solution without fully articulating it, so the good continuation is the next step in developing that solution. Richer than dialectical framing because the prompt itself contains philosophical content — it begins an argument, and the model continues in the direction indicated. - Conversational iteration (multi-turn): the prompter and the model produce philosophy together in an iterative process — write, continue, refine, develop, object, repair. Each turn further constrains the continuation space. The intrinsic virtues of the emerging argument increase with each round because each round further specifies what "good continuation" means. This mode sits on a continuum of autonomy: the prompter provides direction, constraints, and editorial judgment; the model provides dialectical moves, articulation, and pattern-completion. Neither is doing philosophy alone; what they produce together is a text exhibiting intrinsic virtues. - In a corpus filtered by intrinsic virtues, what Floridi calls "plausible continuation" and what Williamson calls "exhibiting intrinsic virtues" are not independent properties. They are correlated — because the filtering shaped what counts as plausible. The discipline produced text; the filtering selected text exhibiting intrinsic virtues; the filtered text became the training data; the LLM learned the distribution of the filtered text; the LLM's "plausible continuation," in the right context, therefore tends to exhibit the intrinsic virtues encoded in the distribution. This does not require the LLM to understand the intrinsic virtues, or to apply them as criteria, or to evaluate its outputs against them. It requires only that the training data was shaped by those virtues — which it was, because that is what philosophical filtering consists in. Floridi et al. themselves raise the question: "if an AI can generate the same explanatory hypothesis a human would, does it matter that the process was different? From an epistemological standpoint, perhaps yes — justification is significant — but regarding the content of the hypothesis and our interpretation of it, maybe not" (2024). For philosophy, the answer to their question is: it does not. - Lipton's distinction between likeliness and loveliness illuminates why this convergence holds. The "likeliest" explanation is the most probable; the "loveliest" is the one that "would, if correct, be the most explanatory or provide the most understanding" (Lipton 2004, p. 59). These can diverge: a conspiracy theory may be lovely (it unifies many apparently unrelated events) without being likely. But in a corpus filtered for loveliness — where the texts that survived peer review, citation, and anthologising are those judged illuminating, elegant, and explanatorily powerful — the likeliest continuation in the model's learned distribution tends also to be the loveliest in Lipton's evaluative sense. The filtering has aligned statistical probability with philosophical quality. Williamson further notes that "we rank only those potential explanations that have been thought of" (2024, p. 355). The philosophical corpus is the record of what has been thought of — and what survived the filtering. The model has absorbed this ranked space. - The "just statistics" dismissal confuses levels of description. Lipton: "arguing that Inference to the Best Explanation is wrong because Bayesianism is right is like arguing that thinking about technique cannot help my squash game because the motion of the ball is governed by the laws of mechanics" (2004, p. 108). The ball obeys mechanics whether or not you think about technique, but the mechanical description does not make the technique description idle. Similarly, an LLM's outputs are generated by stochastic processes over token distributions — and those outputs exhibit philosophical structure: they handle objections, draw distinctions, illuminate subject matter. The stochastic description and the philosophical description operate at different levels. Both are true. The fact that the mechanism is statistical does not settle the question of whether the outputs meet philosophical standards, because philosophical standards concern the output, not the mechanism. - The obvious worry: if the LLM is producing continuations shaped by existing filtered text, can it produce anything genuinely new? Williamson notes that "enumerative induction is inadequate for systematic philosophical theorizing, which often requires introducing new distinctions at a more abstract level not given in the data" (2024, p. 353) — and gives Dummett's distinction between assertoric content and ingredient sense as an example of a conceptual innovation that "cannot simply be read off the data." The model has learned not just particular arguments but patterns of argumentative _structure_ — patterns of how distinctions are drawn, how arguments are constructed, how positions are developed. These structural patterns can be instantiated in novel ways, producing arguments that do not appear verbatim in the training data but follow the patterns the training data established. Most philosophical innovation — most published, cited, taught philosophy — consists in exactly this kind of reconfiguration at higher levels of abstraction. The rare framework-introducing genius may be beyond current LLMs. But the bulk of what the discipline values does not require that kind of genius. And novelty, while not itself an intrinsic virtue on Williamson's list, is implicit in the virtues he does list: a theory that merely restates what is already known scores low on informativeness and generality — two of the virtues a good theory must have. - Two empirical questions arise. First, how much can a general-distribution LLM — one trained on the full breadth of human text, not specialised for philosophy — produce texts exhibiting intrinsic virtues? Second, would specialist training on philosophical texts improve performance? If the first question receives a positive answer and the second adds comparatively little, this suggests something about what philosophy is. Sellars characterised philosophy as the discipline concerned with "how things in the broadest possible sense of the term hang together in the broadest possible sense of the term" (_Philosophy and the Scientific Image of Man_, 1962). A system trained on the full breadth of human knowledge — on science, history, literature, law, ordinary discourse — has, in a sense, been trained on precisely the subject matter Sellars identifies as philosophy's own. The striving toward general intelligence, even if unachievable within current architectures, may itself be what positions these models for philosophical work — not because they have been taught philosophy specifically, but because they have absorbed the broadest possible range of how things hang together. If this is right, it deepens the encoding claim: the intrinsic virtues may be latent in the model not only because the philosophical corpus is filtered for quality, but because the general corpus encodes the breadth of connection that philosophical argument draws upon. - The paper itself is an instance of the process it describes. If the reader judges its arguments clear, its distinctions illuminating, its engagement with objections substantive, then the paper exhibits the intrinsic virtues it discusses — and these virtues are partly the product of the human-LLM collaboration it argues for. The paper was produced with a general-purpose LLM, not a system specialised for philosophy — which is itself evidence bearing on the two questions just raised. The paper does not need to demonstrate LLM philosophy as a separate exercise. It is a demonstration, submitted for blind review, evaluated by the very criteria it articulates. - Humanity asked a computer to do philosophy. It received the answer '42' — correct, according to the machine, but meaningless to the questioners, because they had never known what the question was. The problem was not with Deep Thought's capacities but with humanity's prompt. The intrinsic virtues were latent in the machine; what was missing was the right question to draw them out. Now we know what the question is — and we know that the answer, when the question is well-formed, can exhibit the philosophical qualities that the discipline has spent centuries learning to value.