## Dialectical Saturation and Evaluative Data in Training: A Stress-Test Exploration
### 1. The Pattern/Norm Gap
The sceptical composite runs as follows. Floridi et al. establish that LLMs have a "stochastic core" -- they are "probability-distribution samplers over tokens" whose "process lacks explicit logical rules, deliberate hypothesis testing, or reference to an external world model" (Floridi et al., p. 8). Williamson, on one reading, holds that philosophy proceeds by abduction. Combined: LLMs cannot do philosophy because they cannot genuinely reason abductively.
The dialectical saturation thesis concedes the stochastic core but argues that the philosophical corpora on which LLMs train are so densely structured with argumentative patterns that distributional learning yields something more than Floridi's "mere form." The thesis has a first part -- saturation with move types, move sequences, and success conditions -- and a second, under-explored part concerning evaluative data. Let me take each in turn.
The gap between learning patterns and learning norms is sometimes treated as obvious: learning what typically happens is not the same as learning what should happen. A model that learns "repairs follow counterexamples" has learned a sequential regularity. A model that learns "good repairs follow counterexamples" has learned something normative. But what exactly is the difference between these two kinds of learning, and does it matter?
There are at least three ways to characterise the gap:
**Position A: The gap is real and unbridgeable by distributional means.** Norms are not patterns. Knowing that X typically follows Y is a descriptive fact about text; knowing that X *ought to* follow Y requires understanding *why* X is the right response. A model that has learned the distribution over philosophical moves has learned what philosophers *do*, not what they *should do*. The two happen to correlate, but the correlation is accidental from the model's perspective -- it has no access to the normative ground. This is essentially a Humean is/ought gap applied to training data: you cannot derive a normative conclusion from a distributional premise.
**Position B: The gap is real but narrower than it appears, because philosophical norms are constituted by practice.** This is the line Section 3 of the manuscript develops via Bengson et al. If satisfying the method's requirements "need not be an act of self-conscious adherence" but "can be the upshot of competent engagement in ordinary philosophical activity," then the norms just are the patterns of competent practice. A model that has reliably learned the patterns has, to exactly that extent, learned the norms -- because there is nothing else for the norms to be. The gap between "what philosophers do" and "what philosophers should do" collapses wherever the practice is competent. This position draws support from Bengson et al.'s observation that their endorsed criteria "are familiar from the way many philosophers go about their business" (Bengson et al., p. 107-108). If the norms are enacted rather than consulted, then the text just is the norm-instantiation, and learning the text is learning the norms.
**Position C: The gap exists but evaluative data in the training distribution closes it.** This is the under-explored position, and the one that needs the most development here.
### 2. The Evaluative Data Consideration
Philosophical corpora do not merely contain moves -- they contain evaluations of moves. The training data includes papers that cite other papers approvingly or critically. It includes responses that diagnose where a repair was ad hoc. It includes -- in the case of published collections, handbooks, and review articles -- explicit assessments of which contributions advanced the debate and which did not. Referee reports, editorial decisions, and citation patterns all encode evaluative information about which philosophical moves succeeded and which failed. The question is whether this evaluative dimension, which is genuinely present in the training distribution, closes the gap between learning patterns and learning norms.
#### Arguments for closure
**First**, the model does not just learn that non-ad-hoc repairs follow counterexamples; it learns that non-ad-hoc repairs *get approved* while ad hoc ones *get criticised*. The evaluative response is itself a data point. When a training corpus contains both a philosopher's repair and a critic's diagnosis that the repair was ad hoc, the model learns a distribution over the paired sequence: [repair] -> [criticism of ad-hocness]. When a different repair succeeds and is cited approvingly, the model learns a different paired sequence: [repair] -> [approval/uptake]. Over millions of such pairings, the model develops differential associations between repair types and evaluative responses. This is not the same as learning an explicit rule "ad hoc repairs are bad," but it functions analogously: given a dialectical context in which a repair is needed, the model's probability distribution assigns higher weight to non-ad-hoc repairs because the training data consistently pairs them with positive evaluative continuations.
Walton et al. provide a framework that makes this concrete. Each argumentation scheme comes with "an appropriate set of critical questions" -- the standard challenges that apply to arguments of that form (Walton et al., p. 3). Crucially, in the philosophical corpus, these critical questions are not merely listed; they are *asked and answered*, with the answers evaluated. The model has seen countless instances of a scheme being deployed, a critical question being raised, a response being given, and that response being assessed as adequate or inadequate. The evaluative dimension is built into the dialectical structure.
**Second**, the evaluative data are not separate from the training distribution -- they are part of it. Floridi himself concedes that LLMs have "absorbed patterns of human abductive reasoning as expressed in writing" and that "human-written text in its training data often results from IBE" (Floridi et al., p. 8; p. 10). But the same human-written text also contains evaluations of reasoning. A published philosophy paper does not just present an argument; it positions that argument relative to alternatives, explains why those alternatives fail, and defends its own approach against anticipated objections. These evaluative acts are linguistic acts -- they are expressed in text -- and the model has been trained on them. If Floridi grants that LLMs learn "reasoning structures" from text, then by the same mechanism they also learn evaluative structures from text.
**Third**, the evaluative data create differential reinforcement. I speculate (flagging this as my own inference, not a textual claim) that evaluative data function something like a second-order training signal. The first-order signal is the distribution over move types and sequences. The second-order signal is the distribution over evaluative responses to those moves. A model that has absorbed both levels has, in effect, learned something about which moves are valued and which are not. The resulting competence is not "understanding why X is good" in some deep philosophical sense; it is a reliable disposition to produce moves that would be positively evaluated by the community whose texts the model has ingested. Whether this counts as "learning norms" depends on whether norms can be exhaustively characterised by the evaluative dispositions of a competent community.
The Bengson et al. framework supports this reading. Their Tri-Level Method specifies criteria at three levels -- accommodation and explanation (level one), substantiation and integration (level two), virtues as tiebreakers (level three). The philosophical corpus instantiates these criteria not as abstract prescriptions but as patterns of evaluative practice: a paper is criticised for failing to accommodate a datum; a theory is praised for integrating with background commitments; a proposal is rejected as insufficiently substantiated. These evaluative acts are the criteria in action. A model trained on texts containing such evaluative acts has, in a distributional sense, learned the criteria.
#### Arguments against closure
**First**, evaluative responses in the corpus might themselves be unreliable or biased. Not all published evaluations are correct. Referee reports can be wrong; citation patterns can reflect fashion rather than quality; influential papers can be praised for reasons unrelated to their philosophical merit. If the evaluative data are noisy, then learning from them yields noisy norms -- the model learns what the community *tends to approve*, not what is *actually good*. This is the gap between descriptive consensus and normative correctness.
I think this objection has force but is also commonly overstated. The relevant question is not whether every evaluative signal in the corpus is correct, but whether the aggregate signal -- averaged over millions of evaluative acts across the corpus -- tracks genuine philosophical quality with sufficient reliability. If philosophical practice is even moderately truth-tracking (which any non-nihilist about philosophy must believe), then the aggregate evaluative signal will be correlated with quality, even if individual signals are noisy. The model's advantage is precisely that it aggregates: it is not beholden to any single referee's judgment but has absorbed the evaluative tendencies of the entire community.
**Second**, learning that X gets approved does not guarantee understanding *why* X is good. The model's differential association between non-ad-hoc repairs and positive evaluative responses is -- on a deflationary reading -- just another pattern. The model does not "understand" that ad hoc repairs are bad because they fail to illuminate the underlying phenomenon; it just has a higher probability of producing non-ad-hoc repairs. This is the deepest version of the pattern/norm gap: even with evaluative data, the model's competence might be "right for the wrong reasons" -- it produces good moves because they are statistically associated with approval, not because it grasps the normative ground.
The question is whether this matters. If we accept Bengson et al.'s position that norms are enacted in practice, then "grasping the normative ground" just is having the right evaluative dispositions. A philosopher who consistently produces non-ad-hoc repairs, recognises when a theory fails to accommodate the data, and generates substantiation when challenged is *doing philosophy* regardless of whether they have a meta-theory about why these moves are good. Most working philosophers do not consult their methodology textbook before raising an objection; they have internalised the practice. The question is whether distributional learning from evaluative data yields functionally equivalent internalisation.
**Third**, the evaluative data might be insufficient because of publication bias. Most philosophy in the training corpus is published -- i.e., it has already passed peer review. The model sees more approved than disapproved moves. It has limited access to the full distribution of philosophical attempts, including the failed ones that were rejected. This means the model may have learned what good philosophy looks like but may not have learned what bad philosophy looks like, because bad philosophy rarely makes it into the corpus. This could limit the model's ability to *avoid* philosophical errors, even if it can reliably *produce* philosophical moves.
This is a real concern, but I want to note a complication. Published philosophical texts contain extensive criticism of other published positions. Floridi et al.'s own paper is an example: it argues at length that a certain appearance of reasoning is misleading. The corpus is full of such second-order evaluative acts -- published papers arguing that other published papers are wrong. So the model does see negative evaluations, even if it does not see the original rejected manuscripts. The evaluative data include not just "this was approved" but also "this published view has the following defect."
**Fourth**, there may be a gap between "learning what the community approves" and "learning what is actually good." This is the social constructivism worry. If philosophical norms are constituted by community approval, then learning community evaluative dispositions just is learning the norms. But if there is a norm-independent fact about what counts as good philosophy -- if some community-endorsed work is genuinely bad philosophy despite being approved -- then distributional learning from evaluative data will track community opinion, not truth. The model becomes, at best, a model of the community's philosophical taste.
I take this to be a genuine open question that I cannot resolve here. But I note that the same worry applies to human philosophers: they too are trained (through graduate school, peer review, conference feedback) by absorbing the evaluative dispositions of their community. If the mechanism is suspect when it operates in an LLM, it should be suspect when it operates in a graduate student. The evaluative-data consideration does not give the LLM anything categorically different from what socialisation gives a human philosopher; it gives the same thing by a different route.
### 3. How Much Philosophical Competence Is Captured by Distributional Learning -- In Principle?
Is there a residue of philosophical skill that cannot be learned from text, no matter how rich the corpus? Or is text in principle sufficient?
**Position: Text is in principle sufficient for a large domain of philosophy.** The manuscript's Section 3 argues that philosophy is "grounded in the space of reasons itself" rather than in external entities. Unlike biology (grounded in cells) or physics (grounded in particles), philosophy's objects of study are "logical and inferential relations, not external entities." If this is right, then the symbol-grounding problem that plagues LLMs in empirical domains "is significantly weakened when the domain in question is the system of reasons the model has internalised." Philosophy's verification is largely internal: validity, consistency, dialectical robustness. The text says what it says -- the text *is* the data.
The Bengson et al. framework supports this: the criteria for theory construction and evaluation are "publicly checkable" and "assessable by competent readers." The publicly assessable character of philosophical norms means they are fully expressed in text -- a competent reader can determine from the text alone whether a theory accommodates the data, whether claims are substantiated, whether the theory integrates with background commitments. Nothing beyond the text is needed.
**Counter-position: Some philosophical competence requires extra-textual grounding.** Floridi insists that LLMs "lack grounded semantics connecting words to the physical world or perceptual experiences" (Floridi et al., p. 9). For philosophical domains that depend on empirical observation (philosophy of perception, environmental philosophy, philosophy of action), grounding matters. A model discussing the phenomenology of colour perception without any perceptual grounding may produce text that satisfies formal constraints but misses something important about the subject matter.
I interpret this as a genuine limitation, but a domain-specific one. Pure logic, formal epistemology, much of ethics, aesthetics of arguments, and philosophy of language are domains where textual grounding may be sufficient. Applied ethics, philosophy of mind when it concerns qualia, and philosophy of science when it concerns experimental practice may require more. The saturation thesis is strongest in domains where philosophy's self-grounding character is most pronounced.
### 4. Reproduction vs. Extension
Can distributional learning yield application to genuinely new philosophical territory, or only recombination of existing moves?
The [[Philosophical moves are combinatorial]] note argues that "much day-to-day analytic philosophy consists in a small set of repeatable argumentative manoeuvres" including reframing, distinction-making, model-import, counterexample design, error diagnosis, and synthesis. All of these are "in a literal sense, combinatorial" -- they require sensitivity to patterns, analogies, and argumentative templates.
The optimistic reading: if philosophical novelty is typically "non-obvious recombination under constraints," then a model with a large library of moves and a disposition to combine them in response to dialectical pressure could produce genuinely novel contributions. The model does not need to invent moves ex nihilo; it needs to see which combination of existing moves addresses the current deficit. This is what the [[The obvious move prompting technique]] describes -- the structurally apt continuation that may be rare but follows from the dialectical context.
The cautious reading: there may be a category of philosophical novelty that goes beyond recombination -- paradigm shifts, genuinely new conceptual frameworks, the invention of new philosophical vocabulary that reorganises a domain. These are not cases of applying known moves to known problems; they are cases of changing the rules of the game. Walton et al. acknowledge that their schemes are meant to "cover a large proportion of naturally occurring argument" (p. 39) -- but they do not claim to cover all possible argument. There may be moves outside the scheme repertoire.
Colton and Wiggins offer an instructive parallel from computational creativity. They distinguish between software that generates artefacts within known constraints and software that innovates "at the process level" -- inventing new methods rather than applying existing ones. The analogy suggests that LLMs might currently be competent at within-framework philosophical moves but not at the meta-level innovation that creates new frameworks.
I treat this as genuinely open. The evidence is ambiguous. What would count as evidence that an LLM had extended philosophical territory rather than merely recombining? One criterion: if a competent philosopher encounters the LLM's output and recognises a move that is not classifiable under existing scheme types -- something that does not fit Walton's catalogue or Bengson et al.'s criteria but is nonetheless philosophically valuable. Whether this has happened, or could happen, is an empirical question I cannot settle a priori.
### 5. Whether Saturation Varies by Tradition
Analytic philosophy is highly conventionalised. Its argumentative patterns are relatively explicit: premise-conclusion structures, objection-reply sequences, distinction-making, the deployment of thought experiments. These patterns map well onto Walton's schemes and Bengson et al.'s criteria. The evaluative conventions are also explicit: ad hocness is condemned, clarity is valued, integration with background science is demanded.
Continental philosophy is less conventionalised in this way. Heidegger's hermeneutic circle, Derrida's deconstructive readings, and Deleuze's conceptual creation operate with different -- and less transparent -- patterns. The evaluative norms are harder to extract from the text: what counts as a good deconstructive reading is not assessable by the same publicly checkable criteria that Bengson et al. describe.
This suggests that the saturation thesis gives LLMs stronger competence in analytic philosophy than in continental traditions. The patterns are more learnable, the evaluative data are more explicit, and the norms are more fully expressed in the argumentative texture of the text. For continental philosophy, the "remainder" -- the part of philosophical skill not captured by distributional learning -- is likely larger.
But I want to resist overstating this. Even continental philosophy has recurring patterns: genealogical critique, phenomenological description, the identification of presuppositions, the subversion of dichotomies. These patterns are in the texts. A model trained on Derrida's essays has been exposed to the pattern of identifying a privileged term in a binary opposition and showing how the subordinate term is constitutive of the privileged one. Whether the model has "understood" deconstruction is a different question, but it may have learned the moves.
### 6. The Connection to Floridi's Own Concession
Floridi concedes that "LLMs have effectively absorbed patterns of human abductive reasoning as expressed in writing" (p. 8) and that "the model has absorbed these patterns" of explanation and argument (p. 10). He further grants that "through exposure to billions of words, an LLM acquires a broad range of information about the world" including "common patterns of explanation and argument" (p. 8).
If the evaluative-data consideration is strong, these concessions become more significant than Floridi intends. He treats what has been absorbed as mere form -- structure without content, patterns without normative force. But if the evaluative dimension is part of what has been absorbed -- if the model has learned not just the pattern "repair follows counterexample" but the differential evaluation "non-ad-hoc repair gets approved, ad hoc repair gets criticised" -- then what has been absorbed is more than form. It includes a distributional proxy for normative competence.
Floridi's own question is instructive here. He asks: "if an AI can generate the same explanatory hypothesis a human would, does it matter that the process was different? From an epistemological standpoint, perhaps yes -- justification is significant -- but regarding the content of the hypothesis and our interpretation of it, maybe not" (Floridi et al., p. 12-13). He raises this question but does not develop it. The evaluative-data consideration offers a development: if the model has absorbed evaluative patterns alongside generative ones, then its outputs are not merely stochastically plausible -- they are differentially responsive to the very evaluative standards by which philosophy is assessed. The model does not just produce text that looks like philosophy; it produces text that has been shaped, through training, by the same evaluative pressures that shape human philosophical output.
Whether this constitutes "genuine reasoning" in Floridi's sense is a separate question. But it does suggest that the gap between "stochastic core" and "abductive appearance" may be less clean than Floridi's framework implies. The evaluative data in the training distribution blur the boundary: what the model has absorbed includes not just the patterns of reasoning but the community's evaluative responses to those patterns. The "abductive appearance" is not a surface phenomenon laid over a normatively empty stochastic process; it is the product of a stochastic process trained on normatively saturated data.
---
This exploration deliberately does not reach a verdict. The evaluative-data consideration is the strongest thread in the dialectical-saturation kernel, and the most under-explored. If philosophical norms are constituted by practice, and the practice includes evaluation, and the evaluation is expressed in text, and the model has been trained on the text -- then there is a genuine argument that distributional learning from evaluative data yields normative competence. Whether the argument succeeds depends on commitments about the nature of norms, the reliability of community evaluation, and the relationship between distributional learning and genuine understanding that I have tried to develop honestly from multiple sides.
*Le valutazioni implicite nel corpus filosofico potrebbero funzionare come una bussola normativa silenziosa, orientando le distribuzioni di probabilita verso risposte che la comunita non solo produce, ma approva.*