## The Position and Its Force The constraint-satisfaction thesis, as it appears in the paper, runs roughly as follows. Good philosophy is assessed by whether texts exhibit publicly checkable properties: precision, cost-accounting, non-ad hocness, defeater-sensitivity, fair treatment of rivals, explanatory power, integration with background knowledge. These properties are text-internal. If an output satisfies them, the question of what produced it -- a human brain, a committee, an LLM -- drops out as evaluatively irrelevant. [[Timothy Williamson|Williamson]]'s theoretical virtues (simplicity, strength, elegance, non-ad hocness) are, in his own word, "intrinsic" to theories, not to theorists: > "Apart from its relation to E, the more T has the intrinsic virtues of a good theory, the better (ceteris paribus). It should be elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated." (Williamson, p. 354) [[John Bengson|Bengson]], Cuneo, and Shafer-Landau provide a more articulated version of the same thought. Their Tri-Level Method specifies five criteria -- accommodation, explanation, substantiation, integration, and theoretical virtue -- organised hierarchically, with data-handling at the first level, grounding the theory at the second, and virtues as tie-breakers at the third. They are explicit that these criteria are drawn from ordinary philosophical practice and that satisfying them need not involve self-conscious rule-following: > "We endorse the method not because it makes a philosopher's job easy; indeed, it is quite demanding. Nor are we drawn to its constituent criteria because they revolutionize philosophical thinking; on the contrary, all of them are familiar from the way many philosophers go about their business." (Bengson et al., p. 107-108) > "Whatever philosophical method is, it is something that is friendly to these activities. By this we mean that, in the paradigm case, implementing philosophical method involves engaging in such activities." (Bengson et al., p. 80) The activities include advancing arguments, raising objections, offering replies, providing clarification, developing explanations, and displaying sensitivity to the deliverances of logic, mathematics, science, and common sense. Crucially, Bengson et al. hold that these criteria serve a "dual role": they provide instructions for constructing a theory and standards for evaluating one: > "Methods themselves comprise a set of criteria that serve a dual role: they provide instructions for the construction of a theory, given the data, while also serving as standards by reference to which the merits of theories are evaluated." (Bengson et al., p. 77) This dual-role picture is precisely what the constraint-satisfaction thesis exploits. If the same criteria govern both production and evaluation, then a text that satisfies the evaluative criteria has, by that very fact, satisfied what production was supposed to deliver. Whether the production process was "genuine reasoning" or stochastic pattern-completion becomes irrelevant to the evaluation. The thesis has real force. It is not a trick. It draws on something deeply embedded in philosophical practice: blind review. We do not ask referees to assess the cognitive processes of authors. We ask them to assess texts. The constraint-satisfaction thesis takes this seriously and generalises it. ## Are the Constraints Jointly Sufficient? This is where the position faces its most serious pressure. Four distinct challenges follow, each representing a candidate property that might be missing from the constraint list. ### Significance / Interestingness A text might satisfy every constraint on the list -- it might be precise, cost-accounting, non-ad hoc, defeater-sensitive, fair to rivals, well-integrated -- and still be trivial. It might address a question nobody has reason to care about, or settle an issue whose resolution has no consequences for anything else. Technical competence on an insignificant question is not good philosophy; it is a competent exercise. There is a position that says significance is external to philosophical quality -- it is a sociological property of the question, not of the answer. On this view, evaluating a paper's quality and evaluating its importance are separate tasks. Referees do both, but they are different judgements. A paper on a trivial topic can be a flawless piece of reasoning; its insignificance is not a defect in the reasoning but in the choice of topic. This separation may not hold up. Consider Bengson et al.'s account of what inquiry is for. They argue that a method is "sound" just in case satisfying its criteria positions inquirers to achieve theoretical understanding: > "we propose to call a method 'sound' just in case satisfaction of its criteria thereby positions inquirers to achieve an ultimate proper goal of inquiry." (Bengson et al., p. 27) And theoretical understanding requires a theory that is "robust, answering a multitude of questions about the most important features of the domain under investigation" (Bengson et al., p. 29). The word "important" is doing real work here. A theory that answers only trivial questions about a domain does not yield understanding of it. So on Bengson et al.'s own account, significance is baked into the success conditions for philosophical method. It is not merely sociological. [[Finnur Dellsén|Dellsén]] et al. push a related point. Their account of philosophical progress holds that progress consists in putting people in a position to increase their understanding, where understanding is a matter of better representing the network of dependence relations between phenomena. Progress requires that the contribution make a difference to how well the dependence network is grasped. A technically flawless contribution that maps no new dependence relations, or maps only relations already well understood, does not constitute progress. This suggests significance is partly constitutive of philosophical quality, not merely of its reception. **The counter-position**: But one might insist that significance is a property of contributions relative to a state of the art, not a property of texts as such. The same paper might be significant at one time and insignificant at another. If constraint-satisfaction is a property of the text, and significance is a property of the text-in-context, then they are different kinds of property. The constraint-satisfaction thesis might be right about what makes a text *well-executed* while being silent about what makes it *worth executing*. This distinction may be enough for the LLM question: if we can show that LLMs produce well-executed philosophy on significant questions (where significance is determined by the human prompter), the thesis does its work. I speculate that this is the weakest version of the gap. Significance can plausibly be handled by prompt design -- the human selects the question; the LLM addresses it competently. The constraint list evaluates the addressing, not the selecting. ### Originality This is a harder case. A text might satisfy every constraint and yet be a recapitulation -- a competent restatement of existing arguments, perhaps in a slightly different configuration, but advancing the dialectic not at all. It meets the standards; it adds nothing. One position: originality is a quality-relevant property. Good philosophy must advance the dialectic -- it must make a move that was not already available. On this view, an LLM that recombines existing positions without generating a genuinely new contribution is doing something less than philosophy, however well it executes. A competing position: originality is a sociological property -- a property of a contribution's relation to prior literature, not of its internal quality. A paper that would have been original in 1990 but is published in 2026 does not thereby become a worse piece of reasoning. Its arguments do not weaken. What changes is its novelty relative to the discourse. If we hold the constraint list to be about quality of reasoning, novelty is external. Bengson et al. are suggestive on this. Their Tri-Level Method does not include an originality criterion. Accommodation, explanation, substantiation, integration, and virtue are the criteria. A theory that satisfies all five but happens to have been advanced before is, by the lights of their method, a good theory. The method is about the relation between theory and data, not between theory and prior theories. Dellsén et al. are more helpful to the originality-matters side. On their view, philosophical progress consists in putting people in a position to increase understanding. If a contribution is a recapitulation, it does not put anyone in a position to increase understanding (they already had it). So recapitulation is not progress, even if it is competent. But again, this is about progress, not quality in the narrow sense. A skilled pianist who plays a Chopin etude perfectly is demonstrating quality; she is not advancing music. I interpret this as follows: originality is probably not part of what makes a philosophical *argument* good, but it is part of what makes a philosophical *contribution* valuable. These are different evaluative registers. The constraint-satisfaction thesis operates at the argument level. If the concern is whether LLMs can produce good arguments, the absence of originality from the constraint list is not a defect. If the concern is whether LLMs can make valuable contributions to philosophy, originality matters -- but this is a different, higher-bar question. ### Depth of Insight This is the most elusive candidate and deserves the most careful handling. The intuition: some philosophy displays a quality of "seeing into" the problem -- a depth of penetration that goes beyond technical satisfaction of constraints. Two papers might both be precise, well-integrated, defeater-sensitive, and non-ad hoc, and yet one might illuminate the problem in a way the other does not. The difference is not a matter of meeting or failing to meet identifiable criteria. It is a difference in the quality of thought. **Position A** (the sceptic about depth as separate from constraints): What we call "depth" or "insight" is reducible to the constraint list, but in a way that is not immediately obvious. A "deep" paper is one that satisfies the constraints to an exceptionally high degree -- its explanations are especially powerful, its integration with background knowledge is especially comprehensive, its cost-accounting is especially thorough. We register this as "depth" because we are tracking the degree of constraint-satisfaction, not a separate property. On this view, depth is just excellence along the existing dimensions, not a new dimension. This position gets some support from Bengson et al.'s account of theoretical understanding. Understanding requires full grasp of a theory that is "accurate, reason-based, robust, illuminating, orderly, and coherent" (Bengson et al., p. 28-29). The property of being "illuminating" -- "going beyond a mere description of those features to explain why each exists or is instantiated" -- is already in the framework. Depth might just be a high degree of illumination. **Position B** (depth as irreducible): Some quality of philosophical work resists decomposition into the constraint list. Consider a philosopher who reframes a problem in a way that dissolves what previously appeared to be a genuine tension. The reframing is not a matter of meeting more constraints; it is a matter of seeing the problem differently. The new framing might actually violate some expectations (it might not "accommodate" certain data that were treated as data under the old framing, because it reveals them to be artefacts of a mistaken way of carving the problem). The quality of the reframing is not constraint-satisfaction but something closer to creative perception. This matters for the LLM question because it suggests that there is a kind of philosophical quality that is, in principle, not assessable by checking a list -- however long and sophisticated the list. If depth-of-insight is a real, irreducible property, and it is quality-relevant, then the constraint-satisfaction thesis understates what good philosophy requires. **Position C** (the pragmatist compromise): Even if depth is irreducible in some metaphysical sense, it is text-assessable in practice. A competent reader who encounters a deep philosophical reframing can recognise it as deep. They can point to what makes it deep: "the reframing dissolves the tension by revealing that X and Y are not in competition but address different questions." This is a public, text-internal description of what the depth consists in. So depth might be irreducible to the existing constraint list while still being a *publicly codifiable* property of texts. I speculate that Position C is the most defensible. Depth may not reduce to the five constraints listed, but it may still be a text-internal property. If so, the constraint-satisfaction thesis needs revision: the list is too short, but the principle (philosophical quality is constituted by publicly assessable properties of texts) survives. The revision strengthens the thesis rather than defeating it. ### Understanding Afforded to the Reader Dellsén's account of understanding as dependency modelling introduces a consideration that cuts differently from the others. On his view, understanding a phenomenon consists in grasping a sufficiently accurate and comprehensive model of how it is situated within a network of dependence relations: > "one understands a phenomenon just in case one grasps a sufficiently accurate and comprehensive model of the ways in which it or its features are situated within a network of dependence relations" (Dellsén, 2020, abstract) Understanding is not the same as explanation; one can understand something by learning that it has no explanation, or by learning what it is independent of. The feature relevant here is that understanding is gradable along two dimensions: accuracy and comprehensiveness. Now consider the question: does good philosophy merely satisfy constraints, or does it afford understanding to its readers? That is, does a text need not only to meet structural criteria but also to actually enable a reader to grasp dependence relations they did not grasp before? One position: understanding-affordance is a consequence of constraint-satisfaction, not a separate requirement. If a text accommodates the data, explains them, substantiates its claims, and integrates with background knowledge, then a competent reader who follows the argument will thereby come to understand the phenomenon better. Understanding is not extra; it is what constraint-satisfaction produces. A competing position: understanding-affordance is not automatic. A text can be technically flawless and yet opaque -- not because it is wrong, but because its structure, its choice of examples, its ordering of exposition, do not facilitate the reader's construction of a dependency model. Clarity of exposition, apt illustration, effective structuring -- these are not among the five constraints but they affect whether the text promotes understanding. If Bengson et al. are right that the aim of inquiry is theoretical understanding, and if understanding depends on the reader's grasp, then understanding-affordance is quality-relevant. I interpret Bengson et al. as sensitive to this: their account ties understanding to "full grasp" of a theory, which requires not just that the theory have the six properties but that the inquirer possess that grasp. A text that makes grasping difficult, even while embodying a good theory, is a deficient vehicle for understanding. Whether this counts as a defect in the philosophy or merely in the writing is a question the constraint-satisfaction thesis must answer. ## Are the Constraints Really Codifiable? This is a separate challenge from the sufficiency question. Even if the constraints are jointly sufficient (or close to it), the thesis requires that they be publicly articulable and assessable. The question is whether they actually are. Take "non-ad hocness." In practice, determining whether a modification is ad hoc requires philosophical judgement. A theorist adds a qualification to handle a counterexample. Is the qualification independently motivated, or is it introduced solely to save the theory? This determination often depends on background knowledge about what counts as an independently plausible principle, and that in turn depends on disciplinary training and philosophical sensibility. Williamson himself notes that the criteria are practically applicable even when their ultimate justification is unclear: > "The foregoing sketch leaves it far from clear what makes abduction such a good method. For instance, why should aesthetic criteria such as elegance contribute to the pursuit of truth? Nevertheless, the central role of abduction in the success of the natural sciences provides good reason to think that it is a good method, even though we do not fully understand why." (Williamson, p. 356) This suggests a nuanced picture. The criteria are publicly applicable -- referees can and do apply them -- but they may involve trained perception rather than algorithmic rule-following. A referee's judgement that a move is ad hoc is a judgement, not a computation. It can be articulated ("the qualification has no independent motivation; it is introduced solely to handle this case"), but the assessment of what counts as "independent motivation" is itself a philosophical judgement. **Position A** (the codifiability optimist): The criteria are codifiable in the sense that matters. They can be stated in public language, applied by multiple evaluators, and yield substantial (if not perfect) intersubjective agreement. Perfect algorithmic precision is not required. What is required is that they be public rather than private -- that a reader can check a text against them and report findings to others. This is enough for the constraint-satisfaction thesis because the thesis claims that quality is constituted by text-internal properties, not that quality can be computed by a mechanical procedure. **Position B** (the codifiability sceptic): The criteria are stated in public language but applied by trained judgement, and that trained judgement may not itself be codifiable. "Non-ad hocness" is recognisable by philosophers the way good wine is recognisable by trained tasters -- the recognition is real and intersubjectively reliable, but the basis of the recognition resists full articulation. If this is right, then a model that has learned patterns from text might approximate the surface-level features that correlate with non-ad hocness while missing the deeper property that trained judgement tracks. I interpret Bengson et al. as supporting something like Position A. Their insistence that the criteria are "familiar from the way many philosophers go about their business" and that satisfying them need not be "an act of self-conscious adherence" but can be "the upshot of competent engagement in ordinary philosophical activity" suggests that the criteria are woven into practice in a way that is learnable from exposure to the practice -- which is exactly what happens when an LLM trains on philosophical corpora. ## Necessary versus Sufficient Even if the constraint list is not sufficient for philosophical quality (because significance, originality, or depth are missing), the constraints are almost certainly necessary. No text that fails to be precise, that ignores costs, that is ad hoc, that is insensitive to defeaters, that straw-mans rivals, and that fails to integrate with background knowledge is going to count as good philosophy. How much does this matter for the LLM question? Two readings are available. **The deflationary reading**: Showing that LLMs can satisfy necessary conditions is interesting but not decisive. If sufficiency remains open -- if there are additional requirements that LLMs might not meet -- then the sceptic still has room. The sceptic can say: "Fine, LLMs meet necessary conditions. But they can't produce depth/significance/originality, which are also required." **The inflationary reading**: Showing that LLMs can satisfy all the *publicly articulable* constraints shifts the burden dramatically. The sceptic must now identify a specific property that (a) is quality-relevant, (b) LLM outputs lack, and (c) can be demonstrated from the text. This is a much harder task than simply gesturing at "real understanding" or "genuine reasoning." The burden has been relocated from "prove the LLM can reason" to "identify the textual deficit." And if the sceptic cannot identify such a deficit, then the claim that there must be one amounts to an unfalsifiable faith claim about the necessity of human cognition. I interpret Williamson as supporting the inflationary reading, even though he does not address LLMs. His insistence that theoretical virtues are intrinsic to theories, that theories rank on the abductive scale regardless of who produced them, and that we evaluate theories as "potential explanations" before knowing whether they are true -- all of this supports the idea that evaluation is of artefacts, not of processes: > "We can rank theories (or hypotheses) as potential explanations of our evidence. The point of the qualifier 'potential' is that a false theory is not the actual explanation of the data; in that sense, it does not really explain them. But we need to rank theories as potential explanations before knowing whether they are true, in order then to use the ranking to guide our judgments as to which theory is true." (Williamson, p. 353-354) ## The Bengson et al. Framework and Its Relation to the Constraint List The Tri-Level Method provides a more structured version of the informal constraint list. Here is the mapping, as I interpret it: - **Accommodation** (does the theory render the data likely?) maps loosely to precision and fair treatment of rivals (you must acknowledge what needs accommodating). - **Explanation** (does the theory explain why the data hold?) maps to explanatory power. - **Substantiation** (are the theory's claims defended and supported?) maps to cost-accounting and defeater-sensitivity. - **Integration** (does the theory cohere internally and with background knowledge?) maps to non-ad hocness and integration. - **Virtue** (is the theory parsimonious, elegant, etc.?) maps to Williamson's theoretical virtues. But the Bengson et al. framework adds something the informal list lacks: *hierarchy*. Not all criteria are equal. First-level criteria (accommodation, explanation) take priority over second-level criteria (substantiation, integration), which take priority over third-level criteria (virtue). The Cost-Benefit Method's failure, on their analysis, is precisely that it treats criteria as "roughly equal and aggregative," permitting "illicit trade-offs" where gains on one dimension compensate for losses on another (Bengson et al., p. 98-99). Their critique of this approach is sharp: a theory that achieves parsimony at the expense of explanatory scope "severely diminishes the theory's ability to promote understanding" (Bengson et al., p. 98-99). This hierarchy matters for the sufficiency question. It suggests that constraint-satisfaction is not a matter of tallying points but of meeting criteria in the right order. A text that is beautifully parsimonious but fails to accommodate the data is not good philosophy, no matter how elegant it is. The informal constraint list, stated as a flat list, misses this ordering. Bengson et al.'s framework provides the ordering. Does this give a more precise answer to the sufficiency question? It does, but not a conclusive one. The Tri-Level Method is designed to be sound -- to position inquirers for theoretical understanding -- and Bengson et al. argue that their method, unlike the four alternatives they consider, possesses all the methodological virtues (comprehensiveness, support-requiring, synopticity, multidimensionality, hierarchy, comparativity). If the method is sound, and a text satisfies its criteria in the right hierarchical order, then the text should position readers for theoretical understanding. This is close to sufficiency -- but "positioning for understanding" is not the same as "being good philosophy," because understanding also requires a reader who grasps the theory. The text is a vehicle; whether it reaches the destination depends on the reader too. ## Williamson's Theoretical Virtues and the "Intrinsic" Claim Williamson's claim that theoretical virtues are "intrinsic" to theories deserves careful examination. When he writes that a theory should have "the intrinsic virtues of a good theory," the natural reading is that these virtues are properties of the theory considered in itself -- its structure, its commitments, its explanatory relations -- rather than properties of the theorist's cognitive states or the theory's causal history. This reading supports the constraint-satisfaction thesis directly. If virtues are intrinsic to theories, and theories are (or are expressed as) texts, then the virtues are properties of texts. They can be assessed from the text. Provenance drops out. But Williamson also stresses that simplicity is not merely aesthetic. It is epistemically principled because it protects against over-fitting: > "The restriction helps us avoid mistaking noise for signal, which we do if we fit the current data too closely." (Williamson, p. 368) This raises a subtle question. Simplicity-as-over-fitting-avoidance is a property of the theory's relation to the data -- specifically, a property about how the theory will perform on *future* data, not just the data it was designed to handle. A theory that fits all current data perfectly but does so through ad hoc epicycles will likely fail on new cases. Is this assessable from the text? In one sense, yes: an experienced reader can recognise over-fitting. They can see when a theory has been gerrymandered to handle known cases without principled structure. In another sense, no: the ultimate test of over-fitting is predictive -- how the theory performs on new data. And that test goes beyond what any text can show. I speculate that this is where the constraint-satisfaction thesis meets a genuine limit. The constraints as stated -- precision, cost-accounting, non-ad hocness, defeater-sensitivity, fair treatment of rivals -- are proxies for something deeper: truth-tracking capacity. Williamson's discussion of simplicity makes this explicit: the point of simplicity is to track truth, not to be aesthetically pleasing. A text can satisfy the proxy constraints while failing to track truth, if the proxies are imperfect. The question is whether, for philosophy specifically (as opposed to empirical science), the proxies are good enough. Given philosophy's substantially text-internal character -- the fact that much of its evaluation involves checking properties of arguments and theories as expressed in texts -- the proxies may be closer to the real thing than they are in empirical science. In physics, a theory must ultimately face the tribunal of experiment; in philosophy, a larger share of the evaluative work is done by the text-internal constraints themselves. This is the paper's actual argument -- not that "the map IS the land" (an overclaimed slogan) but that philosophy's evaluative standards are substantially text-internal in a way that distinguishes it from empirical disciplines. ## Assessment The constraint-satisfaction thesis is strong but probably needs refinement rather than abandonment. The constraints listed are almost certainly necessary. They may not be sufficient if we include significance and depth-of-insight as quality-relevant properties. But the principle that underlies the thesis -- that philosophical quality is substantially assessable through publicly checkable properties of texts, rather than through properties of producers -- survives these refinements. What changes is the list of properties, not the location of evaluation. The Bengson et al. framework adds hierarchical structure that the flat constraint list lacks, and this structure matters: it explains why some texts that superficially "satisfy constraints" are still bad (because they satisfy the wrong ones, or satisfy them at the wrong level). The framework also provides a more principled connection between constraint-satisfaction and the aim of inquiry (theoretical understanding), which is exactly what the thesis needs to avoid the charge of formalism. The codifiability question remains genuinely open. The evidence supports the view that the criteria are codifiable enough -- publicly articulable, intersubjectively applicable, learnable from practice -- without being algorithmically precise. Whether this degree of codifiability is sufficient for LLMs to learn the criteria from training data is an empirical question, but the Bengson et al. observation that philosophers satisfy the criteria without self-conscious adherence suggests that implicit, practice-embedded knowledge is the norm even for humans. If the criteria are learnable as practice-patterns rather than explicit rules, then LLMs are well-positioned to acquire them. --- Sources: [[Timothy Williamson]], "Widening the Picture" (from *The Philosophy of Philosophy*); [[John Bengson]], Cuneo, and Shafer-Landau, *Philosophical Methodology: From Data to Theory* (2022); [[Finnur Dellsén]], "Beyond Explanation: Understanding as Dependency Modelling" (2020); Dellsén et al., "What is Philosophical Progress?" (2024). See also [[Notes/Paper-internal constraints as the locus of philosophical evaluation]], [[Notes/From capability to constraint structure]]. *La struttura normativa della filosofia potrebbe essere piu accessibile di quanto pensiamo -- non nascosta nei recessi della mente, ma depositata nei testi stessi, in attesa di chi sappia leggerla.*