## Artefact-Level Evaluation: Developing the Space of Positions ### 1. Setting Up the Kernel The artefact-level evaluation idea can be stated simply: philosophical quality is determined by features of texts, not features of text-producers. When a referee evaluates a submission under blind review, they assess whether the argument is valid, the commitments are clear, the costs are named, the rivals are treated fairly. They do not assess whether the author "really reasoned" or "truly understood." If this is right, then the sceptical argument -- LLMs cannot do genuine abduction, philosophy requires genuine abduction, therefore LLMs cannot do philosophy -- is answering the wrong question. The right question is whether the *output* satisfies the relevant standards. Your paper already develops this idea with considerable force. The question for stress-testing is: *how strong is the artefact-level claim, really?* There are multiple ways to read it, and they differ substantially in what they commit you to and what objections they face. I want to develop five positions, each as if I believed it, and then assess the terrain. --- ### 2. The Practical Reading: Text as Evidence for Something Beyond the Text On this reading, evaluating texts is *our best available method* for detecting philosophical quality, but quality itself is something that outruns the text. We read papers because we cannot read minds. The text is evidence for an underlying intellectual achievement -- genuine understanding, real engagement with the problem space, authentic sensitivity to reasons. When a referee judges a paper as good, what they are really doing (on this view) is forming a justified belief that the author has achieved something epistemically valuable, using the text as their primary evidence. This reading draws natural support from [[John Bengson|Bengson]] et al.'s account of theoretical understanding. They define understanding as a state that "agents possess just when they fully grasp a theory" with six properties: accuracy, reason-basedness, robustness, illumination, orderliness, and coherence (Bengson et al., p. 28-30). The language is explicitly agent-involving. Understanding is something inquirers *achieve*; it is "the state that agents possess." The theory must have certain properties, but understanding consists in the *agent's* full grasp of such a theory. The criteria are properties of the theory, but the *goal of inquiry* -- theoretical understanding -- is a state of the inquirer. If philosophical inquiry aims at theoretical understanding so construed, then the practical reading says: we evaluate texts as proxies for whether understanding has been achieved. A text exhibiting accuracy, reason-basedness, robustness, and so on is good *evidence* that the author grasps a theory with those properties. But the text is not the understanding itself; it is a trace of it. What does this imply for LLMs? On the practical reading, the situation is genuinely uncertain. An LLM output might exhibit all the textual markers of understanding -- precision, cost-accounting, defeater-sensitivity -- while no agent "fully grasps" anything. If quality tracks something agent-level, then the text alone cannot settle the question. The artefact-level evaluator is, on this view, using a reliable but fallible heuristic. And the heuristic was calibrated on human producers; applying it to LLMs is an extrapolation whose reliability is unknown. The practical reading also makes sense of a real phenomenon: philosophers *do* sometimes revise their assessment of a paper when they learn something about its provenance. If you discover that a paper was produced by randomly concatenating sentences from a Markov chain, you would rightly doubt it, even if the text happened to look coherent. The practical reading explains why: the text is evidence, and knowing the production mechanism undercuts the evidential force. A stopped clock is right twice a day, but knowing it is stopped defeats your justification for believing it. I speculate that this is where many working philosophers would land if pressed. They would say: "Of course we evaluate texts. But we evaluate them *as evidence for* good philosophical thinking. If the thinking is absent, the evaluation is undermined." This is a coherent position. Whether it is correct depends on whether philosophical quality is indeed something beyond the text. --- ### 3. The Constitutive Reading: Quality Just IS Text-Internal On the constitutive reading, there is no further fact about philosophical quality beyond what is assessable from the text. The artefact-level evaluation is not our best proxy for something deeper; it *is* the whole thing. Quality is constituted by the text satisfying certain constraints: precision, cost-accounting, non-ad hocness, defeater-sensitivity, fair treatment of rivals. If the text satisfies them, the philosophy is good. Full stop. Section 2 of your paper is most naturally read as advancing this position. The draft states: "philosophy asks for *artefacts that satisfy certain constraints*. Not inner states. Not production mechanisms. Artefacts." And the claim that "for competent philosophical readers, 'looks like good philosophy' in the evaluatively relevant sense just means 'the standards are satisfied in the text'. When they are satisfied, appearance is reality." This reading draws on a particular interpretation of [[Timothy Williamson|Williamson]]. Williamson describes theoretical virtues as "intrinsic" to theories: > "Apart from its relation to E, the more T has the intrinsic virtues of a good theory, the better (ceteris paribus). It should be elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated." (Williamson, p. 354) The word "intrinsic" is doing real work. If theoretical virtues are intrinsic to the theory, then they are properties of the artefact, not relational properties involving the theorist. Elegance, unity, non-ad-hocness, informativeness -- these can all be assessed by reading the theory. You do not need to interview the theorist. Similarly, Williamson's discussion of over-fitting operates at the level of the philosophical community's *output*, not individual minds: > "Strikingly, the philosophical community showed very little aversion to the multiplication of complication. A firmer preference for simplicity and elegance would have warned the community that something was going wrong." (Williamson, p. 369) This is a diagnosis of *papers* -- of what the published record looked like. It is not a diagnosis of individual cognitive states. [[John Bengson|Bengson]] et al. provide further support, though ambiguously. Their account of method states that methods "comprise a set of criteria that serve a dual role: they provide instructions for the construction of a theory, given the data, while also serving as standards by reference to which the merits of theories are evaluated" (Bengson et al., p. 77). The evaluation role is explicitly about "the merits of theories" -- not the merits of theorists. Theories are assessed by reference to criteria; the criteria can be stated; the assessment is of the artefact. Bengson et al. also note that the criteria are "familiar from the way many philosophers go about their business" (Bengson et al., p. 107-108) and that "implementing philosophical method involves engaging in such activities" as "advancing arguments, raising objections, offering replies to these objections, providing clarification, developing explanations" (Bengson et al., p. 80). These are all activities whose products are textual. A text that contains well-constructed arguments, replies to objections, clarifications, and explanations *just is* an implementation of philosophical method -- or so the constitutive reading claims. The strongest form of the constitutive reading amounts to a deflationism about philosophical quality: there is no "hidden variable" that makes philosophy good beyond what is publicly assessable in the text. This is philosophically bold. It implies that blind review is not merely an imperfect approximation of ideal evaluation, but is methodologically ideal: the text is the complete object of evaluation, and knowing the author adds nothing evaluatively relevant. What does this imply for LLMs? If quality is constituted by text-internal features, then provenance is irrelevant in principle, not merely in practice. An LLM output that satisfies the standards is good philosophy, regardless of how it was produced. [[Luciano Floridi|Floridi]]'s own concession points toward this possibility -- he asks whether, "if an AI can generate the same explanatory hypothesis a human would, does it matter that the process was different?" and answers: "regarding the content of the hypothesis and our interpretation of it, maybe not" (Floridi et al., p. 12-13). On the constitutive reading, the answer is a firm "no, it does not matter." But the constitutive reading faces a challenge. Bengson et al.'s account of the *goal* of inquiry is explicitly agent-involving: theoretical understanding is "the state that agents possess just when they fully grasp a theory" with the relevant properties (Bengson et al., p. 28). The criteria for evaluating theories are theory-level, but the *point* of the whole enterprise is agent-level understanding. If you sever the connection between text quality and agent understanding entirely, you lose what Bengson et al. say the enterprise is for. I interpret this tension as genuine. The constitutive reading works well for the evaluation of philosophical *texts* but may not capture everything that matters about the philosophical *enterprise*. The question is whether evaluation and enterprise can be pulled apart. --- ### 4. Intermediate Positions: Splitting the Evaluative Landscape Perhaps some aspects of philosophical quality are text-internal while others are not. This would give us a partition: certain evaluative dimensions are fully capturable from the artefact, while others require something more. A natural partition: *validity*, *internal coherence*, *explicit cost-accounting*, *clarity of commitments*, and *fair treatment of rivals* are plausibly text-internal. You can check these by reading. *Significance*, *depth*, *genuine novelty*, and *sensitivity to the right problems* may be harder to assess purely from the text. Consider significance. Whether a philosophical contribution matters depends partly on the state of the literature -- what has been said before, what problems are live, what moves are exhausted. A referee can assess significance only if they know the field. But note: this is not knowledge about the *author*; it is knowledge about the *dialectical context*. A text's significance is relational (text-to-literature), not producer-dependent (text-to-producer). So significance may be context-sensitive without being provenance-sensitive. An intermediate position could hold that philosophical evaluation is text-plus-context, where "context" means the dialectical landscape, not the author's biography. Depth is trickier. What distinguishes a "deep" philosophical contribution from a merely correct one? One answer: deep contributions reveal previously obscured structural connections, reframe problems in ways that dissolve apparent impasses, or identify hidden assumptions whose removal opens new terrain. These seem assessable from the text -- you can see whether a reframing has occurred, whether an assumption has been identified. But there is another sense of "depth" that gestures at something like the author's long acquaintance with the problem space, their sense for what matters, their capacity to see the forest rather than the trees. This harder-to-articulate quality might resist full codification in text-internal terms. [[Douglas Walton|Walton]] et al.'s framework of argumentation schemes is relevant here. Schemes capture the argument-level dynamics: what counts as a legitimate move, what critical questions apply, what responses are demanded. These are fully codifiable and publicly checkable. As Walton et al. state: > "The method of evaluation of an argument fitting a scheme is that once the argument is put forward by a proponent, it may be defeated if the respondent asks an appropriate critical question that is not answered by the proponent." (Walton et al., p. 3) This evaluation procedure is entirely artefact-level. But schemes operate at the grain of individual arguments. Philosophy also involves theory-level integration -- the long-range coherence of a research programme, the way multiple arguments fit together into a systematic position. Whether this larger-scale coherence is fully text-assessable is less clear. It may require following a theorist's trajectory across multiple works. My speculation: the intermediate position is actually where the strongest defensible claim lives. Almost everything evaluatively relevant is text-internal (or text-plus-context), and the residua that are not text-internal are about *trajectory* and *programme*, not *provenance*. This position still delivers most of what your argument needs, because the sceptical argument targets the *individual paper*, and individual papers are evaluated at the artefact level even on the intermediate view. --- ### 5. How Philosophy Is Actually Practised The artefact-level thesis makes a claim about what philosophical evaluation *should* be, or what it *ultimately* is. But how does evaluation actually work? Blind review is the institutional instantiation of artefact-level evaluation. Referees are not told who wrote the paper; they assess the text. This practice presupposes that the text contains sufficient information for evaluation. If provenance were evaluatively essential, blind review would be irrational -- and yet it is the discipline's gold standard. But there are complications. Referees sometimes recognise authors from writing style, topic choice, or self-citation. When they do, this recognition *does* affect their evaluation, whether they intend it to or not. Studies in the sociology of science suggest that prestige effects are real: famous authors get more favourable readings. Whether this reflects a legitimate epistemic practice (deference to established competence) or a bias (status distortion) is itself debatable. Williamson's observations about academic fashion are pertinent: > "Academic fashions arise because people trained in a discipline have some respect for the judgment of others trained in the discipline as to what is good or fruitful work, worth imitating or following up." (Williamson, p. 429) This is a community-level mechanism that is partly sociological, partly epistemic. It does not straightforwardly track text-internal quality; it tracks perceived fruitfulness, which involves judgments about where a programme is heading -- judgments that go beyond any single text. Your extracted note on [[Provenance is not the right kind of variable in philosophical evaluation]] draws a useful distinction between *justificatory norms* (are the reasons on the page good?) and *triage norms* (given finite time, what deserves attention?). Philosophy's official self-understanding privileges justificatory norms. Triage norms involve provenance -- we read papers from *Mind* before papers from obscure venues, papers by authors we respect before papers by unknowns. But triage is not evaluation. The claim that LLMs short-circuit triage heuristics while satisfying justificatory norms is sociologically interesting but philosophically innocuous. I think the honest assessment is: in practice, philosophical evaluation is *mostly* artefact-level, with provenance operating as a triage filter and occasionally (illegitimately, by the discipline's own lights) leaking into substantive evaluation. The artefact-level thesis is not a radical departure from existing practice; it is an articulation of what philosophical evaluation officially claims to be. The radical move is applying it consistently, including to LLM outputs. --- ### 6. Comparison with Other Domains Mathematical proofs provide the clearest case where artefact-level evaluation is sufficient. A valid proof is valid regardless of who produced it. Automated theorem provers produce proofs that are evaluated entirely by reading them. No mathematician asks whether the computer "really understood" the theorem. The artefact carries the full justificatory load. But mathematical proofs have a property that philosophical arguments often lack: *decidable verification*. You can check each step mechanically. Philosophical arguments involve judgment calls -- whether an analogy is apt, whether a cost has been properly weighed, whether a rival has been fairly characterised. These judgments are publicly articulable but not mechanically decidable. A competent reader can make them, but competent readers sometimes disagree. Code provides an intermediate case. Code can be evaluated by running it -- there is a test suite, a compiler, observable behaviour. But code quality goes beyond passing tests: well-structured code is readable, maintainable, elegant. These qualities are assessed by reading, not running. Software engineering thus has both a mechanical evaluation layer (does it work?) and an aesthetic-structural layer (is it good code?). Both are artefact-level. Philosophy is perhaps closest to the aesthetic-structural evaluation of code. There is no compiler, no test suite. But there are publicly statable criteria -- the kind Bengson et al. and Williamson articulate -- and competent practitioners can apply them with reasonable (if imperfect) agreement. The question is whether the absence of mechanical decidability matters. I think the comparison with mathematics is instructive but potentially misleading. In mathematics, the artefact (the proof) is self-certifying: its validity is in principle checkable by anyone with sufficient competence, and competence is itself checkable (can you follow the steps?). In philosophy, the artefact (the argument) requires *judgment* to evaluate, and what counts as competent judgment is itself contested. This means that "the text is all there is to evaluate" is true in philosophy only if you add: "evaluated by competent judges exercising philosophical judgment." The evaluation is artefact-level, but it is not *mechanical*; it requires a reader who brings something to the text. This brings us back to a subtlety that [[Luciano Floridi|Floridi]]'s "anonymous forum poster" comparison inadvertently highlights: > "one should regard its output more as the opinion of an anonymous forum poster -- possibly correct, possibly incorrect -- rather than an expert." (Floridi et al., p. 19) On the artefact-level view, the anonymous forum poster comparison is exactly right -- and that is a *feature*, not a bug. An anonymous forum poster whose arguments happen to be brilliant should be evaluated on the merits. The post either contains good philosophy or it does not. But Floridi frames anonymity as a defect, implying that not knowing the source undermines the evaluation. This is the practical reading leaking in: the source matters because it is evidence for reliability. On the constitutive reading, the source is irrelevant. --- ### 7. Assessment of the Terrain Having developed these positions, where does the landscape stand? The **constitutive reading** is the most powerful for your paper's purposes, because it makes provenance irrelevant in principle and renders the sceptical argument straightforwardly unsound. But it faces the Bengson challenge: the goal of inquiry is agent-level understanding, not just text-level quality. If you sever quality from understanding entirely, you may win the argument about LLMs at the cost of a deflationism that many philosophers will find too austere. The **practical reading** is the weakest for your purposes, because it preserves a gap between text quality and "real" quality that the sceptic can exploit. But it is probably closer to what most philosophers implicitly believe. The **intermediate position** may be the most honest: almost everything evaluatively relevant is text-internal (or text-plus-context), and the residua that are not text-internal are about *trajectory* and *programme*, not *provenance*. This position still delivers most of what your argument needs, because the sceptical argument targets the *individual paper*, and individual papers are evaluated at the artefact level even on the intermediate view. A strategic observation: your paper may not need to choose definitively among these readings. The argument works against the sceptic on *any* of the three positions. On the constitutive reading, provenance is irrelevant; the text is everything. On the practical reading, the text is our best evidence; if the evidence is good, we should believe the conclusion regardless of priors about the producer (since we have no track record to consult for LLMs, but we do have the text to evaluate). On the intermediate reading, the text-internal dimensions -- which are the ones the sceptic would need to attack -- are assessable from the artefact. The sceptic's move of gesturing at production mechanism ("but it's just statistics!") fails on all three readings, because none of them makes production mechanism the right evaluative variable. The deepest residual worry, I think, is this: if philosophical evaluation is artefact-level, and LLMs can produce texts satisfying the criteria, does this show that LLMs "do" philosophy or that philosophy is easier to simulate than we thought? The constitutive reading says there is no difference between "doing" and "simulating" at the artefact level. The practical reading says there might be, but we cannot tell from the text. The intermediate reading says: for the dimensions that are text-internal, there is no difference; for the (possibly empty) residual dimensions, the question remains open. The honest answer is probably: the question "is this simulation or the real thing?" is either unanswerable from the evidence we have, or it is the wrong question to ask -- and either way, it does not help the sceptic. *Nella valutazione filosofica, il testo si fa giudice di se stesso -- ed e proprio questa autonomia dell'artefatto a rendere instabile ogni scetticismo fondato sulla provenienza.*