%%MATERIAL TO INTEGRATE (moved from Introduction, 4 Apr 2026):
SOKAL FOOTNOTE (for the blind review / "In sum" paragraph):
In 1996, the physicist Alan Sokal submitted a paper to *Social Text*, a cultural studies journal, in which he argued that quantum gravity is a social and linguistic construct. The paper was a hoax — Sokal had written it to test whether a journal would publish an article that, as he later put it, 'sounded good' but whose arguments were nonsensical (Sokal 1996a, 1996b). The journal, which did not practise peer review at the time, published it under Sokal's own name, with his institutional affiliation attached — suggesting that what was being assessed was not so much the reasoning on the page as the person behind it.
TEST CASE (useful framing device for somewhere in the section):
Imagine a philosopher who says "I evaluate published arguments, not people — blind review is correct" (she accepts the text-focused conception) AND "but I think a text produced by an LLM isn't really philosophy, because there's no understanding behind it" (she thinks status within the text-focused conception still depends on pedigree). She is coherent. The introduction handles her first commitment. This section addresses her second.
%%
## The Challenge from Authorship
In this section we address what we might call the _challenge from authorship_: the idea that philosophy is something that only people, or at the very least _minds_, can do. An imperfect comparison would be with _art_: a lot of people, perhaps a majority, would argue that generative AI systems, in virtue of their not being people, cannot make art. This would not be to say that AI outputs cannot be aesthetically pleasing, only that they are not artworks: on this view, art is the product of the right sort of mental activity on the part of its maker — imaginative, expressive — and a system without mental states has no such activity to draw on.
- We might think, for similar reasons, that philosophy is a uniquely human activity, requiring the right sort of mental states to lie behind it. %%%%
We tend to read philosophical texts as evidence of this kind of understanding — as the product of someone who was thinking through a problem — and when we judge the text to be good, part of what we are judging is that the thinking which led to the text was good. **When a philosopher handles an objection well, we take this as evidence that she could see why the objection had force; when she draws a distinction that clarifies the terrain, we take it that she could see why the distinction was needed.** This may be why the history of philosophy is treated as part of philosophy in a way that the history of science is not: understanding a philosophical contribution seems to require understanding the thinking that produced it, in something like the way that understanding an artwork might require understanding what the artist was trying to express or create. **Philosophers study Frege not merely to catalogue his conclusions but to work through his reasoning — to follow the arguments of the *Foundations of Arithmetic* and to see the problem the way he saw it. If the arguments could be stripped away and only the conclusions retained, something philosophically relevant would be lost: not just a route to the result, but the understanding that working through the route makes possible.** If this is right, the question of whether an LLM can do philosophy does not arise. An LLM does not understand problems and cannot think through a difficulty in the relevant sense.
**But does the quality of a philosophical argument depend on the kind of agent that produced it, in the way that an artwork's status might depend on the intentions of its creator?** Putnam was not reporting a previously unnoticed item in the world; he was making a case, by way of thought experiment, that meanings are not fixed solely by what is in the speaker's head. The thought experiment does its work not by pointing to something outside the text — there is no Twin Earth for us to go and inspect — but by constructing a scenario whose internal logic puts pressure on a familiar picture of meaning. A reader who follows the argument does not simply learn that meaning is externally determined; she sees why, through the specific pressure the scenario puts on the assumption that mental life alone fixes what our words mean. That understanding could not be separated from the text that produced it. The philosophical contribution is not something the text reports; it is something the text does. **Someone who had never heard of Putnam, who knew nothing about his career or his reasons for constructing the scenario, would gain the same understanding from the same argument. Philosophical arguments are, in this respect, more like proofs than paintings. A proof is valid in virtue of its structure; nobody needs to consult the mathematician to check.**
Dellsén et al. propose that philosophy makes progress when philosophical research puts people in a position to increase their understanding — where increased understanding is a matter of more accurately or more comprehensively representing the dependence relations in which a phenomenon stands, or fails to stand, to others (2024, pp. 665, 680-81). Understanding, on this account, goes beyond knowing that something is the case. It involves grasping how one phenomenon depends, or does not depend, on another — seeing, for instance, not just that meaning is externally determined, but how the speaker's environment rather than the speaker's psychology fixes what words refer to. One speaker on Earth and another on Twin Earth share every psychological state and yet mean different things by the same word, because their environments differ in ways that bear on reference — a dependence relation of just the kind Dellsén et al. describe. A reader who works through the scenario does not simply acquire the belief that externalism is true; she comes to see why meaning depends on environment, and what features of the case make this so. On Dellsén et al.'s account, enabling that kind of understanding is what philosophical progress consists in. **And progress, so understood, happens "by way of philosophical ideas — theories, arguments, distinctions — becoming publicly available" (p. 679). What is publicly available is the argument. This bears on why history of philosophy is part of philosophy: we return to Frege not to reconstruct his psychology but because the *Foundations of Arithmetic* still puts readers in a position to grasp dependence relations they might not otherwise have seen.**
Not every account of a phenomenon's dependence relations is equally illuminating, however. If philosophical progress consists in enabling understanding, we need a way to distinguish views that genuinely reveal how things depend on one another from views that merely accommodate the data without explaining anything. Lipton distinguishes two ways in which an explanation might count as the best of its competitors:
> "We may characterize it as the explanation that is most warranted: the 'likeliest' or most probable explanation. On the other hand, we may characterize the best explanation as the one which would, if correct, *be the most explanatory or provide the most understanding*: the 'loveliest' explanation. The criteria of likeliness and loveliness may well pick out the same explanation in a particular competition, but they are clearly different sorts of standard. Likeliness speaks of truth; loveliness of potential understanding." (*Inference to the Best Explanation*, p. 59) my italics
Lipton illustrates the contrast with Molière's joke about the dormative virtue of opium. To say that opium sends people to sleep because it has a sleep-inducing power is, in Lipton's terms, the likeliest of explanations — almost guaranteed to be true, precisely because it says little more than that opium sends people to sleep. The explanation repackages the phenomenon without connecting it to anything beyond itself — it maps no dependence relation that the bare statement of the effect did not already contain. A lovely explanation, by contrast, would identify the conditions on which the effect depends, showing what it is about opium that produces sleep.
The same distinction applies in philosophy, though it cuts in a way that is not always recognised. A philosophical view can accommodate the familiar cases and survive the standing objections while doing nothing to connect those cases to their underlying conditions. Such a view handles whatever is put to it — each counterexample met with a new clause, each objection absorbed by a further qualification — but the resulting account, for all its case-by-case accuracy, leaves the reader no wiser about why the cases go the way they do. It is likeliest without being loveliest: defensible without being illuminating. The view survives by becoming more elaborate rather than more revealing, in much the same way that the dormative virtue survives by restating the phenomenon in slightly different words. What Dellsén et al. call philosophical progress requires something different — views that bring dependence relations into view that were not previously visible, views whose loveliness consists in enabling a reader to see how and why the parts of a subject bear on one another.
On the other hand, when likeliness prevails at the expense of loveliness there is no genuine progress. A philosophical view can accumulate ever more elaborate qualifications, each designed to handle a specific objection or accommodate a recalcitrant case, and yet leave the reader no closer to understanding why the cases go the way they do. The view survives by becoming more intricate rather than more revealing, in much the same way that the dormative virtue survives by restating the phenomenon. Williamson draws on Forster and Sober's (1994) work on curve-fitting in statistics to characterise the problem. An equation that passes through every available data point may nonetheless fail to predict new data, because it has mistaken noise for signal; the philosophical analogue is a theory that handles every counterexample by adding complexity, yet grows steadily harder to credit as it does so. Williamson's own example is the post-Gettier literature on knowledge, where each new case prompted a more elaborate analysis without bringing the subject into clearer view. A good philosophical theory, Williamson writes, should be "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated" and should "combine simplicity with strength" (2024, pp. 354, 368-69). These are the marks of a theory that earns its survival through genuine insight rather than through the accumulation of ad hoc qualifications; the marks, in Lipton's terms, of a theory that is lovely rather than merely likely.
Bengson et al. organise these evaluative concerns into a systematic method. Their tri-level framework asks, first, whether a theory accommodates and explains the data in its domain; second, whether the claims that do this explanatory work are themselves substantiated and integrated with one another; and third, whether the resulting theory possesses the relevant theoretical virtues (2022, pp. 108-09). The ordering is not arbitrary. A theory can fit every case and still fail, because the claims doing the explanatory work are poorly supported or because they sit uneasily alongside one another. And a theory can meet the first two levels and still lack the simplicity and coherence that would give it an edge over a rival that does equally well on the data. The third level — theoretical virtue — is where Williamson's desiderata enter: a theory that satisfies Bengson et al.'s first two levels while also combining simplicity with strength has a claim not just to survival but to the kind of progress Dellsén et al. describe.
In sum, the evaluative standards we have assembled all bear on what a philosophical text says and how it argues for it. Lipton's distinction between illumination and mere accommodation, Williamson's desiderata for theoretical virtue, and Bengson et al.'s method for assessing how well those standards are met: **each concerns the argument, not the kind of agent that produced it.** They do not ask how the author arrived at her argument but whether the argument, as it stands on the page, meets the relevant standards. **The practice of blind review in philosophy rests on the same assumption: referees assess what a paper achieves without knowing who wrote it. If the kind of agent behind the argument — whether a senior philosopher or a graduate student, whether a human being or a machine — were relevant to the argument's quality, then blind review would be a defective practice rather than the discipline's standard method of assessment.**
This is not unique to philosophy. Deep Blue, the computer that beat Kasparov in 1997, surveyed vastly more positions than any human could and selected the move most likely to win, what Gaut calls "the epitome of an uncreative way to play chess" (%%REFERENCE: Gaut 2010, fn. 23 — UNVERIFIED, needs checking against source%%). The moves it produced were strong chess all the same. The quality of a move does not depend on the manner of its selection; a move that wins material is a move that wins material, whether it was found by pattern recognition or by brute-force search. If a philosophical text meets the evaluative standards that Williamson and Bengson et al. articulate, if its arguments are lovely in Lipton's sense and its theory combines simplicity with strength, that achievement does not depend on whether it was reached by insight or by search.
The question, then, is what follows when a language model is trained on a corpus that this evaluative apparatus has shaped — a corpus filtered, through peer review, citation, teaching, and anthologising, for the very properties we have been describing. Whether an LLM trained on such a corpus can produce texts that meet these standards is the subject of the next section.[^pigliucci]
[^pigliucci]: We return in Section 3 to the question of worldly starting points and empirical constraint, where it bears directly on the grounding-style objection.