[▶ Live viewer](https://htmlpreview.github.io/?https://gist.githubusercontent.com/nickneek/46ee19235716697d7000780ab3f3867e/raw/generating-philosophy-moves-deck.html) · [▶ Map](https://htmlpreview.github.io/?https://gist.githubusercontent.com/nickneek/46ee19235716697d7000780ab3f3867e/raw/generating-philosophy-moves-map.html) · [▶ Split view](Attachments/tools/moves-split-view.html) · [gist](https://gist.github.com/nickneek/46ee19235716697d7000780ab3f3867e) · source: the talk deck (`Attachments/Generating Philosophy - standalone.html`) + the trailer (`Attachments/generating-philosophy-diagrammatic-offline.html`) # Generating Philosophy with AI — Argument Moves Move-by-move account of §§1–3 for the talk at the [[Sessions/Generating Philosophy|Generating Philosophy]] project's Thursday slot at the 2nd Lingnan–Genoa–Kobe Value Theory Conference (2026-04-23, Zoom). See also [[Generating Philosophy — Talk Hub]]. Moves carrying excessive meta-commentary have been rephrased to make the move rather than narrate it, preserving original wording wherever the meta-commentary wasn't doing work. --- # § 0 — Introduction ## The question - Move 1 — Can LLMs produce philosophy which is worth reading? ## "Worth reading" - Move 2 — "Worth reading" is the concept every philosopher already uses — every time they recommend a paper to a colleague, set one aside after a page, or judge a submission worth sending for review — and the talk takes it as given: a shared working understanding that the discipline trades on, without need for any stipulative definition, and that picks out the same range of cases however one elects to elaborate it. - Move 3 — Journals fit obviously into this picture: they strive to give their readers what is worth reading, though they do not always succeed. --- # §1 — The Challenge from Authorship ## The challenge - Move 1 — Philosophy is a person-only domain: no LLM text can be a work of philosophy, because no philosopher stands behind it — a view we have not seen explicitly articulated in the literature but which captures an intuition a fair number of philosophers seem to hold about what philosophy requires. ## The intuition - Move 2 — The case of art provides an imperfect comparison: many will deny that a purely AI-generated image is an artwork because no artist lies behind its creation, and the thought is that something similar holds for philosophy — no philosopher behind the text, no philosophy. - Sub-move — Both disciplines are often organised around individuals in their teaching and reception: a philosophy undergraduate might take a course on Kant's ethics; a fine-arts student, a course on Turner. ## Davies - Move 3 — The intuition can be made more precise by adapting Davies' performance theory of art: > The work — what the artist achieves — is the process eventuating in that product. Works themselves are neither structures nor objects simpliciter, nor are they contextualized structures or objects [...]. They are, rather, intentionally guided generative performances that eventuate in contextualized structures or objects. > — Davies, *Art as Performance*, p. 98 - Move 4 — Transposed to philosophy, the work is the philosopher's sustained activity in producing the text — her working through of a problem, her formulating and revising of arguments — and the text is what that activity leaves behind: what our engagement is directed at, but not itself what is evaluated. ## Why the transposition fails - Move 5 — The transposition fails because the feature of art that drives Davies' view has no analogue in philosophy: in art, surface underdetermines the work — a canvas-from-a-washing-machine indistinguishable from a Rembrandt, a Danto-style red square that differs in standing from a perceptually identical other, a molecule-identical forgery — but in philosophy no such underdetermination obtains, because two type-identical papers make the same arguments, face the same objections, and admit the same evaluations. - Move 6 — Philosophical evaluation is directed at what the text says: we ask whether premises are defensible, inferences go through, distinctions track real divisions, counter-examples hold against their targets, and conclusions survive the objections the text anticipates — questions whose answers are decidable without knowing who produced the text and that make no reference to any further activity lying behind it. - Move 7 — The achievement-talk objection fails at this point: one might protest that philosophers speak of what an author has *achieved* in a paper, which suggests something beyond the text to be assessed, but achievement-talk in philosophy is parasitic on the text — to say that the author has achieved something is just to say that she has produced a text with such-and-such argumentative properties, not to credit her with any further achievement lying behind them. - Move 8 — The text-focused norm is already institutionalised: philosophy journals strip author information from submissions before sending them to referees, and they do so not as practical convenience but as a matter of principle — what is to be assessed is what the paper says, and information about the author is treated as potentially corrupting that assessment. - Move 9 — Dellsén et al. (2024) make the corresponding normative point: philosophical progress, they argue, consists in putting people in a position to increase their understanding, and this happens paradigmatically through philosophical content — arguments, theories, distinctions, counter-examples — being made publicly available through publication; philosophical progress is, as they put it, a *for-whom* rather than *by-whom* matter, turning on whom the work puts in a position to understand rather than on the cognitive states of whoever produced it. - Move 10 — The Sokal affair is recognisable as a breach of this norm: when *Social Text* published Alan Sokal's 1996 hoax paper without peer review, what went wrong, on the academy's own assessment, was that the journal had assessed Sokal's institutional standing rather than the argument on the page — the recognition of error itself depending on a background norm that the text is what matters. ## §1 payoff - Move 11 — The challenge from authorship is a *constitutive* challenge: it treats the philosopher's activity not as something that produces the work but as part of what the work is, so that even a text indiscernible from a philosophical paper would fail to be philosophy if no philosopher's activity lay behind it — and the preceding argument rejects this: the work consists in the text, and evaluation goes no further than the text. - Move 12 — What remain for §§2–3 are two challenges of a different kind, both *capacity* challenges: each grants that the philosophical work is in the text, and asks instead whether an LLM has the particular mental capacity required to produce a text with the relevant properties — §2 considers the capacity for abductive reasoning, §3 considers the capacity for conscious experience, and the question in each case is whether a system that lacks the capacity can produce a text of the corresponding kind. --- # §2 — Likeliness, Loveliness, LLMs ## The charge - Move 1 — Floridi and colleagues diagnose LLMs with what they call *zeroth-order abduction*: an LLM prompted to explain a phenomenon produces a plausible-looking explanation, but it does so by pattern-matching over learned associations rather than by selecting among competing hypotheses — the appearance of abductive reasoning without the reasoning itself. > LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. > — Floridi et al., "What Kind of Reasoning (if any) is an LLM actually doing?", p. 9 - Move 2 — Williamson argues that contemporary philosophical theorising already proceeds partly by abduction from the armchair: when philosophers defend a view they do so by comparing it to rivals — examining how well each, if true, would explain the evidence — and by weighing what he calls *the intrinsic virtues of a good theory*, roughly simplicity combined with strength (Williamson 2024, pp. 354, 358, 368–69); it is this abductive work that Floridi's diagnosis describes LLMs as unable to perform. - Move 3 — The challenge from abduction is accordingly a *capacity* challenge: good philosophical work requires abductive reasoning; abductive reasoning requires a mental capacity for generating alternative hypotheses and weighing them against one another; humans possess this capacity and LLMs, on Floridi's diagnosis, do not; the objection concludes that an LLM cannot produce philosophical work that genuinely exhibits abductive argument, whatever the surface of its output looks like. ## The page - Move 4 — Floridi himself names the channel by which the patterns of human argument reach the LLM's output — training on philosophical text — and it is worth pressing this point, because what philosophy is conducted in is precisely this text. > This effect is due to the model's training on human-generated texts that encode reasoning structures. > — Floridi et al., Abstract - Move 5 — Philosophy is conducted in writing: a philosopher formulates an argument by writing it down, revises it by reading what she has written, submits the written result to be assessed by other philosophers who respond in further writing — the work of comparing hypotheses, handling objections, and weighing theoretical virtues is done in the prose itself, which is why evaluation consists in reading that prose. - Move 6 — An LLM's training data is this same philosophical prose: the articles, the handbook entries, the published replies and counter-replies that together constitute the written record of philosophical practice — so the training channel Floridi identifies is a channel onto the actual object of philosophical evaluation, not a derivative trace of some prior cognitive state. - Move 7 — The challenge from abduction therefore does not apply straightforwardly: it required that philosophical quality depend on a producer-level mental process an LLM cannot perform, but the properties philosophical evaluation attends to — the comparisons between hypotheses, the handling of objections, the weighing of intrinsic virtues — are properties of the text itself, and the text is what the LLM's training data contains. ## Lipton: likeliness and loveliness - Move 8 — Lipton distinguishes *likeliness* from *loveliness*: the likeliest explanation is the one most probable given the evidence, and the loveliest is the one that, if correct, would provide the most understanding — and the two come apart in principle, because what is probable given the evidence need not be what best explains it. > Likeliness speaks of truth; loveliness of potential understanding. > — Lipton, *Inference to the Best Explanation*, Ch. 4 - Move 9 — Applied to LLMs, this distinction produces the initial worry: an LLM's continuation is the likeliest in a distribution-theoretic sense — whatever is most probable in the distribution the model has learned — and in an arbitrary corpus (advertising copy, product reviews, news aggregation) likeliness of continuation would tell us nothing about the loveliness of what is said, because there is no particular reason the most probable next sentence should be the one that best explains anything. - Move 10 — The philosophical corpus is not such an arbitrary sample: it is the surviving record of philosophers doing inference to the best explanation and ranking candidate theories by the intrinsic virtues — simplicity, strength, accommodation of data — that constitute loveliness; work which fails to meet those standards is less likely to survive in the literature, less likely to be cited, less likely to be assigned to students or included in anthologies. - Move 11 — The corpus is also self-evaluating: every article in it was written by a philosopher reading others, and each was in turn read and engaged with by further philosophers writing back — so the patterns persisting in the corpus are patterns of IBE that have been iteratively tested by subsequent IBE, and the patterns surviving that iteration are the patterns later philosophers have found worth taking forward. - Move 12 — Lipton argues that loveliness can serve as a guide to likeliness — the lovelier explanation is generally the more probable — and the philosophical corpus makes the reverse relation hold as well: because survival in the corpus has itself been filtered by loveliness-tracking evaluation, what is statistically likely in this corpus is what loveliness-tracking evaluation has accepted, and so an LLM's likeliness-tracking over this corpus approximates loveliness-tracking over its content. - Move 13 — An LLM trained on this corpus therefore inherits its *distillate*: pattern-matching over it is not pattern-matching over a static record of good argument but over the surviving product of many rounds of philosophical comparison and ranking — the patterns the model has learned are patterns that have been selected for by the discipline's own evaluative practice. ## The mechanism - Move 14 — Argumentative structure is present in philosophical prose as surface regularity at the level of discourse — the level of clause-to-clause coherence, the handling of objection and reply, the signalling of hypothesis-comparison — not at the level of isolated word or sentence; and regularities at this level are precisely what next-token prediction is built to detect. - Move 15 — Philosophical prose is saturated with discourse markers that encode the comparative work being done in an argument — "however", "the stronger reading is", "one might press the objection that", "consider the cost of denying", "this leaves us with the question of" — and these markers recur with statistical regularity because the work they mark recurs in recognisable forms across philosophers reading and writing back to one another. - Move 16 — An LLM trained on well-formed English acquires syntactic norms by statistical exposure to grammatical text; an LLM trained on philosophical prose acquires argumentative norms — the deployment of these markers and the argument-shapes they mark — by the same mechanism operating over more complex structure, because next-token prediction is indifferent to whether the regularity it is detecting is grammatical or discursive. ## §2 payoff - Move 17 — The §2 upshot is that pattern-matching over a philosophical corpus can carry philosophical quality: the corpus is the surviving record of philosophers doing IBE, it has been iteratively shaped by philosophers doing and evaluating further IBE, and the argumentative patterns it contains are of the kind that next-token prediction can learn. - Move 18 — What §2 has not answered are two further threads: §3 takes up the capacity challenge from phenomenology, the worry that certain properties of philosophical text can be caused only by a human subject with first-person experience of the kind the text's arguments concern; and §4 returns to the fact that an LLM inherits the *distillate* of philosophical practice without itself performing the *distillation* that produced it. --- # §3 — Thought Experiments & Armchair Abduction ## Framing the challenge - Move 1 — Philosophy often draws on conscious experience — on what it is like to see red, to feel time passing, to touch one's own hand — and since an LLM has had no such experience, there is a natural worry that philosophy of this kind is not something it can do. ## Zahavy's worry - Move 2 — Zahavy (2026) argues that some breakthroughs require what he calls *manipulative abduction*, a form of inference that proceeds by simulating the sensory experience of a situation rather than by manipulating symbols — and his paradigm case is Einstein's formulation of the equivalence principle, arrived at via a thought experiment Zahavy describes in the following terms: > Einstein's variation required inventing new axioms based on a physical intuition that did not yet exist in the mathematics. He envisioned a physicist inside an elevator being uniformly accelerated through deep space. Inside this enclosure, the sensory experience reveals a specific pattern: when objects are released, the floor rushes up to meet them. To the physicist, the objects appear to fall with identical acceleration, regardless of composition. Thus, the simulation here was not a permutation of symbols, but a manipulation of perceptual experience. > — Zahavy 2026, §5 - Move 3 — Zahavy's claim about LLMs is that they, operating by manipulating language, have no access to the perceptual experience on which this kind of abduction depends; they are, in his words, "high-dimensional 'Chinese Rooms' (Harnad, 1990), manipulating the language of physics without access to the physical referents that give that language meaning". - Move 4 — It follows, on Zahavy's view, that an LLM trained only on physics text and no other modality could not have run Einstein's reasoning, because that reasoning required the felt experience of acceleration and free-fall — experience an LLM, never having had a body, does not possess. ## The challenge for philosophy - Move 5 — The same worry transfers to philosophy: just as Einstein needed the felt difference between weight-shift and free fall to construct his lift argument, Jackson needed the felt quality of seeing red to construct his Mary argument — Mary has complete physical knowledge of colour but has never seen red, and Jackson argues she learns something new when she first sees red, so framing and evaluating the argument requires a working grip on what seeing red is like, which an LLM without colour experience does not have. - Move 6 — And the challenge is not confined to colour experience: the phenomenology of having an intuition (Bengson 2015), the phenomenology of agency, and the feeling of time passing are all phenomenal modes philosophy has used as starting points for arguments, each of which could yield a Mary-type case that an LLM with no phenomenological experience would be blocked from producing. ## The Pigliucci reframe - Move 7 — Our response begins with an observation Pigliucci makes about how the world figures differently in science and philosophy: in science, the world is what claims are tested against — Einstein's equivalence principle needed Eddington's 1919 eclipse observations and the already-observed precession of Mercury's perihelion to confirm it — whereas in philosophy, the world does not serve this verification role but provides the starting points from which philosophers reason. - Move 8 — Einstein's case illustrates the distinction: the thought experiment gave Einstein the hypothesis — that acceleration and gravity are equivalent — but because physics is in the business of verification, the hypothesis had still to be empirically confirmed before it could be accepted as physics, whereas a philosophical thought experiment's conclusion does not face a further verification demand of that kind. - Move 9 — What philosophy takes from the world, on Pigliucci's account, can therefore be articulated — propositional, describable — rather than raw felt experience, and he puts the point directly: > the basic parameters that philosophers use as their inputs, the starting points of their philosophizing, their equivalent of axioms in mathematics and assumptions in logic (or rules in chess) are empirical data about the world. This data comes from both everyday experience (since the time of the pre‐Socratics) and of course increasingly from the world of science itself. > — Pigliucci, "Philosophy as the Evocation of Conceptual Landscapes," Ch. 6, p. 79 - Move 10 — The challenge accordingly reframes: it is no longer about whether an LLM has had phenomenological experience, but about whether it has access to articulated descriptions of that experience — and since articulated descriptions are exactly what text corpora are made of, the question becomes one the LLM has a real chance of meeting. ## The reply - Move 11 — Sub-disciplines of philosophy work constantly with articulated phenomenology — philosophy of perception on how things look, philosophy of temporal experience on how time is felt, aesthetics on the character of aesthetic response, philosophy of emotion on what emotions are like — and these produce careful descriptions of phenomenal experience that every LLM trained on academic text has processed, so the corpus is saturated with phenomenological description as a demonstrable fact about what the training data contains. - Sub-move — Literary writing adds to this: Proust on the involuntary memory triggered by the madeleine, Virginia Woolf on the texture of ordinary thought in *Mrs Dalloway*, Henry James on the shading of social perception, Nabokov on the visual particulars of rooms and streets — all part of the same training data. ## Running the reply through cases - Move 12 — An LLM has never felt weight-shift or free fall, but on Pigliucci's reframe what philosophical reasoning takes as its input is what the world has had articulated about it, and ordinary English is thick with articulated descriptions of these sensations: every elevator passage in fiction, every description of a roller-coaster drop, every astronaut's memoir of zero-gravity — so an LLM trained on ordinary text has the input that Einstein-style reasoning about weight-shift requires. - Move 13 — An anecdatum of mine: ask a current frontier LLM about design and colour — palettes, complementary colours, contrast and harmony — and it will discuss the phenomenology of colour with a sophistication indistinguishable from a knowledgeable human speaker, which is evidence that the corpus has given it enough of a working grip on colour phenomenology to reason about it. - Move 14 — Wherever philosophical work proceeds from phenomenological starting points that have already been articulated somewhere in text, those starting points are in the corpus an LLM is trained on, and the challenge — that LLMs lack phenomenological experience — is met by the fact that articulated phenomenology is what they have in abundance. ## The limit - Move 15 — There is nonetheless a narrower case the reply does not cover: a philosopher sometimes articulates a *previously undiscovered aspect of phenomenology* — a structural feature of conscious experience that, though present in everyone's lives, has not been explicitly described; this requires sustained first-person attention to one's own experience, and no recombination of existing text can substitute for that attention. - Move 16 — The paradigm is Merleau-Ponty's observation about self-touch: when one fingertip touches another, one finger plays the role of toucher and the other of touched, and the two can reverse roles — but they cannot simultaneously both be toucher, so one's own body is always, at any instant, split between the touching and the touched. - Move 17 — An LLM could not have originated Merleau-Ponty's observation, since there was nothing in prior text to recombine into that insight, and the same holds for the general case: an LLM cannot originate a phenomenological description that is not already, in some form, in its training data — whatever has never been articulated lies outside what the model can produce. ## §3 payoff - Move 18 — The §3 upshot is that the corpus supplies the phenomenological inputs philosophical reasoning takes from the world, wherever those inputs have been articulated in text — which, across most of the discipline (ethics, metaphysics, epistemology, philosophy of mind, philosophy of language), they have. - Move 19 — What §3 leaves to §4 is the further question raised by the persistence of underwhelming LLM output: if the capacity §§2–3 have defended is really present, why are we not seeing its products — and what, if anything, is the contemporary philosopher's work in closing that gap. --- # §4 — A Speculative Coda ## The puzzle - Move 1 — §§1–3 have argued that the in-principle obstacles to LLM-produced philosophy do not hold up: §1 rejected the constitutive challenge, and §§2 and §3 answered the capacity challenges from abductive reasoning and from conscious experience — but an obvious observation presses against these conclusions, namely that we are not in fact seeing LLM-produced philosophy worth reading, at least not at the rate one might expect if the capacity were genuinely there, and the remainder of this section floats some ideas about what might lie in the gap. ## What the prompt asks for - Move 2 — The gap is, on the simplest diagnosis, a matter of what the LLM is being asked for: prompted generically — "write on free will", "explain the Mary argument" — an LLM produces the kind of text most probable in the distribution it has learned, which for topics of philosophical interest is summary text, survey-style, hedged, balanced, non-committal, because the bulk of text written at that level of generality about philosophical topics takes that form. - Move 3 — A paper worth reading is not summary but distinctive argument for a specific conclusion against specific alternatives: a problem is chosen, a position is taken on it, and the position is pressed against its strongest opposition — and when the prompt asks for something that has this structure, the LLM has an articulated target to write toward; when the prompt asks for something more general, what it produces is correspondingly more general, which is not what the practice of philosophy treats as worth reading. ## Distillate without distillation - Move 4 — §2 argued that the philosophical corpus is the *distillate* of many rounds of philosophical criticism — arguments tested by later arguments, with the patterns that persist in the corpus being the ones later philosophers have taken seriously enough to engage with — and that an LLM trained on the corpus inherits this distillate as pattern. - Move 5 — Inheriting the distillate, however, is not the same as performing the *distillation*: the iterated criticism that produced the corpus's patterns operates across time and across many minds, and a single completion by an LLM does not reproduce it — the model has the products of iteration, not the iteration itself. - Move 6 — The philosopher prompting the LLM can perform the iteration in place of the discipline: drafting, pressing the draft against the strongest objection available, redrafting in light of that pressure — and so supplies within the session some portion of the process that the model has absorbed only in its output. ## Generalist, not specialist - Move 7 — A candidate response to all this would be a specialist LLM trained narrowly on philosophical text, but Sellars's characterisation of what philosophy aims at runs in the opposite direction: > The aim of philosophy, abstractly formulated, is to understand how things in the broadest possible sense of the term hang together in the broadest possible sense of the term. Under 'things in the broadest possible sense' I include such radically different items as not only 'cloth, ships, and sealing-wax,' but numbers, duties, possibilities, finger snaps, aesthetic experience, and death. > — Sellars, "Philosophy and the Scientific Image of Man" (1962) - Move 8 — If something like Sellars's characterisation is right, training breadth is not a distraction from philosophy but close to its proper substrate, and the philosophy worth reading from LLMs is more likely to come from generalist models differently prompted than from specialist philosopher-models differently trained — a claim that runs ahead of the evidence currently available to assess it. ## §4 payoff - Move 9 — Putting §§1–4 together: there is no in-principle obstacle to LLM-produced philosophy worth reading, and what is practically required for such philosophy to be produced is what the discipline has always required of its authors — a specific problem taken up, a position committed to, and iteration across objections until the position holds — work which, increasingly, gets done in the prompt rather than in the drafting of a paper.