# Generating Philosophy with AI *Claude's draft v3 — unread by Nick. Sentence-by-sentence rewrite of v2 with paragraph-level depth pass. Triplets, chiasmus, scaffolding clusters, abstract noun-stacks, and meta-narration cut throughout. Two depth-fixes from v2 (Mary-variant and self-touch) now actually run rather than gestured at.* --- Can current large language models produce philosophy worth reading? The question is not about peripheral uses of language models around philosophy — drafting summaries or organising reading lists — but about the philosophical standing of the generated product itself: whether an LLM output, considered as a text, can reward philosophical attention in the way good philosophy does. I want to argue that, in the right conditions, it can. The claim concerns systems we already have, not future artificial general intelligence: current frontier models, trained on large bodies of text and capable of sophisticated generation. The risky thing about it is that the relevant capacity has to be present already, in systems we have used and looked at. By 'worth reading' I do not intend a term of art. Every philosopher knows the standard: a paper either gives us reason to spend time with it as philosophy, or gives us only reason to note its topic and move on. Journals aim at the standard, even when they fail to meet it. A philosophical text worth reading is rarely a bare answer. What matters is that the work of thought be available in the text — that we can follow the route by which a conclusion is reached, or feel the pressure that makes a familiar problem look different. Whether a text repays attention turns on what it makes available for philosophical assessment, not on facts about the producer. The standard is product-centred. Context is not thereby irrelevant: a paper may be worth reading only because of the debate it enters, or because it diagnoses pressure on a position that would otherwise look secure. But the debate it enters is publicly assessable, as is the pressure it diagnoses. Neither requires looking past the text to the producer. Four challenges press against the affirmative answer. First, an authorship challenge: even a text indiscernible from a philosophical paper would fail to be philosophy if no philosopher's activity stood behind it. Second, an abduction challenge: LLMs cannot perform the abductive reasoning on which much philosophical theorising depends, so any apparent abductive structure in their outputs is mere imitation. Third, a phenomenology challenge: LLMs lack conscious experience and so cannot produce philosophy that begins from phenomenology. Fourth, an elicitation challenge: even granting an LLM the capacity to produce worthwhile philosophy, when it does so under prompting from a human user, the human is doing the philosophising, not the model. The first is constitutive — an excellent-looking LLM text would already have failed to be philosophy. The middle two are capacity challenges, doubting that an LLM has what it takes to produce a text of the right kind. The last redescribes whose capacity is being exercised. None of the four, I shall argue, settles the question it raises. ## I. The challenge from authorship Some philosophers hold that no LLM output can really be philosophy, whatever the text in front of them looks like. The complaint is not the familiar one that LLM outputs are often bad. It is that no philosopher has produced them, and so they only resemble philosophical texts from the outside. On this view, the defect lies not in any feature of the paper's argument but in the absence of a philosopher's activity behind it. An excellent-looking LLM output would already have failed before we asked whether its argument worked. The worry can be made more precise by adapting Davies' performance theory of art. Davies writes: > The work — what the artist achieves — is the process eventuating in that product. Works themselves are neither structures nor objects simpliciter, nor are they contextualized structures or objects [...]. They are, rather, intentionally guided generative performances that eventuate in contextualized structures or objects. (Davies, *Art as Performance*, p. 98) On Davies' view, the artwork is not what the artist's activity leaves behind — the painted canvas or the printed score — but the activity itself. A canvas indistinguishable from a Rembrandt but produced by accident lacks the history that would make it a Rembrandt: the work is the doing, the product is the residue. Transpose the view to philosophy, and the philosophical work becomes the philosopher's sustained activity in producing the text: working through the problem, formulating and revising the arguments. The text is what that activity leaves behind. If philosophy were like art in this respect, the authorship challenge would be powerful — a text produced by an LLM might fail to be a work of philosophy in just the way an accidental Rembrandt-looking canvas fails to be a Rembrandt, with the visible product not settling the matter. The transposition does not survive contact with how analytic philosophy is practised. We ask whether premises are defensible and whether conclusions survive the objections the paper anticipates. These are decidable from the text, without reference to any further activity behind it. Achievement-talk seems to push the other way: we say an author has *achieved* something in a paper, suggesting something beyond the text to be assessed. But the achievement-talk is parasitic. To say the author has achieved something is just to say she has produced a text with such-and-such argumentative properties; nothing further is needed for the achievement to be in place. None of this is to say authorship never matters. Credit and responsibility turn on it. But credit-and-responsibility and worth-readingness are different questions, and a paper can answer one well without answering the other. Worth-readingness is a property of the text. Credit, by contrast, depends on who produced it. Consider a duplicate philosophical text. Suppose the same sequence of sentences appears in a human-written article and in an LLM output. The inferential relations available to the reader are the same in both. Whatever the conclusion follows from in one, it follows from in the other.[^1] Causal history affects who should receive credit; it does not, by itself, affect which premises support which conclusions. In this respect philosophy is closer to proof than to painting. A machine-generated proof is not invalid because the machine did not understand it. The analogy should not be pushed too hard, since philosophy is richer than proof. The relevant point survives the difference: what carries assessment is on the page. One might object that an LLM output is not really an argument because no one asserts its premises. There is no agent standing behind the words committed to their truth. But philosophical assessment does not always require sincere assertion by the producer. We assess the arguments that appear in dialogues and reductios without treating every sentence as the author's straightforward commitment. The norms governing argument-assessment and the norms governing assertion are not the same. Anonymous review reflects this: referees assess a paper by what it says, not by who wrote it, and the assessment proceeds in the absence of any authorial information. The Sokal affair is recognisable as a violation of the norm rather than evidence against it. When *Social Text* published Sokal's hoax paper without peer review, what went wrong, on the academy's own assessment, was that the journal had assessed Sokal's institutional standing rather than the argument on the page. What the authorship challenge is left with, after all this, is the more limited claim that LLM publication raises difficult questions about credit and responsibility. So it does. But authorship of that kind does not constitute philosophy. The constitutive challenge — that an LLM output cannot be philosophy at all because no philosopher's activity stands behind it — fails because the work in philosophy is in the text, and the assessable properties of the text are in the text too. The harder worries about LLM-produced philosophy are about whether the system has what it takes to put the right text on the page. They are capacity worries: whether only a system that performs abductive reasoning can produce abductively good philosophy, and whether only a conscious subject can produce philosophy that starts from experience. The next two sections take these in turn. ## II. The challenge from abduction Floridi et al. do not deny that LLMs produce explanation-like answers. They deny that such answers are produced by abductive inference. Their formulation: > LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al., p. 9) The critique becomes serious for the present thesis if Williamson is right that contemporary philosophical theorising already proceeds partly by abduction from the armchair. On Williamson's account, philosophers defend a view by comparing it with rivals and asking which, if true, would best explain the evidence — weighing what he calls the *intrinsic virtues* of a good theory: roughly, simplicity combined with strength. The qualifier 'intrinsic' is doing real work. The virtues are features of the theory itself: it should be elegant and unified rather than ad hoc and messily complicated. They are not features of the theorist's mental processing. If philosophical theorising consists in this kind of comparing-and-weighing, and Floridi is right that LLMs only mimic the surface of such reasoning, then the philosophical character of LLM outputs would be no more than a fluent surface. It would be wrong to respond by arguing that LLMs secretly perform human-style inference to the best explanation. They do not. An LLM does not understand the problem as a problem, and it does not knowingly weigh rivals under a norm of truth. Williamson's qualifier shows where the response should go instead. If the virtues are intrinsic to the theory rather than to the theorist, then whether a theory exhibits them is assessable from the text. The question is not whether an LLM can run the inference but whether a generated text can display abductive philosophical structure — whether it can compare two explanations and show why one handles the pressure better, or trace what would have to be given up if a particular view were rejected. That is a question about the product, not the process. Philosophy externalises much of its abductive work in prose. When a paper says one view explains what a rival cannot, or introduces a distinction to answer an objection, the comparison is on the page rather than hidden behind it. It is part of what the reader assesses. Lipton's distinction between actual and potential explanation bears on this. Inquiry does not begin with explanations already known to be true. It begins with candidates whose explanatory force can be assessed before their truth is settled. An LLM output can offer such a candidate: a way the relevant material would hang together if true. The further distinction Lipton draws between likeliness and loveliness sharpens this. Likeliness concerns warrant; loveliness concerns the understanding an explanation would provide if true. As Lipton writes: "Likeliness speaks of truth; loveliness of potential understanding." A generated answer can be assessed for loveliness before anyone has settled whether it is likely. Loveliness is not a feature of the reasoner's mental life. It shows itself in the way a text makes a problem more intelligible than it was before, by revealing why one explanatory route has more force than another. This kind of intelligibility can be present or absent in an LLM output. When present, the text has the property philosophical readers track when they say a paper repays their attention. Floridi himself points to the channel by which models come to produce such texts. The Floridi paper concedes: "This effect is due to the model's training on human-generated texts that encode reasoning structures." The concession is not incidental. Philosophy is conducted in writing, and the writing is what the LLM has been trained on. The corpus an LLM is trained on is not a heap of sentences about philosophical topics. It is the surviving record of philosophers doing inference to the best explanation and ranking candidates by intrinsic theoretical virtue, as engaged with by further philosophers writing back. Survival in this corpus has itself been filtered by loveliness-tracking evaluation: a paper that fails to satisfy the relevant standards is less likely to be cited or assigned. What is statistically prominent in the corpus is what loveliness-tracking evaluation has accepted, and so a model trained to predict the next token over this corpus is, indirectly, trained to track loveliness over its content. The worry that LLM outputs exhibit only a 'mere surface pattern' moves too quickly. Philosophical prose is saturated with discourse markers that encode comparative work — 'however', 'the stronger reading is', 'one might press the objection that' — and these markers recur with statistical regularity because the work they mark recurs in recognisable forms across philosophers reading and writing back to each other. A model trained on such material is trained on prose shaped by philosophical pressure, and the regularities it picks up are dialectical regularities at the level of clause-to-clause coherence. Floridi's 'zeroth-order abduction' is meant to deflate the output: the model has the linguistic form of explanation without the reasoning. But in philosophy the linguistic form is not detachable from the argument in the way the deflation requires. Explanatory comparisons are how philosophical reasoning appears on the page. They are not surface marks floating free of the inferences they express.[^2] A model that has absorbed such patterns can generate new text in which they are redeployed. It may locate a pressure point or compare two explanations in a way the prompt did not specify. When it does this well, the result is not raw material for a human philosopher to work up but itself a philosophical text worth reading. None of this implies that fluent LLM prose is automatically valuable. LLMs can produce empty philosophy-looking text, just as humans can produce bad philosophy, and the failure of any particular output is shown by reading it, not inferred from the absence of human-style abduction in the process that produced it. LLMs may not perform inference to the best explanation in the way we do. But the abductive structure that bears on philosophical assessment is carried by prose, and a model trained on philosophical prose has been trained on the very thing loveliness-tracking evaluation works on. The producer's inference and the product's structure can be assessed separately, and for philosophy — given that the virtues are intrinsic to the theory rather than to the theorist — only the second is required for the assessment to go through. ## III. The challenge from phenomenology A further capacity worry concerns phenomenology. Even if neither authorship nor abduction rules out LLM-produced philosophy, perhaps conscious experience does. Some philosophy seems to begin from what it is like to see red, or to feel a particular intuition take hold. If LLMs lack conscious experience, perhaps they can only repeat what experiencers have said, and the philosophy that depends on experiential starting points lies beyond them. Zahavy gives this worry a vivid form by way of Einstein's elevator. The passage is worth quoting at length: > Einstein's variation required inventing new axioms based on a physical intuition that did not yet exist in the mathematics. He envisioned a physicist inside an elevator being uniformly accelerated through deep space. Inside this enclosure, the sensory experience reveals a specific pattern: when objects are released, the floor rushes up to meet them. To the physicist, the objects appear to fall with identical acceleration, regardless of composition. Thus, the simulation here was not a permutation of symbols, but a manipulation of perceptual experience. (Zahavy 2026, §5) Zahavy describes this kind of reasoning as *manipulative abduction*.[^3] The thinker varies an imagined experiential situation and attends to what would be experienced within it. If reasoning of this kind depends on simulated experience, LLMs seem blocked from it. As Zahavy puts the point, LLMs operating on text alone are "high-dimensional 'Chinese Rooms' (Harnad, 1990), manipulating the language of physics without access to the physical referents that give that language meaning." They can manipulate descriptions of elevators and gravity, but they have never felt weight or free fall. The same worry transfers to philosophy. Mary's case in Jackson's knowledge argument turns on the apparent gap between knowing all the physical facts about colour and seeing red for the first time. Merleau-Ponty's discussion of self-touch turns on attention to the structure of embodied experience. In each instance, an LLM without the relevant experience might seem unable to engage with the material except by recycling what other writers have already said. The worry is not confined to colour and embodiment: the phenomenology of having an intuition and the feeling of time passing are also phenomenal modes that philosophy has used as starting points, each of which could yield a Mary-style case. Having an experience is not the same as producing philosophy about it. Most people see colours and feel pain, but this does not make them good philosophers of colour or pain. The bare possession of experience is not what matters philosophically. What matters is how experience is articulated and used. Raw phenomenology does not enter an argument directly. It enters once it has been described and turned into a case the argument can address. Articulated phenomenology — phenomenology made public in language — is what does the work. The Einstein analogy works differently in the scientific case from how it would in the philosophical one. In the scientific case, the imagined experience helps generate a hypothesis that has still to be tested against the world. The lift thought experiment gave Einstein the equivalence principle, but the principle had still to be confirmed by Eddington's 1919 eclipse observations and by the precession of Mercury's perihelion. In philosophy, things are different. Once a thought experiment is articulated, its philosophical force lies in what follows from it, not in any further empirical test. What philosophy takes from the world is the input to an argument, not the verifier of a hypothesis, and what gets taken is articulable rather than raw. Mary's case shows what description-level work amounts to. No competent discussant of the knowledge argument has personally undergone her transition. The case works because the scenario is public: Mary knows all the physical facts, has never seen red, and on first seeing red appears to learn something. Once that scenario is on the page, the philosophical work proceeds at the level of articulation. Take Lewis's response. Faced with the conclusion that Mary learns a new fact, Lewis does not need to undergo her experience to reply. He distinguishes propositional knowledge — which Mary already had — from abilities such as recognising red on later occasions or imagining it, and argues that what Mary acquires is an ability rather than a fact. The reply is made entirely on Jackson's articulated description of the case. Whatever one's verdict on Lewis, the move is genuine philosophical work, and the pressure it puts on Jackson is pressure at the level of the public scenario. Human philosophers already rely on articulated phenomenology in this way. They write about blindness and about animal experience without necessarily having had those experiences, drawing on testimony and on the previous philosophical literature. First-person possession is one source of phenomenological material; it is not a universal condition of philosophical work about experience. LLMs lack raw phenomenology, but no text corpus contains raw phenomenology either. What a corpus contains is articulated phenomenology — experience already made available in language — and this is the form in which phenomenology becomes usable in philosophical argument. The training data of any current frontier model is saturated with such material, in academic philosophy of perception and temporal experience and in literary writing of the kind Proust and Woolf produced. A model that has absorbed such material does not merely repeat the descriptions it has seen. To see what a further move looks like, take a variant of Mary: a colour-blind philosopher who reads through Mary's physical knowledge of colour and is then given a colour-corrective implant. Press the Lewis distinction on the variant. On Lewis's reading, what Mary gains is a set of abilities rather than a new fact, and the same should be true of the colour-blind philosopher: she gains the abilities to recognise and imagine red, which is what newly seeing red consists in. So the Jackson argument runs the same way against the variant: someone who knew all the physical facts about colour, including the facts about her own implant and what its activation would bring about, learns something new on first seeing red. The variant therefore makes the Jackson/Lewis disagreement turn on the same point in a slightly different setting, and forcing the question through it can sharpen which side carries the relevant intuition. None of this requires first-person access to a colour experience the discussant lacks. A model trained on the phenomenology-of-colour literature has the materials to construct the variant, and there is no in-principle barrier to its constructing one without prompting. Merleau-Ponty's discussion of self-touch gives the worry its hardest case. When one fingertip touches another, one finger plays the role of toucher and the other of touched. The roles can reverse, but not simultaneously: at any given instant, the body is split between touching and touched. The observation looks like a phenomenological discovery drawn from sustained attention to embodied experience, and an LLM with no such experience cannot have arrived at it from scratch. There is a class of phenomenological observations like this — observations of structure in experience that has not previously been described — and a corpus-trained model is shut out from originating them. But originating an observation and working with one are different things. Take the case in which the touching limb is anaesthetised. The toucher-touched asymmetry, on the original Merleau-Pontian observation, depends on each finger being able to play either role. With one finger anaesthetised, the asymmetry seems to collapse: the anaesthetised finger can be touched but cannot itself touch — the role-reversal becomes unavailable. So either the structure Merleau-Ponty identifies generalises only to non-anaesthetised parts of the body, or the relevant 'touching' is decoupled from sensory feedback in a way the original description does not flag. Either branch puts pressure on what the discovery commits one to. A model trained on Merleau-Ponty and his interlocutors can press the case without itself having touched anything. The pressure it generates is pressure on the articulated structure, and the assessment of the pressure is assessment of how the structure responds to a case the original description did not consider. Two questions need separating. Can an LLM originate a phenomenological description not already present, in some form, in its training data? No. Phenomenology that has not yet been put into words lies outside what a corpus-trained model can reach. Can an LLM produce phenomenology-based philosophy worth reading? Yes, wherever the phenomenological material has been articulated in print — which, across most of the discipline, it has. A model can generate a new description of experience by drawing on the bodily vocabulary and prior descriptions in its training data, and whether the proposal survives reflection is a further question — as for any human phenomenological proposal. ## IV. Authorship redux: the challenge from elicitation A practical embarrassment remains. If current LLMs can produce philosophy worth reading, why are we not surrounded by great LLM philosophical texts? The ordinary experience of using these systems seems to support scepticism. Asked for philosophy, they often produce competent but lifeless exposition — paragraphs that tour a topic without ever applying pressure to it. The diagnosis is not far to seek. A vague topic prompt — 'discuss free will', 'explain the Mary argument' — does not ask for a philosophical intervention. It asks for the most probable kind of text under that topic-label, and in the training distribution that is often hedged survey, because the bulk of philosophical text written at that level of generality takes that form. The prompt is generic, and so is the region of the model's learned space it activates. The diagnosis creates a further worry. If interesting LLM philosophy appears only under careful prompting, perhaps the LLM is not really producing the philosophy after all. Perhaps the human prompter is producing philosophy by using the LLM. The worry is sharpest when the human controls the task and then chooses among the results. The final text may be worth reading, but the value can seem to belong to the human-guided process rather than to the model's production. So 'produced by an LLM' has to mean more than 'appearing in an LLM output window'. A model may output philosophy worth reading without producing the features that make it worth reading. Take the grammar-correction case. A philosopher writes a brilliant argument and asks an LLM only to correct its punctuation. The LLM's response may contain philosophy worth reading, but the model has not produced the argument in virtue of which the text is worth reading. It has output worthwhile philosophy without producing its worth-readingness. So there is a continuum of contribution. At one end, the model merely polishes a human-produced argument; at the other, the human specifies a task and the model generates the philosophical move that makes the output worth reading. The cases that support the present thesis lie towards the latter end. Where the human supplies the argument and the model improves the prose, the LLM has not produced philosophy worth reading. Where the human specifies a problem and the model supplies the objection or distinction that makes the text worth reading, the output is LLM-produced in the sense that matters here. The elicitation challenge also assumes too simple a contrast between autonomous producer and mere tool. Elsewhere I have argued, with Terrone, that generative AI systems of the Midjourney type are best understood as a third category — neither agents nor tools but a new kind of artistic medium, with which the user must grapple under what we call dynamic recalcitrance (Young & Terrone 2025). The same third possibility is available for LLMs. They are not intentional agents and not ordinary tools. Their outputs are partially controllable and dynamically generated. Prompting is best understood as elicitation from such a system: the prompt is not a blueprint that fixes the product in advance but a setting of conditions under which the model generates. The user can constrain and iterate, but cannot determine every relevant feature of what emerges. Recast as elicitation, prompting reveals itself as a different relation. Elicitation is not authorship. A call for papers elicits philosophical answers without authoring them; an interlocutor in conversation elicits arguments from another philosopher without coming to be their author. So a prompt's having elicited an LLM output does not by itself show that the prompter has supplied the philosophical content that makes the output worth reading. Selection is not generation either. A journal selects the papers it publishes but does not thereby produce them, and a reader's selection of a good LLM output is, similarly, an act of assessment rather than of production.[^4] If philosophical corpora encode abductive and dialectical structure in the way I argued earlier, skilled prompting should aim to elicit those structures rather than to request prose about a topic. Skilled philosophical prompting is model-sensitive task specification. It does not ask the model to sound philosophical; it gives the model a dialectical role to play. A prompt can ask the model to defend a thesis against a specific objection, or to identify the explanatory cost of rejecting a claim. Discourse markers are not magic words. Phrases such as 'one might object' matter because they are surface markers of argumentative roles, and a good prompt does not merely insert them — it specifies the role to be filled.[^5] Two forms of such prompting deserve attention. Contrastive prompting asks why one view handles a particular case better than its rival, rather than asking for a discussion of a topic in the abstract — mirroring the contrastive shape philosophy takes when we ask why P rather than Q. Loveliness-sensitive prompting asks not for a conclusion but for a view that, if true, would explain more than its rival. Both target the explanatory virtues that make a philosophical answer worth reading. LLM philosophy improves when the prompt creates dialectical pressure: generic prompts invite generic continuations, and a philosophical prompt should create a space in which some argumentative move is needed, and then leave that move for the model to make. Creativity belongs here, but in a limited role. If the output is elicited, one might wonder whether it can really be creative; the answer depends on where the case sits on the continuum. Grammar correction is not interestingly creative. Generating a new objection or a new distinction may well be. This fits a product-centred approach to creativity. Many accounts require novelty and value, and the present argument need not show that LLMs are creative agents in the fullest sense — only that their outputs can contain novel and valuable philosophical structure. Elicitation does not defeat creativity. Creative work often happens under constraints (a commission, a question put by an interlocutor), and a prompt can set a conceptual space without determining what is found within it. Elicited LLM philosophy can be creative at the level relevant to worth-readingness. A related diagnosis sits behind the practical embarrassment. The corpus an LLM is trained on is the *distillate* of many rounds of philosophical criticism — arguments tested by later arguments, with the patterns surviving the iteration being the ones later philosophers took seriously enough to engage with. Inheriting the distillate is not performing the distillation. The criticism that produced the corpus's patterns operates across time and across many minds, and a single completion by an LLM does not reproduce that. The model has the products of iteration without the iteration itself. The philosopher prompting the LLM, however, can perform the iteration in place of the discipline, alternating composition with criticism and recomposition. So supplied, the model contributes its absorbed patterns to a process the prompter is running, and the output that emerges from the loop is more than what either party would produce alone. It is not, however, made worth reading by the human supplying argumentative content. The human supplies iteration-pressure; the model supplies, when conditions are right, the philosophical move. The absence of many great generic LLM texts is therefore not the verdict it can appear to be. It reflects, at least in part, immature elicitation practices and the prevalence of generic prompting. If philosophy worth reading requires live alternatives and dialectical pressure, vague prompts will rarely elicit it. Elicitation does not defeat the claim that LLMs can produce philosophy worth reading: it shows that such production comes in degrees and depends on task conditions. The question to keep in view is whether the model has generated the philosophical structure in virtue of which the output is worth reading. The answer can be yes. LLMs do not produce worthwhile philosophy merely by being asked for 'some philosophy', and not every LLM-assisted text counts as LLM-produced in the relevant sense. But current models can generate philosophical moves that make a text worth reading, and where they do, the philosophical value is present in the product, and the product was produced by the LLM in the sense that matters here. --- The four challenges treat what an LLM lacks as decisive for what its outputs can be. None of the absences they begin from is trivial. None of them, however, settles whether what appears on the page is philosophy worth reading. Worth-readingness, on the account defended here, is assessed through what a text makes available in its dialectical context. Where an LLM's output carries the relevant comparison and presses the relevant objection, the output is philosophy worth reading, and was produced as philosophy by the system that generated it. What is owed at that point is the assessment philosophy always owes its texts: do the arguments hold? [^1]: The duplicate case is artificial in practice: no human philosopher writes exactly what an LLM happens to produce, and vice versa. The artificiality is doing controlled work. By holding the textual product fixed, the case isolates the question of whether causal history alters inferential structure, and the answer is that it does not. One can run the same case in weaker form: take an LLM output that, after a round of human editing, becomes textually indistinguishable from what a human author might have written unaided. The human-edited version is uncontroversially a philosophical text of the kind under discussion. Whatever distinguishes the edited and unedited versions philosophically must therefore be a textual difference, not a causal-history one. [^2]: A charitable assumption underlies this paragraph: that the philosophical material in the model's training mix is sufficient, and sufficiently weighted, to carry the dialectical regularities the argument relies on. A model trained heavily on shallow material — encyclopaedia summaries, undergraduate term papers, online opinion — will not have absorbed the pressure that good philosophical prose carries. Whether contemporary frontier models meet the relevant condition is an empirical question, and one this paper does not settle. Where the condition is met, the resulting outputs can carry abductive structure; where it is not, they cannot. [^3]: The term originates with Magnani, who introduced it for cases of hypothesis generation through the active construction of mental models. Zahavy uses it in this sense and applies it to Einstein's elevator argument as a paradigm case in physics. [^4]: The journal/reader parallel is not perfect. Selection-and-iteration loops, where the human prompter selects an output, modifies the prompt in response, and selects again, can blur the line: the more the human iterates, the more the resulting text is co-produced. The narrower point survives the disanalogy — selecting a good token is not by itself producing it — but the cleanly LLM-produced cases lie at the lower-iteration end of the continuum. [^5]: A concrete contrast may help. Compare two prompts on the same problem. The generic prompt: "Discuss whether physicalism is true." The contrastive prompt: "Identify the strongest objection that applies to type-A physicalism but not to type-B physicalism, and show what type-A would have to say in reply, given that the standard reply available to type-B is not available." The first asks for a survey and gets one. The second specifies a dialectical role — locate a discriminating objection, run a position through it, and supply a reply that does not collapse the distinction the prompt cares about. A loveliness-sensitive prompt at the same level of specificity might ask for a view that, if true, would explain why type-A physicalism's commitments cluster the way they do, where type-B's clustering looks unmotivated. In each case, what the model is asked for is not a particular conclusion but a particular shape of philosophical move.