# Generating Philosophy with AI *Claude's draft — unread by Nick. Converted from the v5 moveset (Untitled 11.md) into prose. Citations and bibliography deferred. All content from the moveset preserved.* --- Can current large language models produce philosophy worth reading? The question I have in mind is not one about peripheral uses of language models around philosophy — drafting summaries, polishing footnotes, organising reading lists — but one about the philosophical standing of the generated product itself: whether an LLM output, considered as a text, can reward philosophical attention in the way that good philosophy does. The risky part of the claim I want to defend is that it can. I shall not be making a promissory note about future artificial general intelligence. The systems I have in mind are systems we already have: current frontier models, trained on large bodies of text and capable of sophisticated generation. The interesting question is whether such systems already have the relevant capacity, and I want to argue that, in the right conditions, they do. By 'worth reading' I do not intend to introduce a term of art. The phrase names a familiar practical standard in philosophical life: the sense one has, on opening a paper, that the text gives one a reason to spend time with it as philosophy, rather than merely a reason to note its topic and move on. Journals aim at that standard, even when they do not always meet it. A philosophical text worth reading is not normally a bare answer either. What matters is that the work of thought is available in the text — that the reader can see the route by which a conclusion has been reached, or can feel the pressure that makes a familiar problem look different. Worth-readingness in this sense is a feature of what the text makes available for philosophical assessment and uptake. The question is therefore product-centred. This does not mean that context is irrelevant to philosophical interest. A text may be worth reading only because of the debate it enters, or because it diagnoses pressure on a position that would otherwise look secure. But that, too, is a publicly assessable feature of the philosophical product. It is not a private feature of the producer's mental life. The paper proceeds by considering four challenges to the affirmative answer. The first holds that no LLM output can be philosophy because no philosopher stands behind it: an authorship challenge that locates the defect in absent producer-activity rather than in the inferential quality of the paper. The second holds that LLMs cannot perform the abductive reasoning on which much philosophical theorising depends, and that any apparent abductive structure in their outputs is therefore mere imitation. The third holds that LLMs lack conscious experience and so cannot produce philosophy that begins from phenomenology. The final section returns to authorship from another direction: if a human prompt elicits the output, is the philosophy really produced by the LLM, or is the human prompter producing philosophy by using the LLM? My answer to all four is that they fail to settle the question they pose. The authorship challenge fails as a constitutive barrier; the abduction challenge confuses a fact about the generating process with a verdict on the generated product; the phenomenology challenge moves too quickly from the absence of experience in the producer to the absence of phenomenological content in the product; and the elicitation challenge over-simplifies what it is for a system of this sort to produce a text. ## I. The challenge from authorship Some philosophers react to the prospect of LLM-produced philosophy by holding that, whatever the text in front of them looks like, it cannot really be philosophy. The worry I have in mind is not the familiar complaint that LLM outputs are often bad. It is the worry that no philosopher has produced them, and so they only resemble philosophical texts from the outside. The defect, on this view, would not lie in any feature of the paper's argument. It would lie in the absence of a philosopher's activity behind the paper. Even an excellent-looking LLM output would fail before we asked whether its argument worked. The worry gains traction from a comparison with art, particularly with performance theories of the artwork. On Davies' performance theory, an artwork is not merely the physical object left behind, but the intentionally guided performance that eventuates in that object. A canvas that looks exactly like a Rembrandt, but is produced by accident, lacks the history that would make it a Rembrandt. If philosophy were like art in this respect, the authorship challenge would be powerful. A philosophy-looking text produced by an LLM might fail to be a work of philosophy in just the way an accidental Rembrandt-looking canvas fails to be a Rembrandt. The visible product would not settle the matter. The comparison, however, does not transfer. Philosophical assessment is directed at what the text says and does. We ask whether the conclusion is supported, whether the reply meets the objection, whether the distinction does the work it is asked to do. These are features of the public philosophical object, not of a private activity hidden behind it. That is, the assessment we make of a paper of philosophy is in the first instance an assessment of the paper, and the question of whether such assessment depends essentially on a producer's mental performance is precisely what is at issue. This is not to say that authorship never matters. Authorship matters for credit and responsibility, and these are not negligible matters. But credit and responsibility are not the same thing as philosophical worth-readingness. A text can be problematic to publish under someone's name while still containing an argument worth reading; a text can be unproblematically published while containing nothing worth reading at all. Who is responsible for the text and what the text makes available philosophically come apart. Consider a duplicate philosophical text. Suppose the same sequence of sentences appears in a human-written article and in an LLM output. The inferential relations available to the reader are the same in both cases: the conclusion is supported by the same premises, the objection is met by the same reply. The causal history changes who should receive credit. It does not by itself change which conclusion follows from which premises.[^1] In this respect, philosophy is closer to proof than to painting. A machine-generated proof is not invalid because the machine did not understand it. The analogy should not be pushed too hard, since philosophy is richer than proof, but the relevant point survives the difference: public inferential structure is not erased by the absence of a human mental performance. One might object that an LLM output is not really an argument because no one asserts its premises. There is no agent standing behind the words who is committed to their truth. But philosophical assessment does not always require sincere assertion by the producer. We assess the arguments that appear in dialogues and reductios without treating every sentence as the author's straightforward commitment. The norms governing argument-assessment and the norms governing assertion are not the same. A familiar way to see this is to recall the practice of anonymous review. Referees are asked to assess a paper by what it says rather than by who wrote it. This does not prove that authorship is never relevant to philosophical practice, but it does show that philosophical merit is treated, in a substantial portion of philosophical life, as assessable without reconstructing the author's private activity. If this is correct, then the authorship challenge fails as a constitutive barrier. LLMs may not philosophise as persons do, and LLM publication may raise difficult questions about responsibility — questions worth taking seriously in their own right. But authorship alone does not show that an LLM-produced text cannot contain philosophy worth reading. Once that point is in place, the more serious worries concern production capacities rather than authorship as such. Perhaps only a system that performs abductive reasoning can produce abductively good philosophy. Or perhaps only a conscious subject can produce philosophy that starts from experience. These are different objections, and they need different answers. ## II. The challenge from abduction Floridi et al. do not deny that LLMs produce explanation-like answers. What they deny is that such answers are produced by abductive inference. The model does not select an explanation because it best accounts for the evidence; it generates a continuation made probable by training. This becomes a serious challenge to the present thesis if Williamson is right that much philosophical theorising proceeds abductively. Philosophers compare candidate theories and ask what each would explain if true. They prefer a view that handles a wider range of cases, or that integrates better with neighbouring commitments, over a view that does not. If LLMs cannot do that, perhaps the philosophical character of their outputs is no more than a fluent surface. The right response is not to argue that LLMs secretly perform human-style inference to the best explanation. That is unlikely to be the correct hill to die on. An LLM does not understand the problem as a problem, and it does not knowingly weigh rivals under a norm of truth. The issue I want to press is different: it is whether the generated text can display abductive philosophical structure. A text can compare two explanations and show why one handles the pressure better. It can introduce a distinction that turns aside an objection, or trace what would have to be given up if a particular view were rejected. These are properties of the product. They can be present whether or not the process that generated the product was itself human abduction. This matters because philosophy externalises much of its abductive work in prose. When a paper says that one view explains what a rival cannot, or introduces a distinction to answer an objection, the comparison is not hidden behind the text — it is part of what the reader assesses. Lipton's distinction between actual and potential explanation is useful here. Inquiry does not begin with explanations already known to be true. It begins with candidates whose explanatory force can be assessed before their truth is settled. An LLM output can offer such a candidate: a way the relevant material would hang together if true. Lipton's further distinction between likeliness and loveliness sharpens the point. Likeliness concerns warrant — how much reason we have to think the explanation is correct. Loveliness concerns the understanding an explanation would provide if true. A generated answer can be assessed for loveliness before anyone has settled whether it is likely. And loveliness is not a private glow in the reasoner's head. It shows itself in the way a text makes a problem more intelligible than it was before, by revealing why one explanatory route has more force than another. This kind of intelligibility can be present or absent in an LLM output. When it is present, the text has a property that any of us would recognise as a marker of philosophical interest. Floridi's own diagnosis gives us the channel by which models come to produce such texts. If LLMs produce plausible explanations because they have been trained on texts that encode reasoning structures, then philosophical training data matters. Philosophy is, among other things, a textual practice in which explanatory comparison is made public. The philosophical corpus is not a heap of sentences about philosophical topics. It is a record of arguments challenged and replies refined. A model trained on such material is trained on prose that has been shaped by philosophical pressure.[^2] This is why the worry that LLM outputs exhibit only a 'mere surface pattern' is too quick. Next-token prediction is the training task, but the regularities useful for prediction need not be shallow. In philosophical prose, the regularities include dialectical dependencies: the way an objection calls for a reply, the way a conclusion would be too quick if a particular case were not addressed, the way a distinction earns its keep by absorbing a tension. Floridi's 'zeroth-order abduction' is meant to deflate the output: the model has the linguistic form of explanation without the reasoning. But in philosophy the linguistic form is not detachable from the public machinery of the argument in the way that the deflation requires. Explanatory comparisons are how philosophical reasoning appears on the page. They are not ornaments laid over a hidden inferential transaction. A model that has absorbed such patterns can generate new text in which they are redeployed. It may locate a pressure point or compare two explanations in a way not specified by the prompt. When it does this well, the result is not merely raw material for a human philosopher to work up. It is itself a philosophical text worth reading. None of this implies that fluent LLM prose is automatically valuable. LLMs can produce empty philosophy-looking text, just as humans can produce bad philosophy. The point is rather that failure must be shown by reading the text. It cannot be inferred from the absence of human-style abduction in the process that produced it. The challenge from abduction therefore mistakes a fact about the generating process for a verdict on the generated product. LLMs may not perform inference to the best explanation in the way that we do. But because philosophical inference to the best explanation is publicly encoded in prose, current models can produce texts that display abductive philosophical structure. That is enough for the present purpose. ## III. The challenge from phenomenology A further capacity worry concerns phenomenology. Even if neither authorship nor abduction rules out LLM-produced philosophy, perhaps conscious experience does. Some philosophy seems to begin from what it is like to see red, or to inhabit a body. If LLMs lack conscious experience, perhaps they can only repeat what experiencers have said, and the philosophy that depends on experiential starting points lies beyond them. Zahavy gives this worry a vivid form by way of Einstein's elevator. Einstein imagines a physicist inside an enclosed elevator accelerating through space; objects released from the hand appear to fall with the same acceleration. Zahavy's thought is that this reasoning did not proceed by verbal manipulation alone, but by simulated perceptual and bodily experience: the physicist imagines what it would be like to release the object, and reads off the consequences from inside the imagined situation. Zahavy calls this kind of reasoning manipulative abduction. The thinker varies an imagined experiential situation and attends to what would be experienced within it. If reasoning of this kind depends on simulated experience, LLMs seem blocked from it. They can manipulate descriptions of elevators and gravity, but they have never felt weight or free fall. The same worry seems to arise on the philosophical side. Mary's transition from never having seen red to seeing red for the first time, and Merleau-Ponty's discussion of self-touch, both appear to draw on phenomenological access. A system without conscious experience might seem unable to engage with such material except by recycling what other writers have already said. The first thing I want to note is that having an experience is not the same as producing philosophy about it. Most people see colours and feel pain, but this does not make them good philosophers of colour or pain. What matters philosophically is not the bare possession of experience, but how experience is articulated and used. Raw phenomenology does not enter an argument directly. It enters once it has been described, stabilised enough to be referred back to, and turned into a recognisable case to which the argument can be addressed. The phenomenology that does work in philosophy is articulated phenomenology — public, linguistic, repeatable, and open to criticism. This is where the Einstein analogy needs refinement. In the scientific case, the imagined experience helps generate a hypothesis that must then be tested against the world; the imagined elevator is a route to a claim about acceleration that can be checked. In the philosophical case, the thought experiment usually functions differently. Once the scenario is articulated, its philosophical force lies in what follows from it, not in any further empirical test. Mary is the example to keep in mind. No competent discussant of the knowledge argument has personally undergone Mary's transition. The case works because the scenario is public: Mary knows all the physical facts, has never seen red, and then appears to learn something when she first sees red. Once that structure is described, it can be assessed, varied, and resisted, and the philosophical work is done at the level of the description. Human philosophers already rely on articulated phenomenology in roughly this way. They write about blindness, or about animal experience, without necessarily having had those experiences. They rely on testimony, on literary descriptions, and on the previous philosophical literature. First-person possession is one source of phenomenological material, but it is not a universal condition of philosophical work about experience. LLMs lack raw phenomenology, but no text corpus contains raw phenomenology either. What a corpus contains is articulated phenomenology — experience already made available in language. This is the form in which phenomenology becomes usable in philosophical argument, and it is available to models trained on such texts. A model that has absorbed such material need not merely repeat the descriptions it has seen. Once a phenomenological claim has been articulated, there are further moves to make on it: the formulation can be sharpened, an argument can be drawn from the claim to a theoretical commitment elsewhere, or a redescription can be given that brings out a tension between the claim and a neighbouring case. Much of phenomenology-based philosophy consists in moves of this kind. Zahavy's worry therefore shows less than it first seemed to. First-person or sensorimotor imagination may be one route to philosophical discovery. But one route to producing valuable philosophy should not be confused with the philosophical value of the product itself. Merleau-Ponty's discussion of self-touch gives the worry its hardest case. When one hand touches the other, one hand is toucher and the other touched; the roles can reverse, but they do not perfectly coincide. This looks like a phenomenological discovery drawn from attention to one's own embodied experience, and it is hard to imagine arriving at it without that attention. Even here, however, the philosophical value of what Merleau-Ponty offers lies in the articulated structure he leaves behind — the contrast between toucher and touched, the reversibility, the non-coincidence. Once that structure is articulated, it becomes available for criticism and extension. It is no longer merely a private episode of one philosopher's bodily attention. It is a public phenomenological proposal. An LLM cannot check such a proposal by first-person attention. But producing a phenomenological proposal and checking it are different matters. A model might generate a new description of experience by drawing on the bodily vocabulary and prior phenomenological descriptions in its training data. Whether the proposal survives reflection is then a further question — as it is for any human phenomenological proposal too. The challenge from phenomenology therefore moves too quickly from the absence of experience in the producer to the absence of phenomenological value in the product. LLMs lack conscious experience, but phenomenology enters philosophy as articulated content, and articulated content is precisely what a corpus-trained model has access to. Current models can produce phenomenology-based philosophy worth reading. ## IV. Authorship redux: the challenge from elicitation A practical embarrassment remains. If current LLMs can produce philosophy worth reading, why are we not surrounded by great LLM philosophical texts? The ordinary experience of using these systems seems to support scepticism. Asked for philosophy, they often produce competent but lifeless exposition — paragraphs that tour a topic without ever applying pressure to it. This is not surprising once one notices what is being asked. A vague topic prompt does not ask for a philosophical intervention. It asks for the most probable kind of text under a topic-label, and in the training distribution that is often summary or neutral survey. A generic prompt activates a generic region of the model's learned space, and what emerges from a generic region is generic prose. The diagnosis creates a further worry. If interesting LLM philosophy appears only under careful prompting, perhaps the LLM is not really producing the philosophy after all. Perhaps the human prompter is producing philosophy by using the LLM. The worry is at its strongest when the human controls the task and then chooses among the results. In that case, the final text may be worth reading, but it can seem as though the value belongs to the human-guided process rather than to the model's production. We therefore need a more fine-grained account of production. 'Produced by an LLM' cannot mean merely 'appears in an LLM output window'. A model may output philosophy worth reading without producing the features that make it worth reading. The grammar-correction case makes this vivid. Suppose a philosopher writes a brilliant argument and asks an LLM only to correct its punctuation. The LLM's response may contain philosophy worth reading, but the model has not produced the argument in virtue of which the text is worth reading. It has output worthwhile philosophy without producing its worth-readingness. There is, then, a continuum of contribution. At one end, the model merely polishes a human-produced argument. At the other end, the human specifies a task and the model generates the philosophical move that makes the output worth reading. Between these poles lie many mixed forms of elicited production. The relevant cases for the present thesis lie towards the latter end. If the human supplies the argument and the model improves the prose, the case does not support the claim that LLMs can produce philosophy worth reading. If the human specifies a problem and the model supplies the objection or distinction that makes the text worth reading, then the output is LLM-produced in the relevant sense. The elicitation challenge also assumes too simple a contrast between autonomous producer and mere tool. Elsewhere I have argued, with Terrone, that generative AI systems of the Midjourney type are best understood as a third category — neither agents nor tools, but a new kind of artistic medium with which the user must grapple under what we call dynamic recalcitrance (Young & Terrone 2025). A similar third possibility is available here. LLMs are not intentional agents, but they are not ordinary tools either. Their outputs are partially controllable and dynamically generated. Prompting is better understood as elicitation from such a system. The prompt is not a blueprint that fixes the product in advance. It sets conditions under which the model generates. The user can constrain and iterate, but cannot determine every relevant feature of what emerges. This in turn changes how we should think about who produces what. Elicitation is not authorship. A call for papers can elicit a philosophical answer without authoring it; an interlocutor in conversation can elicit an argument from another philosopher without coming to be its author. The fact that a prompt elicits an LLM output therefore does not show that the prompter has supplied the philosophical content that makes the output worth reading. Selection is not generation either. A journal selects the papers it publishes, but it does not thereby produce them. Similarly, a reader's selection of a good LLM output is part of philosophical assessment and uptake. It is not identical to producing the argument selected.[^3] If philosophical corpora encode abductive and dialectical structure in the way I argued earlier, then skilled prompting should aim to elicit those structures rather than merely to request prose about a philosophical topic. Skilled philosophical prompting is model-sensitive task specification. It does not ask the model to sound philosophical. It gives the model a dialectical role to play. A prompt can ask the model to defend a thesis against a specific objection, or to identify the explanatory cost of rejecting a claim. Discourse markers are not magic words. Phrases such as 'one might object' matter not because their presence is sufficient for a philosophical move, but because they are surface markers of argumentative roles. A good prompt does not merely insert such phrases. It specifies the role that needs to be filled. Two forms of such prompting are worth singling out. Contrastive prompting asks why one view handles a particular case better than its rival, rather than asking for a discussion of a topic in the abstract. This mirrors the contrastive structure noted earlier, on which philosophy often becomes sharper when we ask why P rather than Q. Loveliness-sensitive prompting asks not merely for a conclusion, but for a view that explains more than its rival. Both forms aim at the explanatory virtues that make a philosophical answer worth reading. The general point is that LLM philosophy improves when the prompt creates dialectical pressure. Generic prompts invite generic continuations. A philosophical prompt should create a space in which some argumentative move is needed, and then leave that move for the model to make. Creativity belongs in this picture, but in a limited role. If the output is elicited, one might wonder whether it can really be creative. The answer depends on where the case sits on the continuum of contribution. Grammar correction is not interestingly creative; generating a new objection or a new distinction may well be. This fits a product-centred approach to creativity. Many accounts of creativity require novelty and value. The present argument need not show that LLMs are creative agents in the fullest sense. It is enough that their outputs can contain novel and valuable philosophical structure. Nor does elicitation defeat creativity. Creative work often occurs under constraints — a commission from a publisher, a question put by an interlocutor. A prompt can set a conceptual space without determining what is found within it. In that respect, elicited LLM philosophy can still be creative at the level relevant to worth-readingness. The absence of many great generic LLM texts is therefore not the verdict it can appear to be. It reflects, at least in part, immature elicitation practices and the prevalence of generic prompting. If philosophy worth reading requires live alternatives and dialectical pressure, we should not expect vague prompts to elicit it reliably. Elicitation does not defeat the claim that LLMs can produce philosophy worth reading. It shows that such production comes in degrees and depends on task conditions. The question to keep in view is whether the model has generated the philosophical structure in virtue of which the output is worth reading. The paper's answer to that question remains affirmative. LLMs do not produce worthwhile philosophy merely by being asked for 'some philosophy', and not every LLM-assisted text counts as LLM-produced in the relevant sense. But current models can generate philosophical moves that make a text worth reading. When they do, the philosophical value is present in the product, and the product is produced by the LLM in the sense that matters here. --- The four challenges all begin from something LLMs seem to lack: a philosopher behind the text, the abductive activity that drives much philosophical theorising, the conscious experience that some philosophy starts from, and the autonomous authorship that distinguishes a producer from a mere tool. None of these absences is trivial, and none of them should be wished away. But none of them settles the status of the generated text. Philosophy worth reading is assessed through what a text makes publicly available in its dialectical context. If an LLM output makes such material available, it can be philosophy worth reading. [^1]: The duplicate case is artificial in practice — no human philosopher writes exactly what an LLM happens to produce, and vice versa. The artificiality is the point. By holding the textual product fixed, the case isolates the question of whether causal history alters inferential structure, and the answer is that it does not. [^2]: The argument here depends on a charitable assumption about training mix. A model trained heavily on shallow material — encyclopaedia summaries, undergraduate term papers, online opinion — will not have absorbed the dialectical pressure that good philosophical prose carries. Whether contemporary frontier models meet the relevant condition is an empirical question, and one that the argument here does not settle. The point is rather that, where the condition is met, the resulting outputs can carry abductive structure. [^3]: The journal/reader parallel is not perfect. Selection-and-iteration loops, where the human prompter selects an output, modifies the prompt in response, and selects again, can blur the line: the more the human iterates, the more the resulting text is co-produced. The limited point survives the disanalogy — selecting a good token is not by itself producing it — but the cleanly LLM-produced cases lie at the lower-iteration end of the continuum.