## Generating Philosophy with AI — Combined Moves (v5) ## §0 — Introduction I1 — Can present-day LLMs produce philosophy worth reading? The question is not about peripheral uses of language models around philosophy. It concerns the philosophical standing of the generated product itself: whether an LLM output can reward philosophical attention in the way good philosophy does. I2 — The claim is about systems of the sort we already have. It is not a promissory note about future artificial general intelligence. The risky part of the claim is precisely that current frontier models, trained on large bodies of text and capable of sophisticated generation, already have the relevant capacity. I3 — “Worth reading” is not meant as a term of art. It names a familiar practical standard in philosophical life: the sense that a text gives one a reason to spend time with it as philosophy, rather than merely to note its topic and move on. Journals aim at that standard, even though they do not always meet it. I4 — A philosophical text worth reading is not normally a bare answer. What matters is that the work of thought is available in the text: the route by which a conclusion is reached, or the pressure that makes a familiar problem look different. I5 — The question is therefore product-centred. It asks what the text makes available for philosophical assessment and uptake. This does not mean that context is irrelevant. A text may be interesting only because of the debate it enters. But that is still a publicly assessable feature of the philosophical product, not a private feature of the producer’s mental life. I6 — The paper considers three initial worries. The first is that no LLM output can be philosophy because no philosopher stands behind it. The second is that LLMs cannot produce the abductive reasoning on which much philosophy depends. The third is that LLMs lack conscious experience and so cannot produce philosophy that begins from phenomenology. A final section then returns to authorship from another direction: if a human prompt elicits the output, is the philosophy really produced by the LLM? ## §1 — The Challenge from Authorship M1 — Some philosophers react to AI-generated philosophy by thinking that, whatever the text looks like, it cannot really be philosophy. The worry is not simply that LLM outputs are often bad. It is that no philosopher has produced them, and so they only resemble philosophical texts from the outside. M2 — This is a constitutive worry. It says that even an excellent-looking LLM output would fail before we ask whether its argument works. The defect would lie not in the paper’s inferential quality, but in the absence of a philosopher’s activity behind it. M3 — The worry gains plausibility from a comparison with art. On Davies’ performance theory, an artwork is not merely the physical object left behind, but the intentionally guided performance that eventuates in that object. A canvas that looks exactly like a Rembrandt, but is produced by accident, lacks the history that would make it a Rembrandt. M4 — If philosophy were like art in this respect, the authorship challenge would be powerful. A philosophy-looking text produced by an LLM might fail to be a work of philosophy in just the way an accidental Rembrandt-looking canvas fails to be a Rembrandt. The visible product would not settle the matter. M5 — The comparison, however, does not transfer straightforwardly. Philosophical assessment is directed at what the text says and does. We ask whether the conclusion is supported and whether the reply meets the objection. These are features of the public philosophical object, not of a private activity hidden behind it. M6 — This does not mean that the author never matters. Authorship matters for credit and responsibility. But those questions are not the same as philosophical worth-readingness. A text can be problematic to publish under someone’s name while still containing an argument worth reading. M7 — Consider a duplicate philosophical text. If the same sequence of sentences appears in a human-written article and in an LLM output, the inferential relations available to the reader are the same. The causal history changes who should receive credit; it does not by itself change which conclusion follows from which premises. M8 — In this respect, philosophy is closer to proof than to painting. A machine-generated proof is not invalid because the machine did not understand it. The analogy should not be overplayed, since philosophy is richer than proof, but the relevant point remains: public inferential structure is not erased by the absence of a human mental performance. M9 — One might object that an LLM output is not really an argument because no one asserts its premises. But philosophical assessment does not always require sincere assertion by the producer. We assess arguments in dialogues and reductios without treating every sentence as the author’s straightforward commitment. M10 — Anonymous review reflects the same norm. Referees are asked to assess a paper by what it says rather than by who wrote it. This does not prove that authorship is never relevant, but it does show that philosophical merit is treated as assessable without reconstructing the author’s private activity. M11 — The authorship challenge therefore fails as a constitutive barrier. LLMs may not philosophise as persons do, and LLM publication may raise difficult questions about responsibility. But authorship alone does not show that an LLM-produced text cannot contain philosophy worth reading. M12 — Once that point is in place, the more serious worries concern production capacities. Perhaps only a system that performs abductive reasoning can produce abductively good philosophy. Or perhaps only a conscious subject can produce philosophy that starts from experience. These are different objections, and they need different answers. ## §2 — The Challenge from Abduction M1 — Floridi and colleagues do not deny that LLMs produce explanation-like answers. They deny that such answers are produced by abductive inference. The model does not select an explanation because it best accounts for the evidence; it generates a continuation made probable by training. M2 — The challenge becomes serious for philosophy if Williamson is right that much philosophical theorising proceeds abductively. Philosophers compare potential theories and ask what each would explain if true. If LLMs cannot do that, perhaps they can only imitate the surface of philosophical reasoning. M3 — The response should not be that LLMs secretly perform human-style inference to the best explanation. That is unlikely to be the right hill to die on. An LLM does not understand the problem as a problem, and it does not knowingly weigh rivals under a norm of truth. M4 — The issue is whether the generated text can display abductive philosophical structure. A text can compare two explanations and show why one handles the pressure better. These are properties of the product, even if the process that generated the product was not itself human abduction. M5 — Philosophy externalises much of its abductive work in prose. When a paper says that one view explains what a rival cannot, or introduces a distinction to answer an objection, the comparison is not hidden behind the text. It is part of what the reader assesses. M6 — Lipton’s distinction between actual and potential explanation is useful here. Inquiry does not begin with explanations already known to be true. It begins with candidates whose explanatory force can be assessed before their truth is settled. An LLM output can offer such a candidate: a way the relevant material would hang together if true. M7 — Lipton’s distinction between likeliness and loveliness makes the point more precise. Likeliness concerns warrant; loveliness concerns the understanding an explanation would provide if true. A generated answer can be assessed for loveliness before anyone has decided whether it is likely. M8 — Loveliness is not a private glow in the reasoner’s head. It shows itself in the way a text makes a problem more intelligible than it was before: by revealing why one explanatory route has more force than another. That kind of intelligibility can be present or absent in an LLM output. M9 — Floridi’s own diagnosis gives us the channel. If LLMs produce plausible explanations because they have been trained on texts that encode reasoning structures, then philosophical training data matters. Philosophy is a textual practice in which explanatory comparison is made public. M10 — The philosophical corpus is not a heap of sentences about philosophical topics. It is a record of arguments challenged and replies refined. A model trained on such material is trained on prose shaped by philosophical pressure. M11 — This is why the “mere surface pattern” worry is too quick. Next-token prediction is the training task, but the regularities useful for prediction need not be shallow. In philosophical prose they include dialectical dependencies: when an objection calls for a reply, or when a conclusion would be too quick. M12 — Floridi’s “zeroth-order abduction” is meant to deflate the output: the model has the linguistic form of explanation without the reasoning. But in philosophy the linguistic form is not detachable from the public machinery of the argument in the way the objection requires. Explanatory comparisons are how philosophical reasoning appears on the page. M13 — A model that has absorbed such patterns can generate new text in which they are redeployed. It may locate a pressure point or compare two explanations in a way not specified by the prompt. When it does so well, the result is not merely material for a human philosopher. It is a philosophical text worth reading. M14 — None of this implies that fluent LLM prose is automatically valuable. LLMs can produce empty philosophy-looking text, just as humans can produce bad philosophy. The point is that failure must be shown by reading the text, not inferred from the absence of human-style abduction in the process that produced it. M15 — The challenge from abduction therefore mistakes a fact about the generating process for a verdict on the generated product. LLMs may not perform inference to the best explanation. Nevertheless, because philosophical inference to the best explanation is publicly encoded in prose, current models can produce texts that instantiate abductive philosophical structure. ## §3 — The Challenge from Phenomenology M1 — A further capacity worry concerns phenomenology. Even if authorship and abduction do not rule out LLM-produced philosophy, perhaps conscious experience does. Some philosophy seems to begin from what it is like to see red or to inhabit a body. M2 — Zahavy gives this worry a vivid form through Einstein’s elevator. Einstein imagines a physicist inside an enclosed elevator accelerating through space. Objects released from the hand appear to fall with the same acceleration. Zahavy’s thought is that this reasoning did not proceed by verbal manipulation alone, but by simulated perceptual and bodily experience. M3 — Zahavy calls this kind of reasoning manipulative abduction. The thinker varies an imagined experiential situation and attends to what would be experienced within it. If such reasoning depends on simulated experience, LLMs seem blocked from it. They can manipulate descriptions of elevators and gravity, but they have never felt weight or free fall. M4 — The same worry seems to arise in philosophy. Mary seeing red and Merleau-Ponty on self-touch both appear to draw on phenomenological access. If an LLM has no conscious experience, perhaps it can only repeat what experiencers have said. M5 — The first thing to note is that having an experience is not the same as producing philosophy about it. Most people see colours and feel pain, but this does not make them good philosophers of colour or pain. What matters philosophically is not raw possession of experience, but how experience is articulated and used. M6 — Raw phenomenology does not enter an argument directly. It enters once it is described, stabilised, contrasted, or turned into a case. Philosophy works with articulated phenomenological content. That content is public, linguistic, repeatable, and open to criticism. M7 — This is where the Einstein analogy needs refinement. In the scientific case, the imagined experience helps generate a hypothesis that must then be tested against the world. In the philosophical case, the thought experiment usually functions differently. Once the scenario is articulated, its philosophical force lies in what follows from it. M8 — Mary is the central example. No competent discussant of the knowledge argument has personally undergone Mary’s transition. The case works because the scenario is public: Mary knows all the physical facts, has never seen red, and then appears to learn something when she first sees red. Once that structure is described, it can be assessed, varied, and resisted. M9 — Human philosophers already rely on articulated phenomenology. They write about blindness or animal experience without necessarily having had those experiences. They rely on testimony, literature, and previous philosophy. First-person possession is one source of phenomenological material, not a universal condition of philosophical work about experience. M10 — LLMs lack raw phenomenology, but no text corpus contains raw phenomenology either. What the corpus contains is articulated phenomenology: experience made available in language. This is the form in which phenomenology becomes usable in philosophy, and it is available to models trained on such text. M11 — The model need not merely repeat these descriptions. Once phenomenology has been articulated, it can be worked on philosophically: its formulation can be sharpened, its implications drawn out, and its pressure on a theory assessed. This is much of what phenomenology-based philosophy consists in. M12 — Zahavy’s worry therefore shows less than it first seemed to. First-person or sensorimotor imagination may be one route to intellectual discovery. But one route to producing valuable philosophy should not be confused with the value of the philosophical product itself. M13 — Merleau-Ponty’s discussion of self-touch gives the worry its hardest case. When one hand touches the other, one hand is toucher and the other touched; the roles can reverse, but they do not perfectly coincide. This looks like a genuine phenomenological discovery drawn from attention to one’s own embodied experience. M14 — Even here, the philosophical value lies in the articulated structure: touching and touched, reversibility, non-coincidence. Once articulated, that structure becomes available for criticism and extension. It is no longer merely a private episode; it is a public phenomenological proposal. M15 — An LLM cannot check such a proposal by first-person attention. But production and checking are different. A model might generate a new way of describing experience by drawing on bodily vocabulary and prior phenomenological descriptions. Whether that proposal survives reflection is a further question, as it is for human phenomenological proposals too. M16 — The challenge from phenomenology therefore moves too quickly from absence of experience in the producer to absence of phenomenological value in the product. LLMs lack conscious experience, but phenomenology enters philosophy as articulated content. Since articulated phenomenology is public and inferentially usable, current models can produce phenomenology-based philosophy worth reading. ## §4 — Authorship Redux: The Challenge from Elicitation M1 — A practical embarrassment remains. If current LLMs can produce philosophy worth reading, why are we not surrounded by great LLM philosophical texts? The ordinary experience of using these systems seems to support scepticism. Asked for philosophy, they often produce competent but lifeless exposition. M2 — This is not surprising. A vague topic prompt does not ask for a philosophical intervention. It asks for the most probable kind of text under a topic-label, and in the training distribution that is often summary or neutral survey. A generic prompt activates a generic region of the model’s learned space. M3 — But this diagnosis creates a further worry. If interesting LLM philosophy appears only under careful prompting, perhaps the LLM is not producing the philosophy after all. Perhaps the human prompter is producing philosophy by using the LLM. M4 — The worry is strongest when the human controls the task and then chooses among the results. In that case, the final text may be worth reading, but it can seem as though the value belongs to the human-guided process rather than to the model’s production. M5 — We therefore need a more fine-grained account of production. “Produced by an LLM” cannot mean merely “appears in an LLM output window.” A model may output philosophy worth reading without producing the features that make it worth reading. M6 — The grammar-correction case makes this vivid. Suppose a philosopher writes a brilliant argument and asks an LLM only to correct punctuation. The LLM’s response may contain philosophy worth reading, but the model has not produced the argument in virtue of which it is worth reading. It has output worthwhile philosophy without producing its worth-readingness. M7 — There is therefore a continuum of contribution. At one end, the model merely polishes a human-produced argument. At the other end, the human specifies a task and the model generates the philosophical move that makes the output worth reading. Between these cases lie many mixed forms of elicited production. M8 — The relevant cases for this paper lie toward the latter end. If the human supplies the argument and the model improves the prose, the case does not support the paper’s thesis. If the human specifies a problem and the model supplies the objection or distinction that makes the text worth reading, then the output is LLM-produced in the relevant sense. M9 — The elicitation challenge assumes too simple a contrast between autonomous producer and mere tool. The discussion of Midjourney suggests a third possibility. Generative AI systems are not intentional agents, but they are not ordinary tools either: their outputs are partially controllable and dynamically generated. M10 — Prompting is better understood as elicitation from such a system. The prompt is not a blueprint that fixes the product in advance. It sets conditions under which the model generates. The user can constrain and iterate, but cannot determine every relevant feature of what emerges. M11 — Elicitation is not authorship. A call for papers can elicit a philosophical answer without authoring it. Likewise, the fact that a prompt elicits an LLM output does not show that the prompter supplied the philosophical content that makes the output worth reading. M12 — Selection is not generation either. A journal may select a paper, but it does not thereby produce it. Similarly, a reader’s selection of a good LLM output is part of assessment and uptake, not identical to producing the argument selected. M13 — The account of prompting should now return to the argument of §2. If philosophical corpora encode abductive and dialectical structures, then skilled prompting should aim to elicit those structures rather than merely request prose about a philosophical topic. M14 — Skilled philosophical prompting is model-sensitive task specification. It does not ask the model to sound philosophical; it gives the model a dialectical role to play. A prompt can, for instance, ask the model to defend a thesis against a specific objection, or to identify the explanatory cost of rejecting a claim. M15 — Discourse markers are not magic words. Phrases such as “one might object” matter because they are surface markers of argumentative roles. A good prompt does not merely insert those phrases; it specifies the role that needs to be filled. M16 — Contrastive prompting is one promising form of this. Instead of asking for a discussion of physicalism, ask why one view handles a specific case better than its rival. This mirrors the contrastive structure discussed earlier: philosophy often becomes sharper when we ask why P rather than Q. M17 — Loveliness-sensitive prompting is another. Ask not only for a conclusion, but for a view that explains more than its rival. Such a prompt aims at the explanatory virtues that make a philosophical answer worth reading. M18 — The general point is that LLM philosophy improves when the prompt creates dialectical pressure. Generic prompts invite generic continuations. Philosophical prompts should create a space in which some argumentative move is needed. M19 — Creativity belongs here, but only in a limited role. If the output is elicited, one might wonder whether it can really be creative. The answer depends on where we are on the continuum. Grammar correction is not interestingly creative; generating a new objection might be. M20 — This fits a product-centred approach to creativity. Many accounts of creativity require novelty and value. The present paper need not show that LLMs are creative agents in the fullest sense. It is enough that their outputs can contain novel and valuable philosophical structures. M21 — Nor does elicitation defeat creativity. Creative work often occurs under constraints. A prompt can set a conceptual space without determining what is found within it. In that respect, elicited LLM philosophy can still be creative at the level relevant to worth-readingness. M22 — The absence of many great generic LLM texts is therefore not decisive. It reflects, at least in part, immature elicitation practices and the dominance of generic prompting. If worthwhile philosophy requires live alternatives and dialectical pressure, we should not expect vague prompts to elicit it reliably. M23 — Elicitation does not defeat the claim that LLMs can produce philosophy worth reading. It shows that such production comes in degrees and depends on task conditions. The central question is whether the model generated the philosophical structure in virtue of which the output is worth reading. M24 — The paper’s answer remains affirmative. LLMs do not produce worthwhile philosophy merely by being asked for “some philosophy,” and not every LLM-assisted text counts as LLM-produced in the relevant sense. But current models can generate philosophical moves that make a text worth reading. When they do, the philosophical value is present in the product, and the product is produced by the LLM in the sense that matters here. ## Closing payoff The four challenges all begin from something LLMs seem to lack. None of these absences is trivial. But none settles the status of the generated text. Philosophy worth reading is assessed through what a text makes publicly available in its dialectical context. If an LLM output makes such material available, then it can be philosophy worth reading. Generating Philosophy With Ai Draft Iteration V4 ## Generating Philosophy with AI Can current large language models produce philosophy worth reading? I am not asking whether they can be useful around the edges of philosophical work. The question is about the philosophical standing of the generated product itself: whether an LLM output can reward philosophical attention in the way good philosophy does. I shall argue that, in the right conditions, it can. The claim concerns systems of the sort we already have, not future artificial general intelligence. This is part of what makes it risky. It would be much easier to say that some future system, with capacities current models lack, might eventually produce worthwhile philosophy. The stronger claim is that current frontier models, trained on large bodies of text and capable of sophisticated generation, already have the relevant capacity. By “worth reading” I do not intend a term of art. The phrase names a familiar practical standard in philosophical life: the sense that a text gives us a reason to spend time with it as philosophy, rather than merely to note its topic and move on. Journals aim at this standard, even though they do not always meet it. A philosophical text worth reading is rarely a bare answer. What matters is that the work of thought is available in the text: the route by which a conclusion is reached, or the pressure that makes a familiar problem look different. This makes the question product-centred. To ask whether a text is worth reading is to ask what it makes available for philosophical assessment and uptake. This does not make context irrelevant. A paper may be worth reading because of the debate it enters. But that debate is still a public feature of the philosophical product. It is not a private feature of the producer’s mental life. The paper considers four challenges. I begin with the idea that no LLM output can be philosophy because no philosopher stands behind it. I then turn to two capacity worries: LLMs do not perform abductive reasoning, and they do not have conscious experience. The final section returns to authorship from another direction. If a human prompt elicits the output, is the philosophy really produced by the LLM? ## I. The challenge from authorship Some philosophers react to AI-generated philosophy by thinking that, whatever the text looks like, it cannot really be philosophy. The complaint is not the familiar one that LLM outputs are often bad. It is that no philosopher has produced them, and so they only resemble philosophical texts from the outside. On this view, the defect lies not in any feature of the paper’s argument but in the absence of a philosopher’s activity behind it. An excellent-looking LLM output would already have failed before we asked whether its argument worked. The worry can be made more precise by adapting Davies’ performance theory of art. Davies writes: > The work — what the artist achieves — is the process eventuating in that product. Works themselves are neither structures nor objects simpliciter, nor are they contextualized structures or objects \[...\]. They are, rather, intentionally guided generative performances that eventuate in contextualized structures or objects. (Davies, *Art as Performance*, p. 98) On Davies’ view, the artwork is not simply the object the artist’s activity leaves behind. It is the performance that eventuates in that object. A canvas indistinguishable from a Rembrandt but produced by accident lacks the history that would make it a Rembrandt. Transpose the view to philosophy, and the philosophical work becomes the philosopher’s sustained activity in producing the text. The text is what that activity leaves behind. If philosophy were like art in this respect, the authorship challenge would be powerful. A text produced by an LLM might fail to be a work of philosophy in just the way an accidental Rembrandt-looking canvas fails to be a Rembrandt. The visible product would not settle the matter. The transposition is much less plausible once we consider how philosophical texts are actually assessed. We ask whether the argument works and whether the position improves the dialectical situation into which it enters. These questions are answered by attending to what the text publicly makes available. They do not require us to recover a further private activity behind it. Achievement-talk may seem to push in the other direction. We say that a philosopher has achieved something in a paper, and this can make it sound as though the achievement lies behind the text. But in ordinary philosophical contexts the claim is usually parasitic on the text. To say that an author has achieved something in a paper is to say that the paper succeeds in doing some philosophical work. The achievement is not an additional object hidden behind the public argument. None of this is to say that authorship never matters. Credit and responsibility turn on it. But these questions are not the same as philosophical worth-readingness. A text may be ethically problematic to publish under someone’s name while still containing an argument worth reading. Conversely, a text may be impeccably authored and not worth anyone’s time. Consider a duplicate philosophical text. Suppose the same sequence of sentences appears in a human-written article and in an LLM output. The inferential relations available to the reader are the same in both. Whatever the conclusion follows from in one, it follows from in the other. The causal history changes who should receive credit; it does not by itself change which premises support which conclusions. In this respect, philosophy is closer to proof than to painting. A machine-generated proof is not invalid because the machine did not understand it. The analogy should not be pushed too hard, since philosophy is richer than proof. The relevant point survives the difference: public inferential structure is not erased by the absence of a human mental performance. One might object that an LLM output is not really an argument because no one asserts its premises. There is no agent standing behind the words committed to their truth. But philosophical assessment does not always require sincere assertion by the producer. We assess the arguments that appear in dialogues and reductios without treating every sentence as the author’s straightforward commitment. The norms governing argument-assessment and the norms governing assertion are not the same. Anonymous review reflects this product-centred norm. Referees are asked to assess a paper by what it says rather than by who wrote it. This practice does not show that authorship is never relevant. It does show that philosophical merit is treated as assessable without reconstructing the author’s private activity. The authorship challenge is therefore left with a more limited conclusion. LLM publication may raise difficult questions about credit and responsibility. But authorship of that kind does not constitute philosophy. The challenge that an LLM output cannot be philosophy at all because no philosopher’s activity stands behind it fails because philosophical worth-readingness is a feature of what the text publicly makes available. The harder worries concern the system’s capacities. Perhaps only a system that performs abductive reasoning can produce abductively good philosophy. Or perhaps only a conscious subject can produce philosophy that starts from experience. These are different objections, and they need different answers. ## II. The challenge from abduction Floridi et al. do not deny that LLMs produce explanation-like answers. They deny that such answers are produced by abductive inference. Their formulation is direct: > LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al., p. 9) The critique becomes serious for the present thesis if Williamson is right that contemporary philosophical theorising often proceeds abductively. On Williamson’s account, philosophers defend a view by comparing it with rivals and asking which, if true, would best explain the evidence. The relevant virtues are displayed in the theory as presented. We need not inspect the theorist’s psychology before asking whether the proposal is ad hoc or explanatorily powerful. If philosophical theorising consists in this kind of comparison, and Floridi is right that LLMs only mimic the surface of such reasoning, then the philosophical character of LLM outputs would seem to be merely apparent. It would be wrong to respond by arguing that LLMs secretly perform human-style inference to the best explanation. They do not understand the problem as a problem, and they do not knowingly weigh rivals under a norm of truth. The response should go elsewhere. If the virtues by which a philosophical theory is assessed are features of the theory as presented, then the question is not whether an LLM can run the human inference that would normally produce such a theory. The question is whether a generated text can display abductive philosophical structure. A text can compare two explanations and show why one handles the pressure better. That is a property of the product, even if the process that generated the product was not itself human abduction. Philosophy externalises much of its abductive work in prose. When a paper says one view explains what a rival cannot, the comparison is on the page rather than hidden behind it. It is part of what the reader assesses. Lipton’s distinction between actual and potential explanation bears on this. Inquiry does not begin with explanations already known to be true. It begins with candidates whose explanatory force can be assessed before their truth is settled. An LLM output can offer such a candidate: a way the relevant material would hang together if true. The further distinction Lipton draws between likeliness and loveliness sharpens this. Likeliness concerns warrant; loveliness concerns the understanding an explanation would provide if true. As Lipton writes: “Likeliness speaks of truth; loveliness of potential understanding.” A generated answer can be assessed for loveliness before anyone has settled whether it is likely. Loveliness is not a feature of the reasoner’s mental life. It shows itself in the way a text makes a problem more intelligible than it was before. This kind of intelligibility can be present or absent in an LLM output. When present, the text has the property philosophical readers track when they say a paper repays their attention. Floridi himself points to the channel by which models come to produce such texts. The Floridi paper says: “This effect is due to the model’s training on human-generated texts that encode reasoning structures.” The point is not incidental. Philosophy is conducted in writing, and the writing is what the LLM has been trained on. Much of the philosophical corpus is prose shaped by philosophical assessment: arguments answered by later arguments, positions refined under objection. This process is noisy. The corpus is not a purified record of philosophical virtue. But it is not inert either. To the extent that philosophical uptake has been guided by explanatory and dialectical standards, the corpus contains traces of those standards; and a model trained on that corpus can learn regularities shaped by them. The worry that LLM outputs exhibit only a “mere surface pattern” therefore moves too quickly. The relevant regularities are not merely verbal. Philosophical prose has a recognisable dialectical grammar: objections create burdens and replies discharge them. A model trained on such prose can learn these dependencies even though its training objective is next-token prediction. Floridi’s “zeroth-order abduction” is meant to deflate the output: the model has the linguistic form of explanation without the reasoning. But in philosophy the linguistic form is not detachable from the public machinery of the argument in the way the objection requires. Explanatory comparisons are how philosophical reasoning appears on the page. A model that has absorbed such patterns can generate new text in which they are redeployed. It may locate a pressure point in a way the prompt did not specify. When it does this well, the result is not raw material for a human philosopher to work up but itself a philosophical text worth reading. None of this implies that fluent LLM prose is automatically valuable. LLMs can produce empty philosophy-looking text, just as humans can produce bad philosophy. The failure of any particular output is shown by reading it, not inferred from the absence of human-style abduction in the process that produced it. LLMs may not perform inference to the best explanation in the way we do. But the abductive structure that bears on philosophical assessment is carried by prose, and a model trained on philosophical prose has been trained on the kind of public object on which such assessment works. The producer’s inference and the product’s structure can be assessed separately. For the purpose of deciding whether the generated text is philosophy worth reading, it is the second that matters. ## III. The challenge from phenomenology A further capacity worry concerns phenomenology. Even if neither authorship nor abduction rules out LLM-produced philosophy, perhaps conscious experience does. Some philosophy seems to begin from what it is like to see red, or to feel a particular intuition take hold. If LLMs lack conscious experience, perhaps they can only repeat what experiencers have said, and the philosophy that depends on experiential starting points lies beyond them. Zahavy gives this worry a vivid form by way of Einstein’s elevator. The passage is worth quoting at length: > Einstein’s variation required inventing new axioms based on a physical intuition that did not yet exist in the mathematics. He envisioned a physicist inside an elevator being uniformly accelerated through deep space. Inside this enclosure, the sensory experience reveals a specific pattern: when objects are released, the floor rushes up to meet them. To the physicist, the objects appear to fall with identical acceleration, regardless of composition. Thus, the simulation here was not a permutation of symbols, but a manipulation of perceptual experience. (Zahavy 2026, §5) Zahavy describes this kind of reasoning as manipulative abduction. The thinker varies an imagined experiential situation and attends to what would be experienced within it. If reasoning of this kind depends on simulated experience, LLMs seem blocked from it. They can manipulate descriptions of elevators and gravity, but they have never felt weight or free fall. The same worry transfers to philosophy. Mary’s case in Jackson’s knowledge argument turns on the apparent gap between knowing all the physical facts about colour and seeing red for the first time. Merleau-Ponty’s discussion of self-touch turns on attention to the structure of embodied experience. In each instance, an LLM without the relevant experience might seem unable to engage with the material except by recycling what other writers have already said. Having an experience is not the same as producing philosophy about it. Most people see colours and feel pain, but this does not make them good philosophers of colour or pain. The bare possession of experience is not what matters philosophically. What matters is how experience is articulated and used. Raw phenomenology does not enter an argument directly. It enters once it has been described and turned into a case the argument can address. Articulated phenomenology — phenomenology made public in language — is what does the work. The Einstein analogy therefore needs refinement. In the scientific case, the imagined experience helps generate a hypothesis that still has to answer to empirical reality. The lift thought experiment gave Einstein the equivalence principle, but the principle had still to be confirmed. In the philosophical cases at issue here, the role of the thought experiment is different. Once articulated, its force depends primarily on how the scenario bears on a conceptual claim. What gets taken up by philosophy is not raw experience, but an experience as described. Mary’s case shows what description-level work amounts to. No competent discussant of the knowledge argument has personally undergone her transition. The case works because the scenario is public: Mary knows all the physical facts, has never seen red, and on first seeing red appears to learn something. Once that scenario is on the page, the philosophical work proceeds at the level of articulation. Take Lewis’s response. Faced with the conclusion that Mary learns a new fact, Lewis does not need to undergo her experience to reply. He distinguishes propositional knowledge from abilities such as recognising red on later occasions, and argues that what Mary acquires is an ability rather than a fact. Whatever one’s verdict on Lewis, the move is genuine philosophical work, and the pressure it puts on Jackson is pressure at the level of the public scenario. Human philosophers already rely on articulated phenomenology in this way. They write about blindness and about animal experience without necessarily having had those experiences, drawing on testimony and on the previous philosophical literature. First-person possession is one source of phenomenological material; it is not a universal condition of philosophical work about experience. LLMs lack raw phenomenology, but no text corpus contains raw phenomenology either. What a corpus contains is articulated phenomenology — experience already made available in language — and this is the form in which phenomenology becomes usable in philosophical argument. A model that has absorbed such material need not merely repeat the descriptions it has seen. It can work on them philosophically. It can reshape a case and test its pressure on a theory. The important point is not that the model has enjoyed the experience described. It is that, once the experience has entered the public space of reasons, the philosophical work proceeds on the articulated material. Merleau-Ponty’s discussion of self-touch gives the worry its hardest case. When one fingertip touches another, one finger plays the role of toucher and the other of touched. The roles can reverse, but not simultaneously: at any given instant, the body is split between touching and touched. The observation looks like a phenomenological discovery drawn from sustained attention to embodied experience. A model cannot check Merleau-Ponty’s description by attending to its own body. But that does not prevent it from working philosophically on the description once articulated. Take the case in which the touching limb is anaesthetised. The toucher-touched asymmetry, on the original Merleau-Pontian observation, depends on each finger being able to play either role. With one finger anaesthetised, the asymmetry seems to collapse: the anaesthetised finger can be touched but cannot itself touch. So either the structure Merleau-Ponty identifies generalises only to non-anaesthetised parts of the body, or the relevant “touching” is decoupled from sensory feedback in a way the original description does not flag. Either branch puts pressure on what the discovery commits one to. A model trained on Merleau-Ponty and his interlocutors can press the case without itself having touched anything. The pressure it generates is pressure on the articulated structure. This is not to deny the special role of first-person attention in some phenomenological discoveries. An LLM cannot originate a phenomenological description by attending to its own experience. It has no such experience to attend to. But this does not show that it cannot generate a candidate articulation not explicitly present in its training data, by drawing on prior descriptions and conceptual materials. What it lacks is not the capacity to produce a phenomenological proposal, but the capacity to verify it introspectively. Whether the proposal survives reflection is a further question, as it is for human phenomenological proposals too. The challenge from phenomenology therefore moves too quickly from absence of experience in the producer to absence of phenomenological value in the product. LLMs lack conscious experience, but phenomenology enters philosophy as articulated content. Since articulated phenomenology is public and inferentially usable, current models can produce phenomenology-based philosophy worth reading. ## IV. Authorship redux: the challenge from elicitation A practical embarrassment remains. If current LLMs can produce philosophy worth reading, why are we not surrounded by great LLM philosophical texts? The ordinary experience of using these systems seems to support scepticism. Asked for philosophy, they often produce competent but lifeless exposition — paragraphs that tour a topic without ever applying pressure to it. The diagnosis is not far to seek. A vague topic prompt does not ask for a philosophical intervention. It asks for the most probable kind of text under that topic-label, and in the training distribution that is often hedged survey, because the bulk of philosophical text written at that level of generality takes that form. The prompt is generic, and so is the region of the model’s learned space it activates. The diagnosis creates a further worry. If interesting LLM philosophy appears only under careful prompting, perhaps the LLM is not really producing the philosophy after all. Perhaps the human prompter is producing philosophy by using the LLM. The worry is sharpest when the human controls the task and then chooses among the results. The final text may be worth reading, but the value can seem to belong to the human-guided process rather than to the model’s production. So “produced by an LLM” has to mean more than “appearing in an LLM output window”. A model may output philosophy worth reading without producing the features that make it worth reading. Take the grammar-correction case. A philosopher writes a brilliant argument and asks an LLM only to correct its punctuation. The LLM’s response may contain philosophy worth reading, but the model has not produced the argument in virtue of which the text is worth reading. It has output worthwhile philosophy without producing its worth-readingness. So there is a continuum of contribution. At one end, the model merely polishes a human-produced argument; at the other, the human specifies a task and the model generates the philosophical move that makes the output worth reading. The cases that support the present thesis lie towards the latter end. Where the human supplies the argument and the model improves the prose, the LLM has not produced philosophy worth reading. Where the human specifies a problem and the model supplies the objection or distinction that makes the text worth reading, the output is LLM-produced in the sense that matters here. The elicitation challenge assumes too simple a contrast between autonomous producer and mere tool. Elsewhere I have argued, with Terrone, that generative AI systems of the Midjourney type are best understood neither as agents nor as ordinary tools, but as generative systems with which users interact under conditions of partial control. The same lesson applies here, mutatis mutandis. LLMs are not intentional agents, but the ordinary tool model is also too crude. Prompting is not command execution. The prompt is not a blueprint that fixes the product in advance. It sets conditions under which the model generates. The user can constrain and iterate, but cannot determine every relevant feature of what emerges. Recast as elicitation, prompting reveals itself as a different relation. Elicitation is not authorship. A call for papers elicits philosophical answers without authoring them; an interlocutor in conversation elicits arguments from another philosopher without coming to be their author. So a prompt’s having elicited an LLM output does not by itself show that the prompter has supplied the philosophical content that makes the output worth reading. Selection is not generation either. A journal selects the papers it publishes but does not thereby produce them, and a reader’s selection of a good LLM output is, similarly, an act of assessment rather than of production. If philosophical corpora encode abductive and dialectical structure in the way I argued earlier, skilled prompting should aim to elicit those structures rather than to request prose about a topic. Skilled philosophical prompting is model-sensitive task specification. It does not ask the model to sound philosophical; it gives the model a dialectical role to play. The point is not to insert phrases such as “one might object,” as though they were magic words. Such phrases matter only because they mark argumentative roles. A good prompt specifies the role to be filled. Two forms of such prompting deserve attention, though the details are for another occasion. Contrastive prompting asks why one view handles a particular case better than its rival, rather than asking for a discussion of a topic in the abstract. Loveliness-sensitive prompting asks not for a conclusion but for a view that, if true, would explain more than its rival. Both target the explanatory virtues that make a philosophical answer worth reading. LLM philosophy improves when the prompt creates dialectical pressure: generic prompts invite generic continuations, and a philosophical prompt should create a space in which some argumentative move is needed, and then leave that move for the model to make. Creativity belongs here, but in a limited role. If the output is elicited, one might wonder whether it can really be creative. The answer depends on where the case sits on the continuum. Grammar correction is not interestingly creative. Generating a new objection or a new distinction may be. This fits a product-centred approach to creativity. Many accounts require novelty and value, and the present argument need not show that LLMs are creative agents in the fullest sense. It is enough that their outputs can contain novel and valuable philosophical structure. Nor does elicitation defeat creativity. Creative work often happens under constraints; a prompt can set a conceptual space without determining what is found within it. Elicited LLM philosophy can be creative at the level relevant to worth-readingness. A related diagnosis sits behind the practical embarrassment. The corpus an LLM is trained on is the product of many rounds of philosophical criticism. A single completion does not reproduce that history. But a prompt can ask the model to stage, within the generated text, some of the operations that the corpus records across time: an objection pressed, a reply attempted, a distinction introduced under pressure. The human prompt supplies local conditions of elicitation; it need not supply the philosophical move. When the model generates that move, the output is LLM-produced in the sense relevant here. The absence of many great generic LLM texts is therefore not the verdict it can appear to be. It reflects, at least in part, immature elicitation practices and the prevalence of generic prompting. If philosophy worth reading requires live alternatives and dialectical pressure, vague prompts will rarely elicit it. Elicitation does not defeat the claim that LLMs can produce philosophy worth reading: it shows that such production comes in degrees and depends on task conditions. The question to keep in view is whether the model has generated the philosophical structure in virtue of which the output is worth reading. The answer can be yes. LLMs do not produce worthwhile philosophy merely by being asked for “some philosophy”, and not every LLM-assisted text counts as LLM-produced in the relevant sense. But current models can generate philosophical moves that make a text worth reading, and where they do, the philosophical value is present in the product, and the product was produced by the LLM in the sense that matters here. ## Conclusion The four challenges treat what an LLM lacks as decisive for what its outputs can be. None of the absences they begin from is trivial. LLMs do not philosophise as persons do, do not perform human-style inference to the best explanation, do not have conscious experience, and do not normally produce valuable outputs without elicitation. None of this settles whether what appears on the page is philosophy worth reading. Worth-readingness, on the account defended here, is assessed through what a text makes available in its dialectical context. Where an LLM output carries the relevant comparison, presses the relevant objection, or articulates a phenomenological structure worth thinking about, the output is philosophy worth reading. What is owed at that point is the assessment philosophy always owes its texts: do the arguments hold? ## Notes 1. The duplicate case is artificial in practice, but the artificiality is doing controlled work. By holding the textual product fixed, the case isolates the question whether causal history alters inferential structure. It does not. 2. The claim is not that every model trained on any philosophy-adjacent corpus will produce outputs with abductive structure. It is that the process described here is a plausible route by which current frontier models can do so. 3. The term “manipulative abduction” originates with Magnani, who introduced it for cases of hypothesis generation through the active construction of mental models. Zahavy uses it in this sense and applies it to Einstein’s elevator argument as a paradigm case in physics. 4. Iteration complicates attribution, but it does not automatically transfer production to the human. What matters is still whether the human supplies the philosophical move or elicits it. 5. A concrete contrast may help. Compare “Discuss whether physicalism is true” with a prompt asking why one version of physicalism can answer a specific objection that another cannot. The first asks for a survey. The second specifies a dialectical role. What matters is not the wording but the philosophical task imposed.