# I would like you to write me a new draft of section 2 of my generating philosophy draft by impole...
## Skills Used
- deep-writing-mode
- epistemic-discipline
- nick-analytic-voice
- nick-topic-sentences
- no-example-lists
- source-work
- writing-standards
## Retrieval Notes
- Session id: `f3c0aab5-a110-4ee7-b278-91fa01dc5667`
- Last activity: `2026-05-05T08:54:53.776Z`
## My Notes
<!-- Add your notes here. This section is preserved across syncs. -->
## Conversation
### User
/deep-writing-mode
I would like you to write me a new draft of section 2 of my generating philosophy draft by impolementing the plan a the very very end of this converation. you need to read the rest of the conversation for context. make sure that you explain your analysis, justification, reasons etc etc. for your answer BEFORE giving me your final answer.
### Assistant
Let me read the conversation file to find the plan.
### Assistant
Let me find the current Generating Philosophy draft and the original Draft Review Feedback conversation to understand the full context.
### Assistant
Let me look at the transcript files and check for any other plan documents.
### Assistant
Let me search for any section 2 plan documents that might be elsewhere.
### Assistant
Let me look for the original parent conversation that this branched from.
### Assistant
Let me now check the sources that need to be verified for the new draft.
### Assistant
Now let me ground myself in Nick's voice by reading published papers. The deep-writing-mode skill requires this before drafting.
### Assistant
Now let me verify the key Lipton passages from the source extraction.
### User
[Request interrupted by user]
### User
btw, here is the most uptodate draft. this is what you should be referring to # Generating Philosophy with AI 28 Apr 2026
> [!NOTE]
> this is LLM text that I have shouted into existence. currently rewriting it line by line.
The last decade or so has seen the rise of generative artificial intelligence: systems that produce text, images, code, music, video, and other outputs in response to prompts. AI has had success in domains where the value of an output is not exhausted by its superficial fluency. For example, in February 2026, researchers working on gluon scattering amplitudes gave GPT-5.2 worked examples for three, four, five, and six particles and asked it to find the general formula. GPT-5.2 proposed a closed-form expression; another internal model supplied a proof; and the authors then verified the result. The resulting paper argues that single-minus tree-level gluon amplitudes, often presumed to vanish, are non-vanishing in certain half-collinear configurations (Guevara et al. 2026). There are also recent examples in mathematics (Novikov et al. 2025), biomedicine (Gottweis et al. 2025), and materials science (Zeni et al. 2025). In this paper we argue that we should expect similar success in philosophy.
Specifically, we argue that current-generation LLMs are capable of producing philosophical texts that are _worth reading_. This phrase might seem loose, but that is part of its point. We do not want to begin by settling what counts as _good_ philosophy. Instead, we appeal to a distinction that anyone reading this text will recognise. You have read texts that are worth reading, and you have read texts that are not. As you begin reading this article, you likely hope that it is worth reading, in the sense that the time spent reading it will not be wasted. When you write a philosophical text yourself you aim to make it worth readers' while to read it, and whether or not the journal you send it to accepts it, depends on whether or not they agree.
Two clarifications are needed. First, a text’s being worth reading is not the same as its being correct. A text can repay attention even if one rejects its conclusion: it may sharpen a distinction or answer an objection in a way that changes the dialectical situation. Second, the minimal unit we are concerned with is not the bare conclusion of an argument, but the argument itself. If an LLM output consists only in a pronouncement on some philosophical topic ('Direct Realism is correct', 'We should be utilitarians'), it is hard to see why it would be worth reading in and of itself, for the same reason that a bare pronouncement by a human philosopher would not be worth reading.[^1]
The next three sections develop the main argument. Section I rejects the challenge from authorship: the claim that an LLM output cannot be philosophy worth reading because no philosopher lies behind it. Section II turns to abduction and argues that the absence of human-style inference to the best explanation in the producer does not preclude abductive structure in the product. Section III considers phenomenology and argues that the lack of consciousness does not prevent LLMs from producing philosophy grounded in phenomenology.
## I. The challenge from authorship
In this section we address what we might call the _challenge from authorship_: the idea that philosophy is something that only persons, or at least minds, can do. This view has not, to our knowledge, been explicitly defended in just this form, but it gives shape to an intuition that many philosophers may have: philosophy is a person-only domain. An imperfect comparison is with art. One might deny that an image generated by an AI system, at least in the familiar prompt-and-output cases, is an artwork because no artist exercises the relevant kind of intentional control over its production. One might think, for similar reasons, that philosophy can only be done by people. No text produced by an LLM can be a work of philosophy, because no philosopher lies behind it. One might then think, for parallel reasons, that philosophy too requires a philosopher: no text produced by an LLM can be a work of philosophy, because no philosophical agent lies behind it.
Consider also that, like art, the study of philosophy is often focussed on individuals. Philosophy undergraduates take a course on Kant's ethics, or Lewis' metaphysics, and even at more advanced levels one finds specialists, conferences etc. spotlighting the work of specific philosophers. Compare this to the sciences: as a rule, scientific ideas, theories, discoveries etc. are the focus, not the individuals behind them: one does not find scientists who specialise in the work of Newton, or of Einstein; nor do biology departments teach Crick's view of DNA rather than Watson's. %%is this paragraph accurate re: science?%%
We will try now and make this intuition more precise by continuing the comparison with artworks and philosophical works. We shall do this by considering the degree to which Davies' *performance* theory of art can be transposed to philosophy. He writes:
> The work — what the artist achieves — is the process eventuating in that product. Works themselves are neither structures nor objects simpliciter, nor are they contextualized structures or objects [...]They are, rather, intentionally guided generative performances that eventuate in contextualized structures or objects. (p. 98)
On Davies’ view, when a painter paints a picture, the canvas is what we attend to, but it is not the work. The work is the artist’s intentionally guided activity in producing that canvas; the canvas is, in Davies’ terms, the "focus of our appreciative interest in the work" (2004, p. 151). This is why provenance matters to him in a deeper way than it would matter on a view that identifies the artwork with a product plus contextual properties. Facts about how the object came into being help determine what the work is and what is properly appreciated in it. If the same model were transposed to philosophy, an LLM text would fail not because it is badly argued, but because the relevant kind of philosophical performance is missing.
Consider what is involved in attending to a Vermeer. We are not only registering a perceptual surface%%not how i write%%. We are taking that surface as the outcome of a certain painter’s activity, in a certain historical context, with certain resources and limitations. Davies presses this point through cases in which perceptual sameness, or near-sameness, fails to settle artistic identity or appreciation.%%not how i write%% A canvas might emerge by accident from a washing machine and happen to look like a Rembrandt%%is this a danto example? look it up. if it is it needs to be referenced.%%; in that case, there is a Rembrandt-like surface, but no artistic performance of the relevant kind. Or a canvas might be presented as a Vermeer when it was in fact painted by van Meegeren%%famous forger? if so, you need to mention, even if just in a footnote%%; in that case, there is an artistic performance, but not the one the work was taken to make available. The point is not just that provenance gives us extra information. It is that provenance can change what we take the work to be and what kind of achievement we take ourselves to be appreciating. If Davies is right, the surface does not by itself settle the work.
Here is the analogous proposal for philosophy. A philosophical text is not itself the philosophical work. The text is the product of the thinking, writing, and philosophising done by a person or group of persons over time. The text is therefore the focus of our attention, but only as a way of accessing the philosophical performance that brought it into being.
On this proposal, a philosophical work is not identical with the sequence of sentences on the page. The text is the product of someone’s activity of thinking through a problem and giving that activity argumentative form. Reading the text is then a way of engaging with that activity: not just with a conclusion, but with the route by which the conclusion is reached.%%not how i write%% The authorship challenge is therefore not just a worry about missing biography. It is the stronger claim that, if no one has done the relevant philosophising, there is no philosophical work to which the text gives access.
The question is whether this transposition should be accepted. We do not think it should. Davies has a reason to move from product to performance in the case of art: production history can affect which work we are dealing with and what is available for appreciation. A Rembrandt-like surface produced by accident is not a Rembrandt; a van Meegeren presented as a Vermeer is not the work it is taken to be. The philosophical case is different. If two texts contain the same argument, including the same inferential moves, the same considerations count for and against them. Their philosophical merit does not vary with the route by which the words came to be written.
When we assess a philosophical paper, we ask whether the text does philosophical work. Does it introduce a distinction that helps? Does it answer an objection that would otherwise remain pressing? These questions do not require us to look behind the text to the philosopher’s activity. The grounds for the judgement lie in the argument as presented, not in the history of its production.
This is where the analogy with Davies breaks down. Two papers that read identically do not differ in argumentative merit: they make the same moves and face the same objections. In the art case, production history can change what the work is. In the philosophy case, it changes, at most, what we think about the producer or the process by which the text came about.
The organisation of analytic philosophy reflects this. Journals often strip author information from submissions before sending them to referees, and they do so because facts about authorship are treated as possible sources of distortion. The point is not that blind review always succeeds, or that philosophical practice is never interested in authors. The point is narrower: in this central evaluative context, the paper is supposed to be assessed by attending to what it says, not by reconstructing the circumstances under which it was written.
A point from Dellsén et al. (2024) helps to articulate the same thought, although their concern is philosophical progress rather than LLM authorship. On their view, philosophical progress is "for-whom" rather than "by-whom": it consists in putting people in a position to increase their understanding, usually by making philosophical ideas publicly available (2024, p. 679). For present purposes, the useful thought is that philosophy makes its contribution through public materials that others can take up: arguments, theories, distinctions, thought experiments, and ways of framing problems. If this is correct, we should be cautious about locating the philosophical work behind the public text, in the process by which the text came about. The public text is not a dispensable trace of philosophy; it is where the philosophical contribution becomes available.
The challenge from authorship is therefore a constitutive challenge. It treats the philosopher’s activity not merely as something that causes a philosophical work to exist, but as part of what the work is. On this picture, even a text indiscernible from a philosophical paper would not be philosophy if no philosophical activity lay behind it. We have argued that this should be rejected. If a novel philosophical text were produced by the wind blowing sand into a readable pattern, or by a very faulty washing machine, that would not, in and of itself, prevent the resulting text from being worth reading.
What remain are not objections about what philosophy is, but objections about whether LLMs can produce texts with the relevant philosophical properties. That is,
While the authorship challenge argued that text produced by an LLM cannot be philosophy worth reading in virtue of the fact that it was produced by an LLM, these *causal* challenges, on the other hand, can be thought of as claims that LLMs in their current state cannot produce philosophy worth reading, because LLMs lack certain capacities that are required to write worthwhile philosophy, or worthwhile philosophy in certain areas.
In the next section, we consider the challenge from abduction, %%extremely succinct description of section%%Section 3 turns to the parallel concern that some philosophical texts require phenomenal materials available only to conscious subjects.
## Footnotes
1. reference the ai image literature here, and mention that in most cases it is hard to imagine images being created without a human influencing things at least in some way. [↩](#user-content-fnref-1)
2. Note that such a view does not amount to the denial that LLMs can produce beautiful images. We shall return to this point later. [↩](#user-content-fnref-2)
## II. The challenge from abduction
The next challenge concerns abduction.%%not how i write%% Williamson takes inference to the best explanation to be at least part of the methodology of philosophy as actually practised:%%really abrubt change from the topic sentence%% "I still favor inference to the best explanation and an abductive methodology in philosophy",%%the quote does nothing%% he writes, proposing "that philosophy should use a broadly abductive methodology. Indeed, to some extent it already does so" (Williamson 2016). On this picture a philosophical theory does well, qua potential explanation, when it stands in the right relation to the evidence and is intrinsically good qua theory: "It should be elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated. It should be informative and general. In brief, it should combine simplicity with strength" (Williamson 2016). The section's question is whether a system that does not perform inference to the best explanation can nevertheless produce texts containing good abductive arguments. %%this paragraph supposidly introducing abduction and its necessity for philosophy is appalling and needs to be swapped out for something entirely new%%
Floridi et al. press the challenge from the producer side. They write: %%not how i write%%
> LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations… The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025, p. 9)
The challenge separates the appearance of abduction in the output from abductive process in the producer. The text may have the form an abductive reasoner's output would have, while being generated by a process with no grasp of explanatory force at all. %%deeply unclear. look at my published works and you will see i never write something so insubstantial%%
The challenge is not avoided by appealing to so-called reasoning-mode systems.%%not how i write%% Hidden chain-of-thought tokens are still produced by token completion within the same architecture; Floridi et al. treat such systems as advanced uses of token completion rather than escapes from it (Floridi et al. 2025). The challenge therefore applies to the most recent generation of models, not only to plain completion. %%fucking dreaful paragraph, meaningless, inaccurate to the source, needlessly editorialised%%
Granted%%not how i write%%: an LLM does not understand the philosophical problem as a problem, does not know which candidate is true, and does not weigh explanatory rivals under a norm of truth.%%not how i write%% As a description of the producer, Floridi et al.'s claim is correct. But the thesis defended here is not a thesis about the producer; it is a thesis about the _product_.%%not how i write%% The question is whether texts produced by such a system can contain abductive structure that is good as such, and if so, how.
What inference to the best explanation evaluates is a _potential_ explanation.%%not how i write%% Lipton puts it as follows: "Perhaps actual explanations must be true, but the account has us infer to the best potential explanation, a hypothesis that would explain if true. There is no incoherence here" (Lipton 2004). %%not clear and insubstantial%% Williamson characterises the same distinction:
> A potential explanation of the evidence is anything that would explain the evidence if it were true. A theory T is a better potential explanation of evidence E than a theory T* if and only if T would explain E if T were true better than T* would explain E if T* were true – in brief, T would explain E better than T* would. (Williamson 2016)
The point for present purposes is that a potential explanation is something a text can present. The data, the candidate, the comparison with rivals, what the candidate would explain if true %%not how i write fucking fucking lists%%— these are features prose can carry, and they are what is assessed when abductive work is assessed.
Lipton distinguishes the _likeliest_ explanation from the _loveliest_: likeliness concerns the explanation's probability of truth, loveliness the understanding it would provide if true. The two are not separable into truth-talk and ornament-talk:
> we have a kind of feedback between judgments of likeliness and judgments of loveliness. Successful inferences become part of the background, and influence what counts as a lovely explanation and thus influence future inferences. This is as it should be: as we learn more about the world, we not only know more but we also become better inferential instruments. (Lipton 2004)
This blocks the deflationary reading of the explanatory virtues, on which they would be stylistic%%what the actual fuck are you talking about%%. They are evidential.%%not how i write%% Their role in inference to the best explanation is to track likeliness, not to decorate prose. A text exhibiting them at the level of the product is therefore not exhibiting decoration; it is bearing the kind of feature that does abductive work.
The relevant philosophical virtues, in Williamson's compact statement, are that a theory be "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated", and "informative and general", combining "simplicity with strength" (Williamson 2016). %%why is the reader being told this stuff?%%These are features of theories qua presented.%%not how i write%% A philosophical theory is more or less ad hoc as set out in the text; more or less unified in the relations it draws among its claims; more or less informative as it bears on the evidence.%%not how i write%% The reader assesses them by reading.
Lipton's account of inference to the best explanation does not treat candidate-generation as free invention. The hypotheses on offer at any given moment are produced against a structured background:
> the mechanism of explanatory selection plays a role both in the generation of the short list of plausible causal candidates and in the selection from this list. The background beliefs that help to generate the list are themselves the result of explanatory inferences whose function it was to explain different evidence… So Inference to the Best Explanation helps to account for the generation of live candidates because it helps to account for the earlier inferences that guide this process. (Lipton 2004)
That generation, moreover, "favors those that are extensions of explanations already accepted, and so leads towards a unified general explanatory scheme" (Lipton 2004). The point is that candidate-generation operates against a background that is itself the deposit of earlier inferences. This is part of how IBE works in human inquiry.
The philosophical corpus is precisely such a background. Published philosophical writing is not a neutral heap of sentences. It is the residue of philosophical labour, the arguments, objections, replies, distinctions, and refinements that survived philosophical resistance. What survives is not random. It survives because it did relatively well by abductive standards as those standards apply within philosophy: it was less ad hoc than rivals, more unified, more informative. Subsequent philosophical work was generated against the background formed by what survived. The philosophical corpus thus stands to philosophical inquiry as Lipton's structured background stands to inquiry generally — past explanatory selection deposited in textual form.
The mechanism by which next-token prediction expresses this load is ordinary. The model generates text by sampling, one token at a time, from a conditional distribution over what comes next given context. This is its only mechanism; we are not claiming that training on philosophical writing equips it with some additional, non-token-prediction capacity. What training does is fit the conditional distributions themselves. When the training corpus is philosophical, the distributions the model has learned are distributions over continuations of philosophical context, shaped by what tended to come next in surviving philosophical writing. What tended to come next there is not arbitrary. It is the kind of move that survived objection, maintained unity, avoided ad-hocness, and contributed informatively — the same features Lipton and Williamson identify as abductive virtues. The local distribution the model samples from is in this sense abductively loaded at every step. The cumulative effect of locally loaded sampling, over the course of a sustained argument, is a trajectory through textual space that tends to exhibit those virtues at the level of argumentative structure. The system does not select for abductive virtue; it samples from distributions that do. (This is the picture the framework of semiotic physics, developed elsewhere [reference], works out in more detail: the corpus shapes a time-evolution operator that propagates a textual trajectory forward, and that operator is locally biased, at every step, by the regularities of what has survived in philosophical writing.)
Floridi et al.'s own concession licences this reading. They write that LLMs have "absorbed patterns of human abductive reasoning as expressed in writing" (Floridi et al. 2025, p. 9). The phrase should be neither inflated to understanding nor deflated to mere surface mimicry. The patterns absorbed are the patterns that survived past philosophical selection, and they are expressed via the conditional distributions that govern next-token sampling. The capacity to produce texts containing good abductive arguments is therefore the capacity of next-token prediction, when the underlying distribution is philosophical, to produce trajectories bearing the marks of the abductive labour deposited in the corpus.
LLMs do not perform abductive inference. The producer-side claim from Floridi et al. stands. But the thesis defended here is the product-side claim, and it stands too: in virtue of training on the philosophical corpus — itself the residue of past philosophical selection — generation via next-token prediction yields trajectories shaped by that selection. The relevant capacity is therefore present, and the question whether a given output is philosophy worth reading is the question we always ask of a philosophical text: whether it presents a potential explanation that exhibits the philosophical virtues, against the live alternatives, with the contrastive structure the relevant question demands. The next section turns to a different capacity worry — the one concerning phenomenology.
## III. The challenge from phenomenology
### Experiential Axioms
A further capacity worry concerns phenomenology. Few would say that LLMs are conscious, and we will assume the same here; yet this might seem to pose a problem for LLM philosophy, or at least for philosophy grounded in, or making use of, phenomenology. Some philosophy interrogates or refers to what it is like to see red (Harman, 1990), to feel anger (Goldie, 2000), or to have a particular intuition take hold (Chudnoff, 2011). If LLMs lack conscious experience, it seems as if this might hamper their ability to produce worthwhile philosophy which relies on it. This is not to say that all philosophy would be off bounds: large stretches of philosophy of language and modal metaphysics proceed without leaning on the phenomenology of any particular experience.
Zahavy’s discussion of a thought experiment of Einstein's brings out this worry:
> Einstein’s variation required inventing new axioms based on a physical intuition that did not yet exist in the mathematics. He envisioned a physicist inside an elevator being uniformly accelerated through deep space. Inside this enclosure, the sensory experience reveals a specific pattern: when objects are released, the floor rushes up to meet them. To the physicist, the objects appear to fall with identical acceleration, regardless of composition. Thus, the simulation here was not a permutation of symbols, but a manipulation of perceptual experience. (Zahavy 2026, §5)
The thinker imagines[^4] some set of circumstances and attends to what would be experienced within it — in Einstein’s case, that all objects inside the elevator would appear to fall with identical acceleration. That observation becomes the new axiom: a starting point arrived at through experiential simulation rather than formal derivation, from which further reasoning proceeds. If thinking of this kind depends on simulated experience, then it would seem to be out of reach for LLMs. They can provide descriptions of weightlessness or elevators, but they have never felt the sensation of an elevator descending, let alone weightlessness.[^2]
Philosophy also uses experience based thought experiments. Jackson’s Mary case turns on what it is like to see colour, and we might think that as with Einstein's thought experiment, it provides us with an experiential axiom, from which further philosophical reasoning can proceed. The same worry then arises in philosophy: experience based thought experiments seem to require what LLMs do not have.[^3]
### Articulated Phenomenology
Pigliucci offers an account of philosophy on which it is constrained by, but does not aim at, the world as the natural sciences do. He writes:
> This means that the basic parameters that philosophers use as their inputs, the starting points of their philosophizing, their equivalent of axioms in mathematics and assumptions in logic (or rules in chess) are empirical data about the world. This data comes from both everyday experience [...] and of course increasingly from the world of science itself. (Pigliucci, p. 6)
Philosophy begins from worldly materials, but those materials function as starting points for conceptual exploration. They are not used in the same way a physical datum is used to confirm or disconfirm an empirical theory.%%one more sentence, a good one, will make this paragraph substantial%%
Pigliucci elaborates this picture by drawing on Smolin’s account of evocation, taking chess as the paradigm. Positing the rules of a game does not require that they pre-exist; once posited, they generate a structure with rigid properties — a space of consequences that can be explored but not chosen. Once the rules of chess are codified, all the facts about chess become demonstrable, even though chess did not exist before its rules were written down. Pigliucci’s claim is that philosophy operates in this register. Unlike the rules of chess or the axioms of mathematics, however, the starting points of philosophising are constrained empirically. They are constrained by how the world actually is, including by what experience is like.
This is why Pigliucci distinguishes philosophy from fiction: philosophy is not merely the invention of imaginary possibilities. As he puts it:
> Philosophy, I maintain, is in the business of doing empirically informed evoking, not inventing. (Pigliucci, p. 7)
The same picture covers philosophical thought experiments. Even when philosophers explore possible worlds or imagined scenarios, they do so “with an interest in figuring things out as far as this world is concerned” (Pigliucci, p. 7). The thought experiment articulates an axiom — an experiential or empirical starting point — and the philosophical work proceeds within the conceptual landscape that axiom evokes.
This brings out a difference between the elevator and Mary cases. Both are evocations of the kind Pigliucci describes: each posits an experiential axiom and develops what follows from it. What differs is what the evocation is for. In Einstein’s case, the evoked structure yields a hypothesis whose status is then settled by experiment — the elevator gave him the equivalence principle, but the principle’s truth was a matter for empirical confirmation. In Mary’s case, the evoked landscape is itself the object of inquiry; the philosophical question is what the landscape contains, not whether anything outside it corresponds. The role of the evocation, not its presence, is what tracks the disciplinary difference. Evocation is present in both cases; what differs is whether the evoked structure is the means to an external test or is itself the object of inquiry.
No competent discussant of the knowledge argument has personally undergone her transition. Once the case is articulated, work on it is work on the articulation. Responses to Jackson press at the level of the articulated structure, not at the level of any discussant’s experience. Lewis’s reply, for instance, modifies what is taken to follow from Mary’s situation, not what Mary’s situation is taken to be like from the inside.
What allows the Mary case to do philosophical work in public is its articulation: the experiential material it draws on has been made available in language. This is the form in which phenomenology enters philosophy generally. The articulation is what does the philosophical work; the experience the articulation refers to need not be undergone by the people working on it. Philosophers work on the experiences of the blind and on the experiences of non-human animals without first-hand access to either, by working on the articulations the literature has accumulated. The point matters for LLMs in a particular way. They have no raw phenomenology of their own; but no text corpus contains raw phenomenology either. What a corpus contains is articulated phenomenology, and it is in articulated form that phenomenology becomes usable in philosophical argument.
Merleau-Ponty’s discussion of self-touch raises a sharper question — that of phenomenological _discovery_. Suppose the toucher-touched asymmetry was first identified by Merleau-Ponty himself, by sustained attention to his own embodied experience. The asymmetry would then be a phenomenological axiom out of reach of any LLM not trained on Merleau-Ponty or his interlocutors: an axiom an LLM could not have produced for itself, because the system lacks the body and the experience that the discovery requires. When one fingertip touches another, one finger plays the role of toucher and the other of touched. The roles can reverse, but not simultaneously: at any given instant, the body is split between touching and touched. But that does not prevent an LLM from working philosophically on the description once articulated.
What survives, then, is a narrower asymmetry. Even granting that LLMs can work within articulated landscapes, some phenomenological articulations seem to be originated through first-person attention; LLMs have no experience to attend to. First-person attention is one route to an articulation; it is not what gives an articulation philosophical use. What makes an articulation philosophically usable, on Pigliucci’s picture, is not its causal origin but its functioning as an axiom — its capacity to evoke a landscape with rigid properties. An articulation can also be arrived at by working from the articulations a corpus already contains, generating new ones by extension and recombination. Whether a candidate articulation succeeds is a question about what it evokes, and that question is answered the way other philosophical questions are — by the public assessment of the conceptual structure the articulation makes available. It is the assessment any candidate articulation, whatever its origin, must finally meet.
The phenomenology objection rests on a producer-to-product inference: that the absence of experience in the producer must remove phenomenological value from the product. The inference fails. LLMs lack conscious experience, but phenomenology enters philosophy as articulated content. Pigliucci’s account explains why this is not a workaround. Philosophy uses empirical and experiential materials by turning them into constrained spaces for conceptual exploration. Since those spaces are public and inferentially usable once articulated, current models can produce phenomenology-based philosophy worth reading.
[^2]: footnote saying that he calls it manipulative abduction. it should probably also explain why we might think go this as abduction as well as what we talked about in the previous section
[^3]: A nice example in the footnote will be the feeling of understanding that is sometimes used as a way of motivating cognitive phenomenology.
## IV. Authorship redux: the challenge from elicitation
If current LLMs can produce philosophy worth reading, why are we not surrounded by great LLM philosophical texts? The ordinary experience of using these systems seems to support scepticism. Asked for philosophy, they often produce competent but lifeless exposition—paragraphs that tour a topic without ever applying pressure to it. A vague topic prompt does not ask for a philosophical intervention. It asks for the most probable kind of text under that topic-label, and in the training distribution that is often hedged survey, because the bulk of philosophical text written at that level of generality takes that form. The prompt is generic, and so is the region of the model’s learned space it activates.
If interesting LLM philosophy appears only under careful prompting, perhaps the LLM is not really producing the philosophy after all. Perhaps the human prompter is producing philosophy by using the LLM. This version of the worry turns on control. The final text may be worth reading, but if the human fixes the task and chooses the result, the value can seem to belong to the human-guided process rather than to the model’s production. So “produced by an LLM” has to mean more than “appearing in an LLM output window”. A model may output philosophy worth reading without producing the features that make it worth reading.
Take the grammar-correction case. A philosopher writes a brilliant argument and asks an LLM only to correct its punctuation. The LLM’s response may contain philosophy worth reading, but the model has not produced the argument in virtue of which the text is worth reading. It has output worthwhile philosophy without producing what makes it worth reading. So there is a continuum of contribution. At one end, the model merely polishes a human-produced argument; at the other, the human specifies a task and the model generates the philosophical move that makes the output worth reading. The cases that support the present thesis lie towards the latter end. Where the human supplies the argument and the model improves the prose, the LLM has not produced philosophy worth reading. Where the human specifies a problem and the model supplies the objection or distinction that makes the text worth reading, the output is LLM-produced in the sense that matters here.
The elicitation objection assumes too simple a contrast between autonomous producer and mere tool. Elsewhere I have argued, with Terrone, that generative AI systems of the Midjourney type are best understood neither as agents nor as ordinary tools, but as generative systems with which users interact under conditions of partial control. LLMs are not intentional agents, but the ordinary tool model is also too crude. Prompting is not command execution. The prompt is not a blueprint that fixes the product in advance. It sets conditions under which the model generates. The user can constrain and iterate, but cannot determine every relevant feature of what emerges.
Prompting is therefore a form of elicitation. A call for papers elicits philosophical answers without authoring them; an interlocutor in conversation elicits arguments from another philosopher without coming to be their author. So a prompt’s having elicited an LLM output does not by itself show that the prompter has supplied the philosophical content that makes the output worth reading. Selection is not generation either. A journal selects the papers it publishes but does not thereby produce them, and a reader’s selection of a good LLM output is, similarly, an act of assessment rather than of production.
If philosophical corpora encode abductive and dialectical structure in the way I argued earlier, skilled prompting should aim to elicit those structures rather than request prose about a topic. Pigliucci’s account gives this a further formulation. A good philosophical prompt evokes a constrained conceptual space: it makes a problem determinate by placing a contrast under dialectical pressure, but leaves the philosophical move for the model to make. The point is not to insert phrases such as “one might object,” as though they were magic words. Such phrases matter only because they mark argumentative roles. A good prompt specifies the role to be filled.
Contrastive prompting asks why one view handles a particular case better than its rival, rather than asking for a discussion of a topic in the abstract. Loveliness-sensitive prompting asks not for a conclusion but for a view that, if true, would explain more than its rival. Both target the explanatory virtues that make a philosophical answer worth reading. LLM philosophy improves when the prompt creates dialectical pressure: generic prompts invite generic continuations, and a philosophical prompt should create a space in which some argumentative move is needed, and then leave that move for the model to make.
The creativity worry can be handled within the same continuum. Grammar correction is not interestingly creative. Generating a new objection or a new distinction may be. This fits a product-centred approach to creativity. Many accounts require novelty and value, and the present argument need not show that LLMs are creative agents in the fullest sense. It is enough that their outputs can contain novel and valuable philosophical structure. Nor does elicitation defeat creativity. Creative work often happens under constraints; a prompt can set a conceptual space without determining what is found within it. Elicited LLM philosophy can be creative at the level relevant to worth-readingness.
The same point bears on the scarcity of good generic LLM texts. The corpus an LLM is trained on is the product of many rounds of philosophical criticism. A single completion does not reproduce that history. But a prompt can ask the model to stage, within the generated text, some of the operations that the corpus records across time: an objection pressed, a reply attempted, a distinction introduced under pressure. The human prompt supplies local conditions of elicitation; it need not supply the philosophical move. When the model generates that move, the output is LLM-produced in the sense relevant here.
The absence of many great generic LLM texts is not the verdict it can appear to be. It reflects, at least in part, immature elicitation practices and the prevalence of generic prompting. If philosophy worth reading requires live alternatives and dialectical pressure, vague prompts will rarely elicit it. Elicitation does not defeat the claim that LLMs can produce philosophy worth reading: it shows that such production comes in degrees and depends on task conditions. The question to keep in view is whether the model has generated the philosophical structure in virtue of which the output is worth reading. The answer can be yes.
LLMs do not produce worthwhile philosophy merely by being asked for “some philosophy”, and not every LLM-assisted text counts as LLM-produced in the relevant sense. But current models can generate philosophical moves that make a text worth reading, and where they do, the philosophical value is present in the product, and the product was produced by the LLM in the sense that matters here.
## Conclusion
The four challenges treat what an LLM lacks as decisive for what its outputs can be. None of the absences they begin from is trivial. I have not denied those absences; I have argued that they do not, either singly or together, fix the philosophical status of the generated text. None of this settles whether what appears on the page is philosophy worth reading.
A text repays philosophical attention when it changes what can be assessed in a dialectical context. Where an LLM output does that, the output is philosophy worth reading. What comes next is the same question we ask of any philosophical text: does the argument hold?
## Notes
1. The duplicate case is artificial in practice, but the artificiality is doing controlled work. By holding the textual product fixed, the case isolates the question whether causal history alters inferential structure. It does not.
2. The claim is not that every model trained on any philosophy-adjacent corpus will produce outputs with abductive structure. It is that the process described here is a plausible route by which current frontier models can do so.
3. The term “manipulative abduction” originates with Magnani, who introduced it for cases of hypothesis generation through the active construction of mental models. Zahavy uses it in this sense and applies it to Einstein’s elevator argument as a paradigm case in physics.
4. Iteration complicates attribution, but it does not automatically transfer production to the human. What matters is still whether the human supplies the philosophical move or elicits it.
5. Compare “Discuss whether physicalism is true” with a prompt asking why one version of physicalism can answer a specific objection that another cannot. The first asks for a survey. The second specifies a dialectical role. What matters is not the wording but the philosophical task imposed.
[^1]: Consider the following example from the British science fiction novel: *The Hitchhiker’s Guide to the Galaxy*. At one point, humanity builds a computer and asks it for "the answer to the ultimate question of life, the universe, and everything"; after seven and a half million years, the computer returns ‘42’ — correct, apparently, but meaningless, given that mankind neglected to ask what the ultimate question is, they have no context in which to understand the answer. We will return to the question of questioning AI in Section 4, but for now we assume that worthwhile philosophy takes place at the level of the argument.
[^2]: footnote saying that he calls it manipulative abduction. it should probably also explain why we might think go this as abduction as well as what we talked about in the previous section
[^3]: A nice example in the footnote will be the feeling of understanding that is sometimes used as a way of motivating cognitive phenomenology.
[^4]: mention that we assume this is some sort of mental imagery. a conscious experience in itself and derived, at least partly, from perceptual experience
### Assistant
Let me verify Lipton's actual contrastive example before drafting that paragraph.
### Assistant
Let me check if Lipton uses the eggplant example or if I should use a different illustration.
### Assistant
I have enough context now. Let me explain my analysis before producing the new draft.
Let me work through what I've understood and what I'm going to do.
The user wants a new draft of section II of the generating philosophy paper. The plan I am implementing comes from the file `Clippings/nick - Branch · Draft Review Feedback.md`, which contains a "Section II Challenge From Abduction New Iteration" — a complete reformulation of section II produced in a ChatGPT branch. The phrase "the plan at the very very end" refers to this output, which sits at the end of the branch conversation. The branch was branched from an earlier conversation (`nick - Inference to the Best Explanation 2.md`) where Nick discussed in depth how to redeploy Lipton's apparatus to defend LLM-generated philosophy.
I read the new iteration in full. It does several things differently from the existing draft:
1. It develops Lipton's apparatus far more thoroughly (potential explanations, two-filter view, likeliness/loveliness, virtues, contrastive explanation).
2. It uses Bengson, Cuneo, and Shafer-Landau's Tri-Level Method to translate Lipton's framework into philosophy-specific terms (accommodate/explain, substantiate/integrate, objections).
3. It centres the producer/product distinction as the strategic move — granting Floridi's producer-side claim entirely while showing why the product-side criteria are untouched.
4. It does NOT lean on the "next-token prediction over a virtue-filtered corpus" mechanism that the current draft uses. The reply is product-centred, not corpus-centred.
Comparing to the existing section II in the latest draft Nick provided: the existing draft is heavily flagged with `%%not how i write%%` comments, and its strategy is the corpus-filtering thesis. The new iteration's strategy is the product-side reply via Lipton + BCSL. The existing draft also has several voice-flagged moves: an abrupt Williamson opening, a "Granted" construction, list-stacking, meta-commentary like "This blocks the deflationary reading", etc.
Critical voice features I am holding to (drawn from `nick-analytic-voice` and from reading full paragraphs of *Hearing Spaces*, *Growing the Image*, and *Agents of Change*):
- Use "we" throughout (paper is co-authored with Enrico Terrone — flagged as a hard rule in the session file).
- British English, no contractions.
- No announcement phrases ("crucially", "significantly", "importantly").
- No meta-commentary that comments on what the argument or paragraph is doing ("this dissolves", "this establishes", "this demonstrates").
- No triplet flourishes — no "X, Y, and Z" as a sentence-ending list of illustrative examples.
- No banned words: "central", "key", "crucial", "essential", "critical", "fundamental", "primary"; no "textual"/"textually"; no "matter" as a verb; no "pivot".
- Sentences alternate between longer thinking sentences and shorter verdicts. No sequences of short punchy declaratives.
- Interlocutors entered through their actual claims, not via literature-review labels.
- Examples developed properly when they appear; if an example is not going to do work, do not name it.
- Direct claims, not hedged manoeuvres ("it is far from obvious that").
- Engage objections in their strongest form.
What grounds my voice writing concretely:
- *Hearing Spaces* shows Nick's "candidate views laid out then tested" structure: "At first, it might seem that there is an obvious way to support Nudds's position…", "Two things count against this approach. First… Second…". I will use that pattern when laying out alternatives.
- *Growing the Image* (with Enrico) shows the co-authored "we" register: direct interlocutor entry ("As Hertzmann puts it", "However, as Anscomb (2022, p. 25) points out…"), example development ("To see why autonomy is not sufficient for attribution of credit, consider the following example…"), and the use of "There is an obvious alternative to thinking of Midjourney as an agent" as a section opener. That last move is exactly the topic-sentence move I want for the producer/product distinction.
- *Agents of Change* shows clear methodological openings ("My aim here is to show how experience leads to the belief that time passes") and concessive turns ("While I find this line of thought persuasive, I want to argue here that even if we grant…"). I will use that concessive-grant move when granting Floridi's producer-side claim.
I have verified the following quotations against extracted source text:
- Floridi et al. 2025, p. 9 — "LLMs seem to perform a kind of zeroth-order abduction…" (the block quote already in Nick's draft, verified against the extraction).
- Floridi et al. 2025, p. 9 — "absorbed patterns of human abductive reasoning as expressed in writing" (verified).
- Lipton 2004, p. 58 — "start with a souffle" (verified — actually a souffle, not "souffle"; and the phrase is "start with a souffle"; the wider sentence is "Telling someone to infer actual explanations is like a dessert recipe that says start with a souffle.").
- Lipton 2004, p. 59 — "Likeliness speaks of truth; loveliness of potential understanding" (verified).
- Lipton 2004, p. 59 — "live options" used, two-filter structure described (verified).
- Lipton 2004, pp. 59-60 — Newtonian mechanics example (verified).
- Lipton 2004, p. 122 — virtues list: "mechanism, precision, scope, simplicity, fertility or fruitfulness, and fit with background belief" (verified).
- Lipton 2004, ch. 3 — eggplant/sea bass contrastive case is genuinely Lipton's (verified — "If you ask me why I ordered eggplant rather than sea bass (a 'daily special'), I may give the perfectly good answer that I did not know there were any specials").
- Bengson, Cuneo, and Shafer-Landau 2022, p. 108 — Tri-Level Method definition (verified).
- BCSL 2022, p. 110 — accommodation definition (verified).
- BCSL 2022, p. 116 — substantiate language (verified).
- BCSL 2022, pp. 118–119 — integrate definition (verified).
- BCSL 2022, p. 134 — objection definition (verified).
- Williamson 2016 quotations are already in Nick's draft, presumed verified.
Decisions I made about scope and structure:
1. The challenge has to come early. Nick said in the parent conversation "It has to be a response to the challenge as it is presented at the beginning of section two." So I open with a brief setup of abduction's role in philosophy (Williamson) and present Floridi's challenge by paragraph 2. The rest of the section is the reply.
2. The producer/product distinction is the strategic axis. This was the move Nick wanted in the parent conversation when he pushed back against the "modest split between humans and LLMs" framing. I make the distinction the explicit strategic move and let everything else hang on it.
3. Lipton's apparatus is developed as the *specification* of what product-side abductive structure consists in. I do not pretend that LLMs perform Liptonian inference; I use Lipton to say what features in a text count as abductive at the product level.
4. BCSL's Tri-Level Method does the same job, but in philosophy-specific terms. I treat it as the philosophy-specific translation of Lipton's framework.
5. I leave out the corpus-filtering thesis from the existing draft. The new iteration does not need it. The reply is product-centred: a reader can assess these features in the text directly. (I am following Nick's instruction to implement the new iteration's plan, not the existing draft.)
6. I include a mimicry-versus-success contrast at the end, because that is how the new iteration closes and it acknowledges the live worry about empty intelligibility.
7. Section length: I aim for 14–16 paragraphs. The new iteration has 26, which is too long; the current draft has 13, which is too few given the new content. I consolidate where the new iteration repeats itself.
8. I keep the section title "II. The challenge from abduction" rather than the working file's title "Likeliness, Loveliness, LLMs", because (a) it parallels §I's "challenge from authorship" and §III's "challenge from phenomenology", which Nick is using as a unifying frame, and (b) the section now does more than the likeliness/loveliness exposition — it covers Lipton's full apparatus plus BCSL.
9. I attribute Lipton's eggplant case to him (Lipton 2004, ch. 3), since it really is his.
10. I do not include the parenthetical reference to "semiotic physics" from the existing draft — the new iteration goes in a different direction and that machinery is no longer load-bearing.
11. I keep Williamson's quotations from the existing draft because they are already vetted by Nick.
One point I want to flag explicitly: the new iteration does include some illustrative lists ("arguments, objections, replies, distinctions, failed proposals, refined proposals" and "elegance, coherence, illumination" type stretches). I have to police these. The no-example-lists skill says: develop one, or state the general claim, or name the contrast, but do not stack examples as a flourish. I will allow lists where they are exhaustive distinctions doing real work (like the BCSL criteria themselves, which are a closed enumeration), but cut them where they are decorative.
Let me also flag a concern: the new iteration ends with a paragraph that says "The conclusion is limited but not trivial." This is a meta-summary paragraph and risks anti-metacommentary territory. I will end the section with content rather than a recapitulation: a final claim about what the product-side reply allows and what it does not.
Now I will write the new draft. I will keep the draft self-contained — it can replace section II in Nick's manuscript without further edits. After the draft I will note one or two flags for things Nick may want to verify or push on.
# Final Answer
Here is the new draft of section II. It implements the plan from the Branch new iteration: the producer/product distinction as the strategic move, Lipton's apparatus as the specification of product-side abductive structure, and Bengson, Cuneo, and Shafer-Landau as the philosophy-specific translation. The corpus-filtering thesis from the current draft is set aside. Voice is held to nick-analytic-voice; "we" used throughout; quotations verified against extractions.
---
## II. The challenge from abduction
We begin with the methodological picture that gives the challenge its grip. Williamson takes contemporary philosophy to proceed by abduction from the armchair. On this picture, philosophical theorising is not exhausted by conceptual analysis or appeal to intuition. It involves comparing candidate accounts of the data and preferring one because it would, if true, explain the data better than its rivals. Williamson writes that "philosophy should use a broadly abductive methodology. Indeed, to some extent it already does so" (Williamson 2016). The standards by which a theory is preferred carry over from the sciences: a theory should be "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated", "informative and general", combining "simplicity with strength" (ibid.). A philosophical text earns the reader's attention when, judged by these standards, it improves the available account of some phenomenon.
Floridi and colleagues argue that the work involved in producing such a text — the comparison of candidates, the weighing of explanatory virtues — is precisely the work that LLMs cannot do. They write:
> LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. (Floridi et al. 2025, p. 9)
We should be careful about how this passage is read. Floridi et al. allow that a continuation can have the form of a hypothesis or explanation. What they deny is that the form is generated by abductive inference. The model does not ask what would explain the evidence; it samples from a learned distribution. Call this the *producer-side* claim: the process by which the text is produced is not abductive reasoning, however much its surface resembles the surface of abductive reasoning.
Even if we grant the producer-side claim in full, our question is not yet settled. The standards Williamson invokes apply to philosophical theories, not to the cognitive episodes that may have produced them. If the standards are met, they are met by the text; if they fail, they fail in the text. Our question is therefore whether a text generated by an LLM can present a theory that handles the data, supports its claims, and exhibits the virtues we noted; not whether the model has reasoned its way to that theory. The producer-side and the product-side are different questions, and the philosophical evaluation is directed at the latter.
Lipton's account of inference to the best explanation makes this distinction precise. He stresses that what is assessed in such an inference is not an actual explanation but a *potential* explanation: a candidate that would explain the evidence if it were true. The point is partly methodological. If we had to identify the actual explanation before inferring it, we would already have arrived where the inference was supposed to take us; the model would, as Lipton puts it, be like a recipe that tells us to "start with a souffle" (Lipton 2004, p. 58). The notion we need is therefore weaker. An assessable candidate is one that, if it were correct, would explain what we want explained.
The relevant pool of candidates, on Lipton's preferred construal, is not the whole space of logical possibilities. Inquiry runs over what he calls "live options": the serious candidates already in play in the relevant context (Lipton 2004, p. 59). On a two-filter version of the view, a first filter restricts the pool to such candidates, and a second selects from among them. This will be familiar to philosophers. We seldom argue that our preferred view is better than every conceivable rival, including ones nobody has considered; we argue that it is better than the live alternatives in the relevant debate.
Lipton further distinguishes the *likeliest* potential explanation from the *loveliest*. The likeliest is the one most probable given the evidence; the loveliest is the one which would, if true, provide the most understanding. Lipton's compact formulation is that "Likeliness speaks of truth; loveliness of potential understanding" (Lipton 2004, p. 59). The two come apart. A conspiracy theory may be lovely in one respect, since it would unify many disparate events if true, while still being very unlikely. Newtonian mechanics is one of the loveliest explanations in science, even though, after relativistic evidence came in, it became less likely; it remains as lovely an explanation of the older data as it ever was (ibid., pp. 59–60).
If "best" simply meant "likeliest", then inference to the best explanation would say only that we infer whichever explanation we judge most probable. Lipton's stronger thesis is that loveliness is a guide to likeliness: judgements about which candidate would, if true, provide the most understanding feed into our judgements of which is most likely to be true (Lipton 2004, pp. 60–62). On this picture, the features that make an explanation lovely are evidential, not stylistic. They do not decorate the inference; they help to drive it.
Among the explanatory virtues Lipton discusses are "mechanism, precision, scope, simplicity, fertility or fruitfulness, and fit with background belief" (Lipton 2004, p. 122). An explanation with greater scope explains more; a more precise explanation discriminates the phenomenon from nearby alternatives more sharply; a fertile one opens further possibilities of explanation; an explanation that fits with background belief is less isolated from what we already have reason to accept. In scientific cases, mechanism often means a causal mechanism. In philosophy, the analogue will normally be structural — a distinction, a dependence relation, an inferential pattern — that shows why the relevant data hang together. These are features of theories as set out, not of how their authors arrived at them.
Many explanations answer not "Why P?" but "Why P rather than Q?", and the foil determines what is to be explained. To use Lipton's own example: if you ask why someone ordered eggplant rather than sea bass, the answer that they did not realise sea bass was on the menu is a perfectly good explanation of the contrast, even if it does no work as an explanation of why they ordered eggplant simpliciter (Lipton 2004, ch. 3). A great deal of philosophical abduction has this contrastive form. We argue not merely that some distinction can be drawn, but that *this* distinction rather than its rival is the one we need; not merely that some account has a consequence, but that the consequence holds rather than its rival.
This sharpens what we are looking for in a generated philosophical text. A text contains abductive structure in the relevant sense not when it contains the words "explanation" or "best account", but when it identifies what is to be explained, formulates a candidate that would, if true, explain it, places that candidate against the live alternatives, fixes the relevant contrast, and bears the marks — mechanism, scope, simplicity, fit with background — that made the comparison work in the first place. Whether or not the model arrived at any of this by reasoning is a different question.
Bengson, Cuneo, and Shafer-Landau translate this picture into terms tailored to philosophy. On their Tri-Level Method, a theory of a given domain "ought to articulate a set of theses about the domain that (i) accommodate and explain the data, (ii) are themselves substantiated and integrated, and (iii) possess specific theoretical virtues" (Bengson, Cuneo, and Shafer-Landau 2022, p. 108). A theory accommodates a datum when, given the theory, that datum "is likely to hold or be true" (ibid., p. 110); it explains a datum when it gives an account of why the datum holds. To substantiate a claim is to defend and explain it (ibid., p. 116); to integrate a theory is to show that its claims cohere with one another and with our best picture of the world (ibid., pp. 118–19). An objection, on their account, is a consideration that gives reason to think a theory does poorly with respect to one or more of these criteria (ibid., p. 134). These are properties of philosophical work as set out in writing. They are inspectable by readers who want to know whether the work is worth their attention.
The reply to Floridi et al. can now be put precisely. Their producer-side claim, granted in full, leaves the product-side criteria untouched. Whether a generated text formulates a potential explanation, organises live alternatives, fixes the contrast, exhibits explanatory virtues, accommodates and explains the data, substantiates its claims, and integrates them with relevant background — these are properties of the text. A reader can ask whether the explanandum has been correctly identified, whether the alternatives have been fairly represented, whether the candidate would, if true, provide understanding, and whether it does so better than its rivals. None of these questions requires us to know how the text was generated.
Floridi et al. themselves leave room for this reply. They write that LLMs have "absorbed patterns of human abductive reasoning as expressed in writing" (Floridi et al. 2025, p. 9). The phrase should be neither inflated nor deflated. It should not be inflated to the claim that LLMs understand those patterns: nothing in their architecture supports that, and nothing in our argument requires it. But it should also not be deflated to the claim that what LLMs have absorbed are mere verbal templates. Philosophy is conducted, preserved, and criticised in writing. The corpus that an LLM trained on philosophy is exposed to includes arguments, the objections that bear on them, the replies that succeed or fail, and the refinements that survive. To learn the statistical structure of that body of writing is to learn structures that have been shaped by the very evaluative work whose absence Floridi et al. emphasise.
What this allows is also limited. A text generated by an LLM can use the words "the best explanation is" without identifying a genuine explanatory virtue. It can stage an objection without locating any failure of accommodation, explanation, substantiation, or integration. It can introduce a distinction that leaves the same data untouched. These are mimicry cases, in which the public form of philosophy is preserved while the philosophical work goes undone. They are also cases the reader can identify by reading. Our claim is not that LLM outputs are guaranteed to be philosophy worth reading. It is that the question whether they are is settled in the same way it is settled for any text — by attending to what the text contains.
This leaves a recognisable picture. LLMs do not perform inference to the best explanation in the way human philosophers do. They do not understand explanations as explanations, and they do not infer one candidate over another because they judge it lovelier or likelier. What follows from this absence in the producer is exactly what follows from any negative claim about producers: nothing about the product, taken as a product. For the question whether a generated text is philosophy worth reading, the question is not whether the model reasoned abductively, but whether the text presents an abductive structure that is good as such. In some cases, it does. The next section takes up a different capacity worry: the one concerning phenomenology.
---
## Notes for Nick
Three flags worth eyeballing:
1. Williamson 2016 page numbers: I have used the same citations Nick already had in the draft (no page) for the three Williamson quotations. Nick may want page numbers added.
2. Lipton "ch. 3" citation for the eggplant case: I have given a chapter rather than a page because the case appears in Ch. 3 (around p. 33 and again around p. 49 in the same chapter). Nick can swap to a page reference if preferred.
3. The footnote on reasoning-mode systems (the chain-of-thought paragraph from the existing draft) is not in this version. It was flagged as inaccurate to Floridi and editorialised in the existing draft. If Nick wants it preserved, it could go into a footnote rather than the body — the new draft as it stands does not need it to make the producer-side concession.
*L'abduzione, in filosofia, vive nella pagina prima ancora che nella mente: il lettore valuta ciò che il testo offre, non il percorso interiore di chi l'ha scritto.*