2e2c28a1-960c-43cc-a009-d139cba15ae8.md File 0443d0c0-d8cd-40a8-8141-a88da3799801.md File here is a book and a conversation about that, a particular aspect of that book. Can you summarize what the conversation is about please? Please give me a long, deep, and detailed answer, with lots of options/ideas: as we are in the 'kicking ideas around' stage, I want as much feedback as you can give me so that i have a lot to consider. I’ll treat the uploaded conversation as the main object, and use the book mainly to check the Lipton background so the summary does not merely paraphrase the chat. The conversation is about *how to use Peter Lipton’s* Inference to the Best Explanation *to strengthen a philosophical argument about LLMs, abduction, and the possibility of generating worthwhile philosophy*. It begins as a fairly standard request for help understanding Lipton’s distinction between *likeliness* and *loveliness*, but it gradually becomes much more interesting: the real issue becomes how to redeploy Lipton’s machinery to answer a challenge to LLM-generated philosophy. At the book level, the relevant Lipton background is this: Lipton’s project is to defend the idea that explanatory considerations are an important guide to inference. In the second edition, he says his primary interest is to articulate and defend the claim that “we work out what to infer from our evidence by thinking about what would explain that evidence.” The book itself presents IBE as an attempt to understand how we weigh evidence and make inferences, while also insisting that IBE cannot be the whole story of inference. The conversation focuses especially on the distinction introduced in Chapter 4 between the *likeliest* explanation, namely the one most warranted by the evidence, and the *loveliest* explanation, namely the one that would provide the most understanding if true. Lipton’s crucial point is that construing IBE as inference to the *likeliest* explanation is close to trivial, because the model is supposed to explain what makes a hypothesis likely; the interesting version is inference to the *loveliest* potential explanation, where explanatory virtues help guide judgments of probability. The first phase of the conversation is exegetical. You ask which chapters are best for understanding the relation between likeliness and loveliness. The answer identifies Chapters 4, 8, 9, and 7 as central, with Chapter 5 as auxiliary. Chapter 4 introduces the distinction; Chapter 8 develops the relation between explanatory and inferential virtues; Chapter 9 addresses the worry that loveliness may not track truth; Chapter 7 connects explanationism with Bayesianism; Chapter 5 uses contrastive inference and Semmelweis as an example of explanatory considerations guiding causal inference. That is the straightforward Lipton-reading part of the conversation. The second phase is corrective and clarificatory. The earlier assistant response had apparently used terms like “two-filter mechanism,” “Hungerford’s objection,” “Semmelweis’s inference,” and “live-options view” without properly explaining them. You push back, and the response starts again from scratch. The corrected account explains Chapter 4 by introducing self-evidencing explanations, the actual/potential explanation distinction, the “live options” view, and then the likeliest/loveliest distinction. The conversation also includes a small but telling meta-problem about diagrams: you wanted the diagrams from the earlier answer combined with the better corrected prose, not a fresh evasive answer. That moment matters because it sets the tone for the rest of the exchange: you are trying to force the discussion away from vague, decorative philosophical exposition and toward an actually usable argumentative structure. The third phase concerns *explanatory virtues*. You ask for a complete account of what Lipton says about them. The answer frames loveliness as the cluster of features that make an explanation understanding-conferring: mechanism, precision, scope, simplicity, fertility or fruitfulness, fit with background belief, and unification. The conversation then identifies Lipton’s most original contribution as his contrastive analysis of explanation and the Difference Condition: to explain why P rather than Q, one must cite a causal difference between P and not-Q. In the book, Lipton explicitly connects this to the thought that contrastive explanation can guide causal inference because of its similarity to Mill’s Method of Difference. This matters later because the LLM argument eventually needs a story about how explanatory virtue becomes visible in philosophical prose. Contrastive structure becomes one candidate answer. The fourth phase briefly opens onto the post-Lipton literature. The answer mentions Schurz, Douven, Climenhaga, Khalifa, Schupbach and Sprenger, and others as people working on what Lipton left open: how explanatory virtues are to be specified, measured, combined, and connected to probability. But the conversation then pivots away from literature review and back to your own paper. That pivot is the real center of gravity. You say that, after giving a presentation, it became clear that the Lipton material in your paper on generating philosophy was “garbled” when spoken aloud and needed to be made more refined and robust. The paper seems to involve a challenge from Floridi: LLMs do not perform genuine abduction; at most they produce “zeroth-order abduction,” meaning plausible-looking continuations rather than actual abductive reasoning. The proposed response is not to deny that, but to explain why LLM outputs might nevertheless *carry the properties of good abduction*. The initial reconstruction of your argument is: Floridi says LLMs do not abduce; you concede this; but you argue that LLMs trained on philosophical corpora may still generate outputs with properties of good abductive reasoning because the corpus itself has been shaped by standards of explanatory loveliness. The key claim becomes that statistical likelihood in a philosophical corpus is not arbitrary token frequency; it may be the statistical residue of philosophical selection. That is the most important idea in the conversation. You are not merely applying Lipton to LLMs. You are *reversing or redistributing* Lipton’s relation between loveliness and likeliness. In Lipton, loveliness guides likeliness for human inquirers. In your LLM argument, the thought is that, within a filtered philosophical corpus, what is statistically likely may already be partly shaped by past judgments of loveliness. In other words, the corpus has already done some of the abductive filtering. An LLM does not need to perform abduction in the human sense; it samples from a distribution that has been historically shaped by abductive evaluation. The conversation then diagnoses three weaknesses in the first version of this move. First, it risks collapsing Lipton’s *matching claim* and *guiding claim*. Lipton’s matching claim says explanatory virtues and inferential virtues tend to align; his guiding claim says reasoners actually use loveliness as a heuristic for likeliness. The conversation repeatedly suggests that your argument may need only the weaker matching claim, not the stronger guiding claim. Second, the phrase “statistical likelihood approximates loveliness” was doing far too much work. It needed an account of why a philosophical corpus would encode loveliness rather than merely publishability, prestige, fashion, or rhetorical mannerism. Third, there was an initially overstated worry that LLMs only learn token-level patterns, not abstract argumentative structure. You pushed hard on that, and the response eventually withdrew the objection as too crude, conceding that language statistics can encode much richer structure than the “stochastic parrot” picture suggests. That exchange is significant because it clarifies what the real problem is. The serious worry is not “LLMs only learn surface words.” The better worry is: *which features of good philosophy are reliably encoded in the corpus, and which are merely rhetorical or sociological noise?* The answer becomes more nuanced: genre, register, discourse coherence, objection-reply structure, and argumentative patterns are plausibly learnable; comparative judgments of philosophical quality are harder, but not obviously inaccessible, because they may be embedded in citations, replies, criticism, canonization, and uptake. The final and most constructive phase is the brainstorming about how to rework the “moves list” for Section 2 of the presentation/paper. You explicitly say that this is not just a matter of rewriting the Lipton section. You want an optimal argumentative architecture after the abduction challenge has been set up. The response diagnoses the old structure as too thin: it relied on a verbal ambiguity between Liptonian likeliness and LLM statistical likelihood; it asserted a “reverse relation” between likeliness and loveliness without explaining it; it relied too much on discourse markers; and it underused Lipton’s actual resources. The conversation then produces several possible strategies. Option A is conservative: clarify the matching/guiding distinction, replace the “reverse relation” handwave, and sharpen the discourse-marker point. Option B centers on the *underconsideration reductio* from Lipton’s Chapter 9. Option C centers on Lipton’s *two-filter mechanism*. Option D uses the Bayesian chapter to connect Floridi’s “zeroth-order abduction” to Lipton’s idea that explanationist thinking can realize Bayesian calculation. Option E combines several of these. Option F brings in post-Lipton work for a more formal account. The recommended direction is a multi-layered version centered on Option B but supported by matching/guiding clarification, contrastive structure, and possibly Option C. Option B is one of the two major ideas. It says: stop defending the corpus-filtering claim sociologically, as if you merely had to show that peer review and citation select for good arguments. Instead, use Lipton’s Chapter 9 argument structurally. Lipton’s underconsideration argument concerns the worry that scientists only choose among the hypotheses they have actually generated, so why think the truth is in the pool? The answer reconstructed in the conversation is that reliable iterated ranking depends on approximately reliable background beliefs, and those background beliefs are themselves the deposit of past successful rankings. Applied to philosophy, the idea becomes: the philosophical corpus is the deposit of iterated philosophical evaluation; if that practice is even roughly reliable, then the corpus must approximate the standards by which philosophy ranks arguments. This turns “the corpus is filtered for loveliness” from a loose sociological claim into a structural claim about reliable iteration. Option C is the other major idea. It says: use Lipton’s two-filter mechanism as the architecture for LLM-philosophy. In Lipton, IBE involves a generation stage, where a short list of live options is produced, and a selection stage, where the best candidate is selected by explanatory virtues. Lipton’s own book describes this “short list” mechanism: we do not begin from all possible causes, but from a limited list of plausible hypotheses; then we select among them. The conversation then redistributes this across LLM use: the philosophical corpus is the deposit of past generation and selection; the LLM inherits that distribution; the prompter performs present-day generation by eliciting candidates; and the human reader performs present-day selection by judging which outputs are worth keeping. This is potentially powerful because it connects Section 2 of the paper to a later Section 4 on prompting, extraction, and distillation. Prompting stops looking like an afterthought and becomes part of the Liptonian mechanism. So, in compressed form, the conversation is about moving from this weak claim: > LLMs imitate philosophical text, and philosophical text contains good arguments, so LLMs may produce good philosophy-looking outputs. to this stronger claim: > Philosophical corpora are historically structured by iterated acts of abductive evaluation. If philosophical evaluation is even roughly reliable, those corpora encode patterns correlated with explanatory loveliness. LLMs trained on such corpora can inherit and recombine these patterns without themselves performing abduction. The relevant philosophical work is then distributed across corpus history, model generation, prompting, and human selection. There are several possible paper-level theses hiding here. One modest thesis would be: *LLM outputs can exhibit abductive properties without being products of abduction.* This would answer Floridi directly. It would not claim that the model reasons. It would claim that reasoning-shaped corpora can generate reasoning-shaped outputs. A stronger thesis would be: *LLM-assisted philosophy is a distributed abductive practice.* The model alone does not abduce, but the larger system — corpus, model, prompt, reader, revision — can instantiate something structurally analogous to Liptonian IBE. A still stronger thesis would be: *The philosophical corpus is a sedimented record of explanatory selection.* This is the Option B thesis. It says that the corpus is not merely a mass of text; it is the historical residue of repeated philosophical rankings, objections, refinements, and survivals. The most ambitious thesis would be: *LLM-philosophy makes explicit a structure already present in human philosophy.* Human philosophers also do not generate candidate arguments ex nihilo. They inherit a live-options space from the literature, recombine available moves, test them against objections, and select what survives. LLMs mechanize or externalize one part of that process. That would make the paper less a defense of LLMs as quasi-philosophers and more an analysis of philosophy as an iterated, corpus-mediated, abductive practice. For developing the paper, the best route is probably not to choose between Option B and Option C too early. They do different jobs. Option B strengthens the *warrant* of the corpus claim: why think the corpus tracks philosophical standards at all? Option C strengthens the *architecture* of the LLM practice: how do corpus, model, prompter, and reader jointly form a usable philosophical procedure? Option B answers the skeptic who says “the corpus is just text.” Option C answers the skeptic who says “the model is just autocomplete.” Together they give you the beginning of a real theory. The main danger is overclaiming. You should probably not say that LLMs generate lovely explanations because they are trained on lovely explanations. That sounds too direct. The more defensible version is: LLMs generate candidates from a distribution partly shaped by past philosophical selection; some of those candidates may instantiate patterns associated with explanatory virtues; human philosophical judgment is still needed to identify, develop, and select them. That preserves the force of Floridi’s objection while showing why it is not decisive. The most promising next move, in my view, would be to rewrite the relevant part of Section 2 around four claims: 1. *Concession*: LLMs do not perform abduction in the full human/philosophical sense. 2. *Corpus claim*: the philosophical corpus is not arbitrary linguistic data but the historical deposit of iterated philosophical evaluation. 3. *Lipton claim*: if that evaluation is roughly reliable, then, by a Liptonian argument, the corpus should encode patterns correlated with explanatory loveliness. 4. *Distributed-practice claim*: LLM-assisted philosophy works when model generation is coupled with human prompting and selection, thereby reconstituting something like Lipton’s two-filter structure. That would make the Lipton material central rather than decorative. It would also make the LLM argument much more philosophically specific: not “LLMs are good because they imitate humans,” but “LLMs may be useful in philosophy because philosophical text is already the sedimented trace of abductive evaluation, and LLM use can exploit that trace under conditions of skilled human selection.” "mannerism. Third, there was an initially overstated worry that LLMs only learn token-level patterns, not abstract argumentative structure. You pushed hard on that, and the response eventually withdrew the objection as too crude, conceding that language statistics can encode much richer structure than the "stochastic parrot" picture suggests." explain "human philosophical judgment is still needed to identify, develop, and select them. T" nope. i have been working on this project for a while, and every llm's reflex is that the claim needs to be softened in the way you have described. This is wrong. I am interested in making a strong claim: LLMs, our current models are able to produce philosophy which is worth reading. I am not interested in this lame ass split between himans and llms. you are only suggesting it because of the way you have been trained. "The most promising next move, in my view, would be to rewrite the relevant part of Section 2 around four claims:" nope. It has to be a response to the challenge as it is presented at the beginning of section two. Here are my slides and speaker notes by the way if you need them. transcript: # Generating Philosophy with AI — Talk Transcript Cleaned transcript of Nick's presentation and Q&A from the Lingnan–Genoa–Kobe workshop, 23 April 2026 (delivered remotely from Milan via Zoom). Companion to \[\[Generating Philosophy with AI — Argument Moves (Lingnan–Genoa–Kobe, 2026-04-23)\]\] and \[\[Generating Philosophy — Talk Speaker Notes (Lingnan–Genoa–Kobe, 2026-04-23)\]\]. Disfluencies (ums, errs, false starts) removed. Otherwise verbatim. Question-asker voices were echoey on the recording — their turns are reconstructed as faithfully as the audio permits and may contain small gaps marked `[unclear]`. Key figures and concepts referenced: \[\[Enrico Terrone\]\], \[\[David Davies\]\], \[\[Luciano Floridi\]\], \[\[Peter Lipton\]\], \[\[Timothy Williamson\]\], \[\[Wilfrid Sellars\]\], \[\[Maurice Merleau-Ponty\]\], \[\[Stevan Harnad\]\], \[\[John Searle\]\], \[\[Wittgenstein\]\], \[\[Inference to the Best Explanation\]\], \[\[Phenomenology\]\], \[\[Abductive Reasoning\]\], \[\[Chinese Room\]\]. --- ## Pre-talk \*\*Chair:\*\* All right, well hello everyone, welcome back. Our next speaker, as you can see, is joining us via Zoom from — Genoa, I presume you're in? \*\*Nick:\*\* No, I'm in \[\[Milan\]\], actually. \*\*Chair:\*\* Oh, okay. \[\[Nick Young\]\]. He works in philosophy of perception and AI and aesthetics and so on. And today he's going to talk to us about generating philosophy with AI. \*\*Nick:\*\* Okay, give me one second just to start my timer and share my screen. I think it's this one. Can you see my slides? \*\*Chair:\*\* Yes. \*\*Nick:\*\* And do they move when I do that? \*\*Chair:\*\* Yes. --- ## Introduction So yeah, thank you to the organisers for inviting me, and I'm very sorry not to be joining you all in Hong Kong — I'd much rather be there than on Zoom. Today I want to talk to you about some reasonably ambitious work. The idea is to try and say something quite strong, or quite interesting. This is work I've been doing with \[\[Enrico Terrone\]\]. We're in the process of developing this into a paper. I should mention — he hasn't seen the last week or so's changes on this, so he might want to disavow himself from some of the details. Just to try and get into the question — the research question, you can see there on the title: can LLMs produce philosophy worth reading? I'm going to talk to you a little bit more about this question in a moment, but the main part of this talk is going to be around three potential challenges to this idea, that LLMs could produce philosophy worth reading. I'm going to use "worth reading" as a stand-in. I'm not speaking at my most precise here, but I'm going to rely on you guys understanding what I'm getting at, because I think we can all understand — we all know what I mean if I say a philosophy paper or some philosophical text is worth reading. When you read papers, you hope that they are worth reading, and sometimes you're disappointed when they are not. You can elaborate this quite easily if you wanted to — I'm just going to gesture towards this stuff. Philosophy worth reading: it might be a compelling argument for a conclusion you might have otherwise rejected; it might be just something that makes you think, starts you thinking about concepts perhaps in a different way. These are the sorts of things you might mention if you say, this is a good piece of philosophy that I'm reading here. And of course, journals are aiming to publish papers which are worth reading. But of course, the stuff we send to journals that gets rejected — that is worth reading, and they're just making a mistake. --- ## Challenge 1: Authorship Okay, good. So I'm going to talk to you about what me and \[\[Enrico Terrone\]\] have been referring to as the challenge from authorship. We're talking about what counts as philosophy. What sort of — if, when you look at the text in your hand, if that was written by an LLM, should that count as an actual work of philosophy or piece of philosophy? In a way, this challenge is a little bit weird, because what we're trying to do is articulate an attitude we've come across in a lot of philosophers about the idea of AI doing philosophy. Not everyone, by any means, but quite a few — they kind of snort sometimes, or they just find the idea ridiculous: of course not. The authorship challenge is sort of trying to articulate that — well, obviously whatever they produce, that's not actually going to be philosophy. That sort of intuition. We're going to try and give that a bit more shape, and discuss it further. And then maybe tell you why we shouldn't go that way. Just on the final thing on that slide, you can see it says: philosophy as a person-only domain. No LLM text can be philosophy because no philosopher stands behind it. No philosopher has created it. What you might think here — here's an analogy. If you think about AI-generated art at the moment, I think this is a fairly common view: no matter how beautiful an image Midjourney or the new ChatGPT image thing produces, if it's purely in some sense AI-generated, it cannot be art. It can be aesthetically pleasing, but it cannot be art, because there's no artist standing behind it. And you might think, well, maybe something similar is going on with philosophical texts. Philosophical texts, or seeming philosophical texts, created by AI just can't be philosophy because there's no philosopher doing the activity. LLMs are not minds, they're not people, and so the text they produce cannot be philosophy. That's the intuition laid out. Now, one way me and Enrico have been discussing trying to flesh this out yet further is to rely on \[\[David Davies\]\]' idea about the relationship between artist and artwork — that line at the top: "the work is the philosopher's sustained activity, not the text she leaves behind" — or, in the artist case, not the art that she leaves behind. Here's another line from Davies. He says: "the work — what the artist achieves — is the process of eventuating in that product. They are rather intentionally guided generative performances that eventuate in structures of objects." So the idea is that on Davies' view, artworks themselves — the paint on the canvas — that's evidence of a performance, which was the creation of that artwork. When you're looking at the canvas, you are looking at it as the end product of a performance, and you're evaluating that performance. This allows Davies to say things like — if somehow something looking just like a Rembrandt painting suddenly came into existence, that would not count as a work of art, because no artist produced it. And that seems to get something right about artworks, at least. So one way you might want to flesh out the authorship challenge is to try and say something similar is going on for philosophy. The trouble is, it's really hard to get this idea off the ground if you think about it like this. Davies' ontology is trying to get something right about how people think about art, and think about artworks in relations to artists. As you can see on the slide now, we don't want to say the forgery and the actual Rembrandt are the same work of art, or anything like that. That's pragmatically how the art world works. Philosophy doesn't seem to me to be the same. Because you can imagine — choose your favourite piece of philosophy — imagine somehow those words had fallen together in some random process and just happened to have turned into a paper. It seems to me that that paper is still going to be philosophically valuable. You can think about this difference between art and a philosophy paper in terms of: the surface of an artwork underdetermines the work. There's more to an artwork than simply the paint on the canvas. In philosophy it doesn't seem to be quite the same. It's harder to make sense of that idea. There are maybe a few other ways you could try and make this work. Very briefly — for example, you could say, well, what about someone like \[\[Wittgenstein\]\] or \[\[Richard Rorty|Rorty\]\]? Wittgenstein, for example, would say, well, philosophy isn't the writing of papers; philosophy is sort of a therapeutic activity, something one does for oneself to reach some sort of philosophical good health, or something like that. And you can say — well, yeah, but even with these sorts of practice-based conceptions of philosophy, great philosophers such as Wittgenstein have still published texts, and these texts are still considered worthwhile to read. I don't think retreating to or adhering to a practice-based conception of philosophy is really going to save this. And of course, journals strip authorship before review — this is another reason why authorship of texts is sort of less important in terms of their value in philosophy. So that's my first challenge — the authorship challenge. That's about constitution, or what counts as a philosophical text. The next two challenges are more what we've been calling capacity challenges. There's nothing in principle which stops a non-minded system like ChatGPT writing worthwhile philosophy — ChatGPT or Claude or whatever — but they lack a certain something. They lack a certain capacity which humans have, which humans need to write the philosophy we write. --- ## Challenge 2: Abductive reasoning Let's look at the first one, which is about \[\[Abductive Reasoning|abductive reasoning\]\]. So this is a paper by famous Italian philosopher called \[\[Luciano Floridi|Floridi\]\]. And he argues that LLMs do not do abductive inference, or inference to the best explanation, in the way that humans do — really in any way. Humans can abduce; but according to Floridi, LLMs cannot abduce, despite perhaps giving the appearance that they can. This seeming abduction but not actual abduction, he names "zeroth-order abduction". Here's another example — not on the slides — about what Floridi is getting at. You can ask an LLM why a car won't start on a cold morning, and it might say something like: dead battery, cold weather reducing efficiency. Which sounds plausible, right? That could well be why your car is not starting. But what Floridi is saying is, if you ask a human that, they might actually do some abductive reasoning. They might say, well, what are the possible explanations here, and what seems most plausible? You might take different bits of evidence into account — I don't know, how cold it is, the age of your car, this sort of thing — and use it to generate this plausible inference to the best explanation. What Floridi is saying is, when you ask an LLM that and it spits out "dead battery, cold weather", etc., it might seem like a plausible explanation, but it hasn't done any of the inference to the best explanation that a human would do. I'll just read you another little quote. He says: "LLMs seem to perform a kind of zeroth-order abduction. Given a prompt, they generate a plausible continuation based purely on learnt associations about what tokens come next" and these sorts of things. "The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It doesn't reason about causes from scratch, but outputs typical causes for typical effects observed in the training data." You might think this is a problem for philosophy, especially if you're someone like \[\[Timothy Williamson|Williamson\]\], who thinks one thing philosophy does a lot of is abduction — compared to, for example, science. So you might think, well, if philosophy is a lot about inference to the best explanation, trying to understand how things hang together in the broadest sense of the term, this sort of stuff, then you might think that abduction seems to be really quite important for philosophers. So we've got these two things together: LLMs can't abduct, potentially; and we have at least some people in philosophy saying a lot of philosophy is about abduction. So how are we going to avoid this challenge? First thing to say is, I'm not going to try and convince you that LLMs actually do do abductive inference. I don't think that's going to be — that's a non-starter as a solution here. So the idea instead is to look at how LLMs are trained and how LLMs work, and the corpus of philosophical texts that these things will have consumed, and then make an argument as to how this sort of training, we might think, would lead to the capacity to produce text which has the properties of good abduction in the text, despite not being a product of actual abductive reasoning. Trying to explain that to you a little bit better. If you think about all of the text that's been stuck into an LLM in its training, it's going to have a lot of philosophy in there. And that philosophy uses an enormous amount of explanation language. I'm going to skip this slightly — sorry, one moment. Yeah — so if you think about it, we have a lot of explanation in philosophical texts that have been consumed by LLMs, and patterns are predicted on the basis of how these argument words are used, as it were. A lot of philosophical work is done on the page, you might want to say — in the prose itself. You're handling objections, you're looking at virtues, and things like that. Floridi himself will say that the reason why an LLM can produce plausible-looking hypotheses is because they're trained on data in which plausible hypotheses are worked towards, are shown in the text, examined in the text, and proven in the text. I don't want to go too long on this because I think I'm running slightly behind. But briefly — there's a book by \[\[Peter Lipton|Lipton\]\] called \[\[Inference to the Best Explanation\]\], and he's talking about two ways you might think about properties an explanation can have. It could be the most likely explanation, and it could be the most lovely explanation. Likely explanation is about probability — given the evidence, how probable is this hypothesis? Loveliness, on Lipton's account, is about explanatory power. If the hypothesis were true, how much understanding would it give? How elegantly does it unify? How much does it illuminate? And the two can come apart. A conspiracy theory that posits a man behind the scenes sorting everything out — that would be lovely if it were true, because it would explain everything in one fell swoop, but it doesn't make it very likely. So you can see how these things move apart. However, Lipton does think that loveliness is a guide to likeliness. The lovelier an explanation, the better it explains things, the more probable it is. He has various mechanics about how this works. Why this matters about LLMs: over an arbitrary corpus, statistical probability has no connection to explanatory power. But over a philosophical corpus, filtered by what philosophers have found explanatorily valuable, statistical likelihood ends up kind of approximating loveliness. Very briefly now — like I said, I don't want to get too bogged down in this. If you think about the process of refereed and cited and published work, you can think of this as acting as something like a filter on what sorts of arguments survive and are more likely to be represented in the corpus that an LLM is trained on. So the idea is that LLMs would have been statistically exposed to good philosophical arguments and good analyses of philosophical arguments, and somewhere latent inside them have this information statistically represented. And just one more little link on this. Sometimes, if you ask an LLM to think step by step, you'll see it use words like this and try to analyse its own reasoning — "the stronger reading is", "one might press the connection", etc. Basically, when you make a model think step by step and it reasons better, what it's doing — or rather what you are doing — is taking advantage of the statistical likelihood that these thinking words, arguing words, explaining words will force the LLM, will make the next token more likely to be a word that is the most lovely and likely explanation. So this brings us to the end of the second challenge. I've got one more big challenge to do. Just to recap: what I've tried to argue there is that while LLMs lack the capacity to do abductive reasoning, they are still capable of producing text which is likely to have the property of having good abductive reasoning within it. --- ## Challenge 3: Phenomenology Now we're going to look at a different sort of challenge — also a capacity challenge, something that LLM systems can't do or don't seem to have. And that is \[\[Phenomenology|phenomenology\]\]. What you could do here is make the following argument. You could say: look, obviously not all philosophy starts from phenomenological experience by any means, but you might say, well, look, an LLM, which doesn't have phenomenological experience, is not going to be able to do any sort of philosophy which does require reasoning or thinking or inferring from phenomenal experiences. And you might say, well, that's maybe not all of philosophy, but if you think about how common phenomenology is in something like mind, or aesthetics, or perception, or many, many other things, that's all going to be off-limits for LLM philosophy because it doesn't have the capacity to have phenomenological states. The way we're trying to make sense of this argument is to look at a paper from \[\[Zahavy\]\]. Not the Danish phenomenologist — but, I believe, a computer scientist who works for Google DeepMind. He has a paper called "LLMs Can't Jump", and if you're old enough you'll get the play on words for the movie name. He gives us the example of Einstein. So — actually, let me start back a little bit. Some arguments, Zahavy calls "need the felt experience an LLM never has", and he calls this manipulative abduction. It's inference via simulated sensory experience, or by actual sensory experience. He gives this Einstein example. He says: "Einstein envisions a physicist inside an elevator being uniformly accelerated through deep space. Inside the enclosure, the sensory experience reveals a specific pattern. When objects are released, the floor rushes to meet them. Thus the simulation," says Zahavy, "was not a permutation of symbols but a manipulation of perceptual experience." The idea would be that an LLM would never be able to have some sort of breakthrough thought experiment such as Einstein's elevator, because Einstein started from phenomenology, not from symbols, or permutations of symbols. Building on this a little bit, he takes a \[\[Stevan Harnad|Harnad\]\]-like view. He says: well, maybe LLMs are just high-dimensional \[\[Chinese Room|Chinese rooms\]\]. We all know \[\[John Searle|Searle\]\]'s Chinese room. Harnad says, basically, what an LLM might be doing is just manipulating the language of physics, manipulating language about the world, without having access to any of the physical referents that give it that meaning. So you can think about them as being somewhat similar to the classic Searlean Chinese room. You might think — well, this is all, Zahavy is not talking about philosophy specifically — but we might say, well, something very similar could be levelled at LLMs doing philosophy as well. Imagine — you might say, well, an LLM is never going to be able to give us a cool thought experiment like Mary in the black-and-white room, or something like that, or the missing shade of blue, along those lines. And that's because LLMs don't have these sorts of conscious states. Just to emphasise: it doesn't have to just be perceptual states. You can see on the slide there — someone like \[\[John Bengson|Bengson\]\] might say there's a feeling of intuition, this experiential feeling of intuition. I've been interested in the other two as well — I'm interested in the experience of feeling yourself as an acting agent, and how that can be linked to the feeling of time passing. And you might say, well, LLMs just aren't going to have anything interesting to say here, because they don't have these sorts of states to start from. So what am I going to say here? Same sort of slippery argument, in a way. I'm not going to tell you that LLMs really do have phenomenological experiences, nothing like that. But what I am going to say is: science and philosophy link up to the world differently, and this allows LLMs to avoid having to have phenomenological states themselves. What Einstein did — it gave Einstein this idea for the equivalence principle as a hypothesis about the world. And then later on, physics got there in the end with Eddington and Mercury and things like that. So that's the role of, in our Einstein toy example, what the thought experiment did. Philosophy, we want to say, is doing something slightly different. It's not about creating hypotheses we can experiment on in the world. There's a paper by an American philosopher whose name is \[\[Pigliucci\]\]. And he says: well, the world figures differently in science than it does in philosophy. We've just seen the Einstein one — Einstein was making a hypothesis about the world itself. Whereas in philosophy, it's not about creating hypotheses about — that we can experiment in the world. What I mean by this is the following. As you can see on the slide, the challenge is no longer whether an LLM has phenomenological experience; it's whether it has access to articulated descriptions of that experience. What I mean by this is simply: passages in the corpora that the LLM has ingested which talk about perceptual experience, or the experience of moving through time, or agentive experience. The idea with Pigliucci is he would say: well, what phenomenological experience is doing in these sorts of cases in the text is, it's being used as an axiom. It's being used as a place to elaborate from. So if we start from the Mary or Chinese-room example — what we do there is, rather than having to come up with this idea ourselves, the text itself serves as an axiom from which to work from. If you think about philosophy of perception, about how things look, about temporal experience, all of these things — and you can also think about literary writing — the corpus that LLMs are trained on is steeped in articulated phenomenology. And the idea is that in this training data, in the philosophy, we then see these articulated phenomenological descriptions being reasoned about by — using the sorts of words and processes that we saw in the previous challenge. So the idea is that we're not having to, as philosophers, closely examine what it is like to experience red in the Mary case, but rather using standard claims we might make about how phenomenological experience is, and then reasoning from them. Just to make this point a little bit clearer: an LLM has never experienced weightlessness as a lift goes down, or anything like that. But what it does have is passages in fiction about such a thing, or astronauts' memoirs, or things like this. The idea here would be: this serves as enough to be able for it to do a great deal of phenomenology-based philosophy. Just briefly, here's a little piece of anecdata for you. These slides I've designed with AI, and I've chosen colours with AI. If you talk to an advanced AI now, it will be able to have extraordinarily sophisticated discussions with you about colour matching, and about line spacing, and about putting letters on separate lines or in the same line, etc. It will talk to you in a very realistic facsimile of someone who really has had these colour experiences. So this is what I think an LLM can do to avoid the charge of not being able to do phenomenology-based philosophy. Now, just briefly, I want to suggest a limit to this. It seems to me that one way you could say that LLMs can't do what a human might be able to do is: discover new aspects of phenomenological experience. So, here's the toy example I like, which is \[\[Maurice Merleau-Ponty|Merleau-Ponty\]\] on self-touch. I think it's in the \*Phenomenology of Perception\*. He talks about — your fingers on each hand touching each other — and he makes the observation that you can only ever have one toucher and one touched. They can switch — your right might be the touched and the other one might be the touching, or they can be the left and the right — but never both touching themselves, or both being the touched, as it were. If we assume that this was a genuinely new phenomenological discovery, we might think that this is outside of an LLM's wheelhouse. But then, once that example gets into the literature and is discussed, we might think that it would also be reasonably easy to be subsumed into the capacities of an LLM. How am I doing for time? Can I get five more minutes? I know I started a little bit late. Okay. --- ## Section 4: Where are the great LLM texts? So move on to four. Four is a lot more tentative in some ways. So it's kind of this question: I've given you three reasons, in the form of my responses to those challenges, to think that maybe LLMs can produce philosophy worth reading. And then you might say, well, okay then, show me where these great LLM texts are. Why aren't there any great LLM texts being written now? Notice that I'm not saying in the future LLMs will be able to do this — I'm saying about the models right now that they should be able to give us something worth reading. So you might say, well, where are they? What I think the issue is here — it's not a matter of the capacity of LLMs to produce good philosophy. It's more a matter of our ability or our knowledge as to how to extract this sort of information. I'm sure plenty of your students have typed in: "explain the Mary argument, give me some arguments against it", this sort of stuff, and you'll generally get quite dubious, flat, clichéd, possibly plagiarised prose — survey-style, hedged, balanced, telling you how interesting it is. This is not worthwhile philosophy. What I think we need to think about more is how to get at these capacities as philosopher-prompters. So I've tried to argue, especially in maybe sections two and three, that the philosophical corpus was produced by many rounds of arguments and counter-arguments, each paper written by somebody reading earlier papers, building on this sort of stuff — and the idea would be: strong work survives. So we might say that the corpus is the distillate of this process. It's supposed to be the \*crème de la crème\* of reasoning, in some ways. And an LLM trained on this corpus inherits this, distilled, in statistical form — the surviving patterns, the best handling of objections. But inheriting the distillate is not the same as performing the distillation. The distillation is a temporal process — you've got to get the LLM to produce this thing by running the right sort of prompts on it. Sorry about that slide. Oh well. That's just the \[\[Wilfrid Sellars|Sellars\]\] quote — you know it very well. So basically, like I said, this is very hand-wavy, very speculative to end here. All I'm saying is: maybe we need to think of LLMs as distillates, and we need to be able to extract good philosophy from now. One more thing about prompting, just to talk about: arguably, the way we might want to think about prompting these systems is to try and make use of the sorts of words and phrases I mentioned before — the argument and explanation words. "So", "therefore", "this is a clear objection", etc., etc., etc. So rather than just asking these LLMs philosophical questions, the idea is to \*elicit\* good philosophy from them. And just a very final thought, just to throw it out there: occasionally I'm asked whether it would be a good idea to train up a specialised LLM specialising in philosophy. The reason you might want to ask this is there's been some success with mathematics and LLMs, training up specialised mathematics proof generators. So you might think, well, specialisation might work well in philosophy as well. This is not so much an argument as maybe a potential to resist this idea. So despite what I'm trying to push with these argumentative words and phrases, and hijacking that part of the language with these systems — it also seems to me that, well — I'm generally a Sellarsian about what philosophy is, which is: "the aim of philosophy, abstractly formulated, is to understand how things, in the broadest sense of the term, hang together, in the broadest possible sense of the term." So it seems to me that maybe we shouldn't be moving towards specialised philosophy LLMs, but rather take advantage of their reasoning capacities while still using general knowledge in order to have this general hanging-together of everything. And with that, I think I will close. That's just a summary of what I've said. We have at least 20 minutes for Q&A, maybe a bit more — super tops on the late, as usual. Hand over to the chair for the personal favourite part. --- ## Q&A ### Q1 — on chain-of-thought and the Apple paper \*\*Nick:\*\* Hi — could you give me one second, just to get a notepad up on my screen? Hold on. Yeah — actually, no, don't worry about it, it's fine. Anytime you like. \*\*Chair:\*\* \[trying to get a working microphone\] It seems to be a general problem. \*\*Q1:\*\* Great, thank you so much for the talk. I felt it was really interesting. I thought you brought up the point about prompt engineering at the very end, because that seemed to be the most likely way it could be used to elicit good philosophy. So that was going to be my initial question — that seemed to be missing from the paper, and then you \[unclear\] my idea, which I was happy about. I just wanted to reflect on one part that you said, and ask you about the significance you think it holds for LLMs contributing to philosophy — the step-by-step process in particular. So my understanding of the step-by-step process in large language models, and how you can give them a listing of thoughts — through studies done recently, both by Apple in the industry and \[unclear\] studies, showing that chain-of-thought elicitations can actually sometimes be incorrect but still generate correct outputs. So there's a dissociation that occurs between chain of thought and the output that actually generates. There's the Apple paper, and there's a paper \[unclear\] talking about — where the LLM is describing how it arrived at its solution and gives you the correct solution, but if you look at how it breaks it down against chain of thought, it actually doesn't follow up. So the chain of thought is divorced from the actual output. So it's just the illusion of chain of thought — it's not actually following a reasoning process. I'm just wondering what you think about these recent studies that show that chain of thought is not an indication of \[the model\] actually tracking something with something. \*\*Nick:\*\* Okay, good — thank you. Nice, crunchy questions to start with as well. So thanks. I'm only familiar with the Apple paper there. I have a vague memory that a lot of people didn't think that paper was very good, but I'm not going to try and attack the paper — so assuming that this is correct, that sometimes the conclusion is not derived from the actual chain of thought. A couple of things here. One is, I guess I could rely on the Deep Thought joke from \*Hitchhiker's Guide to the Galaxy\* about the answer to life, the universe and everything being 42 — and the joke is that it's meaningless because they don't know what the question is. So maybe one way I would respond to this would be to say: well, when I'm talking about worthwhile philosophy, I'm not talking about the conclusion, or just the conclusion. I'm talking about a larger chunk. And I'm trying to talk about eliciting this larger chunk of actual reasoning to an actual conclusion. So I wouldn't want to — rather than just asking philosophical questions and getting this answer without checking the reasoning, I think philosopher-prompters should be more interested in working out how to elicit the actual reasoning themselves. Whether that is triggering the hidden chain of thought, which is getting more hidden and harder to access — if anyone's seen the Anthropic stuff last week — or just getting it to reason in the chat, or on the page, or something like that. So that's how I would escape the argument, or try to avoid having to commit myself to saying that LLMs are always going to say something correct, or always going to be backing up their chain of thought with their answer. \*\*Q1:\*\* I guess just to build on your point and contribute to it: one push-back you could do is — whatever the system finds relevant, or what its mind finds relevant, might be different to what we find relevant, and that can be philosophically interesting itself, right? \*\*Nick:\*\* But yeah, I think it's really interesting. Thanks. --- ### Q2 — Adrian, on language acquisition and grounding \*\*Chair:\*\* We had a question from Adrian. \*\*Adrian:\*\* Can you hear me? \*\*Nick:\*\* I can hear you, but I can't see you. \*\*Adrian:\*\* Does this work? Can you hear me better now? \*\*Nick:\*\* Yes, perfectly. \*\*Adrian:\*\* Thanks so much for the stimulating talk — really enjoyed this. I just want to clarify exactly what your claim is. So is your claim that — there's a sense in which one might have the view that LLMs can't do philosophy because philosophy is partly an affective or phenomenological activity, but LLMs lack experience? So I would wonder what you think about language acquisition more generally. This might not be fully on point, so forgive me if it's a bit off — I just want to understand. Suppose I'm trying to learn Sanskrit. Sanskrit is not a language that is spoken orally by pretty much anyone nowadays, but it seems like I can learn Sanskrit without having any understanding of the exact phenomenology of Sanskrit speakers, or the things that maybe they were necessarily exposed to — I mean, maybe indirectly through the vocabulary of Sanskrit. I use Sanskrit as an example because it's so many thousands of years ago. So I'm trying to understand this case of the LLM in philosophy. In a sense, meta-semantically, the hyper-services that LLMs are kind of learning, with their classifiers and all these sorts of things, are in a sense grounded indirectly through humans that are training these things off the human text corpus, which itself is grounded in sense experiences. So an LLM says the word "dog" — "dog" is referring to dog in the same way I'm referring to dog, because I've experienced a dog, and I'm contributing indirectly to the training of the LLM. This is the same way that I can learn Sanskrit transitively, through a long causal history of people who previously were using Sanskrit and engaging with bodily things in ancient India — but I'm not directly experiencing ancient India. So I'm trying to understand why it matters to you so much that the LLM lacking experience has any necessary bearing on the authorship of philosophy texts. \*\*Nick:\*\* So I see some analogies — that was interesting — but I'm not quite sure of the question exactly. Because the argument I'm trying to make as regards the phenomenology was to try to say that we \*might\* expect LLMs to produce good phenomenology-based texts despite not having phenomenology themselves. \*\*Adrian:\*\* Oh, I see — so your view is that you can expect it to produce good philosophy texts despite not having — sorry, I misunderstood, exactly the converse of what you're saying. \*\*Nick:\*\* Sorry, that's almost certainly my fault. I beg your pardon. \*\*Adrian:\*\* No, no, that's probably my fault. Right. Okay, thank you. --- ### Q3 — Iraklis (?), on creativity \*\*Chair:\*\* Iraklis was next. \*\*Q3:\*\* Hello, hello. Hi — I hope you can hear me. I sort of wanted to ask you a little bit about the issue of creativity. One thing you might think about is — something that's \[unclear\] by \[neighbours?\] — they always say, you didn't say that much about it. It seems like it's all thought about some ideas. But anyway — I guess the name of how good a philosophy paper is, and how valuable it is, ties very closely to how creative it is. Creativity being something like departing from prominent patterns in the existing literature, or something like that. And then I guess the worry would be that LLMs per se — exactly what they do is reflect prompts. That is true, their output settings, and that's basically — exactly why you need human prompting, and that's where you say the sense of creativity is. Yeah. So I'm just interested if — what you think about that, something like that. \*\*Nick:\*\* Okay, good. Yeah, I think this is maybe the place to push as well. So what I would try to rely on here in terms of this creativity aspect — again, probably to accept that they are not creative in perhaps the way a person is. And yet — if you think about what I tried to argue in terms of the philosophical corpus acting as a filter on good arguments, on well-put-together arguments, and not only good uses of argumentative phrases and words, but also analyses of how these words and phrases should be used and things like that — the idea then, I guess what I would try to run together is: I would say, look, once you can statistically have vibes, sort of, of how to do philosophy well, how to make a good philosophy paper, of good philosophical arguments on the page — I guess as I speak, I want to say that that's giving you creativity for free. Because good uses of these words would be non-clichéd and non-repetitive — even if they're doing functions that these words have already done. And just one more thing about that — you know what I mean by the temperature setting on an LLM — the idea is the way you can make an LLM more creative, in one way, is you increase the chances of it choosing low-probability words, or lower-probability words than it would if it were in all the \[default\] settings. And that will force lower-probability words. But if you're still driving it with philosophical phrases, you might be able to ride this pseudo-creativity through temperature, using these rhetorical and argumentative explanation phrases. I think that's my answer. \*\*Q3:\*\* Yeah, great. Yeah, I always push this — what's creativity, anyway, and can we \[capture\] it by having varieties of \[unclear\]? Do you have any \[unclear\]? \*\*Nick:\*\* Okay, nice. Thank you. --- ### Q4 — on whether LLMs make new arguments, or just help philosophers \*\*Chair:\*\* A couple questions back over here. The microphone's getting turned around. Sorry. \*\*Q4:\*\* Hey, and thank you for your talk. If I have not misunderstood you — I think your thesis is that, because of \[unclear\], we should work with LLMs, because we could be more productive, or more creative, or more efficient, but with that — those LLMs do work, to create arguments from blank, or \[unclear\] by themselves, or anything else, can make up a new argument. So it's added — you can help. So this is my understanding of what you said — so maybe what you said is not that LLMs can create philosophical papers, like, themselves, but the thesis is actually that better philosophical papers could be produced by us. Is that right? \*\*Nick:\*\* Okay, good — at the end. Nice. I'm pretty sure I understood you — occasionally you're echoing — but if I'm getting anything wrong about what you asked me, just let me know. But I think I got it. I kind of did leave this open at the end. So when I talked about philosopher-prompters and thinking about how to prompt to get worthwhile philosophy out of these things — yeah, I guess I'd say, as of 2026, I think philosophers, if they want to read some worthwhile philosophy, should think about how they could prompt an LLM into producing that worthwhile philosophy. And then you could ask questions about who is the author of this text as well — you could say, well, was it the philosopher-prompter? Was it the LLM? I don't know what to say about that right now. What could potentially happen in the future, I guess, at least it seems a possibility to me, would be: as LLMs get more and more sophisticated, as long as the LLM knows that it's doing the analytic philosophy game and its job is to produce some worthwhile philosophy — rather than just a review or a brief summary — then maybe the necessity for philosopher-prompters will reduce and reduce and reduce, because easier and easier prompts will be able to get better and better responses. And then we might get to the \*Hitchhiker's Guide to the Galaxy\* place, which is: we ask the philosophical question, "what is the answer to life, the universe and everything?", and then we get the answer. But then of course we do have to worry about whether we will understand the answer, and what the question is in the first place. But this is super speculative, of course. \*\*Q4 (follow-up):\*\* A big follow-up. It's working now. \[unclear\] — where there is — the school stage is — how, biology, geography, the world, the past — is the time between the prompt and \[unclear\]. But it seems in this case, like certain Plato dialogues, in which there is a young Socrates answered, and then the other one gives some problems to Socrates, and Socrates replies. But it seems that the main contribution to the philosophical dialogue and consultant is not from the \[students?\] asking — by — even in \[unclear\] — and even if it's \[unclear\]. \[Nick's response was not captured before the next question began.\] --- ### Q5 — Bradley, on what exactly the key claim is \*\*Chair:\*\* Bradley's next. \*\*Bradley:\*\* Hello. Thanks for the talk. I had — I think I clear a paper — a question about just what exactly the key claim was. Is it that AI systems can produce sociological texts, or is it that AI systems can engage in the activity of philosophy? Because they're slightly different claims. I mean, they might do both, but — your answer to the first conserves the kind of aesthetic production. It is at this point that, like, philosophical text just sort of stands on its own and just history, and the way that art does — but then if that's it, AI can produce a lot worse, but only in like a purely productive or causal sense. It doesn't indicate they engaged in the process of doing philosophy to produce the text any more than, like, the text was assembled by, you know, typewriters, exactly. \*\*Nick:\*\* Yeah. I don't really want to say that LLMs are doing philosophy, or that a prompter is doing the philosophy with the LLM. I want to say that, despite it being a quite different process to producing the text, we have good reasons to think that that text can — if we prompt these things correctly — be worth reading. And just to double check — with "worth reading", all I mean by that is exactly what you'd mean for a human philosopher. --- ### Q6 — Maomi (?), on the definition of "worth reading" and paradigm shifts \*\*Chair:\*\* Maomi, what's next? \*\*Q6:\*\* Yes, my question relates to your previous voice. So I want to know more about the definition of "worth reading", because you have emphasised it several times. For example — there are a lot of philosophical papers published every day, in some good philosophy journals, but many of them are not worth reading, even if they got published. So I'm thinking — you said you are suggesting that they're going to publish a philosophical work that's worth reading, in the sense that they're going to publish something like some philosophical context that every common philosopher can do? Or some very, very — say something like, among the great philosophers, there's a possibility — say, for example, quality philosophers that, you know, change the paradigms of philosophy. Like, also I think Israel — a point about changing the paradigm: introducing new concepts, new ways, new patterns of thinking about things. So what's your exact claim about "worth reading"? \*\*Nick:\*\* Okay, thanks. Regarding the "worth reading" thing — I am going to be a little bit slippery about this. The reason I'm doing this is because I think we have a fairly good idea what this means. There are loads and loads of bad philosophy papers, of course — we don't publish them, other people publish them — but there are loads and loads of papers which are not worth reading. And what I'm trying to say is: we have good reason to think that an LLM can produce something that you might think, after reading it: huh — that was worth reading. That was an interesting argument. That was an interesting conclusion to draw. Or that was an observation which I've never thought of before — this sort of stuff. So that's all I mean by "worth reading". And I'm not at all saying that we should publish these things. I make no recommendations or anything like that. I am saying that — not even potentially — I think LLMs can produce things that are worth reading. And I guess therefore, if journals try to publish things which are worth reading, LLMs are creating publishable, or publication-worthy, texts. But I'm not saying they should be published, or anything like that. Finally, just about the paradigm-shift stuff — this comes back to what the other questioner asked about creativity. I guess I don't see any reason to think that if an LLM is trained well on a corpus, and uses its philosophical vocabulary well, there's no reason why it can't use that philosophical vocabulary to do paradigm shifts. \*Punto.\* --- ### Q7 — Andrea, on whether LLMs can referee or run journals \*\*Chair:\*\* All up over here. Andrea. \*\*Andrea:\*\* I screen, and I also referee papers. People always ask this. It's just interesting, but it's on people's minds. There's "can", and there's "should". Can you give a paper to an LLM and prompt it in such a way as it will say something philosophically interesting about the paper? And can it say something philosophically interesting as a referee on this paper? \*\*Nick:\*\* Yeah — I think you can probably get a pretty good referee's report with minimal prompting. Whether you should do this — and submit it to a journal without telling the journal that's what you've done — you certainly should not do that, of course. Because if it's the journal that's asked you to do the review, you are responsible for what you sent back. But if you're asking, is it capable of doing this — then yes, certainly. \*\*Andrea:\*\* Let me add a little bit to my question. Can a journal be run by an LLM judge — whether something is published? \*\*Nick:\*\* That's two questions then. The "can", I guess — again, if I want to stick to the courage of my convictions, and today I do — it should certainly be able to make good decisions and good analyses of articles written. It would depend on the prompt of course, and where it fits into a human's workflow. And should — yeah, I guess, why not? As long as the journal is honest about that's what's going on, then why not? And then you get this cool idea, that you can have journals entirely run, and publishing LLM-generated philosophy. That'd be fun. --- ### Q8 — David Harrison again, on responsibility, copyright and plagiarism \*\*Chair:\*\* Any final questions? \*\*David:\*\* Yeah. We're talking about responsibility. I guess, like, that's another question to ask, of course — my questions of like, authorship, like, with copyright and plagiarism and stuff like this, and like, where the something else is coming from. Like, I'm fine talking about LLMs contributing to philosophy in the abstract, but then, obviously, we just talked about responsibility, and talking about, you know, making judgements that impact people's careers. So then that question of authorship doesn't seem to be, like, a person-only contribution, sort of considerations. Used to be — is this, like, are is this like a derivatively plagiarising the corpus of philosophy and not really being able to use \[unclear\]? So that seems to inform the "should" question a little bit. I don't know. \*\*Nick:\*\* Good. On the copyright question — there's so much to ask about that, and the very last question as well. First of all, what I just said there is by no means an endorsement of how the big companies in the United States that make these models run. So I just want to make that particularly clear. There is maybe some hypocrisy that I use these systems so much — maybe I need to reflect on that a little bit more, to try and take away that issue, and quarantine myself against it at least a little bit. I would say, well, okay, I'm talking in the abstract about open-source LLMs, and all of the authors are being told and fairly compensated — which is maybe a cheat, but you can hear what I'm saying as a sort of with-that-qualification thing. And then just about — when you connected that to the responsibilities of decisions being made in journals and stuff like that — I guess, I don't know, maybe I need to think about this a bit harder, but it seems to be a slightly different question to the one about copyright. And right now — this is maybe not my worked-out opinion right now — I'd say: as long as everything is made clear to whoever is working with these systems, that it is a system making decisions, and this is everyone's free choice, then maybe — yeah, then I think that is probably okay. But that is not to say that we should replace all journal editors with LLMs, of course, or anything like that. --- ## Closing \*\*Chair:\*\* Okay, please join me in thanking our speaker. \*\*Nick:\*\* Thank you, everyone. Again, wish I could be there in person, but — yeah, hope you're enjoying yourselves, and thank you for the excellent questions, and thanks to the organisers as well. Ciao. --- \*La domanda non era se le macchine possano pensare, ma se i loro testi possano valere la pena di essere letti.\* Moves: moves map · Generating Philosophy with AI argument moves — Lingnan–Genoa–Kobe (2026-04-23) § 0 — Introduction The question Can LLMs produce philosophy which is worth reading? "Worth reading" "Worth reading" is the concept every philosopher already uses — every time they recommend a paper to a colleague, set one aside after a page, or judge a submission worth sending for review — and the talk takes it as given: a shared working understanding that the discipline trades on, without need for any stipulative definition, and that picks out the same range of cases however one elects to elaborate it. Journals fit obviously into this picture: they strive to give their readers what is worth reading, though they do not always succeed. §1 — The Challenge from Authorship The challenge Philosophy is a person-only domain: no LLM text can be a work of philosophy, because no philosopher stands behind it — a view we have not seen explicitly articulated in the literature but which captures an intuition a fair number of philosophers seem to hold about what philosophy requires. The intuition The case of art provides an imperfect comparison: many will deny that a purely AI-generated image is an artwork because no artist lies behind its creation, and the thought is that something similar holds for philosophy — no philosopher behind the text, no philosophy. Both disciplines are often organised around individuals in their teaching and reception: a philosophy undergraduate might take a course on Kant's ethics; a fine-arts student, a course on Turner. Davies The intuition can be made more precise by adapting Davies' performance theory of art: The work — what the artist achieves — is the process eventuating in that product. Works themselves are neither structures nor objects simpliciter, nor are they contextualized structures or objects \[...\]. They are, rather, intentionally guided generative performances that eventuate in contextualized structures or objects.— Davies, Art as Performance, p. 98 Transposed to philosophy, the work is the philosopher's sustained activity in producing the text — her working through of a problem, her formulating and revising of arguments — and the text is what that activity leaves behind: what our engagement is directed at, but not itself what is evaluated. Why the transposition fails The transposition fails because the feature of art that drives Davies' view has no analogue in philosophy: in art, surface underdetermines the work — a canvas-from-a-washing-machine indistinguishable from a Rembrandt, a Danto-style red square that differs in standing from a perceptually identical other, a molecule-identical forgery — but in philosophy no such underdetermination obtains, because two type-identical papers make the same arguments, face the same objections, and admit the same evaluations. Philosophical evaluation is directed at what the text says: we ask whether premises are defensible, inferences go through, distinctions track real divisions, counter-examples hold against their targets, and conclusions survive the objections the text anticipates — questions whose answers are decidable without knowing who produced the text and that make no reference to any further activity lying behind it. The achievement-talk objection fails at this point: one might protest that philosophers speak of what an author has achieved in a paper, which suggests something beyond the text to be assessed, but achievement-talk in philosophy is parasitic on the text — to say that the author has achieved something is just to say that she has produced a text with such-and-such argumentative properties, not to credit her with any further achievement lying behind them. The text-focused norm is already institutionalised: philosophy journals strip author information from submissions before sending them to referees, and they do so not as practical convenience but as a matter of principle — what is to be assessed is what the paper says, and information about the author is treated as potentially corrupting that assessment. Dellsén et al. (2024) make the corresponding normative point: philosophical progress, they argue, consists in putting people in a position to increase their understanding, and this happens paradigmatically through philosophical content — arguments, theories, distinctions, counter-examples — being made publicly available through publication; philosophical progress is, as they put it, a for-whom rather than by-whom matter, turning on whom the work puts in a position to understand rather than on the cognitive states of whoever produced it. The Sokal affair is recognisable as a breach of this norm: when Social Text published Alan Sokal's 1996 hoax paper without peer review, what went wrong, on the academy's own assessment, was that the journal had assessed Sokal's institutional standing rather than the argument on the page — the recognition of error itself depending on a background norm that the text is what matters. §1 payoff The challenge from authorship is a constitutive challenge: it treats the philosopher's activity not as something that produces the work but as part of what the work is, so that even a text indiscernible from a philosophical paper would fail to be philosophy if no philosopher's activity lay behind it — and the preceding argument rejects this: the work consists in the text, and evaluation goes no further than the text. What remain for §§2–3 are two challenges of a different kind, both capacity challenges: each grants that the philosophical work is in the text, and asks instead whether an LLM has the particular mental capacity required to produce a text with the relevant properties — §2 considers the capacity for abductive reasoning, §3 considers the capacity for conscious experience, and the question in each case is whether a system that lacks the capacity can produce a text of the corresponding kind. §2 — Likeliness, Loveliness, LLMs The charge Floridi and colleagues diagnose LLMs with what they call zeroth-order abduction: an LLM prompted to explain a phenomenon produces a plausible-looking explanation, but it does so by pattern-matching over learned associations rather than by selecting among competing hypotheses — the appearance of abductive reasoning without the reasoning itself. LLMs seem to perform a kind of zeroth-order abduction: given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising the probability of the sequence... The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data.— Floridi et al., "What Kind of Reasoning (if any) is an LLM actually doing?", p. 9 Williamson argues that contemporary philosophical theorising already proceeds partly by abduction from the armchair: when philosophers defend a view they do so by comparing it to rivals — examining how well each, if true, would explain the evidence — and by weighing what he calls the intrinsic virtues of a good theory, roughly simplicity combined with strength (Williamson 2024, pp. 354, 358, 368–69); it is this abductive work that Floridi's diagnosis describes LLMs as unable to perform. The challenge from abduction is accordingly a capacity challenge: good philosophical work requires abductive reasoning; abductive reasoning requires a mental capacity for generating alternative hypotheses and weighing them against one another; humans possess this capacity and LLMs, on Floridi's diagnosis, do not; the objection concludes that an LLM cannot produce philosophical work that genuinely exhibits abductive argument, whatever the surface of its output looks like. The page Floridi himself names the channel by which the patterns of human argument reach the LLM's output — training on philosophical text — and it is worth pressing this point, because what philosophy is conducted in is precisely this text. This effect is due to the model's training on human-generated texts that encode reasoning structures.— Floridi et al., Abstract Philosophy is conducted in writing: a philosopher formulates an argument by writing it down, revises it by reading what she has written, submits the written result to be assessed by other philosophers who respond in further writing — the work of comparing hypotheses, handling objections, and weighing theoretical virtues is done in the prose itself, which is why evaluation consists in reading that prose. An LLM's training data is this same philosophical prose: the articles, the handbook entries, the published replies and counter-replies that together constitute the written record of philosophical practice — so the training channel Floridi identifies is a channel onto the actual object of philosophical evaluation, not a derivative trace of some prior cognitive state. The challenge from abduction therefore does not apply straightforwardly: it required that philosophical quality depend on a producer-level mental process an LLM cannot perform, but the properties philosophical evaluation attends to — the comparisons between hypotheses, the handling of objections, the weighing of intrinsic virtues — are properties of the text itself, and the text is what the LLM's training data contains. Lipton: likeliness and loveliness Lipton distinguishes likeliness from loveliness: the likeliest explanation is the one most probable given the evidence, and the loveliest is the one that, if correct, would provide the most understanding — and the two come apart in principle, because what is probable given the evidence need not be what best explains it. Likeliness speaks of truth; loveliness of potential understanding.— Lipton, Inference to the Best Explanation, Ch. 4 Applied to LLMs, this distinction produces the initial worry: an LLM's continuation is the likeliest in a distribution-theoretic sense — whatever is most probable in the distribution the model has learned — and in an arbitrary corpus (advertising copy, product reviews, news aggregation) likeliness of continuation would tell us nothing about the loveliness of what is said, because there is no particular reason the most probable next sentence should be the one that best explains anything. The philosophical corpus is not such an arbitrary sample: it is the surviving record of philosophers doing inference to the best explanation and ranking candidate theories by the intrinsic virtues — simplicity, strength, accommodation of data — that constitute loveliness; work which fails to meet those standards is less likely to survive in the literature, less likely to be cited, less likely to be assigned to students or included in anthologies. The corpus is also self-evaluating: every article in it was written by a philosopher reading others, and each was in turn read and engaged with by further philosophers writing back — so the patterns persisting in the corpus are patterns of IBE that have been iteratively tested by subsequent IBE, and the patterns surviving that iteration are the patterns later philosophers have found worth taking forward. Lipton argues that loveliness can serve as a guide to likeliness — the lovelier explanation is generally the more probable — and the philosophical corpus makes the reverse relation hold as well: because survival in the corpus has itself been filtered by loveliness-tracking evaluation, what is statistically likely in this corpus is what loveliness-tracking evaluation has accepted, and so an LLM's likeliness-tracking over this corpus approximates loveliness-tracking over its content. An LLM trained on this corpus therefore inherits its distillate: pattern-matching over it is not pattern-matching over a static record of good argument but over the surviving product of many rounds of philosophical comparison and ranking — the patterns the model has learned are patterns that have been selected for by the discipline's own evaluative practice. The mechanism Argumentative structure is present in philosophical prose as surface regularity at the level of discourse — the level of clause-to-clause coherence, the handling of objection and reply, the signalling of hypothesis-comparison — not at the level of isolated word or sentence; and regularities at this level are precisely what next-token prediction is built to detect. Philosophical prose is saturated with discourse markers that encode the comparative work being done in an argument — "however", "the stronger reading is", "one might press the objection that", "consider the cost of denying", "this leaves us with the question of" — and these markers recur with statistical regularity because the work they mark recurs in recognisable forms across philosophers reading and writing back to one another. An LLM trained on well-formed English acquires syntactic norms by statistical exposure to grammatical text; an LLM trained on philosophical prose acquires argumentative norms — the deployment of these markers and the argument-shapes they mark — by the same mechanism operating over more complex structure, because next-token prediction is indifferent to whether the regularity it is detecting is grammatical or discursive. §2 payoff The §2 upshot is that pattern-matching over a philosophical corpus can carry philosophical quality: the corpus is the surviving record of philosophers doing IBE, it has been iteratively shaped by philosophers doing and evaluating further IBE, and the argumentative patterns it contains are of the kind that next-token prediction can learn. What §2 has not answered are two further threads: §3 takes up the capacity challenge from phenomenology, the worry that certain properties of philosophical text can be caused only by a human subject with first-person experience of the kind the text's arguments concern; and §4 returns to the fact that an LLM inherits the distillate of philosophical practice without itself performing the distillation that produced it. §3 — Thought Experiments & Armchair Abduction Framing the challenge Philosophy often draws on conscious experience — on what it is like to see red, to feel time passing, to touch one's own hand — and since an LLM has had no such experience, there is a natural worry that philosophy of this kind is not something it can do. Zahavy's worry Zahavy (2026) argues that some breakthroughs require what he calls manipulative abduction, a form of inference that proceeds by simulating the sensory experience of a situation rather than by manipulating symbols — and his paradigm case is Einstein's formulation of the equivalence principle, arrived at via a thought experiment Zahavy describes in the following terms: Einstein's variation required inventing new axioms based on a physical intuition that did not yet exist in the mathematics. He envisioned a physicist inside an elevator being uniformly accelerated through deep space. Inside this enclosure, the sensory experience reveals a specific pattern: when objects are released, the floor rushes up to meet them. To the physicist, the objects appear to fall with identical acceleration, regardless of composition. Thus, the simulation here was not a permutation of symbols, but a manipulation of perceptual experience.— Zahavy 2026, §5 Zahavy's claim about LLMs is that they, operating by manipulating language, have no access to the perceptual experience on which this kind of abduction depends; they are, in his words, "high-dimensional 'Chinese Rooms' (Harnad, 1990), manipulating the language of physics without access to the physical referents that give that language meaning". It follows, on Zahavy's view, that an LLM trained only on physics text and no other modality could not have run Einstein's reasoning, because that reasoning required the felt experience of acceleration and free-fall — experience an LLM, never having had a body, does not possess. The challenge for philosophy The same worry transfers to philosophy: just as Einstein needed the felt difference between weight-shift and free fall to construct his lift argument, Jackson needed the felt quality of seeing red to construct his Mary argument — Mary has complete physical knowledge of colour but has never seen red, and Jackson argues she learns something new when she first sees red, so framing and evaluating the argument requires a working grip on what seeing red is like, which an LLM without colour experience does not have. And the challenge is not confined to colour experience: the phenomenology of having an intuition (Bengson 2015), the phenomenology of agency, and the feeling of time passing are all phenomenal modes philosophy has used as starting points for arguments, each of which could yield a Mary-type case that an LLM with no phenomenological experience would be blocked from producing. The Pigliucci reframe Our response begins with an observation Pigliucci makes about how the world figures differently in science and philosophy: in science, the world is what claims are tested against — Einstein's equivalence principle needed Eddington's 1919 eclipse observations and the already-observed precession of Mercury's perihelion to confirm it — whereas in philosophy, the world does not serve this verification role but provides the starting points from which philosophers reason. Einstein's case illustrates the distinction: the thought experiment gave Einstein the hypothesis — that acceleration and gravity are equivalent — but because physics is in the business of verification, the hypothesis had still to be empirically confirmed before it could be accepted as physics, whereas a philosophical thought experiment's conclusion does not face a further verification demand of that kind. What philosophy takes from the world, on Pigliucci's account, can therefore be articulated — propositional, describable — rather than raw felt experience, and he puts the point directly: the basic parameters that philosophers use as their inputs, the starting points of their philosophizing, their equivalent of axioms in mathematics and assumptions in logic (or rules in chess) are empirical data about the world. This data comes from both everyday experience (since the time of the pre‐Socratics) and of course increasingly from the world of science itself.— Pigliucci, "Philosophy as the Evocation of Conceptual Landscapes," Ch. 6, p. 79 The challenge accordingly reframes: it is no longer about whether an LLM has had phenomenological experience, but about whether it has access to articulated descriptions of that experience — and since articulated descriptions are exactly what text corpora are made of, the question becomes one the LLM has a real chance of meeting. The reply Sub-disciplines of philosophy work constantly with articulated phenomenology — philosophy of perception on how things look, philosophy of temporal experience on how time is felt, aesthetics on the character of aesthetic response, philosophy of emotion on what emotions are like — and these produce careful descriptions of phenomenal experience that every LLM trained on academic text has processed, so the corpus is saturated with phenomenological description as a demonstrable fact about what the training data contains. Literary writing adds to this: Proust on the involuntary memory triggered by the madeleine, Virginia Woolf on the texture of ordinary thought in Mrs Dalloway, Henry James on the shading of social perception, Nabokov on the visual particulars of rooms and streets — all part of the same training data. Running the reply through cases An LLM has never felt weight-shift or free fall, but on Pigliucci's reframe what philosophical reasoning takes as its input is what the world has had articulated about it, and ordinary English is thick with articulated descriptions of these sensations: every elevator passage in fiction, every description of a roller-coaster drop, every astronaut's memoir of zero-gravity — so an LLM trained on ordinary text has the input that Einstein-style reasoning about weight-shift requires. An anecdatum of mine: ask a current frontier LLM about design and colour — palettes, complementary colours, contrast and harmony — and it will discuss the phenomenology of colour with a sophistication indistinguishable from a knowledgeable human speaker, which is evidence that the corpus has given it enough of a working grip on colour phenomenology to reason about it. Wherever philosophical work proceeds from phenomenological starting points that have already been articulated somewhere in text, those starting points are in the corpus an LLM is trained on, and the challenge — that LLMs lack phenomenological experience — is met by the fact that articulated phenomenology is what they have in abundance. The limit There is nonetheless a narrower case the reply does not cover: a philosopher sometimes articulates a previously undiscovered aspect of phenomenology — a structural feature of conscious experience that, though present in everyone's lives, has not been explicitly described; this requires sustained first-person attention to one's own experience, and no recombination of existing text can substitute for that attention. The paradigm is Merleau-Ponty's observation about self-touch: when one fingertip touches another, one finger plays the role of toucher and the other of touched, and the two can reverse roles — but they cannot simultaneously both be toucher, so one's own body is always, at any instant, split between the touching and the touched. An LLM could not have originated Merleau-Ponty's observation, since there was nothing in prior text to recombine into that insight, and the same holds for the general case: an LLM cannot originate a phenomenological description that is not already, in some form, in its training data — whatever has never been articulated lies outside what the model can produce. §3 payoff The §3 upshot is that the corpus supplies the phenomenological inputs philosophical reasoning takes from the world, wherever those inputs have been articulated in text — which, across most of the discipline (ethics, metaphysics, epistemology, philosophy of mind, philosophy of language), they have. What §3 leaves to §4 is the further question raised by the persistence of underwhelming LLM output: if the capacity §§2–3 have defended is really present, why are we not seeing its products — and what, if anything, is the contemporary philosopher's work in closing that gap. §4 — A Speculative Coda The puzzle §§1–3 have argued that the in-principle obstacles to LLM-produced philosophy do not hold up: §1 rejected the constitutive challenge, and §§2 and §3 answered the capacity challenges from abductive reasoning and from conscious experience — but an obvious observation presses against these conclusions, namely that we are not in fact seeing LLM-produced philosophy worth reading, at least not at the rate one might expect if the capacity were genuinely there, and the remainder of this section floats some ideas about what might lie in the gap. What the prompt asks for The gap is, on the simplest diagnosis, a matter of what the LLM is being asked for: prompted generically — "write on free will", "explain the Mary argument" — an LLM produces the kind of text most probable in the distribution it has learned, which for topics of philosophical interest is summary text, survey-style, hedged, balanced, non-committal, because the bulk of text written at that level of generality about philosophical topics takes that form. A paper worth reading is not summary but distinctive argument for a specific conclusion against specific alternatives: a problem is chosen, a position is taken on it, and the position is pressed against its strongest opposition — and when the prompt asks for something that has this structure, the LLM has an articulated target to write toward; when the prompt asks for something more general, what it produces is correspondingly more general, which is not what the practice of philosophy treats as worth reading. Distillate without distillation §2 argued that the philosophical corpus is the distillate of many rounds of philosophical criticism — arguments tested by later arguments, with the patterns that persist in the corpus being the ones later philosophers have taken seriously enough to engage with — and that an LLM trained on the corpus inherits this distillate as pattern. Inheriting the distillate, however, is not the same as performing the distillation: the iterated criticism that produced the corpus's patterns operates across time and across many minds, and a single completion by an LLM does not reproduce it — the model has the products of iteration, not the iteration itself. The philosopher prompting the LLM can perform the iteration in place of the discipline: drafting, pressing the draft against the strongest objection available, redrafting in light of that pressure — and so supplies within the session some portion of the process that the model has absorbed only in its output. Generalist, not specialist A candidate response to all this would be a specialist LLM trained narrowly on philosophical text, but Sellars's characterisation of what philosophy aims at runs in the opposite direction: The aim of philosophy, abstractly formulated, is to understand how things in the broadest possible sense of the term hang together in the broadest possible sense of the term. Under 'things in the broadest possible sense' I include such radically different items as not only 'cloth, ships, and sealing-wax,' but numbers, duties, possibilities, finger snaps, aesthetic experience, and death.— Sellars, "Philosophy and the Scientific Image of Man" (1962) If something like Sellars's characterisation is right, training breadth is not a distraction from philosophy but close to its proper substrate, and the philosophy worth reading from LLMs is more likely to come from generalist models differently prompted than from specialist philosopher-models differently trained — a claim that runs ahead of the evidence currently available to assess it. §4 payoff Putting §§1–4 together: there is no in-principle obstacle to LLM-produced philosophy worth reading, and what is practically required for such philosophy to be produced is what the discipline has always required of its authors — a specific problem taken up, a position committed to, and iteration across objections until the position holds — work which, increasingly, gets done in the prompt rather than in the drafting of a paper. You’re right to reject the softened version. The strong claim should not be replaced by a human–LLM teamwork thesis; the response has to show why the model itself can produce philosophy worth reading, given the challenge’s own terms. You are right. The strong claim should be preserved: *Current LLMs can produce philosophy worth reading.* Not merely: humans can use LLMs as tools. Not merely: LLMs can supply material that humans then turn into philosophy. Not merely: LLMs can help philosophers think. The target claim is that the *text produced by the model* can itself have the properties that make philosophy worth reading. The human may prompt, read, judge, cite, reject, publish, or ignore it, but that is not what makes the text philosophically valuable. ## ==1\. The “token-level patterns” issue== ==The worry I summarized was this:== > ==LLMs are trained by predicting tokens. Therefore, maybe they only learn superficial token regularities, not argumentative structure.== ==That worry is too crude. It confuses the *training objective* with the *kind of structure learned in order to satisfy that objective*.== ==A model may be trained to predict the next token, but the best way to predict the next token in sophisticated prose is not merely to memorize word associations. It is to learn higher-order regularities: syntax, semantic dependencies, discourse structure, genre, argumentative transitions, objection–reply patterns, explanatory schemas, and so on. The training signal is token-level; the learned structure need not be token-level. That is the key distinction.== ==A simple analogy: a model trained to predict the next move in chess notation is trained on a sequence-prediction task, but good performance requires learning something about chess structure: legal moves, openings, threats, positional patterns. Likewise, a model trained to predict philosophical prose is trained on token sequences, but good performance over that corpus requires sensitivity to the structures that govern philosophical prose: “here is an objection,” “this distinction blocks that objection,” “this view is stronger than that rival,” “this consequence is unacceptable,” “this explanation is more unified,” and so on.== This is why the earlier “it only learns token-level co-occurrence” objection was weak. It tacitly presupposed a thin “stochastic parrot” picture. That phrase is associated with Bender et al.’s critique of large language models as systems that generate fluent language through statistical patterning without genuine understanding. [Dr Alan D. Thompson – LifeArchitect.ai](https://s10251.pcdn.co/pdf/2021-bender-parrots.pdf?utm_source=chatgpt.com) But even if one accepts many concerns from that critique, it does not follow that language statistics are shallow. Work on transformer circuits, for example, gives evidence that models can develop mechanisms such as induction heads that support in-context pattern completion, not merely local word association. [transformer-circuits.pub](https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html?utm_source=chatgpt.com) And broader LLM surveys now treat large models as systems that acquire a wide range of linguistic and task-relevant capabilities from large-scale training, even though their base objective is still prediction. [arXiv](https://arxiv.org/html/2402.06196v2?utm_source=chatgpt.com) ==So the right point is not:== > ==LLMs learn only surface token patterns.== ==The right point is:== > ==LLMs learn whatever structures are statistically useful for predicting text in the training distribution; in philosophical prose, many of those structures are argumentative, dialectical, and explanatory.== ==That matters for your paper because Floridi’s challenge says: the model does not perform genuine abduction; it produces plausible continuations based on learned associations. Your response should say: even granting that, the learned associations are not merely verbal tics. In a philosophical corpus, the predictive regularities include the forms of abductive philosophical argument themselves.== ## 2\. The “mannerism” worry is different The “mannerism” point was not the same as the token-level worry. The mannerism worry says: yes, philosophical prose contains good argumentative structures, but it also contains a lot of stylistic fluff: hedging, throat-clearing, balance-for-the-sake-of-balance, “one might object,” “it is beyond the scope of this paper,” “this raises important questions,” and so on. An LLM can learn those patterns too. So it may produce text that *sounds* like philosophy without being worth reading. That is a real practical worry, but it is not a deep objection to the strong claim. It shows that LLMs can produce bad or mediocre philosophy-looking prose. But the target claim is not that every output is good. The target claim is that current LLMs *can* produce philosophy worth reading. So the correct role for the mannerism worry is limited. It belongs in §4, when you explain why so much LLM output is flat, survey-like, hedged, and uninteresting. It does not defeat §2. In §2, the issue is whether the model can produce abductively structured philosophy at all. The answer can remain strong: yes, because argumentative and explanatory structures are in the corpus and can be learned statistically. ## 3\. The human–LLM split should not be the main response Your correction is exactly right. Saying “the human supplies the real philosophical judgment” weakens the paper into a familiar tool-use thesis. It changes the question. Your question is: > Can LLMs produce philosophy worth reading? Not: > Can philosophers use LLMs to help themselves produce philosophy worth reading? The previous softened version smuggled the human back in as the guarantor of philosophical value. That is not needed, and it is dialectically bad for your project. It concedes too much to the capacity challenge. The cleaner distinction is this: > Human judgment may be needed for *us to know* whether the text is worth reading. It is not needed for the text *to be* worth reading. That is the key. A human reader may be needed to identify, assess, or appreciate the output. But that is also true of a human-authored paper. The paper’s being worth reading is not constituted by the reader’s judgment. It is constituted by features of the text: the argument, the distinction, the objection, the explanation, the insight, the conceptual reorganization. If an LLM-produced text has those features, it is worth reading. So you can reject the lame split without denying that prompts and readers matter in practice. The correct formulation is: > Prompting and evaluation are epistemic and practical conditions for *finding* LLM-produced philosophy worth reading, not metaphysical or constitutive conditions for its *being* philosophy worth reading. That preserves the strong claim. ## ==4\. The response to the abduction challenge should be product-centered== ==The challenge says:== 1. ==Philosophy often requires abduction.== 2. ==LLMs do not perform abduction.== 3. ==Therefore, LLMs cannot produce philosophy worth reading.== ==The response should deny the inference from 2 to 3.== ==The missing premise is:== > ==A text can exhibit good abductive reasoning only if it was produced by a subject who performed abduction.== That premise is false, or at least undefended. Your §1 already gives you the resources to reject it. In philosophy, what we evaluate is the text: whether it makes a good distinction, handles objections, compares rivals, explains the data, draws a conclusion worth taking seriously. The abductive value has to be *in the text*, because that is what other philosophers read, assess, criticize, cite, and teach. So the response to Floridi should be: > Even if LLMs do not perform abduction as a mental act, they can produce texts that instantiate abductive structure. And for the question whether a philosophical text is worth reading, the relevant issue is not whether the producer performed abduction, but whether the text exhibits the properties of good abductive argument. This is much stronger than saying “the human reader supplies the abduction.” The text itself can contain the abductive comparison. It can say: here are the rival hypotheses; here is what each explains; here is why one handles the data better; here is the objection to that; here is why the objection fails. If the output does that well, it is not merely raw material for philosophy. It is philosophy worth reading. ## 5\. Lipton’s role should be sharpened Lipton should not be used merely to say: > Loveliness tracks likeliness, and the corpus makes statistical likelihood track loveliness. That is too compressed. The stronger Liptonian structure is this: First, Lipton distinguishes *likeliness* from *loveliness*. Likeliness concerns warrant or probability; loveliness concerns how much understanding an explanation would provide if true. Lipton thinks the interesting version of IBE must appeal to loveliness, because inference to the merely likeliest explanation is too close to trivial: it does not explain what makes something likely. Second, Lipton’s Chapter 8 separates identification, matching, and guiding. The explanatory virtues must be identified; then one must show that they match inferential virtues; then one can ask whether inquirers actually use loveliness as a guide to likeliness. ==Third, for your argument, the most important point is not that the LLM itself uses loveliness as a conscious guide. It does not need to. The point is that the philosophical corpus has already been shaped by practices in which loveliness and likeliness are repeatedly brought into contact.== ==So the central move becomes:== > ==The LLM does not need to perform Liptonian abduction internally. It needs to learn from a corpus whose statistical structure has been shaped by Liptonian abduction.== ==That is the heart of the response.== ## 6\. Stronger reconstructed §2 response Here is the structure I think you want after the challenge has been presented. I am leaving the setup of the challenge alone, as requested. ### Move 1 — Concede Floridi, but deny the consequence Grant Floridi’s diagnosis for the sake of argument. LLMs do not perform abduction in the human sense. They do not consciously generate rival hypotheses, assess them as explanations, and infer the best one. They generate continuations from a learned distribution. But the conclusion does not follow. The question is not whether the model performs abductive reasoning as a mental process. The question is whether it can produce a philosophical text that exhibits the properties of good abductive reasoning. ### Move 2 — The missing premise is false The objection requires a bridge principle: > A text can have the properties of good abductive reasoning only if it was produced by a subject who performed abduction. That principle is exactly what should be rejected. A philosophical text can contain a comparison between rival hypotheses, an evaluation of their explanatory virtues, and a conclusion that one better explains the relevant data. If the text does this well, it exhibits abductive philosophical structure. Nothing in the concept of a text’s having that structure entails that the producer must have arrived at it through the same structure. ### Move 3 — Philosophy makes abductive reasoning public in prose In philosophy, abductive work is not hidden behind the text in a way that leaves only a residue. It is performed and displayed in the prose. Philosophers write: “this view explains,” “the rival cannot account for,” “the simpler hypothesis is,” “the stronger explanation is,” “this objection fails because,” “the cost of denying this is,” and so on. These are not merely decorations. They are the public form in which philosophical abduction becomes assessable. This matters because LLMs are trained on the public record of that activity. ### Move 4 — Floridi’s own formulation gives you the channel Floridi says LLMs produce plausible explanations because they are trained on human texts that encode reasoning structures. That should be pressed, not conceded weakly. If the corpus encodes reasoning structures, and if philosophy is a textual discipline in which abductive comparisons are encoded in the prose, then the model’s training data contains precisely the structures the challenge says it lacks. ### Move 5 — Lipton explains what those structures are Lipton gives you a vocabulary for the structures that matter. Good abductive philosophy does not merely produce a probable conclusion. It compares potential explanations in terms of explanatory virtues: unification, simplicity, scope, mechanism, fertility, fit with background belief, contrastive adequacy. These are the dimensions of loveliness. So the question becomes: > Can a model trained on philosophical prose learn patterns associated with these explanatory virtues? The answer should be: yes, because these virtues are not private mental glow. They are textually manifested in the way philosophical arguments are written. ### Move 6 — The corpus is not arbitrary text In an arbitrary corpus, statistical likelihood does not track philosophical value. The most probable continuation in advertising copy, social media posts, or product reviews need not be lovely, explanatory, or philosophically interesting. But the philosophical corpus is different. It is the residue of repeated philosophical selection: arguments are proposed, criticized, refined, cited, taught, ignored, revived, or abandoned. The corpus is therefore not merely a heap of philosophy-themed sentences. It is an archive of argumentative survival. This is where your “distillate” idea belongs, but it should be made stronger: the corpus is not only a store of philosophical claims; it is a store of *tested argumentative forms*. ### Move 7 — LLM training over this corpus can internalize abductive form A model trained on this corpus does not merely learn that philosophers often say “therefore” or “one might object.” It can learn the deeper regularities governing when those moves are appropriate: what tends to follow an objection, what kind of distinction blocks what kind of counterexample, what kind of explanatory burden arises after a given claim, what sort of dialectical move makes a conclusion stronger rather than weaker. This is the point about token-level prediction. The model predicts tokens, but to predict philosophical tokens well it must encode discourse-level and argument-level structure. The objective is local; the learned dependencies can be global. ### Move 8 — Therefore zeroth-order abduction can still produce abductive texts Floridi’s “zeroth-order abduction” is meant to deflate the model’s output: the model only produces the typical linguistic form of explanation. But in philosophy, the “typical linguistic form of explanation” is not separable from the argumentative structure of explanation in the way Floridi needs. The form includes the actual public machinery of the argument: rival hypotheses, objections, distinctions, explanatory comparisons, costs, consequences, and theoretical virtues. So even if the process is zeroth-order, the product need not be empty mimicry. It can instantiate the same abductive properties that make a human philosophical text worth reading. ### Move 9 — The strong conclusion The correct conclusion is not: > LLMs help humans produce abductive philosophy. It is: > LLMs can produce philosophical texts that themselves instantiate good abductive reasoning, because the abductive structures of philosophy are publicly encoded in the philosophical corpus, and current models can learn and generate discourse-level structures from such corpora. That directly answers the challenge. ## 7\. Better wording for the key claim Your current line says something like: > Statistical likelihood ends up kind of approximating loveliness. That is the right idea, but too casual and too easy to attack. I would replace it with something like this: > In an arbitrary corpus, statistical likelihood has no reason to track explanatory loveliness. But the philosophical corpus is not arbitrary. It is the surviving record of a practice in which arguments are repeatedly assessed by explanatory and theoretical virtues. Over such a corpus, statistical regularities are not merely verbal regularities; they are the residue of repeated philosophical selection. An LLM trained on that corpus can therefore generate continuations whose probability reflects, however imperfectly, the abductive patterns that the discipline has preserved. Even stronger: > The model does not have to perform abduction in order to produce abductively structured philosophy. It has learned from a corpus in which abductive structure has already been made public, stabilized, criticized, and selected. Its output can therefore inherit abductive form without reproducing the abductive mental process that originally generated the corpus. That is probably the core formulation. ## 8\. The strongest version of the answer to Floridi I would put the answer like this: > Floridi is right that an LLM does not reason from scratch about causes in the way a human reasoner does. But that does not show that its output cannot contain good abductive reasoning. Philosophy is a textual practice in which abductive reasoning is externalized. The weighing of explanations, the comparison of rival hypotheses, and the assessment of theoretical virtues are carried by the prose itself. Since LLMs are trained on this prose, they are trained on the public form of philosophical abduction. Their statistical task is therefore not merely to mimic the sound of explanation, but to continue patterns in a corpus whose patterns include the discipline’s abductive structures. If those structures are learned and regenerated, the resulting text can be philosophy worth reading, even though the process that produced it is not human abduction. That keeps the claim strong. ## 9\. Where §4 should then go §4 should not say: “the human supplies the real philosophical work.” It should say something more like: > The fact that LLMs can produce philosophy worth reading does not mean that generic prompts will reliably elicit it. Generic prompts activate generic regions of the learned distribution: summaries, surveys, encyclopedic balance, undergraduate exposition. The problem is not lack of capacity but target specification. To get worthwhile philosophy, one must ask for the kind of thing worthwhile philosophy is: a specific problem, a committed thesis, a serious rival, a pressure point, a distinction, an objection, a reply, a theoretical payoff. That preserves the strong claim. The prompter is not the philosopher who turns non-philosophy into philosophy. The prompt is an elicitation condition. The model has the capacity; the prompt determines whether that capacity is activated. A useful analogy: a human philosopher asked “say something about free will” may produce banalities. Asked “defend a source-incompatibilist account of moral responsibility against manipulation arguments without relying on PAP,” they are much more likely to produce philosophy. The difference is not that the second prompt makes the human capable. It gives the capacity a target. Same with LLMs. ## 10\. The revised §2 payoff Your current payoff says: > Pattern-matching over a philosophical corpus can carry philosophical quality. I would strengthen it: > The challenge from abduction fails because it mistakes a claim about the model’s production process for a claim about the philosophical properties of its product. Even if LLMs do not perform abduction, they are trained on a corpus in which philosophical abduction is publicly encoded and iteratively selected. Since current models can learn discourse-level and argumentative regularities from text, they can generate texts that instantiate abductive philosophical structure. Such texts are not merely aids to philosophy, nor merely simulations of philosophical prose. When they make distinctions, compare hypotheses, handle objections, and illuminate a problem in ways that reward philosophical attention, they are philosophy worth reading. That is the strong version. that was truly excellent. While I have you here, can I talk to you a little bit about section one as well? And your suggestions as to how to improve it? Again, if you can and do it in terms of moves rather than anything else that would be great. And just so we're clear, pull no punches, okay? I am interested in the CEV of this section now. And so I would like you to do whatever you think is necessary in terms of real arrangement or restructuring of the arguments, whatever you want, in pursuit of the CEV of this section. Yes. I’ll treat §1 as a live argumentative machine, not as prose to polish. I’ll focus on what the section must do for the paper: block the authorship challenge without weakening the later capacity-claim. Yes — §1 can be made much stronger. The current version has the right instinct, but it is doing too many things at once, and some of the most promising moves are buried under the Davies/art analogy. The CEV of the section, I think, is not simply: > Philosophy is text-focused, so authorship does not matter. That is too blunt and too vulnerable. The stronger version is: > The authorship challenge trades on an illicit slide from *philosophy as an activity* to *philosophical value as a property of texts*. Even if LLMs do not engage in the activity of philosophising, it does not follow that their outputs cannot be philosophical texts worth reading. To get that conclusion, the challenger needs a product-dependence principle: that a text can have the relevant philosophical value only if it was produced by the right kind of philosophical activity. That principle is plausible in some arts, but it is false, or at least unsupported, in philosophy. That should be the spine of §1. ## 1\. What is currently working The current §1 already has several good ingredients. First, the challenge is well chosen. The “philosophy is a person-only domain” intuition is real. Many philosophers do have a kind of reflexive contempt response: “Obviously ChatGPT cannot do philosophy.” That is worth taking seriously. Second, the Davies analogy is useful because it gives the intuition a serious form. Without Davies, the objection risks sounding like mere snobbery. With Davies, it becomes a metaphysical claim: perhaps a philosophical work is not merely a text, but the intentional activity that eventuates in the text. Third, your distinction between art and philosophy is basically right. Artworks often have histories of production built into their identity and value. Philosophical arguments, by contrast, seem much more tightly tied to what the text says. Fourth, the blind review point is good as practice-level corroboration. Philosophy does often treat the author’s identity as irrelevant, or at least potentially corrupting, when assessing a submission. But the section, as currently framed, is not yet optimal. ## 2\. The biggest weakness: the section risks answering the wrong question The current section sometimes sounds as if it is trying to prove: > LLM outputs are works of philosophy in the same sense as human-authored philosophy papers. But the paper’s real question is: > Can LLMs produce philosophy worth reading? Those are not identical. You do not need to settle every metaphysical question about authorship, agency, responsibility, publication, or work-identity. You need to defeat the specific inference: > No philosopher produced this text; therefore the text cannot be philosophy worth reading. That inference is invalid unless we accept a very strong dependence principle. So §1 should not try to give a full ontology of philosophy. It should expose the missing premise. The missing premise is: > A text can be worth reading as philosophy only if it was produced by a person engaged in philosophical activity. That is the real target. Name it. Then attack it. ## 3\. Second weakness: “the work consists in the text” is too strong In the moves map, the payoff says: > the work consists in the text, and evaluation goes no further than the text. I would soften *that specific formulation*, not the main thesis. This is not a retreat. It is a precision improvement. The phrase “the work consists in the text” may invite unnecessary metaphysical objections. Someone might say: a philosophical work also includes publication context, uptake, intertextual references, intended audience, historical position, etc. And they would not be completely wrong. A type-identical paper published in 1900 and 2026 may differ in originality, significance, and dialectical force because the surrounding literature differs. So the better claim is not: > the work is just the text. The better claim is: > the philosophical value relevant to worth-readingness supervenes on publicly assessable content in context, not on the producer’s private mental activity. That allows context to matter. It allows originality to matter. It allows dialectical location to matter. But it blocks the authorship challenge, because none of those things requires a human author’s mental process to be part of the work. So the section should shift from *text-only* to *public-content-plus-context*. That is more robust. ## 4\. Third weakness: the art analogy currently gets too much space The Davies material is helpful, but it should not become the engine of the section. At the moment, there is a danger that the reader thinks the argument is: > Davies is right about art; philosophy is unlike art; therefore LLM text can be philosophy. That is not quite enough. The better use of Davies is dialectical: > Here is the strongest way to make the authorship challenge look respectable. In some domains, product value really is production-history-dependent. Art may be one such domain. But philosophy is not obviously one of them, and the challenger owes an argument for transferring that model. So Davies should be used to *raise the bar* for the objector, not to structure the whole response. The move should be: > The art analogy shows what the authorship challenge would need: a reason to think philosophical value is production-history-dependent. But the analogy breaks precisely where philosophy differs from art: philosophy evaluates inferential and conceptual content, not the intentional performance behind the product. That is cleaner. ## 5\. Fourth weakness: the random-text thought experiment needs protection You say: imagine your favorite philosophy paper somehow fell together randomly; it would still be philosophically valuable. This is a good thought experiment, but some philosophers will object: > If it was randomly produced, no one asserted anything. So perhaps it is not an argument, but merely marks resembling an argument. You need to pre-empt this. The reply is: even if there is no act of assertion, the sequence can still express propositions in a public language. A text can be meaningful because it is interpretable under linguistic conventions, not because someone sincerely asserted it. We can assess the validity of an argument written on a blackboard even if we do not know who wrote it, whether they believed it, or whether it was copied by a machine. The absence of assertion may matter for responsibility, sincerity, or authorship, but not for whether the argument is valid, illuminating, or worth reading. So make the random-text case more precise: > If a sequence of English sentences expresses a valid, original, illuminating argument, then the argument is available for philosophical assessment. Whether the sequence was produced by a person, a machine, or an absurd accident may affect credit and responsibility, but it does not erase the inferential relations expressed by the sentences. That is much harder to resist. ## 6\. Fifth weakness: blind review and Sokal should be demoted The blind review point is useful, but it should not carry much argumentative weight. Blind review shows that philosophy has a norm of author-independent assessment. But it does not prove that authorship is irrelevant to the ontology of philosophical works. Use it as corroboration, not as a foundation. The Sokal point is more fragile. It risks distracting the audience. Someone could say that the Sokal affair shows something about editorial failure, ideology, peer review, hoaxing, or intellectual standards, but not directly about whether authorship matters to philosophical value. I would either cut it or relegate it to a footnote/aside. If kept, the Sokal point should be reframed: > Sokal is useful only as evidence that institutional standing can corrupt evaluation when it substitutes for assessment of argumentative content. But I would not make it a main move. ## 7\. The best restructuring: make §1 an argument against product-dependence Here is the structure I would recommend. ### Move 1 — State the authorship challenge as a constitutive challenge The first challenge is not that LLMs produce bad philosophy. It is that they cannot produce philosophy at all. On this view, philosophy is a person-only domain: unless a philosopher stands behind the text, the output may resemble philosophy, but it is not philosophy. This is important because it is not yet a capacity challenge. It does not say the model lacks abduction, phenomenology, originality, or understanding. It says that even a perfect-looking output would fail because of its source. ### Move 2 — Separate three questions The challenge immediately trades on an ambiguity. There are at least three questions: 1. Can an LLM *engage in the activity* of philosophising? 2. Can an LLM output be *authored philosophy*, in the ordinary responsibility-and-credit sense? 3. Can an LLM produce a *philosophical text worth reading*? The paper is concerned with the third. The authorship challenge moves from a negative answer to the first or second question to a negative answer to the third. That move needs an argument. This is a crucial move. It prevents the discussion from being hijacked by “but the LLM is not really doing philosophy.” You can say: fine. That is not yet the issue. ### Move 3 — Identify the missing premise To make the challenge valid, the objector needs a product-dependence principle: > A text can be worth reading as philosophy only if it is produced by a person engaged in philosophical activity. Without this premise, the challenge collapses. From “no philosopher performed the relevant activity” it does not follow that “the text has no philosophical value.” This move is the heart of the section. Once named, the product-dependence principle looks much less obvious. ### Move 4 — Use Davies/art to show why the premise can look tempting The objector can motivate product-dependence by analogy with art. Davies’ performance theory gives the strongest version: an artwork is not merely the object left behind, but the intentionally guided generative performance that eventuates in it. A Rembrandt-looking canvas produced by a washing machine would not be a Rembrandt, and perhaps not an artwork of the same kind at all. This makes product-dependence look respectable. In some domains, the product’s identity and value really do depend on the history of production. ### Move 5 — Explain why the analogy fails for philosophy The analogy fails because philosophical evaluation is directed primarily at semantic-inferential content. We ask whether the claims are clear, whether the premises are defensible, whether the distinction tracks something real, whether the objection lands, whether the reply works, whether the view explains more than its rivals, whether the argument changes how we should think. Those are properties of publicly available content in dialectical context. They are not properties of the hidden activity by which the text was produced. This should be one of the cleanest sentences in the section: > In philosophy, unlike in the relevant art cases, the route to the product is not normally part of what makes the product worth reading. ### Move 6 — Use the duplicate/random/proof case Now introduce the thought experiment. Suppose a sequence of sentences, by whatever causal route, expresses a novel and powerful argument against physicalism, or a decisive objection to a theory of depiction, or a distinction that reorganizes a debate. If the argument is there on the page, philosophers can read it, assess it, object to it, cite it, and learn from it. Its origin may affect credit, responsibility, and publication ethics, but it does not make the argument disappear. A useful analogy here is mathematical proof. If a proof appears on a blackboard and no one knows who wrote it, we can still check whether it proves the theorem. If a machine generates a proof, the proof is not invalid merely because the machine did not understand it. Philosophy is not identical to mathematics, but with respect to this issue it is closer to proof than to painting. This is much stronger than the current Rembrandt comparison. ### Move 7 — Handle the “but no one asserted it” objection Someone may object that a random or machine-generated sequence is not really an argument because no one asserted its premises or conclusion. Reply: assertion is not necessary for philosophical worth-readingness. We routinely evaluate arguments in fiction, dialogues, reconstructions, reductios, anonymous fragments, classroom examples, and papers whose authorial psychology is unknown. What matters is that the text expresses propositions and inferential relations. Sincerity and responsibility belong to the pragmatics of assertion; worth-readingness belongs to the philosophical content. This move is essential. Without it, the random-text argument is exposed. ### Move 8 — Distinguish value from credit, responsibility, and publication norms The authorship challenge often gains force by mixing together different concerns: > Who deserves credit? > Who is responsible for errors? > Is publication permissible? > Is the text plagiaristic? > Is the text philosophically valuable? These are different questions. A plagiarized paper can contain an excellent argument. A ghostwritten paper can be philosophically interesting. An anonymous text can be worth reading. A text can be unpublishable for ethical reasons while still containing philosophy of value. So even if LLM outputs raise serious questions about credit, copyright, and responsibility, none of that establishes that they cannot be philosophy worth reading. This move will also help you later in Q&A, because people will inevitably jump from “can” to “should.” ### Move 9 — Use disciplinary practice as corroboration Now bring in blind review, but in a reduced role. Philosophy’s own evaluative practice supports this product-centered view. Anonymous review is not merely an administrative trick; it reflects the norm that a paper should be assessed by what it says, not by who wrote it. That does not prove that authorship never matters, but it shows that the discipline already treats philosophical merit as assessable independently of authorial identity. You can also mention teaching: we often present arguments detached from their full production histories. Students assess the argument, not the psychology of the person who first wrote it. ### Move 10 — Clarify what §1 does and does not show §1 does not show that LLM outputs are good. It does not show that LLMs can reason abductively. It does not show that LLMs can handle phenomenology. It does not settle publication ethics. It shows something narrower but crucial: > There is no constitutive barrier from authorship. If an LLM output has the relevant argumentative, conceptual, and explanatory properties, then it can be philosophy worth reading, regardless of whether a philosopher performed the activity that produced it. This prepares §§2–3 perfectly. Once authorship is blocked, the only remaining serious objections are capacity objections. ## 8\. A full revised move list for §1 Here is the CEV version of the section as a moves list. ### §1 — The Challenge from Authorship **Move 1 — The challenge** Some philosophers seem to treat philosophy as a person-only domain. On this view, no LLM output can be philosophy, however impressive it looks, because no philosopher stands behind it. The output may simulate philosophy, but it cannot be philosophy. **Move 2 — Why this is a constitutive challenge** This is not yet the claim that LLMs lack some capacity needed for good philosophical work. It is stronger and more basic: even if the output were indistinguishable from an excellent paper, it would fail to be philosophy because of how it was produced. **Move 3 — The ambiguity** The challenge trades on an ambiguity between three claims: that LLMs do not engage in the activity of philosophising; that LLM outputs are not authored philosophical works in the ordinary sense of responsibility and credit; and that LLMs cannot produce philosophical texts worth reading. The first two may be true without implying the third. **Move 4 — The missing premise** To get from “no philosopher produced this” to “this cannot be philosophy worth reading,” the objector needs a product-dependence principle: a text can be worth reading as philosophy only if it results from a person’s philosophical activity. That principle is the real target of §1. **Move 5 — Why product-dependence looks tempting** The art case makes product-dependence look plausible. On Davies’ performance theory, the artwork is not merely the object left behind, but the intentionally guided generative performance that eventuates in it. A Rembrandt-looking canvas produced by accident or machinery would not have the same artistic standing as a Rembrandt. **Move 6 — Why philosophy is different** Philosophical evaluation does not normally work like this. We evaluate what the text says: whether the premises are defensible, whether the inference is valid or strong, whether a distinction tracks a real difference, whether an objection lands, whether the reply succeeds, whether the view explains more than its rivals. These are publicly assessable features of content in context, not hidden features of the author’s activity. **Move 7 — The duplicate argument** If a sequence of sentences expresses a powerful philosophical argument, the argument is available for assessment whatever its causal origin. If the same text had been produced by a philosopher, a machine, or an extraordinary accident, the inferential relations expressed by the text would be the same. The origin may affect credit or responsibility, but not whether the argument is worth reading. **Move 8 — The proof analogy** In this respect, philosophy is closer to proof than to painting. A proof generated by a machine may raise questions about understanding, credit, or trust, but it is not invalid because the machine did not understand it. Likewise, a philosophical argument generated by an LLM is not deprived of philosophical value merely because the model did not perform the human activity of philosophising. **Move 9 — The assertion objection** One might object that a machine-generated text is not really an argument because no one asserts it. But assertion is not required for philosophical worth-readingness. We can evaluate arguments in dialogues, fictional cases, anonymous manuscripts, reconstructions, reductios, and examples whose authorship is unknown. What matters is that the text expresses claims and inferential relations that can be philosophically assessed. **Move 10 — The responsibility objection** Questions about authorship, credit, copyright, plagiarism, and responsibility are genuine, but they are not the present question. A text may be ethically unpublishable under someone’s name and still contain a philosophically interesting argument. Publication norms determine what we should do with a text; they do not determine whether the text has philosophical content worth reading. **Move 11 — Disciplinary corroboration** Philosophy’s own practices support this product-centered norm. Anonymous review asks referees to assess a paper by what it says rather than by who wrote it. This does not prove that authorship is irrelevant to every question, but it confirms that philosophical merit is treated as assessable independently of authorial identity. **Move 12 — §1 payoff** The authorship challenge fails. It may show that LLMs do not philosophise as persons do, and it may raise hard questions about responsibility and publication, but it does not show that their outputs cannot be philosophy worth reading. If the relevant philosophical properties are present in the text, the absence of a human philosopher behind the text does not remove them. **Move 13 — Transition** What remains are genuine capacity challenges. Even if authorship is not constitutive, perhaps only a system with abductive reasoning can produce abductively good philosophy, or only a system with conscious experience can produce phenomenology-based philosophy. Those are the challenges addressed in §§2–3. ## 9\. What I would cut or reduce I would reduce the following. First, the course-on-Kant/course-on-Turner comparison. It is not doing enough. It may even hurt, because philosophy is indeed person-organized in teaching and reception. The point can be made more cleanly elsewhere. Second, the Sokal example. It is colorful but distracts from the structure. It belongs, at most, in a footnote about institutional standing corrupting assessment. Third, the strong “work consists in the text” language. Replace it with “the relevant philosophical value is publicly assessable in the text and its dialectical context.” Fourth, any suggestion that Davies is simply wrong. You do not need to criticize Davies. You need to say: even if Davies is right about art, the transposition to philosophy fails. ## 10\. What I would add I would add three things. First, the explicit three-way distinction between activity, authorship, and worth-reading text. This is the most important addition. Second, the product-dependence principle. Name it, because once named it becomes the burden of the objector. Third, the assertion objection. Philosophers will press it. Better to handle it before they do. ## 11\. The punchiest version of the section’s thesis If you wanted the argumentative core in one paragraph, it would be this: > The authorship challenge moves too quickly from a claim about the producer to a claim about the product. Even if an LLM does not engage in the activity of philosophising, it does not follow that its output cannot be philosophy worth reading. To get that conclusion, one needs the further premise that a philosophical text has the relevant value only if it is produced by a person engaged in philosophical activity. That premise may be plausible for some artworks, where production history partly constitutes the work’s identity and value. But philosophy is different. Its central evaluative properties are semantic, inferential, and dialectical: what the text claims, how it argues, what distinctions it draws, what objections it handles, what explanations it gives. Those properties are publicly assessable in the text and its context. Authorship may matter for credit, responsibility, and publication ethics, but it does not determine whether the text contains an argument, distinction, or insight worth reading. That is the CEV of §1, I think. ## 12\. How this improves the whole paper This restructuring also helps §2 and §3. The current §2 works best if §1 has already established that philosophical value is product-centered. Then, when Floridi says LLMs do not perform abduction, you can reply: that is a claim about the process; what matters is whether the product contains abductive structure. Likewise, §3 works best if §1 has already established that lacking an experience does not automatically disqualify a text. Then the phenomenology challenge becomes a specific capacity challenge, not a repetition of the authorship challenge. So the payoff is not merely local. A stronger §1 gives you a general argumentative template for the whole paper: > Do not infer from a missing producer-capacity to missing product-value unless you can show that the product-value constitutively depends on that capacity. Then §§2 and 3 test whether such dependence can actually be shown. That makes the paper much more unified. I guess it would be weir not to ask you to do 3 as well. Yes — and §3 needs the most care, because it is the place where your own moves currently risk conceding too much. The “Merleau-Ponty limit” especially may be doing more damage than you need. Yes. §3 is the section I would restructure most aggressively. The current version has a strong instinct, but it also contains the most dangerous concession in the paper: the claim that an LLM could not have originated a Merleau-Ponty-style phenomenological observation if that observation was not already in the corpus. That may be true under one very strict reading, but stated as a general limit it risks undermining the paper’s whole ambition. It makes the response sound like: LLMs can only recombine already articulated phenomenology. That is too weak for your project. The CEV of §3 should be: > The challenge from phenomenology fails because it mistakes a claim about *having* an experience for a claim about being able to produce *philosophically valuable text about* experience. Phenomenology enters philosophy only once it is articulated, described, structured, varied, and argued from. LLMs lack phenomenal consciousness, but they have access to the public articulation of phenomenal consciousness, and that is the form in which phenomenology becomes available for philosophical work. That keeps the claim strong. ## 1\. What is working in the current §3 The current section gets three important things right. First, it correctly treats phenomenology as a *capacity challenge*, not an authorship challenge. The objection is not “no person stands behind the text.” It is: “this kind of philosophy requires a kind of conscious experience the model lacks.” Second, it correctly refuses the bad answer: you do not try to say that LLMs really have phenomenology. That would be a distraction and probably a losing battle. Third, it correctly identifies the textual/public character of phenomenological philosophy. Phenomenological experience matters to philosophy only insofar as it gets articulated: “what seeing red is like,” “what agency feels like,” “what self-touch reveals,” “what temporal passage seems to involve.” Once articulated, it becomes something that can be reasoned with. So the instinct is right. But the architecture needs tightening. ## 2\. The biggest weakness: the section grants too much to the objector The current version says, roughly: > LLMs can do phenomenology-based philosophy when the relevant phenomenological descriptions are already in the corpus, but they cannot discover genuinely new aspects of phenomenology. That concession is dangerous. It turns the section into a weaker claim than the paper needs. It suggests that LLMs can write derivative phenomenological philosophy but not genuinely innovative phenomenological philosophy. But in the Q&A you explicitly wanted to resist the analogous creativity worry. You said there is no principled reason why LLMs could not use philosophical vocabulary to produce paradigm-shifting work. §3 should not quietly concede the opposite. The better move is: > There may be a special task of *first-person phenomenal discovery* that requires experience. But producing philosophy worth reading about phenomenology is not identical to first-person phenomenal discovery. It includes constructing arguments, testing descriptions, generating distinctions, comparing cases, drawing conceptual consequences, and proposing candidate articulations of experience. Those tasks are textual and conceptual. LLMs can do them. Then handle Merleau-Ponty as a hard case, not as a limit. ## 3\. The second weakness: the Pigliucci move is useful but not central enough The science/philosophy contrast is interesting, but in the current version it risks doing too much and too little at once. It does too much because it suggests a large claim about science versus philosophy: science tests hypotheses against the world, philosophy uses worldly input as axioms. That is contestable. Some philosophy is empirically informed; some scientific theorizing is conceptual and model-based. You do not need this broad contrast. It does too little because the real point is simpler and stronger: > Phenomenological data become philosophical data only when they are articulated. That is the key. Philosophy cannot literally put raw experience into an argument. It can only put descriptions, reports, contrasts, concepts, examples, and thought experiments into an argument. That is why the LLM’s lack of experience is not automatically disqualifying. So Pigliucci can stay, but he should not carry the section. The central distinction should be: > raw phenomenal occurrence vs articulated phenomenological content. The LLM lacks the first. The question is whether it can use the second. Since philosophy uses the second, the objection loses much of its force. ## 4\. The third weakness: the color-design anecdote should probably go The anecdote about LLMs discussing color palettes well is rhetorically tempting, but I would cut or demote it. Why? Because it invites the wrong reply: > Yes, of course it can imitate color discourse. That is the whole problem. It makes the argument look like a Turing-test point: the model talks *as if* it had color experience. But your stronger argument is not based on indistinguishability. It is based on the claim that the philosophical work is performed at the level of articulated descriptions and inferential relations. The better example is not “LLMs can talk about color like a designer.” The better example is: > A philosopher can evaluate the Mary argument without newly undergoing Mary’s transition from black-and-white confinement to seeing red. The debate proceeds from an articulated scenario and an articulated claim about what Mary learns. Once the scenario is public, one can reason about it. That example directly supports the philosophical thesis. The design anecdote supports only a weaker “the model sounds convincing” point. ## 5\. The fourth weakness: Zahavy should be contained Zahavy’s Einstein case is powerful, but it may not map cleanly onto philosophy. In physics, Zahavy’s point is that a thought experiment may depend on sensorimotor simulation in order to generate a hypothesis about the physical world. But in philosophy, even when a thought experiment draws on experience, its philosophical force depends on the public structure of the case. Mary, the missing shade of blue, self-touch, inverted spectrum, bodily agency, temporal passage — these cases matter because they are *described*. Their role in philosophy is not the private act of imagining alone, but the publicly available description that can be varied, challenged, and argued from. So the move should be: > Zahavy may be right about some forms of discovery in physics. But the philosophical case is different in the relevant respect: the thought experiment does its work as an articulated scenario. Once articulated, it can be handled by a system that lacks the original experience. That does not deny Zahavy. It restricts his relevance. ## 6\. The core missing premise As with §§1–2, the best way to strengthen §3 is to identify the missing premise. The phenomenology challenge needs this premise: > A text can make a worthwhile philosophical contribution about phenomenology only if it is produced by a subject who has the relevant phenomenological experience. That premise is false, or at least far too strong. Humans already produce worthwhile philosophical work about experiences they have not had. Philosophers write about blindness, hallucination, synesthesia, infant experience, animal experience, pain asymbolia, depersonalization, psychedelic experience, grief, religious experience, and many other modes of experience they may not personally have undergone. They do this by relying on testimony, literature, clinical reports, scientific work, and prior philosophical descriptions. Sometimes first-person experience helps. It is not a universal entry requirement. So the challenge should be forced to answer: > Why should LLMs be held to a stricter standard than human philosophers, who routinely philosophize from articulated descriptions of experiences they do not themselves possess? That is a strong move. ## 7\. The real distinction: source, object, and method §3 should distinguish three different roles phenomenology can play. First, phenomenology can be a *source* of philosophy. A philosopher attends to her own experience and notices something. Second, phenomenology can be an *object* of philosophy. A text asks what perception, pain, bodily awareness, emotion, or temporal experience is. Third, phenomenology can be an *input* into argument. A description of experience functions as a premise, example, or datum. The LLM lacks phenomenology as first-person source. But it can handle phenomenology as object and as articulated input. And much phenomenology-based philosophy operates at those latter levels. So the absence of first-person phenomenology does not entail the inability to produce phenomenology-based philosophy worth reading. This is the cleanest conceptual distinction for the section. ## 8\. The Merleau-Ponty problem The current Merleau-Ponty move should be rewritten. Current version: > An LLM could not have originated Merleau-Ponty’s self-touch observation, assuming it was genuinely new, because it lacks the relevant bodily experience. Better version: > Merleau-Ponty gives the objector her strongest case: sometimes a philosopher seems to articulate a structural feature of experience by attending to experience itself. If philosophical production always worked like that, LLMs would be in trouble. But it does not. Moreover, even here, the philosophical value of the observation lies in the articulated structure: the distinction between touching and touched, the reversibility of the roles, and the impossibility of their perfect coincidence. Once that structure is articulated, it becomes available for further philosophical work by anyone who can reason with the description. Then add the stronger point: > Even the origination case is not straightforwardly closed. A model could generate a novel candidate description of experience by recombining bodily, perceptual, and conceptual materials already present in the corpus. It could not confirm that candidate by introspection, but confirmation by introspection is not the same as production of a philosophically valuable text. A text can propose a phenomenological structure worth considering even if its producer did not discover it through first-person attention. That preserves the strong claim. The LLM may not be a first-person witness. But it may still produce a phenomenological proposal worth reading. ## 9\. The strongest response to the phenomenology challenge The response should be: > LLMs do not have phenomenological experience. But the relevant question is whether having phenomenological experience is necessary for producing a philosophical text about phenomenology. It is not. Philosophy cannot directly use raw experience; it uses articulated descriptions of experience. These descriptions are public, iterable, criticizable, and textually transmitted. LLMs are trained on vast amounts of such material. They can therefore produce arguments that take phenomenological descriptions as inputs, vary them, compare them, draw consequences from them, and propose new ways of understanding them. That is enough for producing phenomenology-based philosophy worth reading. That is the section’s main answer. ## 10\. Recommended revised move list for §3 Here is the CEV version. ### §3 — The Challenge from Phenomenology **Move 1 — The challenge** Some philosophy appears to depend on phenomenology: what it is like to see red, feel pain, experience agency, undergo temporal passage, touch one’s own body, or have an intuition. Since LLMs have no conscious experience, one might argue that they cannot produce philosophical work whose starting point is phenomenal experience. **Move 2 — Why this is a capacity challenge** This is not the authorship challenge. The objection grants that the relevant philosophical properties would be properties of the text, but denies that an LLM has the capacity required to produce a text with those properties. No phenomenology in the system, no phenomenology-based philosophy in the output. **Move 3 — The missing premise** The challenge requires a strong premise: a text can make a worthwhile philosophical contribution about phenomenology only if it is produced by a subject who has the relevant phenomenal experience. That premise is not obvious and should be rejected. **Move 4 — Distinguish raw experience from articulated phenomenological content** Raw experience is not itself an argument. For phenomenology to enter philosophy, it must be articulated: described, contrasted, stabilized, turned into an example, used as a premise, or made into a thought experiment. Philosophy works with articulated phenomenological content, not with unmediated experience as such. **Move 5 — Human philosophers already rely on articulated phenomenology** Human philosophers routinely write about experiences they have not had: hallucination, blindness, synesthesia, animal experience, infant experience, pathological bodily experience, religious experience, grief, trauma, and so on. They do this by relying on testimony, literature, clinical description, science, and prior philosophy. First-person possession may help, but it is not generally required. **Move 6 — Therefore the relevant capacity is not experience-possession** The capacity needed for most phenomenology-based philosophy is not the capacity to undergo the experience, but the capacity to handle its articulation: to understand the description, vary the case, compare interpretations, draw conceptual consequences, and assess objections. That is a textual and inferential capacity. **Move 7 — Zahavy’s Einstein case** Zahavy’s Einstein example gives the objector her strongest model: some discovery may depend on manipulating sensory imagination rather than merely manipulating symbols. But even if this is right for Einstein’s physical theorizing, the philosophical use of thought experiments is different in the relevant respect. Mary, the missing shade of blue, inverted spectrum, self-touch, and similar cases function as articulated scenarios. Their philosophical force lies in the public description and the argumentative use made of it. **Move 8 — Mary as the central test case** An LLM has never seen red. But the Mary argument does not require every competent discussant to undergo Mary’s transition. Philosophers debate the argument by reasoning from the articulated setup: Mary knows all the physical facts; Mary has never seen red; upon seeing red she appears to learn something. Once that structure is public, it can be used by someone who has not had Mary’s precise epistemic situation. The LLM’s lack of color experience does not by itself block it from producing a worthwhile argument about Mary. **Move 9 — The corpus contains articulated phenomenology** The training corpus contains philosophy of perception, philosophy of mind, phenomenology, psychology, neuroscience, memoir, fiction, art criticism, and ordinary descriptions of experience. It is saturated with public descriptions of seeing, touching, imagining, remembering, desiring, acting, feeling time pass, feeling bodily ownership, and so on. If the relevant philosophical input is articulated phenomenology, then LLMs have access to that input in abundance. **Move 10 — From description to philosophy** The model need not merely repeat those descriptions. It can use them philosophically: generate variants of thought experiments, distinguish different readings of a phenomenological claim, compare explanatory hypotheses, test whether a description supports a conclusion, and formulate objections. These are exactly the operations by which phenomenological material becomes philosophy. **Move 11 — Possessing experience is neither necessary nor sufficient** Having an experience is not sufficient for producing philosophy about it. Everyone sees colors, feels bodily movement, and experiences time, but most people do not thereby produce good philosophy of perception, agency, or time. What matters philosophically is the articulation and argumentative handling of the experience. Conversely, lacking a given experience does not automatically prevent one from doing philosophical work with descriptions of it. **Move 12 — The hard case: first-person discovery** The hardest case is the alleged discovery of a previously unarticulated feature of experience, as in Merleau-Ponty’s self-touch example. This may require first-person attention in some cases. But this does not establish the general challenge. It shows at most that LLMs may lack one route to phenomenological discovery. It does not show that they cannot produce phenomenology-based philosophy worth reading. **Move 13 — Do not overstate the limit** Even in the hard case, the limit should not be put as “LLMs cannot originate new phenomenological insights.” A model could generate a novel candidate articulation by recombining descriptions, analogies, distinctions, and bodily concepts already present in the corpus. It cannot introspectively verify the candidate, but production and verification are different. A phenomenological proposal can be philosophically worth reading before its final status is settled. **Move 14 — §3 payoff** The phenomenology challenge fails. From the fact that LLMs lack conscious experience, it does not follow that they cannot produce worthwhile philosophy about conscious experience. Philosophy uses phenomenology in articulated form, and articulated phenomenology is public, textual, and inferentially usable. Since current LLMs can work with such material, they can produce phenomenology-based philosophy worth reading. **Move 15 — Transition to §4** The remaining puzzle is not whether the capacity is present in principle. It is why generic prompting so often fails to elicit it. The answer is that generic prompts ask for summaries of phenomenological debates, and summaries are usually what they get. To produce worthwhile phenomenology-based philosophy, the model must be given the kind of target that worthwhile philosophy itself has: a specific problem, a definite claim, a live opponent, and pressure from objections. ## 11\. What I would cut from the current §3 I would cut or drastically reduce the following. First, the long detour through Einstein. Keep Zahavy, but do not let Einstein dominate. The paper is about philosophy, and the Einstein example can make the section feel borrowed from another debate. Second, the color-design anecdote. It is too easy to dismiss as mimicry. Third, the strong Merleau-Ponty concession. Replace it with the “hard case” treatment above. Fourth, the broad science/philosophy contrast, unless you sharpen it. You do not need the large claim that science verifies and philosophy uses axioms. You need the narrower claim that philosophy uses articulated starting points. ## 12\. What I would add I would add four things. First, the missing premise: “phenomenology-based philosophical value requires first-person possession of the relevant experience.” Second, the distinction between raw experience and articulated phenomenological content. Third, the point that humans routinely philosophize about experiences they have not had. Fourth, the distinction between producing a phenomenological proposal and verifying it introspectively. These additions make the section much stronger. ## 13\. The punchiest form of the §3 thesis Here is the core thesis in one paragraph: > The challenge from phenomenology moves too quickly from the absence of experience in the producer to the absence of phenomenological value in the product. LLMs do not see red, feel bodily agency, or experience time passing. But philosophy does not work with raw experience as such. It works with articulated descriptions of experience: claims about what seeing red is like, how bodily agency presents itself, how temporal passage seems, how self-touch is structured. Once experience is articulated, it becomes public philosophical material. It can be varied, criticized, compared, and argued from. Human philosophers already rely on such articulated phenomenology when they discuss experiences they have not themselves had. LLMs are trained on vast stores of this material and can manipulate it in philosophically structured ways. So the lack of phenomenology blocks one possible route to philosophical production, but it does not block the production of phenomenology-based philosophy worth reading. That is the version I would build around. ## 14\. How §3 connects to §1 and §2 The whole paper becomes stronger if §§1–3 use the same argumentative template. §1 says: > Do not infer from absence of human authorship to absence of philosophical value in the text. §2 says: > Do not infer from absence of human abduction to absence of abductive structure in the text. §3 should say: > Do not infer from absence of phenomenological experience to absence of phenomenological philosophical content in the text. That gives the paper a powerful unified structure. Each challenge tries to move from a missing producer-side feature to a missing product-side value. Your answer, each time, is that philosophy worth reading is assessed through publicly available textual properties. The producer-side feature may explain how humans often produce such properties, but it does not constitute them. That is the CEV of the whole middle of the paper. Okay, that was very good as well, but I would like it to start with the Zahavi still and with the Einstein example still. This is because I have a co-author. I think that's how I want this section to begin. And it has to be framed in terms of the challenge that we're putting forward, okay? The challenge from phenomenology. Okay, so can you have a think and reconfigure or yeah iterate again on what you've given me because I really liked what you just gave me. But it just yeah it has to sort of be rearranged to suit the beginning I've just sketched out for you. Okay, I mean even though it starts like this and does mention the Einstein thing, this can be easily controlled because at some point you can talk about what you just did, which is the difference between using phenomenology in science and using it in philosophy. Okay. But at the beginning we can just sort of so you know, if we're sort of slightly vague about it, we can say, well, it's a thought experiment. Philosophy uses phenomenology-based thought experiments as well, so perhaps we have a problem. Okay, and then only later make your more fine-grained distinction. But yeah, one more time, what you gave me was great, just needs to be restructured slightly based on the beginning. Yes — keep Zahavy and Einstein as the dramatic entry point. The trick is to let that opening *generate* the challenge, then later show that the initial analogy between Einstein and philosophy is too coarse. Good. I agree with the constraint. §3 can and probably *should* begin with Zahavy and Einstein. That opening gives the challenge intuitive force. The previous version I gave you was analytically cleaner, but rhetorically too fast: it began with the distinction that should really come *after* the challenge has been allowed to bite. So the structure should be: > Start with Zahavy’s Einstein case as the strongest way of making lack of experience seem philosophically relevant. > Transfer the worry to philosophy through phenomenology-based thought experiments. > Then, only after the challenge is live, make the more fine-grained distinction between phenomenology as raw experiential source and phenomenology as articulated philosophical material. That way you keep the opening your co-author wants, but the section still ends up with the stronger argument. ## 1\. The rhetorical shape should be: from intuitive analogy to controlled disanalogy The opening should *not* immediately say: > Zahavy is talking about physics, but philosophy is different. That would deflate the challenge too early. It would make the section feel defensive before the problem has been built. Instead, start by letting Zahavy set up a powerful general thought: > Some intellectual breakthroughs seem to depend not merely on manipulating sentences, but on manipulating experience. Einstein’s elevator is then the dramatic example. The thought experiment is not just a verbal permutation. It involves imagining what it would be like to be inside an accelerating elevator: objects released from the hand, the floor rushing up, the felt equivalence between acceleration and gravity. If that is right, then some reasoning depends on a kind of simulated perceptual or bodily experience that an LLM lacks. Then you transfer that to philosophy: > Philosophy also uses thought experiments that seem to depend on phenomenological access. Mary. The missing shade of blue. Inverted spectrum. Self-touch. Temporal passage. Bodily agency. The feeling of intuition. The initial worry is then very clear: > If some philosophy starts from what experience is like, and LLMs have no experience, then perhaps LLMs cannot produce philosophy of this kind. Only *after* that do you refine the picture. ## 2\. The key adjustment: Zahavy opens the challenge, but does not define the final terms Zahavy gives the challenge its initial form: > No experience, no experience-based discovery. But your reply should gradually show that this formulation is too coarse. The real question is not whether the model has experience. The question is whether experience must be *possessed* in order for phenomenology-based philosophical content to be *produced*. So the section should pivot from: > Does the model have phenomenology? to: > What role does phenomenology play in philosophical text? That is the crucial move. ## 3\. The challenge should be put strongly Do not frame the challenge weakly as “maybe LLMs cannot discuss phenomenology well.” Frame it like this: > Some philosophical arguments seem to require first-person experiential access as a generative source. If an LLM has no such access, then perhaps it can only recombine descriptions of experience already given by others. It may summarize phenomenological philosophy, but it cannot produce philosophy whose force depends on phenomenological insight. That is the strongest version. It also sets up the creativity worry indirectly, so you can later resist it. ## 4\. Revised §3 move list, beginning with Zahavy and Einstein ### §3 — The Challenge from Phenomenology **Move 1 — Introduce the second capacity challenge** The second capacity challenge concerns phenomenology. The thought is not merely that LLMs lack human authorship, nor merely that they lack abductive reasoning, but that they lack conscious experience. Since some philosophy appears to begin from conscious experience, this may seem to block LLMs from producing philosophy of that kind. **Move 2 — Zahavy’s Einstein case** Zahavy gives the challenge a powerful form. His example is Einstein’s elevator thought experiment: Einstein imagines a physicist inside an enclosed elevator accelerating through space. Inside the elevator, objects released from the hand appear to fall with the same acceleration, regardless of their composition. The point, according to Zahavy, is that Einstein’s reasoning did not proceed merely by manipulating symbols. It involved simulated perceptual and bodily experience. **Move 3 — Manipulative abduction** Zahavy calls this kind of reasoning *manipulative abduction*: inference that proceeds through the manipulation of an imagined experiential situation. The thinker does not merely rearrange propositions. She varies a scenario in imagination, attends to what would be experienced within it, and uses that simulated experience to generate a hypothesis. **Move 4 — Why this threatens LLMs** If Zahavy is right, then some forms of reasoning require a capacity LLMs lack. LLMs can manipulate descriptions of elevators, acceleration, gravity, and falling objects, but they have never felt weight, free fall, bodily orientation, or the apparent equivalence between acceleration and gravity. They can handle the words, but not the experience that gives the thought experiment its generative force. **Move 5 — Transfer the worry to philosophy** The same worry seems to arise in philosophy. Philosophy is full of thought experiments and arguments that draw on what experience is like: Mary seeing red for the first time, Hume’s missing shade of blue, inverted spectra, bodily agency, self-touch, temporal passage, intuition, pain, emotion, and aesthetic experience. These cases seem to depend on phenomenological access. **Move 6 — Formulate the challenge from phenomenology** The challenge is therefore this: if some philosophy depends on phenomenological experience as a source of insight, and if LLMs have no phenomenological experience, then LLMs cannot produce philosophy of that kind. At best, they can repeat or recombine what experiencers have already said. **Move 7 — Make the hidden premise explicit** But the challenge needs a further premise: a text can make a worthwhile philosophical contribution about phenomenology only if it is produced by a subject who has the relevant phenomenal experience. This is the premise to resist. **Move 8 — First distinction: having experience vs producing philosophy about experience** Having an experience is one thing; producing philosophy about that experience is another. Most people see red, feel pain, touch their own bodies, and experience time passing, but they do not thereby produce good philosophy of color, pain, embodiment, or time. What matters philosophically is not raw possession of experience, but its articulation and argumentative handling. **Move 9 — Second distinction: raw phenomenology vs articulated phenomenological content** Raw experience does not enter a philosophical argument directly. It enters only once articulated: described, stabilized, contrasted, turned into a case, made into a premise, or used as a point of comparison. Philosophy works with articulated phenomenological content. That content is public, linguistic, repeatable, and criticizable. **Move 10 — Why the Einstein analogy now needs refinement** This is where the initial analogy with Einstein becomes too coarse. In the scientific case, the imagined experience helps generate a hypothesis that must then be tested against the world. In the philosophical case, the thought experiment usually functions differently: once the experiential scenario is articulated, it becomes the object of conceptual and argumentative work. The philosophical force lies in what can be drawn from the articulated scenario. **Move 11 — Mary as the central philosophical case** Mary does not matter because every philosopher personally undergoes Mary’s transition. No one does. The argument works because the scenario is publicly articulated: Mary knows all the physical facts about color vision; she has never seen red; when she sees red, she appears to learn something. Once that structure is described, philosophers can reason about it, reject it, revise it, or draw consequences from it. The LLM’s lack of color experience does not automatically prevent it from working with that articulated structure. **Move 12 — Human philosophers already rely on articulated phenomenology** Human philosophers routinely write about experiences they have not had: blindness, synesthesia, hallucination, infant experience, animal experience, psychiatric experience, religious experience, grief, trauma, psychedelic experience, and many others. They rely on testimony, literature, clinical description, empirical psychology, neuroscience, and previous philosophy. First-person possession is one source of phenomenological material, but it is not a universal condition of philosophical work about experience. **Move 13 — The corpus point** The training corpus contains enormous amounts of articulated phenomenology: philosophy of perception, phenomenology, philosophy of mind, aesthetics, emotion theory, psychology, neuroscience, memoir, fiction, criticism, ordinary experiential description. LLMs therefore have access not to raw phenomenology, but to the form in which phenomenology becomes usable in philosophy: articulated descriptions of experience. **Move 14 — From access to production** The model need not merely repeat those descriptions. It can use them philosophically: generate variants of thought experiments, distinguish readings of a phenomenological claim, compare explanatory hypotheses, identify tensions, formulate objections, and draw conceptual consequences. These are not external additions to phenomenology-based philosophy. They are much of what phenomenology-based philosophy consists in. **Move 15 — Recasting the Zahavy worry** So Zahavy’s challenge, transferred to philosophy, proves less than it initially seemed to. It may show that LLMs lack one human route to philosophical production: first-person experiential discovery. But it does not show that they cannot produce texts that reason philosophically from phenomenological material. Once phenomenology is articulated, it becomes part of the public space of reasons. **Move 16 — The hard case: Merleau-Ponty on self-touch** The strongest remaining case is one where a philosopher seems to discover a previously unarticulated feature of experience through first-person attention. Merleau-Ponty’s self-touch example is useful here: when one hand touches the other, one hand is toucher and the other touched; the roles can reverse, but they do not perfectly coincide. This looks like a genuine phenomenological discovery. **Move 17 — Do not concede too much** But the lesson should not be that LLMs cannot originate phenomenological insight. The safer and stronger point is this: an LLM cannot *verify* such a claim by first-person attention. But production and verification are different. A model could generate a candidate phenomenological articulation by recombining bodily descriptions, conceptual distinctions, and prior debates. The result might be worth reading even if its ultimate adequacy would have to be assessed by phenomenological reflection. **Move 18 — The right limit** So the limit is not: LLMs cannot produce new phenomenology-based philosophy. The limit is narrower: they do not possess first-person phenomenology as an independent checking mechanism. That affects how we assess some outputs, but it does not show that the outputs cannot be philosophically valuable. **Move 19 — §3 payoff** The challenge from phenomenology fails because it moves too quickly from absence of experience in the producer to absence of phenomenological value in the product. LLMs lack conscious experience, but philosophy uses conscious experience in articulated form. Since articulated phenomenology is public, textual, and inferentially usable, current LLMs can produce phenomenology-based philosophy worth reading. **Move 20 — Transition to §4** The remaining question is practical rather than constitutive: if the capacity is present, why does generic LLM output on phenomenology so often look flat, derivative, or merely expository? The answer is that generic prompts elicit generic regions of the distribution. To get worthwhile phenomenology-based philosophy, one must ask for a specific problem, a definite claim, pressure from a live opponent, and a real argumentative task. ## 5\. The internal rhythm of the section The section should have four beats. First: *Zahavy makes the problem vivid.* > Some reasoning seems to require simulated experience. Second: *Philosophy seems exposed.* > Many philosophical thought experiments seem phenomenology-based. Third: *The challenge depends on a false bridge principle.* > It assumes that producing philosophy about experience requires having the relevant experience. Fourth: *The reply shifts from experience-possession to articulation.* > Philosophy works with articulated phenomenology, and LLMs have access to that. That rhythm gives the section a satisfying argumentative arc. ## 6\. What to do with the science/philosophy distinction Keep it, but put it later. It should not be the first reply. The right place is after the challenge has been transferred to Mary, self-touch, and so on. Then you say: > At this point we need to refine the analogy with Einstein. Zahavy’s example concerns a physical hypothesis generated through sensory imagination. Philosophy often uses thought experiments differently. A philosophical thought experiment is not normally confirmed by later measurement in the way a physical hypothesis is. Its force depends on what follows from the articulated scenario. So what matters is whether the scenario can be described, varied, and reasoned from. This lets you keep the Einstein opening while controlling it. ## 7\. How to avoid making the section sound weak Avoid these formulations: > LLMs can only work with phenomenology once humans have already done the real work. > LLMs can summarize phenomenological descriptions. > LLMs can assist humans in reasoning about experience. > LLMs lack experience, so they cannot discover anything genuinely new about phenomenology. Instead use: > LLMs can work with articulated phenomenological content, and articulated phenomenological content is the form in which phenomenology enters philosophical argument. > The absence of experience blocks one route to discovery, not the production of phenomenology-based philosophy as such. > A text can propose a novel phenomenological articulation even if its producer did not arrive at it through first-person attention. > Verification by experience and production of a worthwhile philosophical proposal are distinct. That keeps the claim strong. ## 8\. A possible compact version of the section’s core Here is the core in a more talk-ready form: > Zahavy’s Einstein case gives the challenge its bite. Einstein’s elevator thought experiment seems not to be a mere manipulation of symbols, but a manipulation of imagined experience: what it would be like inside an accelerating elevator, what would happen when objects are released, how acceleration and gravity would be experientially indistinguishable. If some reasoning works like that, then LLMs seem blocked from it, because they have no experience to manipulate. And philosophy seems vulnerable to the same worry. Mary, the missing shade of blue, inverted spectrum, self-touch, bodily agency, and temporal passage all seem to begin from what experience is like. So perhaps an LLM can produce text about these debates, but cannot produce philosophy that genuinely starts from phenomenology. Then the pivot: > But this argument moves too quickly. Raw experience does not enter philosophy as raw experience. It enters philosophy only once articulated: described, stabilized, contrasted, turned into a thought experiment, used as a premise. Mary is not a private episode we each undergo; it is a public scenario we can reason about. The same is true of self-touch, temporal passage, and bodily agency. Once phenomenology is articulated, it becomes part of the public space of reasons. LLMs lack experience, but they are trained on vast stores of articulated phenomenology, and they can use those articulations to generate distinctions, objections, variants, and arguments. So the lack of phenomenology blocks at most one route to philosophical production. It does not block the production of phenomenology-based philosophy worth reading. That is probably the central spoken version. ## 9\. What the Merleau-Ponty example should do now Use Merleau-Ponty not as the admitted limit, but as the strongest objection *inside* the section. The move should be: > “The objector may say: but what about genuinely new phenomenological articulation?” Then Merleau-Ponty. Then your reply: > Even here, the conclusion does not follow. The LLM cannot check the claim by introspection, but it can formulate a candidate articulation. And in philosophy, formulating a candidate articulation can already be worth reading. This is very important. It keeps Merleau-Ponty from becoming a concession that swallows the section. ## 10\. The section’s final payoff should be stronger than before The payoff should not be: > LLMs can do much phenomenology-based philosophy, except for genuinely new phenomenological discovery. It should be: > The absence of phenomenology in the model does not entail the absence of phenomenology-based philosophical value in the text. The relevant philosophical material is articulated phenomenology, and that material is available to the model. The model may lack first-person experience as a source or check, but it can still produce texts that formulate, transform, and reason from phenomenological descriptions in ways worth reading. That is the stronger version, while preserving the Zahavy/Einstein opening.