# User: you must invoke the following skills BEFORE DOING ANYTHING * Skill contemplate * Skill nick... ## Skills Used - epistemic-discipline - nick-analytic-voice - nick-philosophical-prose - source-work - twork - writing-standards ## Retrieval Notes - Session id: `95a96a6d-2492-4db4-9f27-8ac6698da12e` - Last activity: `2026-03-05T21:07:53.601Z` ## My Notes <!-- Add your notes here. This section is preserved across syncs. --> ## Conversation ### User User: you must invoke the following skills BEFORE DOING ANYTHING * Skill contemplate * Skill nick-analytic-voice * Skill nick-philosophical-prose * Skill twork * Skill source-work * Skill epistemic-discipline * Skill writing-standards so here is a very early stage draft of the paper me and my coauthor are writing together. I can think of a few places it needs improvement , but for now I want to focus ona different problem. I have a very ambitious aim for the section coming immediately after how the paper currently ends. it will be practical, it will involve writing about prompt creaft and how we can get LLMs to produce good philosophy. Now, I have gotten LLMs do produce very good philosophical ideas, but it is lame just to present bits and pieces and say 'look', I am not sure why, but i guess I am not sure what this section 4 needs to consist in. Basically, and I know the preceding sections are far from perfect but they include all the ideas i want to talk about even igf not propley fleshed out or present below is a transcript of my coauthor presenting the paper in its current form today, followed by both of us asking questions, because the recording is from my computer speakers i suspect the transcript will be pretty damn bad but do the best you can with context to work out what enrico is saying. . it will hopefully fill in a few of the spaces that the paper leaves out. there are still plenty more though. in particular how things like theoretical virtues, or phenomenological data are embedded in the training data of the llm in such a way that it can generate novel philosophy which keeps in accordance with theoretical virtues. (and when required be able write about phenomenology which rings true) . Now I thonk about it these are too big questions to deal wirth and they are not really worked out in any detaikl yet in secs 0 - 3 , but i think you can at least see where i am coming from. i think i would ikus to focus only on the theoreticaly virtues -how arguments can be *drawn out* from llms, that seems to be the thing we have to say in this paper and i am fine with that, lets leave the phenomenolgoy right now and just focus on this. think about the most obvious way to argue that theoretical virtues are some how latent in an llm (or sometihng similar) and how a skillful prompter might draw out arguments which embody them. Tell you what, could you work out the cev of this idea about theoretical virtues (or however we ended up phrasing thngs in the paper, i am not sure that is the right label, but maybe there is a more general term as to what i am after, see how you go, and check what is best. please draw on any resources you have available regarding the project notes, previous chats, and sources. this is a big job. transcript: Hey, what's the difference? So it seems that Florida. It's again, a problem with the process uh is putting too much emphasis on the process. Okay. And then, uh, yeah. So the main objection, the only objection that it doesn't, uh, in which the connection to the process is really relevant because it's, it's the idea that the process, uh, put constraints on, on the, on the text. So, the, the task cannot be valuable. Uh, but then we have these arguments that this arguement that, yeah. Uh, that's a priority to physics. Yeah. Okay. Uh one thing I haven't done yet, as well is talk about benzianism. That still hasn't been quite finished off yet. Yeah, but this is just an orange suspicion. Maybe also for the sake of clarity. Also the The Einstein principle of equivalence maybe, but that's all. I think it's already that zombie paper, so it's not a problem for us to, to explain the principle. So yeah, I know most substantial objection may be related to the role of intuitions in philosophy. I put in the charter Mercedes book, which is usually considered About the meta philosophy and it's quite influential. Even though I don't know how relevant, it can be to what we say, but maybe if at least for this. And I was, yeah, I wondering also whether There may be areas of philosophy in which appeal to experience, even introspection, which is not the kind of experience that is relevant to physics. Maybe but especially if philosophy of mine philosophy of Consciousness in which the very subject matter is something that seems to require acquaintance. So, probably one, we may also I wonder whether because we doesn't care so much about that because it's not the kind of philosophies interested in, it's more of your doing language metaphysics, uh model logic, but perhaps in some more mathematical like philosophy. And so one way, one possible objection is that this Supply only to to a certain areas of philosophy. Also uh evaluative philosophy like Aesthetics so Experience like like the Einstein experience, it's very hard to make a shift for this one may apply the The the the Einstein case to towards of influential paper on category so far, in which he realised how important historical category like impressionism and R in our engagement with words of art or or, uh, the other more uninfluential paper on. Oh, it's water again, in a sense, the imaginative resistance, again, to introduce the concept of a system, you have to feel some resistance gay certain fictions. And so that that is the way of of extending uh the the objection to to certain area of philosophy. And in that thing, maybe the the Us, people usually don't talk about what they feel in a date or so what. But in in those cases it seems that this kind of experience. It's already sedimented in the data set. Maybe not the philosophy data set, but there are history data set or the good diary autobiography data set. Absolutely, yeah, good. Now when you're right, this is a nice way of broadening stuff out. But yeah, exactly. The and yeah, cuz we could say there are various different data sets, which will have different levels of abstraction about experience in different ways. Good. Because I I think may let me know if you disagree. I think maybe the best way of dealing with this objection. Is maybe not trying to present what you've just said about stuff being in the Corpus as being 100. A solution to this objection but at least showing we can probably get quite a lot further than you would imagine. I'm tempted to sort of not be so strong in saying, we've completely solved the phenomenology aspect with the Corpus. Well no no. But it seems possible because one way to go it's it's the one we have now is basically phenomenology doesn't care. Just care, just scare manipulating concept at a higher level of abstraction and that's that's already in the philosophical text. They are selected Mean that for a certain area of philosophy, indeed, also, Sort of even the cream cake is because cricket seems sort of experience of using language, intuitions about So um, yeah. One mayor yards. The. The value objection at least on certain cases the role of intuitions in philosophy and I think that's what the, the Machery book it should be about. And but then when we say, oh yeah, but maybe this. So this is another way of using the data setting, say not only as as a repository of Concepts and arguments but also as a repository of experience, second hand experience is uh it may be that this experience is, I'm not uh, it's not that it, it can solve everything but at least there are things that the, the the can use to, to rely on some experiences. That's really, yeah, a repository of experiences is really nice way of thinking about it actually. Yeah, or second. The hand actually is good. That is, can I ask so, I don't know the book you're mentioning about intuition, what is an intuition on that account? Is it a? Is it a feeling? Is it is Girl, who brought their case it is. It's not that There should be something on. Usually digested intuitions are like perception, but an intellectual level. Okay. So, they are not in financial, it's sort of grasping like, like a vision. But instead of grasping, uh, perceptor content, you grasp ideas, The propositions that's to me the the most And that's also why some some the philoso denied that Indonesia really exist precisely because talk cannot be grasped in a perceptual uh straightforward non-mediated way. But other things they did if they care. For instance the actions in mathematics and the principles not in transition is something that these things we just grasp it without the need of of deducing it from counterpences. One way of. I think there is also something on the debate for instance on conceptability, the intuition that Zombies. Chinese zombies is something. We cannot perceive them, but there is, we have the sense that they can like us but they don't have Consciousness. But it's, it's a messy debate, but there should be something and, and may be relevant. So, yeah, it seems that that's also a way of reaching the paper because I don't know if you had other things in mind to make it for. I think, for the presentation for the tomorrow presentation, what we have, it's perfect, uh, so if you can just, um, make the the PowerPoint out of that, maybe adding also these things about, um, if you want also to say something about uh, this case of uh, second hand. Uh yeah that's for for tomorrow is perfect. Then we we can start how long is that? It seems quite short as a beaker now, right? It's something like, let me, let me check. I have that the wall file just here. It's a well, yeah, it's not so short. It's a 4 000 War, so maybe already analysis paper. Yeah, I mean, it would be nice to get this done quick, wouldn't it? Let's see what, let's see, how it goes tomorrow. Yeah, probably Thing that so not for tomorrow. But the only other thing I was thinking with the extension is, Someone's going to say, okay, in that case, prove it Show us. Show us an llm doing good philosophy, okay? And then there's two ways to go with any section there. One would be An actual demonstration, which I think is somewhat interesting to try but another one would be to at least show how we could get get these things to do. Good philosophy by talking about prompting techniques or talking about these systems themselves and how you do it. Because I'm I feel like we should do something at some point because otherwise people are going to say if they can Why aren't they? Yeah. Yeah, I remember in the first version it was fun because you you presented that as a self-proving say oh this paper is. So if you think this paper is good. Yeah yeah. I mean, it's kind of the same to be honest. That's that's another, at least we can discuss it. I don't know if we have to, okay, at least we can we can discuss it and because I mean this is lesson, this is less important but it would also Go. If we did talk about prompting at some point, it would actually fit very nicely with The Hitchhiker's Guide Galaxy joke. At the beginning. Yeah. It would be nice to, to have a sort of payoff of, you know, the setup at the beginning. But another way which we can add something is also exactly about prompting and autonomy. So the two different issues that that also was in some previous version of the paper and can can do philosophy. Can mean, two things can do philosophy while in in collaboration with human philosophers or can you philosophy on their own? Yeah, yeah. Okay. It's another important thing. Certainly I I it kind of goes back to that Continuum I've mentioned earlier. And then the less interesting, the claim becomes the further along you get towards the prompter having to do all the work basically. Yeah. And and also the the the the the the scene or the character of bronze because you're, if the brand is just please uh, solve the the Mind Body problem. That's not the thought it seems that. In that case, you may say that the llm is doing philosophy. The problem is just giving a problem to solve on the other hand. If the ground is tree, is considered that and then there is interaction is not just one, prompt bada conversation, there are adjustments, there are then order is an initial idea and developing the so that's another kind of prompting. So we may probably distinguish between uh, yeah, one shot prompt and just probably which is just a question or a problem to solve. For tomorrow. But there's other interesting things to say as well about it's not even simply asking questions as well. So you remember with The semiotic physics ideas, they're always talking about sort of good continuation. So another interesting thing to think about prompting is well you really need to be doing is writing a prompt the good continuation of which will be good. Um but yeah, maybe another distinctually be uh, problem oriented and the solution oriented Brands. So the primary identity just uh a prompt that uh in a sense. Also a question is starting a good continuation of which is uh is is the answer A problem or you, you write the problem and say please solve it and a good continuation. But in this case, it's very differential on the other hand, you may write something, which there are already ideas that point toward the solution, and the good continuation, is another step toward the solution but then you can add another beat and a good continuation. So, there is a, a more obvious and shallow sense of good continuation, which is just the idea, uh, for which the, the, the, the the answer is a good continuation of the question. And then, there is a more interesting uh, sense in which the good continuation. The basis for the good continuation is not just a question, but a sort of of rough Uh, gesturing towards the solution and then the llm make the solution much more robust, good. Are you recording the conversation? Oh yeah, yeah, don't worry is everything. So, and I was just also, I've got a prototype presentation. So if you go in the chat, And I'm gonna send you a link. Yeah. And I'll send you a password as well. And this is the first, this is the first chance, but you'll see I've just been doing this. Well, I just said, make this while we're talking. Uh, it asked me a task for that in the chat as well. Ah, sorry. Yeah. Is the next message in the chat? Um Anna is the zoom chat another, the WhatsApp chat. Sorry, I was in the brown chat. Okay, got it. You see it? Yeah. Yeah. I can see it pressed down as well and you'll see. It moves in a very nice way as well. You see. Yeah. Yeah. So, Obviously, we can change things but it's yeah, quick to make these sorts of things. It look nice now. Okay. Anyway, you don't have to look through all of that now, but, um, And this is, The grant, the the typefaces are all based on stuff. I've chosen for my I've made my own note-taking app now with Claude code. So you can now just invent apps. I can say, make me a calendar app, make me a note-taking app and you get it. Exactly. And now I've told it to so now I have a nice house style. If you look at my website it's the same style as well. Cool. Yeah, cool. Eh Um, but yeah, I'll of course, make a better version before you talk tomorrow, okay. Yeah, and and then we will arrange it because also, I, I don't know. These guys are not philosophers, but they are working on AI a lot and they, they seems quite serious. So, let's see. Also, whether the other tours can be interesting for you, we will ranch away or YouTube to attend, the course, the other dogs. And yeah, I mean maybe take that microphone uh, that we use for the. Where is the the conference tomorrow? Eh in Sinaloa Madzini which is good microphone if the one with the sort of yeah a microphone on each table but maybe about today write an email to to look and franchise to to record and do we not you won't be able to come and solve please uh help you to to organise uh the online thing. Okay, great, I can do that. I don't know. Is there anything else we need to talk about right now? Or do you want? I don't know where. Now, this seems Already. It seems to we have good stuff or for stuff for for developing the paper in a yeah longer form or something like 6000. Maybe no more than 8,000 if if we can because in this way. Yeah it it depends because the paper is very ambitious. So we may also try the the super big journals like philosophical review or mind or Journal of philosophy. Okay, but we can also look for just for for another I don't know, but let's see. I think we can just keep on writing till the paper, looks unitary and go here. And and I don't think we need in principle. But by the end of the month, maybe we can, or maybe we can wait at the Hong Kong, uh, conference, uh, To do that. I need to deal with this. Um the income is the best place to to have that because they are very good. That just basically that then just doing philosophy for AI and they're not doing anything else. So and then there are smart people. Uh, so that's really a place where we can have very, very useful feedback. Okay. I still need to get my Administration fixed with that, but it should be fine. Now, I'll book the table, I'll book the return very, very soon. And I need to talk to Agatha about the As well. It'll be fine though. Okay, okay. Some, okay, so I think it's And just just, uh, keep on, keep on working on that. But for tomorrow, I, I, I, I, I, I feel quite quite quite confident it just maybe, maybe if you, um, And it's possible from. I can't see that from this website to download. Um, it's not this was just a preview. I will, I will send you a proper presentation. Yeah, at a certain point uh by by in the afternoon, you can send me up in this a presentation that I I can print. Uh, okay this afternoon, Well, so far, it's no problem. It's funny. Yeah. Also I map you also in the current version the one you you just a readable version in this way. I can I can prepare the talk just by by looking at the data on paper writing my notes on paper. So even now if you want it because I'm I'm really happy with this one for tomorrow. I think it's a it's it's then if you can add all the things we have say. Maybe if you can add something about we just discussed and then say bye bye. Five after guy would like to print it after Gaya guys. Have to go to teach. Okay, but in this way I have I have that I can tomorrow morning before I can. Okay, it's also true that we have a lot of time because we have the last of the day. Yeah. Yeah, but I I, I, I know that usually, then there are many things during the conference and then so I I would like just to have a printer, uh, draught printed version, uh, You. If you can, when you send it to me as a PDF, if you can switch change the background having a white white background. Otherwise favourite. Yeah, yeah. The the printer just go. How to think exactly bankrupt, the university by league. So yeah. Okay, in that case. So I've got two hours before Gaia, I'll just do two hours more work on this and probably yeah I'll send you something just as guy as things starts or something like that. Okay. Okay, great, cool. See you a little bit. See you later. Bye! Um, let's try to picky. I didn't mean to, I was just Not gonna be like, in my cupboard. I need some tissues. Tissue coffee water. DRAFT: Generating Philosophy Without Artificial Intelligence Nick Young & Enrico Terrone 0. Introduction “Forty-two,” said Deep Thought, with infinite majesty and calm. It was a long time before anyone spoke. Out of the corner of his eye Phouchg could see the sea of tense expectant faces down in the square outside. “We’re going to get lynched aren’t we?” he whispered. “It was a tough assignment,” said Deep Thought mildly. “Forty-two!” yelled Loonquawl. “Is that all you’ve got to show for seven and a half million years’ work?” “I checked it very thoroughly,” said the computer, “and that quite definitely is the answer. I think the problem, to be quite honest with you, is that you’ve never actually known what the question is.” — Douglas Adams, The Hitchhiker’s Guide to the Galaxy In The Hitchhiker’s Guide to the Galaxy, humanity asks an AI to do some philosophy. A computer named Deep Thought is constructed and instructed to provide “The Answer to the Ultimate Question of Life, the Universe, and Everything.” Humanity builds this computer only to receive the answer ‘42’—an answer which, while apparently correct, means next to nothing at all due to humanity’s failure to know what the Ultimate Question in fact is. In 2026, humanity has reached a position in which it can actually ask machines philosophical questions. One reason for optimism is that AI has had considerable success in other domains. In February 2026, researchers working on gluon scattering amplitudes gave GPT-5.2 worked examples for three, four, five, and six particles and asked it to find the general formula. The model conjectured a formula, completed a formal proof, and overturned a forty-year-old assumption (Guevara et al. 2026).1 The question is harder to answer than it might seem, because philosophy does not have uncontroversial success conditions. What counts as a contribution depends on 1Other AI-assisted breakthroughs include protein structure prediction, which won the 2024 Nobel Prize in Chemistry (Hassabis and Jumper, AlphaFold); solving a 30+ year challenge in quantum error correction (Google Quantum AI, Willow chip); and discovering new symmetries in black hole event horizon equations (Lupsasca with GPT-5). what philosophy is, and conceptions of philosophy differ in ways that matter for the question about AI. Some conceptions locate philosophy in texts. Dellsén, Firing, Lawler, and Norton (2024) argue that philosophical progress consists in putting people in a position to increase their understanding—what they call the for-whom rather than by-whom account—in which public utility of the work, not the internal states of whoever produced it, determines whether progress has occurred. Bengson, Cuneo, and Shafer- Landau (2022) characterise philosophical inquiry as theory construction evaluated by criteria: accommodation of data, explanatory power, integration, theoretical virtue, all of which are assessable by examining the theory itself. Williamson (2024) defends an abductive methodology judging theories by their simplicity, elegance, and explanatory power. On any of these accounts, a philosophical contribution is a text exhibiting certain properties; the question who or what produced it does not enter the evaluation.2 Other conceptions locate philosophy in the practitioner. For Hadot (1995), philosophy is a practice of self-transformation; for the later Wittgenstein (1953), it is a form of therapy; for Merleau-Ponty, it requires us to “slacken the intentional threads which attach us to the world” (1945, p. xv) in order to examine them. What these views share is a commitment: philosophy requires being a certain kind of subject, capable of self-transformation, or therapy, or phenomenological observation—and machines are not such subjects. On Nietzsche’s account (as Sorgner reads it), philosophers are creators of values, expressing drives and psychophysiology that LLMs lack. To a proponent of any such view, the question of whether LLMs can do philosophy is closed before it opens.34 I do not take a position here on which conception of philosophy is correct. This paper assumes the text-focused conception. If philosophical evaluation concerns proper2Pigliucci offers a related formulation: philosophy “attempts to clarify things, or to analyze in order to bring about understanding, not really to discover new facts, but rather to evoke rational conclusions arising from certain ways of looking at a given problem or set of facts.” Whether such evocation requires a human evoker is the question at issue. 3On transformative conceptions, what makes an activity philosophical is something that happens in the practitioner rather than anything assessable in what she produces (Hadot 1995; cf. late Wittgenstein on philosophy as therapy). Transcendental and phenomenological approaches presuppose having experience (Kant 1781/1787; Merleau-Ponty 1945). World-view conceptions require the philosopher to live a human life (Dilthey; see Overgaard, Gilbert & Burwood 2013: ch. 8). Jones (2006) holds that philosophy requires entering an identity-conferring conversation within a community; Sorgner reads Nietzsche as requiring biology and psychophysiology. 4The distinction between text-focused and practitioner-focused conceptions maps imperfectly but suggestively onto the analytic/continental divide: analytic philosophy tends to emphasise texts and arguments as the locus of evaluation, while continental traditions more often locate philosophical activity in lived practice or self-transformation. ties of texts—coherence, handling of objections, illumination of subject matter —then whether LLMs can do philosophy is a question about the texts they produce. Even within the text-focused framework, some argue that LLMs cannot produce texts exhibiting the right properties. These objections do not concern what philosophy is; they concern what philosophical reasoning requires. Floridi, Nobre, and Taddeo (2024) argue that genuine abductive reasoning is beyond LLMs’ capacities; Zahavy (2026) argues that they cannot make the creative leaps needed to propose new theoretical frameworks. If either argument succeeds, LLMs cannot do philosophy regardless of what we think philosophy is. I argue that LLMs can produce philosophy exhibiting the relevant properties. Section 1 argues that philosophical evaluation concerns text-internal criteria. Section 2 presents objections from Floridi et al. and Zahavy. Section 3 responds. Section 4 considers what a demonstration would look like. 1. Philosophy in the Text When Watson and Crick published their paper on the double helix in 1953, they announced what they had found: a particular arrangement of nucleotides, with two strands running in opposite directions, held together by hydrogen bonds between complementary base pairs. The double helix existed before they described it; their paper reported what was already there. Had Rosalind Franklin announced it first, the finding would have been the same—the same arrangement, the same base pairings —just differently attributed. Quine’s “Two Dogmas of Empiricism,” published two years earlier, is not like this. Quine made arguments: against the coherence of the analytic/synthetic distinction, against reductionism about meaning. The arguments are the contribution. There is no arrangement of facts the paper reports, no prior reality that someone else might have found instead. A different philosopher reaching the same conclusions would have had to make arguments. If the arguments differed, so would the contribution.5 What, then, are we evaluating when we evaluate a piece of philosophy? Not whether a text accurately reports something prior to it, since there is nothing prior that it reports. We evaluate the arguments themselves. But for what? 5Literature has a similar character: when we evaluate a novel, we assess prose, pacing, tension—features internal to the text. Philosophy shares this constitutive character but differs in evaluative criteria: literature is assessed aesthetically, philosophy for argumentative virtues. Consider Lipton’s distinction between likeliness and loveliness. The likeliest explanation is the most probable. The loveliest is the one that would provide the deepest understanding if it were true. Lipton’s point is that we can assess loveliness independently of likeliness. Semmelweis investigated why women in the First Division of the Vienna maternity hospital died at higher rates than those in the Second. He considered several potential explanations: differences in birthing position, differences in the route taken by the priest administering last rites, differences in exposure to cadaveric matter. Each was a potential explanation—something that would explain the difference if true. And Semmelweis could assess their loveliness without yet knowing which was correct. The cadaveric explanation was lovelier because it unified the phenomenon with known facts about infection. The priest explanation, even if true, would leave it mysterious why the priest’s presence caused death. Philosophical evaluation has this character. We assess arguments for properties that can be judged from the arguments themselves, without first establishing that their conclusions are correct. These properties include elegance and unity—a good theory explains much with little, and its parts hang together rather than being a collection of separate claims. Williamson notes that these are the same theoretical virtues that guide theory choice in science, but in philosophy they must be weighed without direct empirical test. He compares the philosophical cycle of analysis, counterexample, and revised analysis to overfitting in statistics. Each repair to accommodate a new counterexample risks making the theory more ad hoc, more gerrymandered to the cases at hand. Deep Blue plays good chess. Its moves respond effectively to threats, secure positional advantages, and contribute to coherent strategic plans. But Deep Blue does not play creatively—it searches exhaustively rather than intuiting the best move. Gaut observes that this shows creativity and domain-specific excellence can come apart. A move is good chess or it is not, regardless of whether it was found by creative insight or brute computation. Whether a philosophical argument handles objections well, draws distinctions at the right places, or illuminates its subject matter is assessable in the same way. The question is what properties the argument has, not how it came to have them. Dellsén and colleagues argue that philosophical progress consists in putting people in a position to increase their understanding. Suppose a scientist publishes an important finding and then dies. Everyone who read the paper also dies, or forgets what they read. Has the progress been lost? No. Progress occurred when the publication made it possible for someone to understand, whether or not anyone actually did. The materials remain publicly available; that is what matters. The same applies to philosophy. A published argument constitutes progress if it enables understanding, regardless of who or what produced it, and regardless of whether anyone currently grasps it. Blind review operationalises this: referees assess whether an argument handles objections and illuminates its subject matter without knowing who wrote it. The practice treats authorship as irrelevant to evaluation. If philosophical evaluation concerns properties of arguments—elegance, coherence, illumination of subject matter—and these properties are assessable by reading the arguments, then the production process is not evaluatively relevant. The question is whether a text exhibits these properties, not what brought it into existence.6 2. LLMs and Abduction I want to argue that the objections to LLM philosophy presented in the previous section do not apply to the kind of philosophical work that Williamson describes. Floridi et al. and Zahavy identify capacities that LLMs lack—hypothesis evaluation in the one case, embodied simulation in the other—but these capacities are not required for what Williamson calls philosophical abduction. Consider first how LLMs produce their outputs. An LLM predicts the next token in a sequence based on probability distributions learned from training data. When prompted to explain why a car might not start on a cold morning, it generates text that exhibits explanatory structure: it identifies a hypothesis (the battery), provides a reason (cold weather reduces battery efficiency), and presents the explanation with the connectives and qualifications that explanations typically have. But the LLM does not select this explanation by comparing it with alternatives and judging it best. It outputs the most probable continuation given its training. Floridi et al. put the point this way: Given a prompt, they generate a plausible continuation (a hypothesis or explanation) based purely on learned associations. In reality, their operation is driven by maximising 6Different theorists articulate these criteria differently. Williamson emphasises elegance, unity, and non- ad-hocness (2024, pp. 152–3). Bengson, Cuneo, and Shafer-Landau organise evaluative criteria into five levels: accommodation, explanation, substantiation, integration, and virtue (2022). Dellsén and colleagues cash out philosophical progress in terms of representing dependence relations accurately and comprehensively (2024). The vocabularies differ, but all concern properties assessable from theories themselves. the probability of the sequence… The model does not understand what an explanation is, but it produces text that follows the typical phrasing and structure of explanations. It does not reason about causes from scratch but outputs typical causes for typical effects observed in the training data. Floridi et al. call this zeroth-order abduction. The phrase marks an absence: what is missing is the comparative evaluation that genuine abduction involves. In genuine abduction—what Floridi et al. call strong abduction—one generates multiple hypotheses, compares them, and selects the best. LLMs do not do this. They generate a plausible continuation without evaluating whether that continuation is better than alternatives they did not generate. This matters for some questions. It matters, for instance, if we want to know whether LLMs reason in the way humans reason. Floridi et al.‘s answer is that they do not: the mechanism is stochastic, not inferential. But it matters less if the question is whether LLMs can produce outputs that meet philosophical standards. For philosophical evaluation concerns the output—whether the theory is elegant, unified, and handles the evidence—not the process that generated it. Floridi et al. themselves note this: “if an AI can generate the same explanatory hypothesis a human would, does it matter that the process was different? From an epistemological standpoint, perhaps yes—justification is significant—but regarding the content of the hypothesis and our interpretation of it, maybe not.” Zahavy’s objection cuts differently. His concern is not that LLMs fail to evaluate hypotheses but that they cannot generate certain hypotheses at all. His paradigm case is Einstein’s formulation of the equivalence principle. Einstein did not have data sufficient to infer general relativity inductively; Newtonian mechanics faced no empirical crisis, and the anomaly of Mercury’s perihelion was attributed to an undiscovered planet rather than a flaw in Newton’s laws. Nor could Einstein deduce the equivalence principle from prior axioms—it was itself a new axiom, something that had to be formulated before deduction could begin. How, then, did Einstein arrive at it? Zahavy’s answer is manipulative abduction: generating hypotheses through embodied simulation rather than symbolic manipulation. Einstein imagined himself inside a falling elevator. He simulated the sensations of an observer in that scenario—objects released from the hand appearing to hover, the floor rushing up to meet falling things—and abduced from that simulated experience that gravity and acceleration must be the same phenomenon. The thought experiment was not a logical exercise conducted in symbols but a sensory one conducted in imagination. LLMs, Zahavy argues, cannot do this. They operate entirely in the domain of symbols —tokens, vectors, probability distributions—without access to the physical referents those symbols represent. Zahavy quotes Harnad’s phrase: LLMs are “high-dimensional ‘Chinese Rooms’, manipulating the language of physics without access to the physical referents that give that language meaning.” They can derive consequences from axioms once those axioms are given in symbolic form, but they cannot make the leap from sensory experience to new axioms that Zahavy takes to be constitutive of scientific invention. Zahavy limits his argument to physics: “this proposal is specifically tailored to the physical sciences, where the object of study is external material reality.” But the argument structure extends to phenomenological experience more broadly. If LLMs lack subjective experience altogether, then philosophy that relies primarily on phenomenological observation will be difficult for them. Similar considerations apply to intuitions (Machery 2017) and to aesthetic experience. In the next section I argue that these considerations are less damaging than they appear. 3. Thought Experiments and Armchair Abduction Williamson characterises philosophical abduction in terms general enough to encompass both science and mathematics. Theories are ranked as potential explanations of a body of evidence, and the ranking depends on two things: how well the theory fits the evidence, and how well it scores on what Williamson calls the intrinsic virtues of a good theory: Apart from its relation to E, the more T has the intrinsic virtues of a good theory, the better (ceteris paribus). It should be elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated. It should be informative and general. In brief, it should combine simplicity with strength. (Williamson 2024, p. 354) The evidence base for philosophical abduction is unrestricted. Williamson writes that “nothing in this account requires the evidence propositions, the explananda, to be of some special kind. Any known truths will do” (p. 355). In philosophy, this evidence consists in arguments, counterexamples, thought experiments, and the distinctions and results that prior inquiry has established—in short, the accumulated textual record of the discipline. The explanations philosophy offers are typically constitutive rather than causal: accounts of what something consists in, how concepts relate, what follows from what. Williamson draws an explicit analogy with mathematics: “mathematics is a precedent for a successful discipline with an ‘armchair’ methodology that still has a key role for abduction. Thus it would be myopic to assume that an abductive methodology for philosophy implies its assimilation to the experimental sciences” (p. 358). Philosophical abduction can be conducted from existing knowledge, without new empirical observation, and evaluated by examining the theory itself rather than comparing it to mind-independent facts. This matters for the question of whether LLMs can produce good philosophy. The philosophical corpus—the body of philosophy that has survived peer review, been taught, been cited, and been anthologised—exhibits the intrinsic virtues Williamson identifies. This is not accidental. Peer review is a filter: papers that fail to handle objections, or draw arbitrary distinctions, or offer no illumination of the subject matter, are rejected. What survives to be published and taught is a sample of what the discipline judges good, where the standards of goodness track exactly the intrinsic virtues Williamson describes. An LLM trained on this corpus has learned the distribution. It has learned what makes a philosophical explanation score well—not through explicit instruction, but through exposure to a body of text that has been filtered by those standards over centuries. When Floridi et al. describe LLMs as “engines of generative plausibility,” they are describing systems that have absorbed, from the corpus, the evaluative standards that philosophical abduction employs. Floridi et al.‘s diagnosis—that LLMs produce plausible outputs without evaluating alternatives—is correct at the level of mechanism. But what counts as “plausible” in philosophy is exactly what scores well on Williamson’s virtues; and what scores well on those virtues is exactly what the corpus encodes. Turn now to Zahavy’s argument. Zahavy claims that scientific invention requires a leap from sensory experience to formal axioms—what he calls the E→ A Jump— and that this leap involves embodied simulation rather than symbolic manipulation. Einstein imagined the sensations of an observer in a falling elevator, and from that simulated experience abduced the equivalence of gravity and acceleration. LLMs, lacking access to physical referents, cannot make this leap. Does philosophy require something similar? Consider how philosophical thought experiments actually work. Take Putnam’s Twin Earth case. Putnam asks us to imagine a planet where the clear liquid in the lakes and rivers is not H2O but a different chemical compound, XYZ, which is superficially indistinguishable from water. Oscar, on Earth, and Twin Oscar, on Twin Earth, both use the word “water” to refer to the clear liquid in their environments. They are molecule-for-molecule identical in their internal states, yet—Putnam argues—they mean different things by “water.” Oscar means H2O; Twin Oscar means XYZ. The conclusion: meaning is not determined by what is in the head. Notice what this thought experiment does not require. It does not require anyone to simulate the sensations of being on Twin Earth or drinking XYZ. The thought experiment is articulated entirely in language, recorded in text, and does its intellectual work at the level of concepts and propositions. Readers evaluate it by asking whether the scenario is coherent, whether the conclusion follows, whether the argument illuminates something about meaning—and all of these questions can be answered by examining the text. The same is true of Jackson’s Mary, Searle’s Chinese Room, Parfit’s teleporter, and every other philosophical thought experiment in the literature. They are textual objects, and the work they do is textual work. Zahavy’s model of creative invention—sensory experience, embodied simulation, formal axioms—fits Einstein’s physics, where the object of study is external material reality and the axioms must connect to that reality through grounded concepts. But philosophical thought experiments do not make this demand. They are already articulated in language; they enter the record as text; and their intellectual contribution consists in the arguments they embody. Whatever private experiences philosophers have in arriving at thought experiments, the thought experiments themselves—as they enter the literature and do philosophical work—are linguistic objects. But what of the broader objection—that LLMs lack phenomenological experience, intuitions, and aesthetic response? We do not claim that LLMs have these capacities. What they have is a training corpus containing extensive descriptions of human experience. This is not first-hand access to experience but access to descriptions—and descriptions of experience are what philosophical argument typically works with. This bears on the question of novelty. Zahavy’s argument, if it worked, would suggest that LLMs cannot produce genuinely new theories—that they are limited to recombining existing materials. But what does philosophical novelty consist in? Williamson notes that “enumerative induction is inadequate for systematic philosophical theorizing, which often requires introducing new distinctions at a more abstract level not given in the data” (p. 353). Dummett’s distinction between assertoric content and ingredient sense, for instance, is not derived from data by induction; it is a conceptual innovation, a new way of organising existing materials. Kripke’s framework of rigid designators similarly reconfigured existing materials about reference, identity, and modality. Lewis’s modal realism took the apparatus of possible-world semantics and proposed treating it with full ontological seriousness. These are innovations, but they are conceptual innovations—new ways of carving up the problem-space, new distinctions at a higher level of abstraction. They do not involve leaps from bodily sensation to formal axioms. An LLM trained on a corpus that documents both the conceptual materials and the standards for combining them has, in principle, the resources to produce similar reconfigurations. Whether it will in fact do so is an empirical question. But the arguments that it cannot—Zahavy’s arguments about embodied simulation and physical grounding—do not straightforwardly apply when the materials and the standards are textual. Lipton makes a related point about the relationship between different levels of description: If these suggestions are along the right lines, then arguing that Inference to the Best Explanation is wrong because Bayesianism is right is like arguing that thinking about technique cannot help my squash game because the motion of the ball is governed by the laws of mechanics. (Lipton 2004, p. 108) The ball’s motion is governed by mechanics, and an LLM’s outputs are generated by stochastic processes. But the mechanical description of the ball does not make thinking about technique pointless, and the stochastic description of the LLM does not settle the question of whether its outputs meet philosophical standards. Floridi et al. correctly identify the mechanics; philosophical evaluation concerns a different level—the level at which we ask whether the arguments are clear, whether they handle objections, whether they illuminate the subject matter. These are questions about the output, not about the process that generated it, and they are the questions that matter for determining whether what the LLM has produced is good philosophy. References Frankish, K. (2024). What are large language models doing? In A. Strasser (Ed.), Anna’s AI Anthology: How to live with smart machines? (pp. 55–78). Xenomoi. Assistant: I need to invoke all the required skills first, then read the relevant project materials before engaging with this substantial request. [Tool Skill input: skill=contemplate status=completed] User: you got cut off, please continue from where you left off. ### Assistant Let me continue invoking the remaining skills and then get to work. ### Assistant Let me start by reading the project materials, knowledge files, and source texts in parallel. ### Assistant Now let me search for past sessions, read the source texts, and check the current longform manuscript files in parallel. ### Assistant Good, all five agents are working through the materials. Let me also read some additional resources while waiting — the Lipton extraction and the Bengson source, plus the CEV skill itself. ### Assistant The CEV skill file wasn't found at that path. Let me find it, and also read the Lipton and Floridi extractions from where they actually are. ### Assistant Good, the extractions exist. Let me find the CEV skill and read the Williamson extraction (the agents are handling Floridi and Zahavy already). ### Assistant Good, this is the beginning of Williamson's chapter. The passage on theoretical virtues and abduction in philosophy is likely later in the text. Let me read the section where he discusses abductive methodology and theoretical virtues specifically. ### Assistant Excellent — I now have the Williamson text on theoretical virtues and abduction. The passage at lines 1622-1625 is the one your draft already quotes. Let me now also find the CEV skill and check on the background agents. ### Assistant Good — the session search found the relevant sessions. The most substantive are the March 3 "Map moves" (33 turns) and "Analyze co-author discussion" (54 turns) sessions, plus the March 4 "Analyze co-author transcript" (13 turns). Let me read the CEV command and then start deep-reading the most relevant past sessions while waiting for the other agents. ### Assistant Good, I have the CEV format. Now let me wait for the remaining agents to complete — particularly the source texts, manuscript files, and vault notes agents, which are doing the heaviest lifting. ### Assistant The Williamson agent is done. Let me now check on the remaining agents and collect their results. ### Assistant Excellent. The Floridi/Zahavy agent, manuscript agent, and vault notes agent have all completed. I now have a comprehensive picture. Let me update the todos and wait for the two session-reading agents to complete, then produce the CEV.