# Generating Philosophy with Artificial Intelligence ## Cleaned transcript *Editorial note: This transcript has been cleaned against the accompanying draft and slides. Repetitions, false starts, filler words, and obvious transcription errors have been removed. The pre-session conversation and technical troubleshooting have been omitted. The substance and oral character of the talk and discussion have been retained. Where a questioner’s name could not be established confidently from the recording, a descriptive label is used.* ## Introduction **Chair:** Our speaker is currently a postdoctoral researcher on Enrico Terrone’s ERC project, *The Philosophy of Experiential Artifacts*, at the University of Genoa. Before that, he was a postdoctoral researcher at the Centre for Philosophy of Time at the University of Milan. His research covers a wide range of topics, including the philosophy of perception, art, and, most recently, technology. Today’s talk is based on joint work with Enrico Terrone and is called “Generating Philosophy with Artificial Intelligence.” I’ll hand over to Nick. ## Talk **Nick:** Thanks for the introduction and for the invitation. It is nice to be speaking to you all. I am going to begin with something I do not normally include in presentations, which is a brief personal introduction. I started working with Enrico in early 2022, which was more or less the moment when generative AI began appearing prominently in the news and becoming something that many people actually used. I have been fascinated by it from the start. I am not only writing about philosophy and AI. By now, apart from reading, I do nearly all my work with LLMs. I would describe what I do as writing philosophy with an LLM. As we will see toward the end of the talk, there is an interesting question about how much of the resulting work is mine and how much is the model’s. I am reasonably confident that most of it is still coming from me, but you will see what I mean. Because I use these systems so much and am so interested in them, I am aware that I can sound slightly evangelical. Many people in the humanities dislike this technology, and there are many good reasons for concern, including serious moral questions. Nevertheless, I do not think the technology is going away. What interests me in the background is how we might use it in a nourishing way, one that genuinely helps us. I am not going to discuss the morality of using these systems today. Those questions are obviously important, but there are other interesting questions as well, and one of them is the research question of this paper. Part of the motivation is simply genuine enthusiasm. Over the last few years, I have become much more confident intellectually as a philosopher. Working with LLMs has exercised my philosophical abilities in a way that earlier technologies did not. I also have a slightly contrarian streak. I rather enjoy liking something that many other people in the humanities dislike. I am also interested in the thought, suggested by the familiar sequence of blows to human self-importance, that part of what makes AI so unsettling is that it threatens to knock us off our intellectual perch, and perhaps our creative perch as well. AI has already contributed to interesting discoveries and is achieving increasing success in physics, mathematics, biomedicine, materials science, and other disciplines. Another way of framing this paper, then, is to ask whether philosophy should expect similar successes. The claim I will defend is this: > Current LLMs can produce philosophical texts that are worth reading. By “worth reading,” I am trying to avoid settling large metaphilosophical questions about the ultimate purpose or value of philosophy. I want to appeal to a practical distinction that everyone in this room understands. When we read a philosophy paper or book, or listen to a philosophy talk, we hope it will be worth our time. We hope it will be philosophically nourishing in some way. That is what I am suggesting current LLMs can produce. I do not mean systems that might exist in the future. I mean systems available now. A text’s being worth reading does not mean that it is correct. We have all read philosophy that we regard as valuable even though we disagree with some or all of it. Nor is a bare philosophical pronouncement usually worth reading. This is not *The Hitchhiker’s Guide to the Galaxy*, where we build a computer and it announces the answer. A book consisting only of unargued pronouncements, whether written by a person or a model, would fall on the wrong side of the distinction. What matters is argument. My claim is that LLMs can produce arguments of sufficient quality and novelty to be worth reading. I will defend this claim against six challenges. Some will take longer than others, and the section on abduction will be the longest. ### 1. The challenge from authorship The first challenge is the challenge from authorship. Sometimes, when I speak to philosophers about this topic, I sense an intuition behind what they say. Even if an LLM produced a text that looked exactly like a philosophical argument, they would still refuse to count it as philosophy. The thought seems to be that there is no person behind the text. No philosophizing has occurred and then accumulated in the output. This can lead to a fairly blank refusal to accept that such systems might have anything useful to contribute to philosophy. An imperfect comparison can be made with art. If I generate an image with a minimal prompt, some people will say that it is not an artwork because there is no artistry behind it. We can ask whether something similar is true in philosophy. There is some initial support for this thought in the way philosophy is organized. Undergraduate and graduate courses are often devoted to particular philosophers or to what a particular philosopher thought about a particular subject. We do not see the same thing to anything like the same degree in scientific disciplines. I assume that physics undergraduates do not take many courses simply on Isaac Newton as a person. To put some flesh on the challenge, consider David Davies’s performance theory of art. Davies says that the work, what the artist achieves, is the process that eventuates in the product. Works are intentionally guided generative performances that culminate in contextualized objects or structures. On this view, when a painter produces a painting, the canvas is what we attend to, but it is not the work. The work is the activity that led to the canvas. Someone might try to transpose this view to philosophy. The philosophical work would be the philosophizing, and in the LLM case no philosophizing has taken place. Davies uses a case in which a storm blows pigment across a canvas and creates something indistinguishable from an abstract painting. He would say that there is no artwork there because no artistic activity led to it. The trouble is that this does not seem to transpose successfully to philosophy. We do not judge a philosophical text on the basis of the activity that preceded it. Peer review is partly designed to separate the origin of a text from its content. Imagine that the same storm traces an argument in English in the sand, perhaps an argument against naïve realist theories of perception. If it were coherent, I would still want to read it, and we could still judge it on its merits. What makes a philosophical text worth reading is not who wrote it, or even whether anyone wrote it. What matters is what is on the page, or, in the case of a spoken argument, what is said. That is the response to the first challenge. The next three challenges are different. They concern capacities that an LLM might lack. Think of a parrot. If a parrot somehow produced an interesting philosophical argument, it would be worth recording and considering. But parrots do not produce novel philosophical arguments. They lack capacities needed to do so. Perhaps LLMs are in the same position. Perhaps nothing excludes their outputs from philosophy by definition, but they lack something that worthwhile philosophy requires. ### 2. The challenge from abduction The first capacity challenge concerns abduction. It says that LLMs cannot properly perform abductive inference, or inference to the best explanation. Suppose I go downstairs and find water on the kitchen floor. There are several possible explanations: a leaking tap, a wet mop, a burst pipe, or rain through an open window. If the water is directly beneath an open window and it rained during the night, rain coming through the window is the most plausible explanation. The evidence does not deductively entail it, but it is the best explanation available. One way of ranking competing explanations is in terms of “loveliness,” a term I take from Peter Lipton. The loveliness of a theory is the understanding it would provide if it were true. When I infer that rain came through the window, nothing rules out an unlikely chain of events that produced the same evidence. Given the available materials and candidate explanations, however, the rain hypothesis is the most elegant, unified, and non-arbitrary. Williamson and many others argue that abduction plays a major role in philosophy. If an LLM cannot perform abductive inference, we therefore have a serious problem. Floridi and his colleagues argue that LLMs cannot do it. They describe these systems as having “a stochastic core and an abductive appearance.” LLMs are not trained to aim at truth. They are trained to predict the next probable word or token. Nevertheless, their outputs sometimes appear abductive. Floridi and his colleagues explain this by saying that LLMs have absorbed patterns of human abductive reasoning as expressed in text. Their example concerns a car that will not start on a cold morning. A model may say that the battery is weak, that the oil has thickened, or that there is some other problem, before concluding that the battery is the most likely explanation. The output contains the kinds of words a person might use while making an abductive inference. But, Floridi argues, the system has not actually performed such an inference. It is functioning as a brainstorming assistant. It throws out explanation-like material without filtering it, and the human user must identify the good explanation. The difficulty for this view is that the model’s answer in the car case is often good. Floridi and his colleagues acknowledge that it may give the same explanation a human would choose, perhaps even the optimal answer by inference-to-the-best-explanation criteria. So the output is plausible, but they still insist that no abduction has occurred. Their reply is that the system is fitting common patterns. It knows what explanations of a car’s failure to start on a cold morning normally look like. The façade will crack, they suggest, in uncommon or genuinely novel cases. There is a mixed empirical literature on this. Some benchmarks find LLMs relatively strong at abduction, while others find that they perform worse on abduction than on deduction. There is also disagreement about what these benchmarks actually test. I will leave the empirical question to one side for now and begin my response with Stephen Wolfram’s discussion of grammar. An LLM learns grammatical rules without being given an instruction sheet explaining English grammar. It is exposed to a huge quantity of text and acquires the ability to produce well-formed sentences. It does not construct those sentences in the way we do. Its linguistic capacities are realized differently from ours. But nobody says that it produces only a veneer or façade of grammaticality. Its sentences really are grammatically correct. Perhaps something similar is possible with abductive inference. The model need not abduct in the way a person does in order to acquire the capacity to produce text that contains an inference to the be st explanation. The obvious objection is that loveliness is not grammar. Grammatical competence gives us no reason to expect good explanations. That might seem to be comparing apples with oranges. But Wolfram points out that LLMs do not ordinarily produce sentences such as “Inquisitive electrons eat blue theories for fish.” That sentence is grammatically perfect, with nouns and verbs in the right places, but it is meaningless. Similarly, when you ask why a car will not start on a cold morning, a model does not usually say, “There are no squirrel tracks, so freak Arctic winds must have blown into the exhaust pipe.” That has the superficial form of an explanation but is obviously nonsensical. Instead, the model gives plausible possibilities such as a weak battery or thickened oil. Wolfram suggests that models develop what he calls a “semantic grammar”: rules governing what can sensibly be said of what. He does not develop the idea very far, but he also suggests that patterns such as syllogistic logic can be learned from the corpus and used to produce correct inferences. My suggestion is that an LLM can infer, in the relevant output-level sense, insofar as its semantic grammar captures which claims can sensibly be made about which things and which conclusions follow from which considerations. It learns the connective vocabulary of inference, words such as “therefore,” “however,” and “despite,” together with constraints on their sensible use. That may give it a capacity to produce abductive text even though it does not perform human-style abduction internally. This leads directly to the next challenge. Even if the model can produce something abductive in flavor, it may seem unable to apply that capacity to the world. It takes text in and produces text out. ### 3. The challenge from detachment This is the challenge from detachment. If I really have a car that will not start, the model has no direct access to my particular car. It cannot know the relevant facts about it or check its hypothesis against the world. LLMs are supposedly disconnected from the world, and therefore unable to make or test abductive inferences about actual cases. In 2026, this claim requires qualification because LLMs can use web search and other tools that provide some connection to the world. But I will put those tools aside and focus on the language model itself. Tomer Zahavy, a computer scientist rather than the Danish phenomenologist Dan Zahavi, develops a version of this worry in relation to science. He discusses Einstein’s elevator thought experiment and the development of the equivalence principle. Einstein imagined an elevator accelerating through space and considered what would happen to objects inside it. He then advanced a principle that could be tested against the world. On Zahavy’s account, Einstein needed a connection with the world at both ends. He needed experience of the world in order to imagine the physical situation, and the resulting principle then had to be checked against reality. Zahavy sometimes moves between the model’s lack of worldly access and its lack of conscious experience. I think the central issue in his argument is access to the world. I will treat conscious experience separately in the next section. My response turns on a difference between the ways science and philosophy relate to the world. Philosophy is not cut off from the world, but it normally relates to it differently from empirical science. Massimo Pigliucci offers a useful way of thinking about this in terms of evocation from empirical data. The model is similar to mathematics or chess. Once the rules of chess are in place, we can work out an enormous number of possible games and objective truths about them. We are not simply inventing whatever we please. We are constrained by the starting rules. The same is true of mathematical axioms. We begin from an established structure and work out what follows within it. Pigliucci suggests that philosophy also proceeds by evocation, but its starting points are constrained by empirical facts about the world. I am setting experimental philosophy aside and thinking mainly about ordinary analytic philosophy. Analytic philosophers still begin from empirical assumptions, usually commonsense truths that do not require fresh investigation: there are people, objects ordinarily look colored, and so on. We start from things we take to be true of the world and then evoke a constrained conceptual structure. This differs from invention. A science-fiction writer can decide that the laws governing a fictional world are different. In philosophy, once the starting assumptions are fixed, the structure has rigid properties. What follows is no longer simply up to the author. We can therefore accept Zahavy’s claim that the model lacks a direct connection with the world while observing that its corpus contains an enormous amount of writing about the world. Through its semantic grammar, it acquires a good grasp of the commonly accepted starting points from which philosophical arguments develop. The empirical materials philosophy ordinarily uses already reach it in linguistic form. ### 4. The challenge from experience The third capacity challenge is the challenge from experience. LLMs have no phenomenology. They have no conscious or subjective experience. Perhaps this does not rule out philosophy in general, but it might rule out work in which phenomenology or consciousness plays a central role. Consider Mary in the black-and-white room. How could a system that has never seen red engage philosophically with the question of what it is like to see red? I am willing to accept that current LLMs have no phenomenology. But we can make a move similar to the one made in the previous section. Although LLMs do not have experiences, the corpus contains a vast quantity of articulated phenomenology, that is, human descriptions of experience set down in words. This includes philosophical descriptions, such as those in J. L. Austin’s *Sense and Sensibilia*, but also memoirs, novels, diaries, interviews, psychiatric reports, and countless other texts in which people describe what experience is like. The model has never felt weightlessness, but it has read descriptions written by astronauts. It has never seen red, but it has read a great deal about color experience. Philosophical texts that concern phenomenology rarely require the reader to inspect every fine-grained detail of their own experience while reading. A paper in the philosophy of perception usually presents articulated claims about experience, and the philosophical work proceeds by reasoning from those claims. The model has access to them in the same way any reader does. There is, however, an important limitation. An LLM cannot produce a genuinely novel phenomenological observation by attending to an experience of its own. My tentative example is Merleau-Ponty’s observation that, when one hand touches the other, the roles of toucher and touched can alternate, but cannot be occupied simultaneously in the same way. If he was the first person to notice and articulate that feature of experience, then I do not see how an LLM could have made the observation independently. It has no hands and no tactile phenomenology. So there may be a limit here. The model can reason from articulated phenomenology, but it cannot be the original source of a novel observation that requires having the experience. ### 5. The challenge from observation We now come to the fifth challenge. The first challenge asked whether an LLM output could count as philosophy at all. The next three concerned abduction, connection with the world, and phenomenology. The final two are not capacity challenges in quite the same way. The challenge from observation asks an obvious question: if everything I have said is correct, where is all the worthwhile LLM philosophy? Ask a model about the meaning of life and you are likely to receive a bland survey of familiar positions, a Monty Python joke, or a reference to *The Hitchhiker’s Guide to the Galaxy*. Something must be preventing these systems from producing worthwhile philosophy, because they do not seem to be doing so. My response is that we have not yet learned, in any general way, how to prompt these systems so that they produce it. Even people who understand how LLMs work are tempted to treat them as persons or as oracles. If you ask a factual question, especially now that models can search the web and hallucinations have diminished, you may immediately receive a correct answer to something quite obscure. It is therefore tempting to ask the oracle a philosophical question and expect a philosophical answer. But LLMs produce reasonable continuations of texts. They generate words that are probable or reasonable given all the words that have come before. The likely continuation of “What is the meaning of life?” or “What is the true theory of perception?” is not a perfect analytic philosophical argument. In ordinary life, even a professional philosopher confronted with such a question would not instantly produce one. The issue is not simply what question we should ask, but what input we need to provide in order to make philosophical argument a reasonable continuation. One obvious method is to request careful, rational, step-by-step argument. Another is to provide a substantial body of philosophical material. You can give a model several books or a long draft and ask it to decide which position is better, or to argue for one account over another on a specific question. Doing so makes argumentative vocabulary and argumentative relations a more probable continuation. It supplies concepts and commitments that the response must connect. In slogan form, we need to put philosophy in to get philosophy out. But that immediately generates the final challenge. ### 6. The challenge from instrumentality If a model can produce philosophy only when a philosopher gives it a philosophical prompt, who is really producing the worthwhile philosophy? Perhaps the philosopher is doing all the work, while the LLM merely rearranges or remixes what the prompt already contains. Saying that the model produced the philosophy might then resemble saying that a ventriloquist’s dummy produced the speech. At one level, this is trivially correct. I can paste a worthwhile section of a paper into a model and ask it to correct the spelling. If it returns a typo-free version, then in one very weak sense it has produced worthwhile philosophy. But the philosophy was already present. I wrote it; the model merely tidied the text. The mistake would be to react by claiming that the LLM always does all the philosophy and the prompt contributes nothing. That is obviously false. The right answer depends on what the prompt contains and what the continuation adds. There is a movable line between the prompter’s contribution and the model’s. To illustrate this, I will borrow an analogy Enrico and I use in our work on AI image generators. Consider the relation between a gardener and a garden. It differs from the relation between a painter and a canvas. A painter can, in principle, be responsible for nearly every mark on a surface and decide exactly where each one goes. A gardener does not have that degree of control. The gardener plants, prunes, directs, coaxes, and marshals, but the garden also develops under its own biological and physical constraints. The balance varies. A formal French garden displays a high degree of control, with geometrical beds and carefully shaped plants. An English country garden may preserve a degree of overgrowth and allow nature more freedom. In each case, the line between what the gardener controls and what the garden contributes is different. Something similar applies to image generation and to LLMs. A prompt is not exactly a seed, but it is an articulated starting point. A good philosophical prompt elicits an evoked philosophi cal structure through the model’s semantic grammar. Depending on how much work the prompt does and how capable the model is, responsibility for the resulting philosophy may lie more heavily on one side or the other. A very detailed prompt may contain most of the philosophical work, just as a highly formal gardener controls much of the garden. A more minimal prompt, given to a more capable model, may leave the model responsible for much more of the resulting development. The point of the final section is to unite evocation and semantic grammar. By marshaling the model’s semantic grammar into the service of philosophy, a prompter can cultivate a philosophical development in a particular direction. The resulting division of responsibility is not captured by the typewriter or ventriloquist analogy. The garden is a better model. Let me conclude by returning to the six challenges. First, authorship does not determine whether a philosophical text is worth reading. Second, LLMs do not perform inference to the best explanation in the same way humans do, but they can produce abductively structured text through the semantic grammar they acquire from training. Third, philosophy’s relation to the world differs from that of empirical science, and the worldly starting points philosophy uses ordinarily reach it in linguistic form. Fourth, LLMs have no conscious experience, but they have access to a vast body of articulated phenomenology. They cannot make novel first-person phenomenological observations, but they can reason philosophically from observations that have been articulated. Fifth, the present scarcity of worthwhile LLM philosophy does not show that the systems lack the capacity. It may show that we do not yet know how to elicit it reliably. Finally, an LLM is not merely a typewriter or a ventriloquist’s dummy. Its relation to the prompting philosopher is more like the relation between a garden and a gardener, and this gives us a more plausible way to describe their respective contributions. Thank you. ## Questions and discussion ### The practice of philosophy and the educational use of LLMs **Chair / first questioner:** Thank you. That was extremely interesting, especially given the way universities are currently responding to AI. It seems clear to me that AI could produce something worth reading. One possible advantage is that readers might not become preoccupied with assessing the person behind the text. Take students who do not respond well to feedback from tutors and experience it as a personal attack. Feedback generated without a personal author might allow them to focus more directly on the argument and on where the essay went wrong. But I also wondered whether part of the value of philosophy is that it is a practice we engage in with one another. Who do you argue with? An AI might produce a good synchronic piece of philosophy, but philosophy also involves an author defending a position over time, responding to objections, and participating in an ongoing exchange. An LLM could perhaps examine what it previously said and answer criticisms, but it would not be a philosopher in quite the same sense. I am not saying that its output could not be philosophy. My concern is that it might not be participating in the practice or “the game” of philosophy. **Nick:** There is a lot there. On the point about students, feedback that is not presented as coming from a person might indeed be easier to take less personally. It could allow the student to get directly to the content. I also think the imperfection of LLMs can be an educational advantage. Once students realize that they cannot simply trust the model, they may learn to push back against it. An LLM has infinite patience. I can tell it, “That is an outrageously stupid idea; stop wasting my time,” and it will continue working with me. That means one can sometimes cut directly to the intellectual issue without managing another person’s feelings. I know that is not exactly your point, but I do think there are productive educational uses of this kind. On philosophy as a practice, I agree that something important is at stake. I do not want to stop doing philosophy and be replaced by an LLM. Mathematicians may already be confronting a version of this crisis as more problems are solved quickly by automated systems. I do not know where philosophy will be in five or ten years, and I would not want it to become an LLM-only activity. At the same time, I do think these systems will change how we write. I doubt that in ten years philosophers will simply sit in Microsoft Word and write in exactly the way they do now. I may have moved away from your precise question, but I agree that the practice matters, even if I am uncertain whether that leads to an optimistic or pessimistic conclusion. ### Authorship, lookalike texts, and formalism **Dan:** Thank you, Nick. I wanted to return to the authorship argument, but I have two preliminary observations about the gardening analogy. First, there are other media in which we already find a movable line between the author’s contribution and something that is simply there. Photography is an obvious example. The qualities of a photograph arise partly from the photographed world and partly from the photographer’s activity. Tom McClelland has a short paper on generative AI and the Aeolian harp. An Aeolian harp is played by the wind. His suggestion is that generative AI is trained on vast bodies of human-produced music and images, so its outputs partly reflect the culture contained in those datasets. We may not say that AI channels the divine, as people sometimes said of an Aeolian harp, but perhaps it channels the Zeitgeist. That offers another way of thinking about authorship and about standing on the shoulders of giants. My main question concerns your claim that a philosophical text is not worth reading merely because it was written by a particular person. I wonder whether that needs sharpening. Suppose a n unknown manuscript by Wittgenstein were discovered. We would want to read it because it was by Wittgenstein. Perhaps you would say that this gives us a historical rather than a distinctly philosophical reason to read it. Is that your position? **Nick:** In the Wittgenstein case, I would want to read it because his name gives us evidence that the text may be good. Compare a newly discovered philosopher from a hundred years ago, someone nobody has heard of but who is reportedly brilliant. I would want to read that person too. Wittgenstein’s name is a stamp of probable quality rather than something that independently adds philosophical value. I realize that may sound slightly cold. I am simply less moved than you are by the fact that the manuscript would be Wittgenstein’s. **Dan:** There may be a useful way to sharpen your claim by using the standard lookalike cases in aesthetics. Davies thinks perceptually indistinguishable objects can differ in artistic value. A surface made by wind may look exactly like a painting but lack the same value because of its production history. Your claim, by contrast, could be that two textual outputs with the same content must be identical in philosophical value. **Nick:** Yes. That is exactly what I want to say. That is a very helpful way of putting it. **Dan:** In that case, formalism, which is often unattractive in aesthetics, may be exactly what you want for philosophy. **Nick:** Yes. I want to be a formalist about philosophical texts. The Aeolian-harp example is also useful. I had not considered it. Pollock may provide another comparison. In producing a drip painting, he relies on momentum, gravity, and other physical forces. He does not painstakingly place every mark. He allows nature to play a role. That seems somewhat similar to what occurs with an image generator and perhaps with an LLM. The “shoulders of giants” thought is also suggestive. Enrico and I have played with the image of earlier thought decomposing into mulch from which new ideas grow, or earlier images decomposing and then supporting the growth of new ones. These analogies may help us make sense of the contribution. ### Education, intentionality, peer review, and the art analogy **Costas:** Thank you, Nick. You have given me a lot to criticize. I will limit myself to three or four points, and perhaps we can continue over coffee another time. First, I am also a teacher, and I regard AI as extremely dangerous in education. I am trying to develop particular skills in students. If a student thinks they can use AI to produce something that receives a passing mark, I have failed as a teacher. Students need to understand that relying on AI may prevent them from developing the skills the course is meant to teach. Second, I am not persuaded by the comparison between philosophy and art. Plato, for example, insists on a strong distinction between what philosophers and artists do. Saying that philosophy is like art misses a large part of the philosophical tradition. Third, I am not sure that the argument written in the sand is really an argument. It may reveal a thought, but there is no thinker making an inference. Fourth, the point about peer review does not establish complete independence from persons. I submit to a peer-reviewed journal because the editors and referees are people with philosophical education, experience, and a history of doing philosophy. They are my peers. We enter a shared practice on the assumption that we are persons with related forms of experience and training. The same issue arises with abduction. Many philosophers will say that “abduction-flavored” text is not abduction. An inference requires a person performing the inference. Intentionality may therefore provide the strongest argument against LLMs producing worthwhile philosophy. Philosophers have intentions and a philosophical terminology that LLMs do not have. We also want students to acquire that intentionality and terminology. **Nick:** Let me divide that into education, peer review, and the final point about intentionality and philosophical practice. On education, I completely understand the position you are in. It is extremely difficult to teach undergraduate and master’s students when these systems make laziness so easy. I am not underplaying the problem. My response is that the systems are not going away. I am pessimistic about our ability to prevent students from using them. We therefore need to work out how students can use LLMs productively while still developing the relevant skills. I have personally become much better at philosophy through using LLMs since 2022. They have put rocket boosters on my philosophical capacities. But I am importantly different from a beginning student because I spent many years doing philosophy without them. The educational challenge is to put students in a position where they can obtain the benefits without losing the underlying abilities. I do not yet know how to do that. I do not think prohibiting their use will work. AI detection will become a continuing arms race in which models improve, detectors improve, and the models improve again. I do not think academia will win that war. On peer review, I agree that we want experts to assess the paper. But a person outside academia could still submit a paper. Imagine a visionary genius who had never held an academic job, perhaps a figure like Wittgenstein. The journal could send the paper to experts and they could judge it. My point was only that peer review is designed to remove autobiographical detail and focus assessment on the argument as presented. **Costas:** But submitting a paper is part of a collective thinking process. I participate in good faith on the assumption that I am a person and that the reviewers are people with relevant experiences and histories. The reviewers address me on the same assumption. I doubt that an LLM without experience or history can participate in that transaction or take what it is supposed to take from the exchange. And I still object to the comparison between philosophy and art. **Nick:** I am actually arguing against the transposition from art to philosophy. My claim is that, unlike artistic value on Davies’s view, philosophical value is all on the page. That connects with the peer-review point. As Dan put it, I am a formalist about philosophical texts. **Costas:** I do not understand that. If you are a formalist about philosophy, do you discard Hegelian or Continental philosophy because it is not formalist? **Nick:** The philosophers themselves do not need to be formalists. I am using “formalism” in the aesthetic sense Dan introduced. In aesthetics, a formalist says that everything relevant to aesthetic or artistic value is present in the perceptible object, rather than depending on the artist’s biography or process. My analogous claim is that everything relevant to the philosophical value of a text lies in the text. It does not matter whether the text is analytic, Hegelian, Continental, or something else. I care about what it says, not who wrote it. **Costas:** But that still leaves much to explain. Think of Plato and Russell. Both wrote texts, but Plato deliberately leaves much unsaid and invites the reader to think beyond the text, whereas Russell often wants the reader to attend closely to what is explicitly stated. **Nick:** I see what you mean. I do not think I agree, but I understand the challenge. Thank you for the bracing questions. ### Phenomenology, mystical states, and ineffability **Bob:** Thank you for a very good talk. My interests are phenomenology and the paintings of Salvador Dalí. I do not think AI can produce art for the same reason that I do not think it can do phenomenology. One might say that LLMs possess a kind of experience because they have access to an enormous database of texts. They use prior information in something like the way Anil Seth describes the Bayesian brain: we begin with priors, update them in light of evidence, and assess how well the result fits the model. AI can do that very well, and perhaps it can do some forms of philosophy. But I do not think it can do phenomenology. Following Robin Carhart-Harris, I think consciousness is not exhausted by Bayesian processing. Dalí claimed to be a mystic, and I am exploring that claim. Artists may need access to mystical or otherwise expanded states of consciousness of the kind described by William James. Some aspects of experience cannot be written down. They exceed the propositional. To make art, one may need to move into those nonordinary states, and a computer cannot do that. A pocket calculator may perform extremely sophisticated Bayesian operations, but it cannot enter a mystical state. This is why I think phenomenality itself is necessary for phenomenology. In some traditions, if one asks, “What is the Dharmakaya of the Buddha?” an apparently nonsensical answer may nevertheless be philosophically or spiritually apt. Such an answer may make no sense within Bayesian logic but make sense within another mode of thought. I would be interested in your reaction. **Nick:** I think I did not explain my position on phenomenology clearly enough in the talk, so let me reconstruct it. I am happy to say that LLMs have no phenomenology, Bayesian or otherwise. What they do have is access to an enormous body of text in which human beings describe their phenomenology. This includes philosophy, novels, memoirs, psychiatric reports, and many other forms of writing. My claim is that this gives an LLM enough material to produce most philosophical arguments about phenomenology. The limit you are describing through mystical states is close to the limit I was trying to acknowledge. People have attempted to articulate mystical states, but some experiences may be genuinely ineffable. I have read Akiko Frischhut’s work on timeless consciousness in meditation, for example. If a philosophical argument depends on an experience that cannot be articulated at all, that may lie beyond the systems I am discussing. My claim applies more confidently to ordinary phenomenological experience, which is represented extensively in the corpus. The model can talk convincingly and in great detail about what experience is like even though it has no experience of its own. It does so through autoregressive token prediction. A simple example is color. If you ask a good model about color combinations and aesthetic effects, it can speak as though it knows exactly which shades work together. It may tell you that a particular combination will make something “pop,” and often it is correct. It can discuss color experience at a fine-grained level despite having no color experience. That is all I need for a great deal of philosophy. **Bob:** I agree that AI-generated images can display a sophisticated grasp of color and shading. My concern is that mystical experience may be so unlike ordinary experience that a person who has had it can look at other people’s descriptions and dismiss them as gibberish. Sometimes art may require a kind of “holy gibberish,” or altered states that make the work suddenly come alive. **Chair:** Thank you. We have several more people in the queue and limited time, so I will move us on. ### “Worth reading,” utility, and moral concerns **Henry:** Thank you. My background is psychology, so I feel slightly like an impostor in a room of philosophers. I was interested in your decision not to provide a substantial philosophical definition of good philosophy, but instead to say that a text must be “worth reading,” so that the time spent on it is not wasted. That initially sounded to me like a utilitarian argument: if reading it is not a waste, it is useful. A great deal of academic opposition to AI is moral. There are worries about replacement, laziness, human suffering, and whether philosophy should help us understand or alleviate such suffering. I wonder why you do not position the philosophical use of LLMs in relation to those moral arguments. **Nick:** Let me make sure I understand. Are you asking whether I should pay more attention to moral reasons not to use these systems at all? **Henry:** Partly. I am also asking whether your claim that philosophy should produce something useful commits you to a utilitarian position, and how that relates to duty-based objections to AI. **Nick:** I would resist the move from “worth reading” to utilitarianism. I introduced the phrase precisely to sidestep disputes in metaphilosophy about the purpose of philosophy. I am appealing to something familiar from our ordinary practice. When we begin reading a paper, we hope it will be worth our time. We have all read excellent work, average work, and work that felt like a complete waste of time. The claim is only that an LLM can produce philosophy that does not waste its reader’s time. It is not a claim about usefulness to society at large. The analogy with the scientific examples may help. If an AI-assisted paper solves a long-standing mathematical problem, then it is worth reading for mathematicians. That is the modest sense I intend. On the moral questions, I will be somewhat evasive because they are not the subject of the paper. At a personal level, I do not believe I am doing anything especially wrong by using these systems, or at least nothing obviously more wrong than many ordinary forms of participation in compromised institutions and technologies. I am not persuaded by every environmental claim made about AI, and I also find some copyright-based condemnations rather selective given ordinary practices of copying and piracy. But I do not want to pretend that this settles the moral questions. My argument today is about capability, not the all-things-considered permissibility of using that capability. **Henry:** So the argument is not based on utility. It appeals instead to a shared practical understanding of what it is like to read good philosophy. **Nick:** Yes, something like that. An everyday shared understanding rather than a theory of value. ### Bias in training data and “ethically grown” LLMs **Audience member (historian):** I want to echo Henry’s concern. I am a historian rather than a philosopher, and I am interested in the moral dimension of AI as it relates to authorship. LLMs are developed within patriarchal and racialized societies, and their training data are produced within those societies. If we are thinking about literature, philosophy, and attempts to decolonize the academy, where do LLMs fit? Does the fact that they consume texts produced in sexist, patriarchal, and racialized contexts affect their outputs, and how could that be changed? **Nick:** I understand the question. The corpus used to train these systems presumably contains a large quantity of material written from patriarchal and otherwise objectionable points of view. It includes texts from many societies and historical periods, together with a vast amount of internet material. I am afraid I cannot give a very satisfying answer because we no longer know in detail what the major laboratories train their models on. They are secretive about the composition of the datasets. We also know that artificial or synthetic data, text generated by other models, now forms part of some training processes. Laboratories also perform quality control and filter some categories of material. That does not remove the problem, because we then need to ask who decides what is acceptable. What does Sam Altman, Elon Musk, or any other powerful figure think should be included or excluded? That is a real issue, but it is also contingent. In principle, one could imagine an ethically sourced or “ethically grown” LLM, perhaps a free-range LLM rather than one raised on horrible material. Of course, we would then have to decide which ethics should govern its cultivation. There is also a second stage of training. The first stage trains the model to predict the next token from a corpus. A raw text predictor would simply continue producing text. A later stage trains it to behave as an interlocutor, typically as a helpful assistant responding to a user. Those assistant models are trained very heavily not to produce offensive or harmful material. They are not foolproof, but this creates a second line of resistance. The weakness is that this layer is open to manipulation. A well-known example involved changes to Grok that caused it to inject claims about anti-white racism in South Africa into many unrelated answers. A system prompt or other intervention can distort outputs in that way. So the systems are open to abuse. Some of the worst material is probably filtered during one or both stages of training, but the more satisfactory long-term response may require thinking seriously about how LLMs are cultivated and whose ethical judgments shape them. ### Philosophical pleasure, style, biography, and the point of producing more philosophy **Philip:** Thank you for a very intriguing talk, and thanks to the previous speakers. Their questions have helped me narrow down my concern. For present purposes, it is not primarily about ethics or phenomenology. I am afraid I cannot let you avoid metaphilosophy entirely. What makes philosophy worth reading? My suggestion, which you may not like, is that the philosophy I find worth reading gives me pleasure. That standard rules out a great deal of peer-reviewed academic work. Much contemporary philosophy is written in a desiccated manner because people must publish in order to gain or retain academic positions. Too much is published, and much of it gives very little pleasure. How does that fit with the prospect of LLMs producing even more material of a similar kind, perhaps better than some human work but still not pleasurable? What would that tell us about philosophy as presently practiced and about the quality of philosophical writing in peer-reviewed journals? I also want to add a po int connected with Costas and Bob. When I wrote philosophy, I knew when to stop. That is one respect in which good philosophy resembles good art. Philosophical decisions are made by particular embodied people under the contingencies of their lives and biographies. Those contingencies may have little to do with the explicit question being studied, but they shape where one stops and what one produces. That is part of what makes interesting philosophy interesting. It is not only what is on the page. In Wittgenstein and many other figures, biographical features contribute to our interest in the philosophy. **Nick:** Let me begin with pleasure. You seemed to suggest that part of the pleasure comes from biographical details and from the person writing the work. **Philip:** That is an additional source of interest. The central pleasure comes from a well-crafted philosophical text. **Nick:** In that case, the disagreement may be narrower. I think LLMs can craft coherent, elegantly argued philosophical texts. I have elicited arguments from models that I regarded as elegant and well argued, and I experienced pleasure in reading them. Is your concern that LLMs cannot produce that kind of craft, or that knowing the text came from an LLM diminishes the pleasure? **Philip:** My central question is simply: what is the point? We already have more human philosophers, more fellow medium-sized language models, producing philosophy than most people want to read. What is the motivation for producing even more? If this is partly an enthusiastic exploration of how remarkable LLMs are, that is fine, but I am not sure there is a large market for additional philosophical text. **Nick:** That is a good question. I admit that the project is partly an enthusiastic exploration, but not entirely. We face many questions about what to do with LLMs in academia, philosophy, and society. One question that bears on the rest is whether they can produce worthwhile philosophy. My paper is about that capability. You are then asking: even if they can, so what? I suppose I am part of the market. I want to read good philosophy. I want philosophy that increases my understanding and allows me to learn something new. I have argued that LLMs can produce this and will become increasingly capable of doing so. I therefore see them as a possible source of more good philosophy. Of course, what each of us regards as good philosophy varies. LLMs could be used to produce the desiccated kind of philosophy you dislike, but they could also be used to produce the kind you enjoy. I am not sure that fully answers the challenge, but I think the basic motivation is that I want more good philosophy, and I think these systems may help create it. **Chair:** I think we are out of time. Thank you, Nick, for a very interesting talk, and thanks to everyone for the discussion. **Nick:** Thank you all for the questions and for the invitation. Please feel free to email me if you want to continue the discussion.