# Generating Philosophy with Artificial Intelligence ## Abstract We argue that current-generation large language models (LLMs) are capable of producing philosophical texts that are worth reading, in the same way that a good piece of human-written philosophy can be. We defend this claim against various challenges. First, one might think that a text produced by an LLM cannot count as a work of philosophy because no person lies behind it, just as one might think that an image produced by AI cannot count as an artwork. Second, one might think that LLMs lack a particular capacity, or capacities, that producing worthwhile philosophy requires: the capacity to weigh rival explanations, to connect with the world, or to draw on experience. Finally, one might ask why, if we are right, no philosophy worth reading has yet come from these systems, and whether the philosophical content of any text produced under a philosopher's prompting should be credited to the philosopher rather than to the model. We argue that none of these challenges succeeds. Two texts that differ only in how they came to be written do not differ in whether they are worth reading; the distinction between genuine and merely apparent abduction cannot be located in the texts LLMs actually produce; and the absence, so far, of worthwhile LLM-written philosophy reflects how these systems are used rather than what they can produce. --- ## 0. Introduction The last decade or so has seen the rise of generative artificial intelligence: systems that produce text, images, code, music, video, and other outputs in response to prompts. AI has had success in domains where the value of an output is not exhausted by its superficial fluency. There have been recent AI-assisted discoveries in physics (Guevara et al. 2026), mathematics (Novikov et al. 2025), biomedicine (Gottweis et al. 2025), and materials science (Zeni et al. 2025). In this paper we argue that we should expect similar success in philosophy. Specifically, we argue that current-generation LLMs are capable of producing philosophical texts that are _worth reading_. We do not want to begin by settling what counts as _good_ philosophy. Instead, we appeal to a distinction familiar to anyone who reads philosophy: between texts that repay the time spent reading them and texts that do not. As you begin reading this article, you likely hope that it is worth reading, in the sense that the time spent reading it will not be wasted. When you write a philosophical text yourself you aim to make it worth readers' while to read it, and whether or not the journal you send it to accepts it depends on whether or not they agree. Note that a text's being worth reading is not the same as its being correct: we take many of the philosophical texts we read to be mistaken in their conclusions, and few of them to have wasted our time. Note also that the mere statement of a philosophical conclusion is unlikely to be worth reading. A text consisting only of bare pronouncements — that direct realism is correct, that we should be utilitarians — would not repay anyone's attention, and this is as true of a human philosopher's pronouncements as of an LLM's.[^1] What would repay attention is argument, and it is the capacity of LLMs to produce philosophical arguments, rather than philosophical pronouncements, that concerns us in what follows. We will argue for this claim by considering six challenges that might be raised against it. **The first challenge is constitutive, the next three concern capacities the model is said to lack, and the last two ask why, if we are right, no philosophy worth reading has yet appeared, and whose it is when a philosopher's prompting draws it out.** Section 1 rejects the challenge from authorship: the claim that an LLM output cannot be philosophy worth reading because no philosopher lies behind it. Section 2 turns to abduction and argues that the absence of human-style inference to the best explanation in the producer does not preclude abductive structure in the product. Section 3 turns to the challenge from detachment, and argues that the materials philosophy takes from the world reach it **already articulated in language**. Section 4 takes up the challenge from experience, and argues that the **subjective** experiences philosophy argues about reach it in the same way. Section 5 turns to the challenge from observation — that if all this is right, philosophy worth reading should already be coming from these systems, and is not — and argues that this reflects how they are used rather than what they can produce. Section 6 addresses the challenge from instrumentality — that where a philosopher's prompting draws out such a text, the philosophy is the philosopher's — and argues that a prompt articulates a starting point whose development it does not fix. [^1]: TODO: add footnote. --- ## 1. The Challenge from Authorship Philosophy might be thought to be something that only persons, or at least minds, can produce. This view has not, to our knowledge, been explicitly defended in just this form, but it gives shape to an intuition that many philosophers may have: philosophy is a person-only domain. An imperfect comparison is with art: one might deny that an image generated by an AI system is an artwork, because no artist exercises the relevant kind of intentional control over its production.[^2] One might think, for similar reasons, that philosophy can only be done by people: no text produced by an LLM can be a work of philosophy, because no philosopher lies behind it. Call this the _challenge from authorship_. There is some initial support for this thought in the way philosophy is studied. Like the study of art, and unlike the study of physics, the study of philosophy is organised around individuals: undergraduates take courses on Kant's ethics or Lewis's metaphysics, and reading an important philosopher's own words is held to be of value in a way that reading Newton's is not. A physics student is taught Newtonian mechanics from a current textbook, and the course loses nothing if the _Principia_ is never opened; a course on Kant's ethics that never opened the **second _Critique_** would scarcely count as one. In the sciences, that is, what a text contributes can be carried entirely by other texts, while in philosophy the contribution and its original presentation are harder to prise apart. One might take this as evidence that a philosophical work is bound to the activity of the person who produced it, in a way that a scientific result is not. **We will now try to make this challenge from authorship more precise, by considering how far Davies' _performance_ theory of art transposes to philosophy.** Davies writes: > [T]he work – what the artist achieves – is the process eventuating in that product. Works themselves are neither structures nor objects simpliciter, nor are they contextualized structures or objects. They are, rather, intentionally guided generative performances that eventuate in contextualized structures or objects (or events, as we shall see) – performances completed by what I am terming a focus of appreciation. (2004, p. 97) On Davies' view, when a painter paints a picture, the canvas is what we attend to, but it is not the work. The work is the artist's intentionally guided activity in producing the canvas; the canvas is "the focus of our appreciative interest in the work" (2004, p. 150). What we appreciate in a painting, on this account, is an achievement, and achievements are individuated by the activities that bring them about: the same surface, reached by some other route, would be a different achievement. Facts about how the object came into being thus do more than supply context; they help determine what the work is, and what is properly appreciated in it. Suppose, to adapt one of Davies' examples (2004, p. 102), that a storm were to blow pigment across a stretched canvas, leaving a surface indistinguishable, mark for mark, from an abstract painting. A theorist who identifies the painting with its surface must either count the storm's product as a painting or explain why it is not. There is a surface with no painting of a picture behind it, and the corresponding case for philosophy — a text indistinguishable from a philosophical argument, with no philosophising behind it — is the one an LLM presents. If Davies is right, the surface does not by itself settle the work. Transposed to philosophy, Davies' proposal would run as follows: a philosophical text is not itself the philosophical work; the text is the product of a person's philosophising, and reading it is a way of engaging with that prior activity. The activity, on this proposal, is part of what the work is, so that there is a philosophical work only where the philosophising has taken place, which makes the challenge a constitutive one. If no one has philosophised, there is no work to which the text gives access, however the text reads — an LLM text would stand to philosophy as the storm-made canvas stands to painting. We do not think the transposition should be accepted. What makes Davies' view plausible in the case of art is that indiscernible surfaces really do seem to differ in artistic value: the storm-made canvas is worth nothing as a painting, while a brushed one may be worth a great deal. Whoever transposes the view to philosophy is therefore committed to the corresponding claim, that two texts with the same contents can differ in whether they are worth reading according to how each came to be written, and it is difficult to see what such a difference could consist in. Whether an argument is valid, whether its premises are plausible, and whether the objections to it have been answered are questions about a text's contents, and two texts with the same contents receive the same answers to them. A difference in the route by which two texts came to be written, without any difference in their contents, is therefore no difference in whether they are worth reading.[^dan][^putnam] Peer review proceeds on the same assumption: the grounds for accepting or rejecting a paper lie in the argument as presented and in what it adds to the literature, not in the history of its production.[^3] ~~It might still be insisted that where there has been no philosophising there is no work, whatever the resulting text contains. We need not resist this, because our thesis concerns texts worth reading, and a text can be worth reading without being a work in Davies' sense. Suppose a desert wind traced out, in the sand, a sound argument against enactivist theories of perception. Nobody would deserve credit for the argument, and there would be no performance for the text to give access to; but a reader who worked through it would still encounter a thesis and the considerations advanced in its favour, and would be in a position to answer or to extend it. 'Work' may, if the performance theorist insists, be reserved for texts with performances behind them. What cannot be so reserved is being worth reading, since everything that judgement answers to is on the page.~~ That no one lies behind a text does not, then, settle whether it is worth reading. Whether an LLM can actually _produce_ a text worth reading is a further question, and it is to this that we now turn. [^2]: This is not to deny that systems of this kind can produce beautiful images; we return to image generation in Section 4. [^3]: Similar ideas can be found in the literature on philosophical progress. According to Dellsén et al. (2024), philosophical progress consists in 'putting people in a position to increase their understanding' (p. 679). The usual way of putting people in such a position, they note, is to make philosophical ideas publicly available in journals and books (pp. 679–680). The fact that a contribution is here defined by what readers can take up, and never by the activity which produced it, shows that this literature, too, locates philosophy in the public text rather than in its author. [^dan]: According to Levinson (1980), the aesthetic properties of a musical work are partly determined by the context in which it was composed: two works identical note for note, composed in different contexts, can differ aesthetically. Transposed to philosophy, the claim would be that whether a text is worth reading is partly determined by the context in which it appears. Imagine _Sense and Sensibilia_ written today rather than in 1962. Its arguments would be no worse, but its readers would already have them, and the text would no longer repay their attention. A difference of this kind is made by the context in which a text is read, not by the route by which its words came to be written. We thank Dan %%surname%% for the point and the example. [^putnam]: It might be objected, following Putnam (1981), that marks produced by nothing with intentions do not represent anything at all — the trail an ant happens to trace in the sand is no caricature of Churchill, however closely it resembles one — and hence that a text produced by an LLM has no contents in which its worth could reside. But a reader who works through such a text follows its argument exactly as she would had a philosopher written it; and it is to what its reader can find in it, not to what its marks represent in themselves, that a text's being worth reading answers. --- ## 2. The challenge from abduction In the previous section, we argued against the idea that LLMs cannot produce philosophy worth reading simply because they are not human. In this and the two sections that follow we shall consider a different form of challenge: even if LLMs cannot be ruled out of producing philosophy worth reading, they lack particular _capacities_ that are required to produce it. If our arguments regarding authorship are correct, then a novel philosophical argument produced by a parrot would deserve to be taken just as seriously as one produced by a human being. No such argument will be forthcoming, however: a parrot can only reproduce sounds it has already heard. One might suspect that LLMs are in the parrot's position — that while nothing rules their outputs out of being philosophy worth reading _tout court_, they lack some capacity that producing it requires. **In this section we address one capacity challenge, which we will call _the challenge from abduction_. In **Sections** 3 and 4 we shall look at two more: the _challenge from detachment_ and the _challenge from experience_.** Abductive inference is reasoning from a body of evidence to the hypothesis that would best explain it. In a deductive argument the premises fix the conclusion: if all men are mortal and Socrates is a man, then Socrates is mortal, and that is that. Now, imagine walking into your kitchen one morning and finding that part of the floor is wet. What has happened? The wet floor does not determine the answer in the way the two premises determined Socrates' mortality: a burst pipe would have left the floor wet, and so would a spilled bucket. But, given that the window is open, the water is under the window, and it rained last night, rain coming through the window seems the most plausible answer. To reason in this way, inferring the best explanation for a set of facts, is common in the sciences as well as everyday life. Williamson argues that philosophy should use a broadly abductive methodology (2007; 2021, §9.2). In philosophy, as in science, there are data that a candidate theory must accommodate — intuitions about cases, and the phenomena of the domain itself — and rival theories that would each accommodate them at different costs. The theory to prefer is the one that would, if true, best account for the data, where what makes one theory's account better than another's is its explanatory virtue: a good philosophical theory is, in Williamson's words, "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated", and should "combine simplicity with strength" (2021, §9.2). Nor is this view of theory choice confined to those who, like Williamson, take philosophy to be methodologically continuous with the sciences: Bengson, Cuneo and Shafer-Landau (2022) argue that the assessment of rival theories by their explanatory and unifying merits is a constraint on sound philosophical method as such.[^4] We shall assume it in what follows.[^5] Lipton (2004, p. 59) distinguishes two things that "the best explanation" might mean: the _likeliest_ explanation, the one most warranted by the evidence, and the _loveliest_, the one that would, if true, provide the most understanding. As he puts it, "Likeliness speaks of truth; loveliness of potential understanding" (2004, p. 59). The two can come apart. Return to the wet kitchen floor: that water has fallen on it is the likeliest explanation of its being wet, and it is scarcely possible that it is false, yet it affords no understanding at all of how the floor came to be in that state: of course it is wet because water fell on it, but how, and which water? The explanatory virtues by which, on the methodology we have assumed, rival philosophical theories are weighed are virtues of loveliness: elegance, unity, and simplicity concern the understanding a theory would provide were it true, not the probability that it is true. Where a philosophical text turns on the weighing of rival explanations — not all philosophical writing does — whether the text is worth reading and whether it weighs its rivals well go together. If a capacity for abduction is required to produce worthwhile philosophy, do LLMs possess it? Floridi et al. (2025) argue that they do not: > We argue that such LLMs generate text based on learned associations rather than performing abductive inferences. […] LLMs produce plausible hypotheses, simulate commonsense reasoning, and provide explanatory answers without grounding them directly in truth, semantics, verification, or understanding, and without any abductive reasoning. (Floridi et al. 2025, p. 1) An LLM, on this account, has "a stochastic core and an abductive appearance" (2025, p. 2); the further charge, that these systems cannot "genuinely validate" their explanations against reality (ibid., p. 6), we take up in the next section. Models are trained to predict which words are likely to follow which, and to produce the continuation that their training makes probable; nothing in the training directs them at truth. The appearance comes from what the training data have passed on: models have "absorbed patterns of human abductive reasoning as expressed in writing" (p. 9) — how explanations are typically phrased, which causes are typically offered for which effects. [^6] Floridi et al.'s own example is a car that will not start on a cold morning. Asked why not, a model replies that a weak battery is one possibility, since cold reduces a battery's efficiency; that thickened engine oil is another, since a cold engine is harder to turn over; and that, "[b]ased on your description, the battery is the most likely explanation" (2025, p. 10). On their reading, the LLM is not actually reasoning about causes from the case before it; it is the statistical reproduction of the causes such explanations typically cite (p. 9). What they deny is that the model weighs the candidates — that it "entails choosing the best explanation among alternatives" (2025, p. 3). The verdict that settles on the battery does no weighing; it reproduces how explanations of this kind conventionally end. And where the output marks a genuine difference between the two — some consideration that would tell the battery from the oil — it is one already drawn in the explanations the model learned from, not one worked out afresh for the case in hand. Because LLMs produce no more than a "veneer of explanation" (ibid., p. 20), Floridi et al. conclude, they can play only a supporting role in intellectual work: > In essence, LLMs function like brainstorming assistants that toss out ideas without filtering for quality. After all, they work like statistical interfaces to an enormous amount of data accumulated for millennia by generations. A cautious human collaborator can sift through and assess them. (2025, p. 11) A device for brainstorming seems a far cry from something which might produce worthwhile philosophy. If you were presented with a text and told that it contains a number of philosophical ideas, none of which have been filtered for quality, it is unlikely you would think it is worth your while reading it. The filtering these systems are said to lack is, in Lipton's terms, filtering for loveliness: **what a model cannot do,** on Floridi et al.'s account, is prefer one candidate explanation to another on the grounds of the understanding it would afford. Floridi and his colleagues allow that, in ordinary cases such as their own cold-morning car, the model's answer is a good one: "the same explanation a human reasoner would likely choose" (2025, p. 10), one that may be "even optimal by IBE criteria" (2025, p. 19), since such systems "echo the obvious, common explanations" (2025, p. 10). The complaint cannot then be that the explanation is poor, which leaves it hard to say what, in such an answer, is supposed to be merely apparent. What their account points to is the uncommon case: "in less common situations, LLMs can falter" (2025, p. 10), and on inputs "that go beyond their training" "the facade can crack" (2025, p. 9), the success on familiar cases being "a sign of overfitting to common patterns" (2025, p. 15). **A model overfits when it memorises the patterns of the examples it was trained on rather than the regularity behind them: its success then carries over only to cases resembling those examples.** Once the question of its truth is set aside, this is what the facade reduces to: the claim that the competence shown on common problems is overfitting that would give out in less common ones. It is an empirical question as to whether the competence shown on common problems does give out on less common ones. Various benchmarks which set these systems abductive tasks have been put forward in recent years. A recent survey by Salimi et al. (2026) gathers the assessments these benchmarks have so far yielded, some finding abduction the easiest of the three classical forms of reasoning for a language model, and others finding that models struggle with abductive tasks despite strong performance on deductive ones. As Salimi and colleagues point out, there is no shared definition of what an abductive task requires, making the results of different benchmarks difficult to compare. While these benchmarks tell us how often a model names the keyed culprit or ranks the right diagnosis first, they still score the answer a text arrives at rather than the weighing that Floridi and his colleagues say is merely apparent. According to Salimi and colleagues, indeed, these metrics 'only assess the end-result answers to the task, completely bypassing the actual reasoning trace' (2026). To find out whether the competence does give out, then, we might need to look beyond the benchmark scores altogether and consider the explanations these systems actually produce. We do not disagree with Floridi et al.'s characterisation of how LLMs function: these systems do not weigh and choose among alternatives in the way that humans do. As we saw in the previous section, however, a difference in how a text came to be written, without a difference in its contents, is no difference in whether it is worth reading. To see how what something is can come apart from how it came to be, consider the artificial images made by generative AI: a picture of a cat created in this way is no less a picture of a cat, yet it is produced quite differently from any picture a human would make by painting or photography.[^milliere] Similarly, it might be possible for LLMs to produce text which displays abductive reasoning, without any actual abductive reasoning behind it. Despite their stochastic core, LLMs are perfectly capable of producing grammatically correct text: without being given any specific rules, the system "implicitly 'discovers' them—and then seems to be good at following them" (Wolfram 2023). Does this mean that the texts LLMs produce have merely the appearance of being grammatically well-formed? Clearly they do not: LLM sentences _are_ grammatically well-formed, **regardless of their stochastic roots**. A stochastic core, then, need not mean that the best an LLM can manage is a veneer of abductive inference. It may be, rather, that the core is marshalled to produce text exhibiting actual abductive inference, in just the way it is marshalled to produce actual grammatical correctness. It might be objected that the grammar analogy will not stretch this far. Grammar, the objector would say, is a system of rules a model can follow without understanding anything, and explanatory loveliness is not — it is not a system of rules you follow to get the right answer, so the model's grammatical competence gives no reason to expect it to produce lovely explanations. If Floridi et al. are right that the output has the form of an inference to the best explanation but not its substance, then there should be cases where the form is present and the substance is absent: cases where the model produces something that looks like an explanation but is empty or nonsensical, just as it could produce something that is grammatically faultless and says nothing. Wolfram's example of the latter is "Inquisitive electrons eat blue theories for fish" (2023) — impeccably grammatical, and meaningless. The model **almost never produces** such strings. The abductive equivalent would be a passage with the form of an inference to the best explanation, grammatically faultless, and yet senseless — asked why a car will not start on a cold morning, there are no squirrel tracks, so it must be freak arctic winds blowing into the exhaust pipe. That has the form of evidence weighed towards a conclusion, and it is nonsense; and it is nonsense the model **almost never produces**. Put the question to it and it offers the weak battery and the thickened oil, plausible explanations rather than arctic winds. Wherever the facade is to be located, it cannot be located there: what the model produces are, **for the most part,** plausible explanations. Moreover, the answer Floridi and his colleagues describe is not quite the answer these systems give. Their account asks us to picture a confident verdict laid over a hollow core, yet what one finds in practice is closer to hedging. Asked why a car would not start on a cold December morning, a current model will say that the battery is the most likely culprit, set out other plausible contributors — thickened oil, fuel-system problems, ignition faults — and end by observing that if the car started once the day had warmed, the battery is almost certainly the primary cause, though a load test would confirm it.[^7] The model's avoidance of senseless explanations, like its avoidance of senseless sentences, points to a semantic competence picked up in training, over and above syntax. A syntactic grammar settles only how the parts of speech may be combined, and not whether what a sentence says makes sense: avoiding strings like the inquisitive electrons takes a grasp of what can sensibly be said of what — which predicates go with which subjects, which causes with which effects. Wolfram calls such a grasp a semantic grammar: > to deal with meaning, we need to go further. And one version of how to do this is to think about not just a syntactic grammar for language, but also a semantic one. (Wolfram 2023) A model trained on enough text has, on his account, come by one: > From its training ChatGPT has effectively "pieced together" a certain (rather impressive) quantity of what amounts to semantic grammar. (Wolfram 2023) What the training pieces together extends beyond the fit of predicate to subject. The arctic winds owed their senselessness as much to the inference as to its parts — nothing in the absence of squirrel tracks bears on what enters an exhaust pipe — and Wolfram suggests that inference is picked up by the same route: just as Aristotle might have arrived at syllogistic logic by working through many examples of rhetoric, a model in training can "discover syllogistic logic" in the text it reads, and can be expected to produce text containing "correct inferences" (2023).[^8] It is this, rather than any contact with the case in hand, that keeps the arctic winds out of the model's explanations: a feel for what, in a working model of the world, can hang together.[^9] But a feel for how things hang together is gathered from the text the model has read. It amounts to a grasp of how failures of this kind are explained in general — that cold weather weakens batteries, say — and it gives the model no means of finding out which explanation is true of the particular car whose fault is in question. It does, however, underwrite the weighing itself: the model's answer sets out the candidate explanations and fits its confidence to the evidence it has been given, and neither of these requires access to the car. What requires such access is settling which candidate is correct. In much philosophical abduction there is no analogue of the car: what a thought experiment commits us to, or which of two theories carries the lighter explanatory cost, is settled from what is already set down in writing, and a corpus already holds evidence of that sort. It is now difficult to see where the facade should be located. The explanations the model produces are plausible, and the confidence they express is fitted to the evidence it has been given. What the model lacks is access to the particular case, and that is a limitation concerning its relation to the world rather than its capacity for abduction. What that costs a philosophical text is a matter for the next section. [^4]: TODO: add footnote. [^5]: Understanding-based accounts of philosophical progress converge on the same standard: if progress in philosophy consists in placing people in a position to increase their understanding (Dellsén et al. 2024, p. 679), then rival theories are rightly weighed by the understanding they would, if true, provide. [^6]: TODO: add footnote. [^7]: The exchange, with Kimi k2.67 (%%add date of test%%), is reproduced in full in the Appendix. This is not yet to say that the facade charge is refuted: a hedged answer can still be, on Floridi et al.'s account, a statistical reproduction of how hedged explanations typically read. [^8]: Wolfram's expectation comes with a caveat: on "more sophisticated formal logic" he expects the model to fail, for the same reasons it fails to match parentheses across long sequences (2023). The caveat concerns extended derivation, and the weighing of explanations is not derivation: as the objector above puts it, loveliness is not a system of rules one follows to get the right answer. [^9]: We use "model of the world" in Wolfram's informal sense: a body of implicitly learned regularities concerning what goes with what. This is weaker than the technical sense now disputed in machine learning, on which a world model is an internal representation of an environment that supports prediction and counterfactual reasoning. Nothing in our argument requires deciding whether LLMs possess world models in the technical sense; how they connect to the world is the matter of §3. **[^milliere]: For an account of how images of this kind are produced, see Millière (2022).** --- ## 3. The Challenge from Detachment A second capacity challenge, the _challenge from detachment_, concerns the model's relation to the world. In addition to the argument considered in the previous section, Floridi et al. object that an LLM stands in no relation to it: its words rest on no perception of anything, and a hypothesis, once produced, is never tested against how things are (2025, pp. 6–7). A discipline whose theories answer to how things are, the challenge runs, cannot be advanced by a system with no access to how things are, and a text produced by such a system gives its reader no reason to think it worth reading. Zahavy (2026) raises a worry of this kind about scientific discovery. A model can carry out the deductive part of discovery, working out the consequences of premises it has been given; what it cannot do, he holds, is produce the premises — **make the leap from sense experience to new first principles.[^11]** His case is the thought experiment that gave Einstein the equivalence principle: > Einstein's variation required inventing new axioms based on a physical intuition that did not yet exist in the mathematics. He envisioned a physicist inside an elevator being uniformly accelerated through deep space [...]. Inside this enclosure, the sensory experience reveals a specific pattern: when objects are released, the floor rushes up to meet them. To the physicist, the objects appear to fall with identical acceleration, regardless of composition. Thus, the simulation here was not a permutation of symbols, but a manipulation of perceptual experience. (Zahavy 2026, §5) **Einstein imagined** a set of circumstances and **attended** to what would be experienced within them: everything released inside the elevator appears to fall with identical acceleration. On Zahavy's reconstruction the simulation supplies an observation, and from that observation the new axiom is inferred — the simulated experience of acceleration was indistinguishable from the remembered experience of gravity, and Einstein concluded that the two are one phenomenon.[^12] A model has no access to that observation. It can produce descriptions of elevators and of weightlessness, both present in its corpus, but it has undergone neither, and a discovery whose premises are fixed by simulated experience is beyond a system that, in Zahavy's words, lacks the capacity he calls sensory agency. **While Zahavy often uses phrases which suggest that it is a model's lack of phenomenology that is the principal stumbling block — Einstein arrived at the equivalence principle by "simulating the physical feelings" of his observer, guided by "the sensation of gravity" (2026, §5) — on closer inspection it is actually better thought of as an argument about a model's access to the world. If a model is to make the leap, Zahavy argues, it must be equipped with an interactive world model, a system that simulates a physical world within which the model can act and observe what follows (2026, §5). The fact that such a system contains no phenomenology, yet would supply what the model is missing, shows that what is missing is access rather than phenomenology. Having said that, one might **still** ask whether a lack of phenomenology does in fact preclude an LLM from doing at least some sorts of philosophy. How can a system which is not conscious produce worthwhile philosophy about consciousness? We will examine this question in the next section.** We argue that the challenge fails because the materials philosophy takes from the world reach it already articulated in language. We can grant that the model perceives nothing, and that nothing it produces is put to the test against the world. Whether these concessions disqualify it depends on how philosophy itself meets the world. Pigliucci (2017), writing on how philosophy makes progress, holds that philosophy is constrained by the world without investigating it as the natural sciences do: > This means that the basic parameters that philosophers use as their inputs, the starting points of their philosophizing, their equivalent of axioms in mathematics and assumptions in logic (or rules in chess) are _empirical_ data about the world. This data comes from both everyday experience [...] and of course increasingly from the world of science itself. Philosophy, I maintain, is in the business of exploring the sort of conceptually evoked spaces that Smolin is talking about, where the evocation is the result of whatever starting assumptions are made by individual philosophers working within a particular field and, crucially, of the constraints that are imposed by our best understanding of how the world actually is. (2017, pp. 79–80) Unlike the axioms of mathematics or the rules of chess, philosophy's starting points are constrained by how the world actually is. Being constrained by the world, however, is not the same as investigating it. The scientific data Pigliucci has in mind reach philosophers already set down in language: a philosopher of physics works from published results, not from having run the experiments. Everyday experience enters philosophising in the same way, as claims about the world that anyone may state and no one need establish — that objects look smaller from far away, that people act on what they believe. Philosophy's starting points are constrained by the world, then, in virtue of being claims about how it is, and such claims are what an LLM's corpus contains in enormous quantity. A system that perceives nothing can gather none of this for itself, but the empirical materials philosophy ordinarily uses reach it in linguistic form. Pigliucci takes the term _evocation_ from Smolin (Unger and Smolin 2015).[^10] Evoked truths are neither discovered, in the sense of corresponding to mind-independent states of affairs, nor invented, in the sense of being arbitrary constructs. Smolin's illustration is chess: before the rules of the game were codified there were no facts about chess, but once they were, a great many facts about the game became demonstrable — objective facts, in the sense that anyone who demonstrates one demonstrates the same fact as anyone else (Unger and Smolin 2015, p. 423; quoted at Pigliucci 2017, p. 78). On Pigliucci's proposal, philosophy is in the business of ascertaining evoked truths. Pigliucci's contrast case is a science-fiction writer who describes the same planet across three different timelines, keeping every description within the laws of physics and of logic. The writer is exploring conceptual spaces of a kind, but he is inventing rather than evoking: his worlds have no rigid properties, because even the constraints he keeps to could have been otherwise — he could as easily have imagined planets with a different physics, or a different logic. Philosophy, on the other hand, "is in the business of doing empirically informed evoking, not inventing", and its objects of study accordingly have rigid properties (2017, p. 80): once a philosophical starting point has been articulated, what holds within the structure it opens is no longer up to anyone, any more than what follows from the rules of chess is up to anyone. **In a thought experiment,** the philosopher sets up an imagined scenario, but explores it "with an interest in figuring things out as far as this world is concerned" (2017, p. 80), and the philosophical work then proceeds within a structure whose properties the philosopher does not control. **A philosophical thesis is not tested in the way the equivalence principle was, by experiments that could have refuted it.** Whether it holds is a question about what obtains within an evoked structure, and it is settled in the way questions about chess are settled: by working out what the starting points commit us to. This working-out is carried out in writing, in the literature's back-and-forth of argument and reply, and it is work of a kind Section 2 argued a model's text can display. A model's inability to run experiments therefore costs it nothing that a philosophical text needs: the testing philosophy does is done on the page. ## 4. The Challenge from Experience A third capacity challenge, the _challenge from experience_, holds that some philosophy depends on **subjective** experience in a way that a system which has never had any cannot meet. Few would say that current LLMs are conscious, and we assume here that they are not. Does Mary, released from her black-and-white room knowing every physical fact about colour vision, learn something when she first sees red (Jackson 1982)? To settle the question one must consider what the experience of seeing red is like, and the argument proceeds from that verdict. Here experience enters as the starting point of an argument, and a system that has never experienced anything seems unable to supply it. Experience enters philosophy as subject matter too, in work that asks what it is to see red (Harman 1990), to feel anger (Goldie 2000), or to have an intuition take hold (Chudnoff 2011).[^13] If starting points of the first kind, and subject matters of the second, can be handled only by a being that has the experiences, then much of the philosophy of mind lies beyond a model's reach. Experience does not, however, enter the knowledge argument in the way the challenge assumes. The facts about colour experience that the argument **draws on** are general ones, articulated many times in the literature: that colour is experienced as a quality of the things seen, and that one cannot imagine a colour one has never seen — which is why Mary's room must be black and white. The texts produced in response to Jackson's paper work with claims of this kind and with his description of the case, contesting what they commit us to — whether what Mary gains on her release is knowledge of a fact at all. **None of this requires having seen red for oneself:** a reader of Jackson's paper need not introspect the fine details of their own experience of red to follow his argument, or the literature that discusses it. Philosophy does not usually begin from raw encounter; it works on what has already been articulated. Even where a case does begin in its author's own experience — as Jackson's may have — what enters the argument is the description he makes of it. Descriptions of this kind — in which what experience is like has been set down in words, and can be read and reasoned from like any other claim — are what we will call _articulated phenomenology_. Consider, for example, Austin's discussion of colour phenomenology in _Sense and Sensibilia_, where aspects of what it is like to see colour are carefully described and considered (e.g. how a pointillist painting's appearance changes depending on how far one stands from it, or how an object's colour changes under different illumination; 1962, pp. %%page%%, 66). Philosophers are not the only ones in an LLM's corpus to articulate their phenomenology in text; we might also consider novels, diaries, interviews — anything in which someone describes _what it is like_ to have an experience. An LLM has never felt what Zahavy's physicist feels inside the elevator, but it has read the memoirs of astronauts who have. As with the scientific results of the previous section, a model has undergone none of these experiences, but the descriptions of them are in its corpus, and it is the descriptions that philosophy works from. What models do lack is the capacity to make _novel_ phenomenological observations. An example of this sort of novelty might be Merleau-Ponty's (1945 %%page%%) description of how, when one hand touches the other, the roles of toucher and touched can alternate but cannot be occupied at once. The experience is a familiar one, but the asymmetry it contains seems not to have been described before. Noticing it required having hands, and tactile experience of them; a model has neither, and so could not have been the first to set it down. However, philosophical arguments which rely on novel phenomenological observations are rare indeed: the descriptions of experience that philosophy works with have, for the most part, long been in the literature. And once an observation of this kind is written into the corpus, everything philosophical that can be done with it — what it shows about bodily self-awareness, what follows once the asymmetry is granted — is there on the page for any reader, and none of it depends on being able to have made the observation oneself. Nor is first-person attention the only source of descriptions the literature lacks: a new description can also be reached from those it already contains, by drawing out what they have not been taken to imply, or by combining them as no one has, and this a model can do. There is, then, one route to new starting points that a model cannot take. Whether the route it can take yields descriptions the literature lacks is the question of novelty, which Section 6 takes up. A model can, then, do the philosophy that turns on experience, with one exception: it cannot be the source of a novel phenomenological observation. Whether experience enters as an argument's starting point or as its subject matter, it reaches the model as articulated phenomenology, which it can work through as well as any reader. [^10]: _The Singular Universe and the Reality of Time_ is jointly authored, but its second part, which contains the discussion of evocation, was written by Smolin alone, as Pigliucci notes (2017, p. 77); we follow him in attributing the view to Smolin. [^11]: Zahavy too calls this leap abduction, but the word picks out something other than it did in Section 2. There, with Floridi et al., abduction was the weighing of rival explanations, and the charge was that a model only mimics it; here it is the generation of new first principles from experience, and Zahavy's claim is that a model cannot make the move because it has had no experience to move from. The present challenge rests on that second claim, about experience, and not on any verdict about the weighing. [^12]: Zahavy, following Magnani, calls the process _manipulative abduction_: hypothesis generation through the manipulation of a model — here a simulated experience — rather than of symbols (Magnani 2009; Zahavy 2026, §5). It is abduction in Section 2's sense: the equivalence principle is inferred as the best explanation of the simulated observation, the simulation supplying an explanandum that no search over existing text would have produced. What experience contributes, on this picture, is not the inference but its starting point. [^13]: The materials need not be sensory: the feeling of understanding something is sometimes used to motivate the claim that thought itself has a phenomenology (Pitt 2004). --- --- ## 5. The Challenge from Observation The challenge from observation begins with an obvious question: if all of these arguments are correct, where is all the LLM philosophy?[^14] If you ask an LLM the answer to the hard problem of consciousness or the meaning of life, the response you receive will not typically contain much in the way of philosophical argument. More likely you will get a somewhat bland survey of various philosophical positions on the topic, or perhaps just an irreverent joke.[^15] One might think that whatever the merits of the arguments we have put forward so far, _something_ must be preventing them from producing philosophy worth reading because they _don't_ produce philosophy worth reading. We want to suggest that this is a failure of prompting rather than evidence that LLMs lack the capability to produce philosophy worth reading. **It is tempting to treat a chatbot as an oracle. Ask it the capital of a country, or about the news of the day, and a correct answer comes back at once; a straight philosophical question then seems just one more thing it will answer.** As we saw in Section 2, **however,** these systems are trained to produce reasonable continuations of the text they have been given (Wolfram 2023), before being trained further to converse as helpful assistants. In the text these systems are trained on, a question like "what is the meaning of life?" is typically met with a summary of familiar positions or with a joke, but very rarely with philosophical argument. The assistant training pushes the same way: an assistant is trained to answer whatever it is asked, and a summary of the familiar positions answers a philosophical question. How, then, can we elicit philosophy from these systems? What counts as a reasonable continuation depends on what is being continued, and a bare question is rarely continued by philosophical argument. To elicit philosophy from an LLM, we suggest, is to give it a text of which philosophical argument is a reasonable continuation. One obvious way is simply to _ask_ for careful, rational arguments about a philosophical question: a request for careful reasoning about a stated problem is continued by text in which that reasoning is carried out. One can refine the request by giving the model a problem or set of facts and asking for the explanation that would, if true, provide the most understanding of them — that is, by asking it for an inference to the best explanation. We saw in Section 2 that a model's continuations are produced under the semantic grammar it has acquired — a grasp not only of what can sensibly be said of what, but also of which inferences may correctly be drawn. The claims an argument advances and the objections it raises stand in the text being continued, and whether a further step is reasonable is settled in the light of them: a continuation that contradicted what the prompt conceded, or passed over a pressure it raised, would not be a reasonable continuation of it. A prompt that contains an argument therefore fixes more than a subject matter; it fixes what the continuation must answer to. Nor need the continuation be merely argument-shaped. The senseless passage that merely looks like reasoning is what the model almost never produces — that was the lesson of the arctic winds — and the inferences its grammar licenses are, for the most part, correct ones. **We saw in Section 3 that an articulated starting point evokes a structure whose facts become demonstrable, though almost none of them are stated. A claim, once advanced, has consequences, and the consequences are there whether or not anyone has drawn them; a bare question names a subject without advancing any claim about it, and so evokes little. A prompt containing an argument, in contrast, articulates a position and the pressures upon it, and thereby evokes, and what it evokes is there for the continuation to develop.** The developing is done one step at a time. At every step, what the model continues is the whole text so far — the prompt together with everything it has itself already written (Wolfram 2023) — so each inference the continuation draws becomes part of **the text its next step continues**. Early moves open commitments that later moves are made under, and the further a continuation runs, the more of what it is continuing was never in the prompt at all. **And at** any point there are several ways of going on that would each be reasonable (Wolfram 2023), **so** a prompt constrains its continuation without fixing it. What the steps produce, where the prompt is philosophical, are the developments themselves: consequences drawn from the starting point, rivals weighed against one another. How much of a starting point a prompt articulates is a matter of degree, and the more it articulates, the more there is for a continuation to develop. Nothing restricts evocation to a single form of prompt: several hundred pages of draft manuscript make starting assumptions and accept constraints as surely as two paragraphs of thought experiment do, and which forms of prompt evoke most productively is an empirical question which we shall not try to settle. It is in the continuations of such prompts that the question of these systems' philosophical capacities would be decided.[^16] If a philosopher must articulate the starting point, then it will be said that whatever philosophy the resulting text contains should be credited to the philosopher, and that the model is merely its instrument. This is the challenge of the next section. [^14]: Strictly speaking, the observation is difficult to verify. To know that no LLM-written philosophy has appeared in the journals over the last half-decade, one would need to know that every author who has declared that they did not use AI in preparing their manuscript was telling the truth. **Most journals, moreover, do not at present require any such declaration, so an LLM-written text could have appeared in them without any author having lied.** [^15]: TODO: add footnote. [^16]: TODO: add footnote. --- --- ## 6. The Challenge from Instrumentality One might object to the account of prompting we have just put forward on the following grounds: if the prompter must put so much philosophy into the prompt to elicit a text worth reading, then whatever philosophy the resulting text contains should be credited to the prompter rather than to the model. On this picture, the model is a tool that a philosopher uses to write worthwhile philosophy. That is, it does not generate any worthwhile philosophy of its own; rather, it rearranges and remixes the philosophy contained in the prompts it receives. To believe the model responsible for the valuable properties of a philosophical text is akin to believing that it is the ventriloquist's dummy which is doing the talking. Call this the _challenge from instrumentality_. At a certain fineness of grain, the challenge is trivially true. Consider a philosopher who puts a section of a worthwhile paper into an LLM and tells it to produce one without any spelling errors or typos. If the LLM performs this task correctly, then in a certain sense we might think that it has produced worthwhile philosophy. But the philosophy in the corrected text is the philosophy that was pasted in, and we would not think that the model deserves any credit for it in any interesting sense. A model that corrects a typescript and a model that develops an argument are doing the same thing: each produces a continuation of the text it has been given. If the operation is the same in both cases, then the difference between the two outputs cannot be a difference in what the model has contributed. Dependence of this kind is, however, the normal condition of philosophical writing. A reply is elicited by the objection it answers, and nobody counts the reply as the objector's contribution; what the objector contributed is the objection. Whether the challenge succeeds therefore depends on what the prompt contributes and on what the continuation adds. We saw in the previous section that a prompt which articulates a position and the pressures upon it evokes, and that the continuation is produced from this starting point under the semantic grammar the model has acquired; what it states are facts of the evoked structure. There are, then, three contributions to distinguish: the articulated starting point, which is the prompter's; the structure it evokes, whose facts belong to no one; and the text that develops them, which is the model's. A prompt settles only a starting point, and the consequences the model's text goes on to state were stated by no one beforehand. We saw in Section 3 that once a starting point has been articulated, what holds within the structure it opens is no longer up to anyone — and that includes the person who articulated it. Smolin observes that whoever lays down the rules of a game is afterwards in the same position as everyone else: exploring the game feels like exploring a pre-existing territory, because at each point there is little or no choice about what its facts are, and the territory holds surprises for its own inventor (Unger and Smolin 2015, p. 422). The same holds for a philosopher and the starting point she articulates, and for the same reason. Articulating it makes its developments demonstrable, but it demonstrates none of them: a demonstration is the working-out of what the starting point commits us to, and an articulation does no such work. What a prompt states, then, is a starting point; the developments it makes demonstrable are stated nowhere in it. A model given such a prompt has nothing to rearrange but the starting point itself, and no rearrangement of a starting point is a demonstration of its developments; a continuation that states them has demonstrated them. Nor does the prompt select among the demonstrations available: several continuations of it would each be reasonable. It might still be insisted that the developments were contained in the starting point all along, as theorems are contained in the axioms that entail them, and that in drawing them out the model adds nothing of its own. Even in mathematics, however, the containment does not do the demonstrating: the axioms make a theorem provable, and the proof still has to be produced. And philosophical developments do not stand to their starting points as theorems stand to axioms in any case. We saw in Section 2 that where a philosophical text weighs rival explanations, what decides between them is loveliness — the understanding each rival would, if true, provide — and that loveliness is not a system of rules one follows to get the right answer. A theorem can be recovered from its axioms by derivation; a verdict that one rival carries a lighter explanatory cost than another cannot be recovered from the prompt by any procedure, because it is reached by weighing, and weighing is not derivation. What the prompt contains is the problem; the weighing is done in the continuation. Suppose, then, that the starting point is the philosopher's own — a case or an objection the literature does not contain. The structure it evokes is new, and until the prompt was written there were no facts about it to state; whatever the continuation states of it has been stated by no text before it. Novelty of this kind is measured against the prompt, and it is what the instrumentality challenge turns on: the prompter did not state these developments, and so they are not her contribution. Whether they are new to philosophy is a further question, for arguments repeat across texts — two texts can contain the same argument — and a continuation may state what nothing has stated about this case while containing no argument the literature lacks. A starting point the literature does not contain makes the stronger novelty available; whether a continuation achieves it is settled by reading it against the literature, as the narrower question was settled by reading it against the prompt. And whether what it states was worth stating is judged as any philosophical text is judged, by whether the considerations it cites genuinely tell between the positions and by the understanding they would, if correct, provide. One kind of starting point remains beyond prompting. We saw in Section 4 that a model cannot be the source of a novel phenomenological observation, and prompting does not change this: a prompt supplies text, not experience, and an observation of this kind must be noticed before it can be set down. Whether philosophy worth reading ever requires such observations among its starting points is a question we leave open. The person writes the prompt, decides which continuations to pursue and when to stop, and without her there is only the response to the point-blank question; this much of the challenge is correct. A request for careful reasoning, however, states none of the reasoning that answers it. And although in choosing between continuations she decides which of them the text goes on from, she lacks direct control over exactly what any continuation says. Nor can she make the developments hers by enriching the prompt. A richer prompt is a larger starting point, and it asks more of the continuation: what would add something to two paragraphs of thought experiment may add nothing to several hundred pages of manuscript. But however much the prompt states, the consequences the output draws were not among the statements; and we saw in the previous section that at every step the model continues the whole text so far, so that the further a continuation runs, the more of what it is continuing was never in the prompt at all. Whether anything has been added is settled by reading the two together. Where a continuation merely rephrases its prompt, the model has done no more than rearrange, and the philosophy is the prompter's; **where it develops the starting point in ways the prompt neither states nor fixes, the development is the model's.** To sum up, a philosopher who prompts an LLM states the way a problem is to be looked at, and the conclusions that arise from looking at it that way are stated only in the model's continuation. Philosophy, Pigliucci writes, attempts "to evoke rational conclusions arising from certain ways of looking at a given problem or set of facts" (2017, p. %%page%%). A prompt supplies a starting point without fixing what follows from it, and the philosophical standing of any output is settled by what its continuation adds — first to the prompt that occasioned it, and then to the literature it would join. ## 7. Conclusion %%Conclusion to be written by NY%% ## Appendix: maybe this should go somewhere in Sec. 2 > User: Last December, I tried to start my car early in the morning when it was very cold. It wouldn't start. Why not? > > Kimi k2.67: The most likely culprit is the battery. In very cold weather, a battery's chemical reactions slow dramatically, reducing its available capacity by up to 50%. If the battery was already marginal (older than three years, or not fully charged), freezing temperatures can push it below the threshold needed to turn the starter motor. You may have heard rapid clicking or nothing at all. > > Other plausible contributors: > > - Thickened engine oil — Cold increases viscosity, making the engine physically harder to crank and demanding even more current from the already weakened battery. > - Fuel system — Condensation in the fuel lines can freeze, blocking flow. Diesel fuel can gel; petrol can struggle to vaporise properly in extreme cold. > - Spark/ignition — Moisture in the distributor cap or on spark plugs can freeze, preventing a proper spark. > > If it started fine once temperatures rose later in the day, the battery is almost certainly the primary cause. A load test would confirm whether it needs replacement or just a longer drive to reach full charge. --- ## References Austin, J. L. (1962). _Sense and sensibilia_ (G. J. Warnock, Ed.). Oxford University Press. Bengson, J., Cuneo, T., & Shafer-Landau, R. (2022). _Philosophical methodology: From data to theory_. Oxford University Press. Chudnoff, E. (2011). What intuitions are like. _Philosophy and Phenomenological Research_, 82(3), 625–654. Davies, D. (2004). _Art as performance_. Blackwell. Dellsén, F., Firing, T., Lawler, I., & Norton, J. (2024). What is philosophical progress? _Philosophy and Phenomenological Research_, 109(2), 663–693. https://doi.org/10.1111/phpr.13067 Floridi, L., Morley, J., Novelli, C., & Watson, D. (2025). What kind of reasoning (if any) is an LLM actually doing? On the stochastic nature and abductive appearance of large language models. _arXiv_. https://doi.org/10.48550/arXiv.2512.10080 Frankish, K. (2024). What are large language models doing? In A. Strasser (Ed.), _Anna's AI anthology: How to live with smart machines?_ (pp. 55–78). Xenomoi. Goldie, P. (2000). _The emotions: A philosophical exploration_. Oxford University Press. Gottweis, J., Weng, W.-H., Daryin, A., Tu, T., Palepu, A., Sirkovic, P., Myaskovsky, A., Weissenberger, F., Rong, K., Tanno, R., Saab, K., Popovici, D., Blum, J., Zhang, F., Chou, K., Hassidim, A., Gokturk, B., Vahdat, A., Kohli, P., ... Natarajan, V. (2025). Towards an AI co-scientist. _arXiv_. https://doi.org/10.48550/arXiv.2502.18864 Guevara, A., Lupsasca, A., Skinner, D., Strominger, A., & Weil, K. (2026). Single-minus gluon tree amplitudes are nonzero. _arXiv_. https://doi.org/10.48550/arXiv.2602.12176 Harman, G. (1990). The intrinsic quality of experience. _Philosophical Perspectives_, 4, 31–52. Jackson, F. (1982). Epiphenomenal qualia. _The Philosophical Quarterly_, 32(127), 127–136. Lipton, P. (2004). _Inference to the best explanation_ (2nd ed.). Routledge. Magnani, L. (2009). _Abductive cognition: The epistemological and eco-cognitive dimensions of hypothetical reasoning_. Springer. %%single-authored — in-text "Magnani et al. 2009" in footnote 12 needs aligning%% Merleau-Ponty, M. (2012). _Phenomenology of perception_ (D. A. Landes, Trans.). Routledge. (Original work published 1945) Millière, R. (2022). Deep learning and synthetic media. _Synthese_, 200, 231. https://doi.org/10.1007/s11229-022-03739-2 Novikov, A., Vũ, N., Eisenberger, M., Dupont, E., Huang, P.-S., Wagner, A. Z., Shirobokov, S., Kozlovskii, B., Ruiz, F. J. R., Mehrabian, A., Kumar, M. P., See, A., Chaudhuri, S., Holland, G., Davies, A., Nowozin, S., Kohli, P., & Balog, M. (2025). AlphaEvolve: A coding agent for scientific and algorithmic discovery. _arXiv_. https://doi.org/10.48550/arXiv.2506.13131 Pigliucci, M. (2017). Philosophy as the evocation of conceptual landscapes. In R. Blackford & D. Broderick (Eds.), _Philosophy's future: The problem of philosophical progress_ (pp. 75–90). Wiley-Blackwell. Pitt, D. (2004). The phenomenology of cognition: Or what is it like to think that P? _Philosophy and Phenomenological Research_, 69(1), 1–36. Salimi, M., Adim, S., Parnian, D., Alighardashi, N., Jafari Siavoshani, M., & Rohban, M. H. (2026). Wiring the 'why': A unified taxonomy and survey of abductive reasoning in LLMs. _arXiv_. https://doi.org/10.48550/arXiv.2604.08016 Unger, R. M., & Smolin, L. (2015). _The singular universe and the reality of time: A proposal in natural philosophy_. Cambridge University Press. Williamson, T. (2007). _The philosophy of philosophy_. Blackwell. Williamson, T. (2021). _The philosophy of philosophy_ (2nd ed.). Wiley-Blackwell. Wolfram, S. (2023). _What is ChatGPT doing … and why does it work?_ Stephen Wolfram Writings. https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-doing-and-why-does-it-work/ Zahavy, T. (2026). _LLMs can't jump_ [Preprint]. PhilSci-Archive. https://philsci-archive.pitt.edu/28024/ Zeni, C., Pinsler, R., Zügner, D., Fowler, A., Horton, M., Fu, X., Wang, Z., Shysheya, A., Crabbé, J., Ueda, S., Sordillo, R., Sun, L., Smith, J., Nguyen, B., Schulz, H., Lewis, S., Huang, C.-W., Lu, Z., Zhou, Y., ... Xie, T. (2025). A generative model for inorganic materials design. _Nature_, 639, 624–632. https://doi.org/10.1038/s41586-025-08628-5