# Abstract In this paper, we argue that current-generation LLMs are capable of producing philosophical texts that are worth reading. We do not begin by settling what counts as good philosophy; instead we appeal to a distinction any reader will recognise, between texts that repay the time spent reading them and texts that do not. We defend this claim against a series of challenges. The first, the challenge from authorship, holds that philosophy is something only persons can produce. Drawing on Davies' performance theory of art, we argue that a text's being worth reading answers entirely to what is on the page: two texts containing the same argument cannot differ in philosophical merit. The second, the challenge from abduction, holds that LLMs cannot weigh rival explanations but only reproduce the appearance of doing so. We argue that the distinction between genuine and merely apparent abduction cannot be located in the texts these systems actually produce. The third, the challenge from connecting to the world, holds that a system with no perception cannot reach philosophy's starting points or test its claims; the fourth, the challenge from experience, holds that a system that has never experienced anything cannot do the philosophy that begins from or concerns experience. We argue that the materials philosophy takes from the world and from experience reach it already articulated in language, which a model can work on as any philosopher does. The fifth, the challenge from observation, asks why, if all this is correct, worthwhile LLM-written philosophy is nowhere to be found. We argue that this reflects how these systems are used rather than what they can produce. The final challenge, from instrumentality, holds that where a philosopher's prompting elicits a text worth reading, the philosophy is the philosopher's and the model merely its instrument. We argue that a prompt articulates a starting point which underdetermines its development, and that what the development states beyond the prompt is not the prompter's. --- # 0. Introduction The last decade or so has seen the rise of generative artificial intelligence: systems that produce text, images, code, music, video, and other outputs in response to prompts. AI has had success in domains where the value of an output is not exhausted by its superficial fluency. For example, in February 2026, researchers working on gluon scattering amplitudes gave GPT-5.2 worked examples for three, four, five, and six particles and asked it to find the general formula. GPT-5.2 proposed a closed-form expression; another internal model supplied a proof; and the authors then verified the result. The resulting paper argues that single-minus tree-level gluon amplitudes, often presumed to vanish, are non-vanishing in certain half-collinear configurations (Guevara et al. 2026). There are also recent examples in mathematics (Novikov et al. 2025), biomedicine (Gottweis et al. 2025), and materials science (Zeni et al. 2025). In this paper we argue that we should expect similar success in philosophy. %%this needs to be replaced with the recent maths discovery%% Specifically, we argue that current-generation LLMs are capable of producing philosophical texts that are _worth reading_. We do not want to begin by settling what counts as _good_ philosophy. Instead, we appeal to a distinction that anyone reading this text will recognise. You have read texts that are worth reading, and you have read texts that are not. As you begin reading this article, you likely hope that it is worth reading, in the sense that the time spent reading it will not be wasted. When you write a philosophical text yourself you aim to make it worth readers' while to read it, and whether or not the journal you send it to accepts it, depends on whether or not they agree. Two clarifications are needed. First, a text’s being worth reading is not the same as its being correct. A text can repay attention even if one rejects its conclusion: it may sharpen a distinction or answer an objection in a way that changes the dialectical situation. Second, the minimal unit we are concerned with is not the bare conclusion of an argument, but the argument itself. If an LLM output consists only in a pronouncement on some philosophical topic ('Direct Realism is correct', 'We should be utilitarians'), it is hard to see why it would be worth reading in and of itself, for the same reason that a bare pronouncement by a human philosopher would not be worth reading.[^1] The next six sections develop the main argument. Section 1 rejects the challenge from authorship: the claim that an LLM output cannot be philosophy worth reading because no philosopher lies behind it. Section 2 turns to abduction and argues that the absence of human-style inference to the best explanation in the producer does not preclude abductive structure in the product. Section 3 turns to the challenge from connecting to the world, and argues that the materials philosophy takes from the world reach it already set down in words, which an LLM can work on as any philosopher does. Section 4 takes up the challenge from experience, and argues that the experiences philosophy argues about reach it in the same way. Section 5 turns to the challenge from observation — that if all this is right, philosophy worth reading should already be coming from these systems, and is not — and argues that this reflects how they are used rather than what they can produce. Section 6 addresses the challenge from instrumentality — that where a philosopher's prompting draws out such a text, the philosophy is the philosopher's — and argues that a prompt articulates a starting point whose development it does not fix. --- # 1. The Challenge from Authorship In this section we address what we might call the _challenge from authorship_: the idea that philosophy is something that only persons, or at least minds, can produce. This view has not, to our knowledge, been explicitly defended in just this form, but it gives shape to an intuition that many philosophers may have: philosophy is a person-only domain. An imperfect comparison is with art: One might deny that an image generated by an AI system is an artwork, because no artist exercises the relevant kind of intentional control over its production.[^1] One might think, for similar reasons, that philosophy can only be done by people: no text produced by an LLM can be a work of philosophy, because no philosopher lies behind it. Similar to the study of art, the study of philosophy is often organised around individuals: undergraduates take courses on Kant's ethics or Lewis's metaphysics, and conferences are devoted to, the work of particular philosophers. Physics students, on the other hand, are taught Newtonian mechanics from a current textbook, and the course loses nothing if Newton's own writing is never looked at. In the sciences, then, what a text contributes can be carried by other texts. In philosophy, the contribution and its original presentation are harder to prise apart, and we might take this as evidence that a philosophical work is bound to the activity of the particular person who produced it, in a way that the sciences are not. We will now try to make this challenge from authorship more precise, by considering how far Davies' _performance_ theory of art transposes to philosophy. Davies writes: > [T]he work – what the artist achieves – is the process eventuating in that product. Works themselves are neither structures nor objects simpliciter, nor are they contextualized structures or objects. They are, rather, intentionally guided generative performances that eventuate in contextualized structures or objects (or events, as we shall see) – performances completed by what I am terming a focus of appreciation. (2004, p. 97) On Davies' view, when a painter paints a picture, the canvas is what we attend to, but it is not the work. The work is the artist's intentionally guided activity in producing the canvas; the canvas is "the focus of our appreciative interest in the work" (2004, p. 150). What we appreciate in a painting, on this account, is an achievement, and an achievement is individuated by the activity that brought it about: the same surface, reached by some other route, would be a different achievement. Provenance, on this view, does more than supply context: facts about how the object came into being help determine what the work is and what is properly appreciated in it. Davies makes the case with examples in which indiscernible surfaces differ as works. In one kind there is no performance at all: an instance of the verbal structure of _Kubla Khan_ might be generated by desert wind, or a monkey at a typewriter. A theorist who identifies the poem with its verbal structure must then either count these as instances of Coleridge's work or explain why not (2004, p. 102). There is a surface indistinguishable from the poem with no writing of a poem behind it, and the corresponding case for philosophy — a text indistinguishable from a philosophical argument with no philosophising behind it — is the one an LLM presents. Questions of misattribution, where one performance is taken for another, do not bear on it. If Davies is right, the surface does not by itself settle the work. Here is what the analogous proposal for philosophy would be. A philosophical text is not itself the philosophical work: the text is the product of a person's philosophising, and reading it is a way of engaging with that prior activity. The activity, on this proposal, is part of what the work is, so that there is a philosophical work only where the philosophising has taken place, which makes the challenge a constitutive one. If no one has philosophised, there is no work to which the text gives access, however the text reads — an LLM text would stand to philosophy as the wind-made _Kubla Khan_ stands to poetry. Should the transposition be accepted? We do not think it should. In the art case Davies' move is licensed by an evaluative fact: indiscernible surfaces can differ in artistic value, the wind-blown surface worth nothing as a painting where a brushed one may be worth a great deal. The transposition therefore commits its defender to the corresponding claim about philosophy: that two texts containing the same argument could differ in philosophical merit. We can find no difference for the merit to consist in. If two texts contain the same argument, including the same inferential moves, the same considerations count for and against them: whether the argument is valid and whether the objections are answered are questions about the texts' contents, and two texts with the same contents receive the same answers. Their philosophical merit does not vary with the route by which the words came to be written. The discipline's evaluative practice is built on the same denial. Journals strip author information from submissions before review because facts about authorship are treated as potential sources of distortion; if texts with the same contents could differ in merit, anonymising would discard evaluatively relevant information, and review would not be designed this way. The grounds for the judgement lie in the argument as presented, not in the history of its production.[^3] Nor is the author-centred teaching noted earlier in tension with this. That ethics is taught through Kant's _Groundwork of the Metaphysics of Morals_ rather than a digest of its conclusions reflects what a reader gains in understanding by working through Kant's own arguments, and is consistent with holding that the philosophical merit so gained is a feature of the text rather than of its author. Suppose the performance theorist holds firm: where there has been no philosophising there is no work, whatever the resulting text contains. This may be allowed, because the thesis in question concerns texts worth reading, and a text can be worth reading without being a work in Davies' sense. Suppose the desert wind assembled not _Kubla Khan_ but a sound argument against enactivist approaches to perception, an argument for which no one would deserve credit. A reader who worked through it would nonetheless meet a thesis and the considerations advanced for it, and would be in a position to answer or extend it. Whether such a text is a _work_ may then be reserved for texts with performances behind them; what cannot be reserved is the text's being worth reading, since everything that judgement answers to is on the page. Nothing in a text's having had no one behind it, then, settles whether it is worth reading. Whether an LLM can produce a text worth reading is a further question, and a doubt of a different kind bears on it. [^1]: This is not to deny that systems of this kind can produce beautiful images; we return to image generation in Section 4. [^3]: The same location of philosophy in the public text is reached by accounts of philosophical progress. Dellsén et al. (2024) hold that progress is "for-whom" rather than "by-whom": it consists in putting people in a position to increase their understanding, usually by making philosophical ideas publicly available (p. 679). Being put in such a position requires something one can take up and think through, and what is available to be taken up is the text. A philosophical contribution so understood is constituted by what the public text makes available, which a view that locates the philosophy in the antecedent private activity mislocates. --- # 2. The challenge from abduction In the previous section, we argued against the idea that LLMs cannot produce philosophy worth reading simply because they are not human. In this section and the next we shall consider a different form of challenge: even if LLMs cannot be ruled out of producing philosophy worth reading tout court, they lack particular _capacities_ that are required to produce it. If our arguments regarding authorship are correct, there is no reason that a novel philosophical argument produced by a parrot should be taken any less seriously than one produced by a human. Yet parrots _cannot_ produce such complex strings of sounds as their powers are mimetic, rather than productive. One might think the same is true for LLMs. They just don't have a capacity, or capacities, required to produce worthwhile philosophical argument. In this section we address one capacity challenge, which we will call _the challenge from abduction_. In the two sections that follow we shall look at two more: the _challenge from connecting to the world_ and the _challenge from experience_. Abductive inference is reasoning from a body of evidence to the hypothesis that would best explain it. In a deductive argument the premises fix the conclusion: if all men are mortal and Socrates is a man, then Socrates is mortal, with no wriggle room whatsoever. Now, imagine walking into your kitchen one morning and finding that part of the floor is wet. What has happened? The wet floor does not determine the answer in the way the two premises determined Socrates' mortality: a burst pipe would have left the floor wet, and so would a spilled bucket. But, given that the window is open, the water is under the window, and it rained last night, rain coming through the window seems the most plausible answer. To reason in this way, inferring the best explanation for a set of facts, is common in the sciences as well as every day life. A scientist chooses one theory over another when it explains the same data more simply: Copernicus's model of the solar system was preferred to Ptolemy's because it explained the observed planetary motions without the elaborate epicycles the older model required. Williamson argues that philosophy should use a broadly abductive methodology (2007; 2021, §9.2). In philosophy, as in science, there are data that a candidate theory must accommodate — intuitions about cases, and the phenomena of the domain itself — and rival theories that would each accommodate them at different costs. The theory to prefer is the one that would, if true, best accounts for the data, with 'best' being cashed out in terms of explanatory virtue What makes one theory's explanation better than another's, on this account, is a matter of explanatory virtue%%repeat%%: a good philosophical theory is, in Williamson's words, "elegant and unified, not arbitrary, gerrymandered, ad hoc, or messily complicated", and should "combine simplicity with strength" (2021, §9.2). This need not rest on a science-modelled conception of philosophy: Bengson, Cuneo and Shafer-Landau (2022) argue that the assessment of rival theories by their explanatory and unifying merits is a constraint on sound philosophical method as such.[^holders] We shall assume it in what follows. If a capacity for abduction is required to produce worthwhile philosophy, do LLMs possess it? Floridi et al. (2025) argue that they do not: > We argue that such LLMs generate text based on learned associations rather than performing abductive inferences. […] LLMs produce plausible hypotheses, simulate commonsense reasoning, and provide explanatory answers without grounding them directly in truth, semantics, verification, or understanding, and without any abductive reasoning. (Floridi et al. 2025, p. 1) We will split this quite comprehensive critique of LLMs' abductive capacities into two parts. In the rest of this section we will consider the charge that these systems do not genuinely infer to the best explanation but simply generate text based on learned associations. In the next, we will focus on the idea that LLMs are not connected to the world in a way that would allow them to "genuinely validate" (ibid. p. 6) their explanations against reality. An LLM, on this account, has "a stochastic core and an abductive appearance" (2025, p. 2). Models are trained to predict which words are likely to follow which, and to produce the continuation their training makes probable — to aim at the likely continuation rather than truth. The appearance comes from what the training data have passed on: models have "absorbed patterns of human abductive reasoning as expressed in writing" (p. 9) — how explanations are typically phrased, which causes are typically offered for which effects. [^1] Floridi et al.'s own example is a car that will not start on a cold morning. Asked why not, a model replies that a weak battery is one possibility, since cold reduces a battery's efficiency; that thickened engine oil is another, since a cold engine is harder to turn over; and that, "[b]ased on your description, the battery is the most likely explanation" (2025, p. 10). On their reading, the LLM is not actually reasoning about causes from the case before it; it is the statistical reproduction of the causes such explanations typically cite (p. 9). What they deny is that the model weighs the candidates — that it "entails choosing the best explanation among alternatives" (2025, p. 3). The verdict that settles on the battery does no weighing; it reproduces how explanations of this kind conventionally end. And where the output marks a genuine difference between the two — some consideration that would tell the battery from the oil — it is one already drawn in the explanations the model learned from, not one worked out afresh for the case in hand. The fact that LLMs can produce no more than a "veneer of explanation" (ibid. p.20), Floridi et al. argue, means that LLMs can only ever play a supporting role in intellectual work: > In essence, LLMs function like brainstorming assistants that toss out ideas without filtering for quality. After all, they work like statistical interfaces to an enormous amount of data accumulated for millennia by generations. A cautious human collaborator can sift through and assess them. (2025, p. 11) A brainstorming assistant seems a far cry from something which might produce worthwhile philosophy. If you were presented with a text and told that it contains a number of philosophical ideas, none of which have been filtered for quality, it is unlikely you would think it is worth your while reading it. The filtering Floridi et al. have in mind is not a matter of selecting the most likely continuation — the stochastic core already does that. It is a matter of preferring what Lipton (2004, p. 59) calls _lovelier_ explanations to merely _likelier_ ones: the likeliest is the one most warranted by the data, the loveliest the one that would, if true, provide the most understanding. As he puts it, "Likeliness speaks of truth; loveliness of potential understanding" (PAGE REF). The two come apart. Consider again the wet floor in the kitchen. A very likely, almost certainly true, explanation is that the floor is wet because water has fallen on it, yet offering that as an explanation would be met with exasperation: of course it is because water fell on it, but how, and which water? Banally true explanations do little for a person's understanding. On Williamson's abductive methodology, weighing rival explanations is a matter of explanatory virtue, and the explanation to prefer is the one with a virtue the others lack (2021, §9.2). To prefer an explanation on that ground is to prefer the lovelier rather than the likelier, since explanatory virtue is a matter of the understanding an explanation would afford if true, not of its probability. Not all philosophical writing turns on this kind of weighing, but where a philosophical text does turn on it, whether the text is worth reading and whether it weighs its rivals well go together. And it is just this kind of weighing that Floridi and his colleagues say a system that does no more than continue text cannot do. We do not disagree with Floridi et al.'s characterisation of how LLMs function: these systems do not weigh and choose among alternatives in the way that humans do. However, we should not be too quick to jump from this to the conclusion that LLMs cannot _produce text_ which exhibits abductive reasoning. A pocket calculator does not have the capacity to do arithmetic in the way a person does, but does have capacity to produce the correct answer to sums which are entered into it. Similarly, it might be possible for LLMs to produce text which displays abductive reasoning, despite it not being grounded in any actual abductive reasoning. Floridi et al. are not merely saying that the process by which an LLM produces its output is stochastic rather than abductive. They are saying that the output itself is only apparently abductive — that what looks like inference to the best explanation is, in their words, a "compelling illusion of genuine and structured inferential reasoning" (2025, p. 2). Trained on a great deal of writing in which explanations are offered and weighed, such a system absorbs the forms this writing takes and, prompted to explain, reproduces them, following "the typical phrasing and structure of explanations" and offering "typical causes for typical effects" rather than "reason[ing] about causes from scratch" (2025, p. 9). On this view the output has the form of an abductive explanation but not the substance — the shape of a weighing of explanations, taken over from the writing the model has digested, and not a genuine weighing of the case in hand. Floridi and his colleagues allow that, in ordinary cases such as their own cold-morning car, the model's answer is a good one: "the same explanation a human reasoner would likely choose" (2025, p. 10), one that may be "even optimal by IBE criteria" (2025, p. 19), since such systems "echo the obvious, common explanations" (2025, p. 10). The complaint cannot then be that the explanation is poor, which leaves it hard to say what, in such an answer, is supposed to be merely apparent. What their account points to is the uncommon case: "in less common situations, LLMs can falter" (2025, p. 10), and on inputs "that go beyond their training" "the facade can crack" (2025, p. 9), the success on familiar cases being "a sign of overfitting to common patterns" (2025, p. 15). Once the question of its truth is set aside, this is what the facade reduces to: the claim that the competence shown on common problems is overfitting that would give out in less common ones. Whether the competence gives out on the uncommon case is an empirical question, and the most recent survey of abductive reasoning in language models seems at first to bear the conjecture out (Salimi et al. 2026). Current models do markedly worse on abductive tasks than on deductive ones: where their median accuracy on deductive tasks is near eighty per cent, on abductive ones it is some forty-two and a half, with a spread "extending down to near-zero accuracy", and the survey reports that "strong deductive performance does not reliably imply strong abductive performance". This is close to Floridi's own diagnosis — an answer that reproduces a common pattern instead of reasoning to it — now voiced from within the field that builds the systems. Taken at face value, it tells in Floridi's favour. Our response to the challenge from abduction begins by considering what text having an abductive appearance actually amounts to. Consider first that, despite their stochastic core, LLMs are perfectly capable of producing grammatically correct text. Despite not being given specific rules, LLM training means that the system "implicitly 'discovers' them—and then seems to be good at following them" (Wolfram 2023). Does this mean that the texts LLMs produce have merely the appearance of being grammatically well-formed? Clearly not. LLMs sentences _are_ gramatically well-formed despite their stochastic roots. This suggests that a stochastic core need not mean that the best an LLM can do is produce a veneer of abductive inference. It may be, rather, that the core is marshalled to produce text exhibiting actual abductive inference, in just the way it is marshalled to produce actual grammatical correctness. It might be objected that the grammar analogy will not stretch this far. The model has picked up the shape of abductive explanation, but it is not actually performing abduction. The preceding paragraph suggested that because LLMs implicitly discover and follow the rules of grammar, they might in the same way be marshalled toward genuine abductive inference. But grammar, the objector would say, is a system of rules a model can follow without understanding anything, and explanatory loveliness is not — it is not a system of rules you follow to get the right answer, so the model's grammatical competence gives no reason to expect it to produce lovely explanations. Floridi's claim is that the model's abductive output is merely apparent — that it has the form of an inference to the best explanation but not the substance. If that is right, then there should be cases where the form is present and the substance is absent: cases where the model produces something that looks like an explanation but is empty or nonsensical, just as it could produce something that is grammatically faultless and says nothing. Wolfram's example of the latter is "Inquisitive electrons eat blue theories for fish" (2023) — impeccably grammatical, and meaningless. The model does not produce such strings. The abductive equivalent would be a passage with the form of an inference to the best explanation, grammatically faultless, and yet senseless — asked why a car will not start on a cold morning, there are no squirrel tracks, so it must be freak arctic winds blowing into the exhaust pipe. That has the form of evidence weighed towards a conclusion, and it is nonsense; and it is nonsense the model does not produce. Put the question to it and it offers the weak battery and the thickened oil, plausible explanations rather than arctic winds. Wherever the facade is to be located, it cannot be located there: what the model produces are plausible explanations, and it is not yet clear what, in a plausible explanation, is supposed to be merely apparent. Moreover, the answer Floridi and his colleagues describe is not quite the answer these systems give. Their account asks us to picture a confident verdict laid over a hollow core, yet what one finds in practice is closer to hedging. Asked why a car would not start on a cold December morning, a current model will say that the battery is the most likely culprit, set out other plausible contributors — thickened oil, fuel-system problems, ignition faults — and end by observing that if the car started once the day had warmed, the battery is almost certainly the primary cause, though a load test would confirm it.[^kimi] The reply does not announce the answer; it fits its confidence to the little it has been told, and marks the point past which it will not go without knowing more of the particular car. This is not yet to say that Floridi's facade charge is refuted — a hedged answer can still be, on his account, a statistical reproduction of how hedged explanations typically read. But it does mean that Floridi et al.'s own example does not quite fit the charge. The case they describe, in which a confident surface verdict is laid over a hollow core, is not the case their own example presents: the confidence these systems express is already answerable to their evidence, and a confidence so answerable is not the facade the objection has in view. That the model keeps clear of senseless explanations, as it keeps clear of senseless sentences, points to a core of semantic competence picked up in training, over and above syntax. Wolfram calls it a semantic grammar. Syntax, he notes, settles only how the parts of speech may be combined: > to deal with meaning, we need to go further. And one version of how to do this is to think about not just a syntactic grammar for language, but also a semantic one. (Wolfram 2023) A model trained on enough text has, on his account, come by one: > From its training ChatGPT has effectively "pieced together" a certain (rather impressive) quantity of what amounts to semantic grammar. (Wolfram 2023) Wolfram's notion of a semantic grammar is meant to capture what a model picks up over and above syntax. A syntactic grammar settles only how the parts of speech may be combined — what counts as a well-formed sentence. It does not settle whether what a sentence says makes sense. "Inquisitive electrons eat blue theories for fish" is grammatical and says nothing, and it is the absence of such strings from the model's output that tells against treating its syntax as appearance alone. Avoiding such strings takes more than syntax: it takes a grasp of what can sensibly be said of what, which predicates go with which subjects, which causes go with which effects. ~~Wolfram calls this a semantic grammar~~, and his claim is that a model trained on enough text has come by one — has, in his words, "pieced together" a quantity of what amounts to semantic grammar (2023). It is this, rather than any contact with the case in hand, that keeps the arctic winds out of the model's explanations: a feel for what, in a working model of the world, can hang together. But a feel for how things hang together is gathered from the text the model has read, and a model of the world is not the world.[^wm] What a semantic grammar supplies is a sense of what would sound right, not a line to how things actually stand. The same detachment that keeps the model's explanations sensible is what leaves it unable, on its own, to reach the car on the drive. A semantic grammar is a grasp of how such failures are explained in general — cold slows batteries, thickens oil, fouls ignition — and it gives no hold on the particular car whose fault is in question. Pressed for the cause of this failure to start, the model can go only so far before it needs what it cannot get for itself: some purchase on the actual car in the actual world, the lights tried, the turn of the key heard. Set that aside, and what is left is, so far as the text goes, abduction itself: a weighing of explanations that respects sense, fits its confidence to the evidence, and stops where the evidence stops. How much the missing purchase costs depends on the kind of abduction at issue. The car is a hard case, because its answer waits upon the world; only the car itself can tell the weak battery from the frozen line. Much philosophical abduction does not wait upon the world in that way: what a thought experiment commits us to, or which of two theories carries the lighter explanatory cost, is settled from what is already set down, where the discriminating evidence is the sort a corpus already holds. The case that shows the model at its most hobbled is thus the world-bound one, and not the case philosophy most often presents — a matter we take up in the next section. What is left for this one is the empirical question of whether, on harder cases, this competence in fact gives out. Returning to the survey by Salimi et al. (2026), the scores measure whether the model arrived at the right answer, not the weighing that got it there. On the harder tasks the scores are low, and on long mysteries with their clues strewn through the text the best models fall just short of the average human solver; taken at face value, the numbers count against the model. But the score is a score for the answer — whether the named culprit was the keyed one, whether the right diagnosis came first — and not for the weighing that reached it; such a score, Salimi says, "completely bypass[es] the actual reasoning trace". So when a model misses the keyed culprit, the number marks the miss, and not the comparison it set out on the way — which explanations it canvassed, and why it came down on one. And the survey's hardest tasks are built from low-prior, non-stereotypical outcomes, where several explanations may be reasonable; there, matching an answer to a single reference "underestimates explanation quality", so a low score on such tasks does not show that the model cannot weigh, only that it did not land on the keyed answer. What a one-shot score leaves unscored is the weighing behind it, and that is what a fuller prompting would have to draw out. It is now difficult to see where the facade should be located. The explanations the model produces are plausible rather than senseless, the confidence they express is fitted to the evidence it has been given, and the benchmark scores that seemed to confirm the diagnosis measure answers rather than the weighing behind them. What the model lacks is purchase on the particular case: a semantic grammar is gathered from text, and so allows the model to say how cold-morning failures are explained in general, but not to tell what has gone wrong with a particular car. This is a limitation concerning the model's relation to the world rather than its capacity for abduction. How much it costs a philosophical text is a matter for the next section; what a one-shot prompt leaves undrawn, and how it might be drawn out, is a matter for Section 5. [^wm]: We use "model of the world" in Wolfram's sense; whether such systems possess world models in the sense now disputed, and how they connect to the world, is the matter of §3. Example of real exchange with kimi 2.6 high > **User:** Last December, I tried to start my car early in the morning when it was very cold. It wouldn't start. Why not? > > **Kimi k2.67:** The most likely culprit is the battery. In very cold weather, a battery's chemical reactions slow dramatically, reducing its available capacity by up to 50%. If the battery was already marginal (older than three years, or not fully charged), freezing temperatures can push it below the threshold needed to turn the starter motor. You may have heard rapid clicking or nothing at all. > > Other plausible contributors: > > - Thickened engine oil — Cold increases viscosity, making the engine physically harder to crank and demanding even more current from the already weakened battery. > - Fuel system — Condensation in the fuel lines can freeze, blocking flow. Diesel fuel can gel; petrol can struggle to vaporise properly in extreme cold. > - Spark/ignition — Moisture in the distributor cap or on spark plugs can freeze, preventing a proper spark. > > If it started fine once temperatures rose later in the day, the battery is almost certainly the primary cause. A load test would confirm whether it needs replacement or just a longer drive to reach full charge. --- # 3. The Challenge from Connecting to the World %% Note to self: disambiguate what I mean by world models in this section at some point; see World models in AI - terminology note for Section 3. %% In this section we address a second capacity challenge, the _challenge from connecting to the world_. In addition to the argument considered in the previous section, Floridi et al. object that an LLM stands in no relation to the world: its words rest on no perception of anything, and a hypothesis, once produced, is never tested against how things are (2025, pp. 7–9). A discipline whose theories answer to how things are, the challenge runs, cannot be advanced by a system with no access to how things are, and a text produced by such a system gives its reader no reason to think it worth reading. We argue that the challenge fails because the materials philosophy takes from the world reach it already articulated in language, and a model can work on those words as any philosopher does. We can grant that the model perceives nothing and that it cannot put what it produces to the test against the world. Whether either concession bears on the texts it produces depends on what philosophy does with the world — on where a philosopher's starting points come from, and on what becomes of a philosophical claim once it is made. Pigliucci (2017) addresses both. He holds that philosophy is constrained by the world without investigating it as the natural sciences do: > This means that the basic parameters that philosophers use as their inputs, the starting points of their philosophizing, their equivalent of axioms in mathematics and assumptions in logic (or rules in chess) are _empirical_ data about the world. This data comes from both everyday experience [...] and of course increasingly from the world of science itself. Philosophy, I maintain, is in the business of exploring the sort of conceptually evoked spaces that Smolin is talking about, where the evocation is the result of whatever starting assumptions are made by individual philosophers working within a particular field and, crucially, of the constraints that are imposed by our best understanding of how the world actually is. (2017, pp. 79–80) _Evocation_ is a term Pigliucci takes from Smolin (Unger and Smolin 2015),[^4] for truths that are neither discovered, in the sense of corresponding to mind-independent states of affairs, nor invented, in the sense of being arbitrary constructs, and his example is chess: when the rules of a game are codified, a whole bundle of facts about it becomes demonstrable — objective facts, in that anyone who can demonstrate one demonstrates the same fact as anyone else — although chess did not exist before its rules were written down (Unger and Smolin 2015, p. 423; quoted at Pigliucci 2017, p. 78). Pigliucci's proposal is that philosophy ascertains evoked truths in this sense, with an addition that separates it from mathematics and chess alike: its starting points are constrained by how the world actually is. The addition is also what separates philosophy from fiction, on his account. A novelist's worlds are invented rather than evoked — nothing about them is rigid, since even the constraints the novelist adopts could have been otherwise — whereas philosophy "is in the business of doing empirically informed evoking, not inventing", so that its objects of study have rigid properties (2017, p. 80). A thought experiment is itself a case of such evoking: the philosopher sets up an imagined scenario but explores it "with an interest in figuring things out as far as this world is concerned" (2017, p. 80), so that what it evokes has the rigid properties Pigliucci means, and the philosophical work proceeds within the structure it opens. Philosophy's starting points, then, are empirical, and a system that perceives nothing cannot reach them on its own. But the data Pigliucci describes comes from everyday experience and, increasingly, from science, and the scientific kind reaches working philosophers already articulated — already set down in language, available to be read rather than undergone. A philosopher of physics works from published results, not from having run the experiments. The same holds for everyday experience: what the discipline retains of it, it retains as the literature's accumulated descriptions of how things seem. For the worldly materials philosophy actually uses, written access is the profession's normal condition rather than a deficiency, and a corpus is such access. On this point the model stands where every philosopher already stands with respect to nearly all of the empirical data they use. The other objection was that the model never checks what it produces against the world. But a philosophical claim is not the kind of thing that gets checked that way. Compare Einstein's equivalence principle with Jackson's (1982) Mary, who has spent her life in a black-and-white room and knows every physical fact about colour vision. The equivalence principle, once Einstein had it, faced a tribunal of measurement: the experiments might have gone against it, and then it would have been dropped. Mary's case faces no such tribunal. The question Mary raises — whether she learns something new on first seeing red — is a question about what follows within the scenario Jackson has set up, and that is settled in the way any question about chess is settled: by working out what the set-up commits us to, something any competent party can do and none can decide by fiat. This is the only checking a philosophical thesis gets, and it happens in the literature, in the back-and-forth the previous section described. So a model's inability to run experiments costs it nothing a philosophical text needs: the testing that philosophy does is the working-out of what a scenario commits us to, and that is done on the page. A model's lack of perception and experiment therefore leaves untouched both the starting points a philosophical text works from and the checking its claims receive: the former arrive already set down in language, and the latter consists in working out what they commit us to, which is done on the page. It might be doubted, however, that experience can be treated in the same way, since a description of what it is like to see red is not obviously an adequate substitute for seeing it, and the next section addresses a challenge built on this doubt. [^4]: _The Singular Universe and the Reality of Time_ is jointly authored, but its second part, which contains the discussion of evocation, was written by Smolin alone, as Pigliucci notes (2017, p. 77); we follow him in attributing the view to Smolin. # 4. The Challenge from Experience In this section we address a third capacity challenge, the _challenge from experience_. The challenge holds that some philosophy depends on experience in a way a system without experience cannot meet: experience supplies the starting point of some philosophical reasoning, and is itself the subject matter of some philosophical inquiry. Few would say that current LLMs are conscious, and we assume here that they are not. We argue that the challenge fails for the same reason as the challenge from connecting to the world: the experiences philosophy argues about reach it already articulated in language, and a model can work on those descriptions as any philosopher does. Zahavy (2026) raises a worry of this kind about scientific discovery. A model can carry out the deductive part of discovery, working out the consequences of premises it has been given; what it cannot do, he holds, is produce the premises — make the move from sense experience to new first principles. On the picture he takes from Einstein, that move is a leap, and it is the leap[^5] that gives a theory its axioms. His case is the thought experiment that gave Einstein the equivalence principle: > Einstein's variation required inventing new axioms based on a physical intuition that did not yet exist in the mathematics. He envisioned a physicist inside an elevator being uniformly accelerated through deep space [...]. Inside this enclosure, the sensory experience reveals a specific pattern: when objects are released, the floor rushes up to meet them. To the physicist, the objects appear to fall with identical acceleration, regardless of composition. Thus, the simulation here was not a permutation of symbols, but a manipulation of perceptual experience. (Zahavy 2026, §5) Einstein imagines a set of circumstances and attends to what would be experienced within them: everything released inside the elevator appears to fall with identical acceleration. On Zahavy's reconstruction the simulation supplies an observation, and from that observation the new axiom is inferred — the simulated experience of acceleration was indistinguishable from the remembered experience of gravity, and Einstein concluded that the two are one phenomenon.[^2] A model has no access to that observation. It can produce descriptions of elevators and of weightlessness, both present in its corpus, but it has undergone neither, and a discovery whose premises are fixed by simulated experience is beyond a system that, in Zahavy's words, lacks the capacity he calls sensory agency. We might think that philosophical thought experiments depend upon experience in the same way. Does Mary, released from her black-and-white room, learn something when she first sees red? Settling the question requires considering what the experience is like, and the argument proceeds from the verdict; this is experience entering as the starting point of an argument, and a system that has never experienced anything appears unable to supply it. Experience enters also as subject matter, in philosophy that asks what it is to see red (Harman 1990), to feel anger (Goldie 2000), or to have an intuition take hold (Chudnoff 2011).[^3] Were these claims correct, much of the philosophy of mind would lie beyond a model's reach. Take the knowledge argument itself. Nobody who debates it has been through what Mary goes through — released into colour after a lifetime of black and white — and the debate does not suffer for it. The participants have in front of them Jackson's description: a scenario set down in words, with a claim about what it is meant to show. Lewis's reply (1988) works on that description, changing what we should say follows from Mary's release while saying nothing about what her first sight of red is like from the inside. This is the usual way experience figures in philosophy. A philosopher need not have had an experience to argue about it: Nagel (1974) asks what it is like to be a bat without supposing he could find out, working instead from a description of the bat's situation. A model is no worse off here than these philosophers are. It has had no experiences of its own, but it has not needed them, because what its corpus supplies, and what the philosophy works on, is experience already put into words — the same descriptions that, in Section 2, carried displayed reasoning into its texts. Merleau-Ponty noticed something about touching one hand with the other that was not, so far as we know, already written down anywhere. At any moment one hand is the toucher and the other the touched; the roles can switch, but they cannot both hold at once (Merleau-Ponty 1945 %%page%%). Suppose he came to this only by attending to his own body, to what the touching was like from the inside, and not from anything already in the literature. Then it is a starting point a model could not have reached on its own, since reaching it took a body and a first-person view that a model does not have. Once Merleau-Ponty has written the description down, a model can take it up and argue about it as well as anyone. It could not, though, have been the one to set it down first. What a philosopher does with Merleau-Ponty's description, though, has nothing to do with who first arrived at it. The description earns its place by what can be drawn out of it — what it shows about the body, what follows once the asymmetry is granted — and that is there for any reader to work through, whether or not they could have come to the asymmetry on their own. First-person attention is not, in any case, the only way new descriptions come about. A description the literature does not yet contain can also be reached from the descriptions it does contain, by drawing out what they have not been taken to imply, or by putting two of them together as no one has — and a model can do this. Whether today's models in fact produce descriptions the literature lacks is the question of novelty, which Section 6 takes up. A model can do the philosophy that turns on experience, save for the one point already granted: it could not be the first to set down a description that only first-person attention could yield. The experiences philosophy argues about reach it as descriptions, which it can work through as well as any reader. In Section 2 abduction entered philosophy as reasoning displayed in a text, and in Section 3 the world entered it as starting points set down in one; here experience enters it the same way. What a model produces is read as any philosophy is read, and how it was produced settles nothing in advance. [^2]: Zahavy, following Magnani, calls the process _manipulative abduction_: hypothesis generation through the manipulation of a model — here a simulated experience — rather than of symbols (Magnani et al. 2009; Zahavy 2026, §5). It is abduction in Section 2's sense: the equivalence principle is inferred as the best explanation of the simulated observation, the simulation supplying an explanandum that no search over existing text would have produced. What experience contributes, on this picture, is not the inference but its starting point. [^3]: The materials need not be sensory: the feeling of understanding something is sometimes used to motivate the claim that thought itself has a phenomenology (Pitt 2004). [^5]: Zahavy too calls this leap abduction, but the word picks out something other than it did in Section 2. There, with Floridi et al., abduction was the weighing of rival explanations, and the charge was that a model only mimics it; here it is the generation of new first principles from experience, and Zahavy's claim is that a model cannot make the move because it has had no experience to move from. The present challenge rests on that second claim, about experience, and not on any verdict about the weighing. --- --- # 5. The Challenge from Observation [this section needs the most work] In this section we address what we will call the _challenge from observation_. In Section 1 we argued that LLMs should not be disqualified from producing worthwhile philosophy tout court. In Sections 2, 3 and 4 we argued that, although LLMs neither perform abductive inference, nor have experience, nor are connected to the world, there is still reason to think that they are capable of producing text which exhibits good quality abduction, and works with articulated axioms about experience and the world. The challenge from observation begins with an obvious question: if all of these arguments are correct, where is all the worthwhile LLM-written philosophy? If you ask an LLM the answer to the hard problem of consciousness, or the meaning of life,1 you will not receive _the correct answer_, but instead a competent but unopionated survey of the field if you are lucky, or a less accurate but equally bland survey if you are unlucky. The observation is accurate, and it reports less than it seems to: it reports what models produce under one use — a bare question, put once, answered in one pass. How these systems are built explains why that use yields what it does. A model is first fitted to a vast general corpus and trained to continue text, so its response to a bare philosophical question is the likely continuation of such a question in writing at large, and the likely continuation of "what is the meaning of life?" in a general corpus is not an analytic tract. It is the sort of text that follows the question at large: a survey of views, a consoling generality, a joke. The model is then further shaped to converse as a helpful assistant, and the shaping presses the same way, since a person employed to be helpful to all comers would not answer the question with a tract either. The survey is not a ceiling the systems have hit; it is the likely continuation of exactly what was given them. The use that generates the observation treats the model as an oracle: a system whose answers are its measure, so that asking is all the eliciting there is.2 The empirical record tells against the assumption. The survey of abductive benchmarks discussed in Section 2 runs every test with a single fixed instruction and scores the answer, while cataloguing, in the same pages, methods that alter what models produce — prompts that separate the stages of a task, pipelines in which an answer is criticised and revised over several passes (Salimi et al. 2026). What a model returns depends on what it is given, and the observation samples one point in that space, the bare question. It therefore cannot discriminate between the two hypotheses at issue — that the capacity defended in the preceding sections is absent, and that it has not been elicited. Both predict the observed record, and an argument against this paper needs the first; the observation supports it no better than the second. --- # 6. The Challenge from Instrumentality In this final section we address the _challenge from instrumentality_, which arises as a natural escalation of the challenge addressed in the previous section: if philosophy worth reading comes out of these systems only when a philosopher directs the process — supplies the framing, sets the constraints, presses for development — then the philosophy, it will be said, is the philosopher's. The model is an instrument in the production, as a typewriter is, and crediting it with the result is crediting the dummy with the ventriloquism. Section 1's challenge held that a model's text is not philosophy tout court; what stands here is narrower, that the philosophy in such a text is not the model's. Whether the escalation succeeds depends on what prompting a model involves: what a prompt supplies, and what the model's continuation adds to it. Section 3's account of starting points says the first; Section 2's account of continuing text says the second. A prompt articulates a starting point, as a thought experiment does. A prompt that sets out a position and the rivals it must beat stands to the model as Jackson's two paragraphs stand to the profession: a starting point handed over for development. What an articulated starting point does, on the account already in place, is evoke a structure with rigid properties — there are facts about what holds within it, demonstrable by anyone and chosen by no one, and they outrun whatever has been stated, just as the facts about chess outran the rules the moment the rules were written down. Most of what a starting point evokes, no one has ever said. What the model contributes is the development, and the mechanics are the ones Section 2 drew from Wolfram: a model produces a reasonable continuation of the text it has been given, where what counts as reasonable is relative to the corpus it was fitted to (2023). A prompt is part of the text the model has been given. An articulated starting point therefore changes what there is to continue — the reasonable continuation of a stated position under stated constraints is not the reasonable continuation of a bare question — and the model makes use of what the prompt states in everything that follows: tell one of these systems something once, Wolfram observes, and it is used thereafter (2023). The continuation that results states consequences of the starting point that the starting point does not state. Section 2 said what it is for such a text to go well — the comparison it displays cites differences that tell between the positions, and would, if correct, give understanding — and whether a given continuation goes well is read off the continuation. Nothing in this makes the development a transcription. An evoked structure contains more than any text states: the rules of chess settle every fact about chess, and do not settle which theorems get written down, in what order, or to what depth, so that two writers working from the same rules produce different books, both correct, neither dictated by the rules. The mechanics mirror the structure, since the same prompt, run twice, yields different continuations (Wolfram 2023). The starting point underdetermines the development, and the gap between them is where the model's contribution lies: were there one text the prompt fixed, the output would transcribe what the person had already settled, and the instrument description would be true. The gap also leaves room for error. A development can state what does not hold in the evoked structure — a chess writer can publish a false theorem, a philosopher can misdraw the consequences of their own thought experiment, and a model can do both, along with its characteristic failure of stating fluently what nothing supports. The errors are found on the page. And an error is attributable only to a developer: no one blames the rules of chess for a false theorem, and no one's typewriter has ever made a mistake of content. Three contributions, then, and three owners: the articulated starting point is the person's; the structure it evokes, and the facts that hold there, are no one's; the text that develops them is the model's. Much in the instrument picture is true. The person writes the prompt and the prompt is authored; the person chooses which continuations to pursue and when to stop; without the person, there is the survey. What the picture adds to these truths is a description of the model — a device, like the typewriter, that fixes only what its user has already settled — and the description is what the account above denies. Every word of the novel was the author's before the typewriter touched it; the consequences a model's text states were nobody's before the text stated them. What the user of a typewriter settles is the text; what the writer of a prompt settles is a starting point. The account invites an obvious enrichment of the prompt. State the position, name the rivals, list the objections and the lines along which they are to be met, and at some point, it will be said, the prompt contains the philosophy and the model is expanding what the person wrote — so that where a model's output is good, one should suspect a prompt rich enough to have done the work. But enriching a prompt enlarges the starting point without converting it into the development. A game with more rules is a bigger game, not a book of its theorems, and however much the prompt states, the consequences the output draws were not among the statements. There is a genuine limiting case — a prompt that states the comparison and the verdict, so that the continuation only rephrases — and it is identified the way everything in this paper is identified: set the output against the prompt and ask what the text states that the prompt did not. A text that states nothing beyond its prompt is a paraphrase, and owed to the person; a text that states what the prompt left unstated is a development, and the unstated part is not the person's. Which of the two a given output is, is settled by reading them together. Even if all this is granted, we might still ask whether such a text can do more than handle well the positions a literature already contains and make a distinction that literature lacks, and so be creative in the stronger, public sense in which a human philosophical text is creative. We want to be careful here, because this is quite different from the modest sense in which an output is novel only relative to its prompt; what is at issue is the stronger claim of saying something the literature had not yet said. We do not need to settle this question here. It would be settled as the rest has been, by setting the output not only against the prompt but against the literature, and asking what it says that the literature had not. This is not meant to settle the matter either way, but it does give us a reason to leave the question open, for further work. We can now return to the challenge from the previous section, where the ordinary blandness of what these systems produce when given a bare question was taken to show that they have nothing to contribute to philosophy. The outputs to such questions do tend towards the empty and the thin; but that bears only on whether bare questions are good tests of philosophical capacity, since what these systems do is continue the context they are given, and a context with no shape of argument can only draw from them an output with no shape of argument. This is no reason to think that a context which hands the system a position, its rivals, and the pressures bearing on each of them cannot draw a development of its own; and whether such a development is worth reading is settled not by looking at the system but by reading the continuation first against the prompt that occasioned it and then against the literature it means to add to. To sum up, these systems are not oracles; they are continuation systems, and philosophy worth reading needs a dialectical context which a prompt can supply without thereby fixing the development, so that the philosophical standing of any output depends on what the continuation itself adds. ## Footnotes 1. While preparing this paper we asked GPT-5.5 for a detailed overview of the positions an analytic philosopher might take on the meaning of life. What came back was a competent, hedged survey of the field; what did not come back was an argument for any position in it. %%add date of test%% ↩ 2. That these systems are mischaracterised as oracles — with the corollary that no benchmark of single-pass answers should be expected to probe the upper limits of what they can produce — has been argued from inside the practitioner literature (Janus 2022). ↩ --- # References # References Frankish, K. (2024). What are large language models doing? In A. Strasser (Ed.), _Anna's AI Anthology: How to live with smart machines?_ (pp. 55–78). Xenomoi. # Artifact, Forest, Person. The Chimera Aesthetics of Generative AI ## Introduction In recent years, aestheticians and philosophers of art have turned their attention towards generative AI, investigating whether AI systems can be authors or co-authors, and whether AI-generated work has any aesthetic merit at all (Anscomb 2022; Wojtkiewicz 2023; Cross 2025). Still, there is another interesting philosophical question that has not been properly addressed so far, namely, whether and how AI systems themselves can be objects of aesthetic appreciation. This paper aims at answering this question by focusing on the paradigm case of LLMs. We argue that there are three main ways in which the aesthetics of LLMs might be built. The first, most straightforward way, is within the frame of the aesthetics of design, treating LLMs simply as designed artifacts. Still, LLMs's aesthetically relevant features — the patterns in their outputs, a particular model's characteristic 'feel' — emerge from training and from the way generated continuations develop from context, rather than being specified by designers themselves. The text the appreciator reads is the outcome of a process the trained system has gone on from the prompt — under conditions design has set, but in a direction it has not. LLMs, indeed, are not only designed but also trained. Olah (2024) describes the relation between what design fixes and what training produces as a type of growth. > I think one useful way to think about neural networks is that we don't program, we don't make them, we grow them. We have these neural network architectures that we design and we have these loss objectives that we create. And the neural network architecture, it's kind of like a scaffold that the circuits grow on. It starts off with some random things, and it grows, and it's almost like the objective that we train for is this light. And so we create the scaffold that it grows on, and we create the light that it grows towards. But the thing that we actually create, it's this almost biological entity or organism that we're studying. Given the way in which an LLM appears to us, as a subject able to engage in conversation, one might be tempted to identify this almost biological entity or organism with a sort of person. This brings us to the second candidate frame for the aesthetics of LLMs: the aesthetics of persons. Users indeed talk about different models' personalities, and it is natural to respond aesthetically to these apparent traits. Still, LLMs lack the temporally extended life, stable dispositions, and projects that underwrite person appreciation. This suggests that the aesthetic of persons can capture at most the way LLMs appear, not what they are. There is, instead, we argue, a third kind of aesthetic appreciation that fits LLMs' deep nature. That is the aesthetic appreciation of natural objects and environments such as mountains, rivers, and forests. Carlson (REF) argues that we appreciate nature by attending to patterns produced by natural forces, guided by knowledge — geology, ecology, and the like — that makes those patterns visible. He calls this mode of aesthetic engagement "order appreciation" and he contrasts this with "design appreciation", which is, instead, the mode of aesthetic engagement appropriate to artifacts. We argue that LLMs are special artifacts that call not only for design appreciation but also for order appreciation: attention to patterns produced by trained continuation systems, guided by knowledge of how generated text develops from context under learned constraints. We conclude that, in a sense, all the three basic modes of appreciation are relevant to the aesthetics of LLMs, which thus reveal themselves to be especially peculiar aesthetic objects that merit a distinctively threefold appreciation. Specifically, design appreciation accounts for LLMs as designed systems, order appreciation accounts for LLMs as systems that are not only designed but also, as Olah puts it, "grown", and person appreciation accounts for how LLMs appear, that is, the experiential effects they elicit from human users. Still, we contend, order appreciation is the key to the aesthetics of LLMs, the crucial central piece that bridges the gap between the appreciation of their designed structure and the appreciation of their personish behavior. We will proceed as follows. Section 1 sets out Carlson's distinction between design appreciation and order appreciation, and considers why person appreciation has to be discussed alongside it. Section 2 asks whether design appreciation can guide the appreciation of LLMs, and gives only a partially positive answer since LLMs are best understood as trained continuation systems that transcend design. Section 3 shows that order appreciation is the key to the aesthetics of LLMs, and argues that it is based on knowledge of how trained continuation systems develop text from context. Section 4 considers the role of person appreciation in our engagement with LLMs, arguing that it can shed light on their conversational function but it rests upon design-directed knowledge and especially order-directed knowledge, which enable us to properly appreciate the structure in virtue on which that function is fulfilled. We conclude by showing how this framework guides appreciation of an LLM as a whole. --- # 1. Design Appreciation, Order Appreciation and… Person Appreciation Both our criticism of artefactual and agentive conceptions of LLMs and our positive account will draw from Carlson's environmental aesthetics, as laid out in his 2000 book Aesthetics and the Environment. In particular, we adopt Carlson's general recommendation for aesthetic appreciation: take things as what they are, and look at them in the light of the right kind of knowledge. He applies this to the appreciation of the natural environment thusly: > First, that, as in our appreciation of works of art, we must appreciate nature as what it in fact is, that is, as natural and as an environment. Second, it recommends that we must appreciate nature in light of our knowledge of what it is, that is, in light of knowledge provided by the natural sciences, especially the environmental sciences such as geology, biology, and ecology. The natural environmental model thus accommodates both the true character of nature and our normal experience and understanding of it. (Carlson, 2000, p. 6) This captures something intuitive about how we appreciate nature versus art. Appreciating mountains and cliff faces as the work of a divine artisan, rather than of natural forces, would be wrong-headed (cf. Carlson, 2000, Chapter 8); so would appreciating a Rembrandt as if it were the product of natural forces slopping paint together (cf. Danto 1974, p. 140). In both cases, appreciation is undermined by a failure to recognise what the object really is. Carlson argues that artworks and everyday objects call for design appreciation. With paradigmatic artworks, we recognise them as objects whose features are, as Gombrich (1950, p. 13) puts it, each "the result of a decision by the artist". We appreciate such works by seeing how well the result realises the artist's design. The same approach extends to designed artifacts more generally. Carlson (2000, p. 188) is explicit that functional objects are properly appreciated by seeing how their forms answer to what they are for: > With anything functionally designed, not only its form, but much of its aesthetic interest and merit, 'follows function'. So, on this account, a chair, or a bridge invite the same style of attentive appraisal as a painting – guided by knowledge of ends, materials, constraints, and the fit between purpose and realisation. For the natural environment, in contrast, Carlson recommends a different mode: order appreciation. Appreciation of things like trees or valleys cannot be grounded in considerations of how well a designer managed to realise her intentions, because they are not designed objects. Instead, Carlson recommends that the knowledge grounding appreciation of the natural world is knowledge of how the order we find has been shaped by natural forces: > On the assumption that order appreciation provides the correct model for the appreciation of nature, such appreciation has the following general form: An individual qua appreciator selects objects of appreciation from the things around him or her and focuses on the order imposed on these objects by the various forces, random and otherwise, that produce them. Moreover, the objects are selected in part by reference to a general nonaesthetic and nonartistic story that helps make them appreciable by making this order visible and intelligible. (Carlson, 2000, p. 119) Order appreciation is not confined to nature, however. Carlson finds it already required by unconventional works whose patterns are produced rather than designed: Pollock's action painting, where the pattern arises from the behaviour of the paint and the painter's unplanned movements, and the Dada experiments with chance, where Tzara drew the words of poems from a hat and Arp placed cut-outs by chance (Carlson 2000, pp. 110–113). Such works have no initial design against which they could be judged, but they do have an ordered pattern, produced by a combination of forces: the properties of the materials, chance, and an artist who figures as one force among the others and whose remaining role is to select which results are kept. In this respect they are closer to trees and valleys than to the chair or the bridge; Arp wanted works of this kind to "remain anonymous and form a part of nature's great workshop as leaves do, and clouds" (quoted in Carlson 2000, p. 117). In both modes, knowledge guides acts of aspection — ways of attending to an object that partly constitute its appreciation (Carlson, 2000, pp. 41–42). But the character of this knowledge differs. In the case of design, we need functional and technical understanding: what the designer intended and what constraints they faced. In natural cases, we need an account of the processes that produced the order we perceive — knowledge that lets us see natural structures as effects of forces (Carlson, 2000, pp. 60–61). Once the relevant account is in play, some cases will exhibit its order more clearly and interestingly than others; still, Carlson holds that all of nature is more or less equally appreciable (Carlson, 2000, pp. 118–120). So far so good but, we contend, Carlson's approach to aesthetics overlooks another important category of object of appreciation: people. In ordinary life we not only admire landscapes and artifacts; we admire people too – their wit, their manner, their steadiness. Some philosophers have taken this practice seriously, investigating the aesthetic appreciation of personality – sometimes termed "beauty of character" – and asking whether traits such as kindness, wit, or courage can be aesthetically as well as morally valuable (Gaut 2007; Paris 2018). Carlson's recommendation seems naturally extendable here: appropriate aesthetic appreciation of persons will depend on the right kind of person-directed knowledge – familiarity with a life (real or fictional) and a sense of the values and dispositions that organise it. A stranger's single act can of course strike us as kind. The object of beauty-of-character appreciation, however, is a trait, and a trait is a standing disposition which a single act is not sufficient to establish: the same piece of behaviour might be an expression of generosity or of calculated self-interest, and which of these it is depends on how it fits into the rest of the agent's life. Appreciating kindness as a feature of someone's character therefore requires knowing the pattern of their generous responses given who they are and what they have faced. As Parsons stresses, such knowledge is typically built up through direct interaction, careful biography, or the more precarious route of gossip (Parsons 2023, 297–299). Although Carlson does not consider person appreciation, it is not difficult to imagine ways in which his framework might be extended or modified to accommodate it. A first option is to treat the appreciation of character as a special case of order appreciation: we focus on the psychological, social, and biographical forces that shape a life, much as we attend to geological and ecological forces in a landscape. A second option would be to emphasise the ways in which personalities are, at least in part, self-shaped, and to appreciate them as self-designing projects – a thought that has obvious attractions for existentialist traditions. Nothing in what follows requires settling the issue between these options. On the principle implicit in Carlson's account, moreover, modes of appreciation are individuated by the kind of knowledge that guides them. Knowledge of a life is knowledge of a person's reasons and values, and this is not reducible either to knowledge of a designer's intentions or to knowledge of the forces operating on an object; the philosophers of character cited above already treat its appreciation as a practice of its own. We will accordingly speak of persons as a third category alongside natural items and artifacts, with their own distinctive mode of appreciation anchored in their status as subjects, leaving open how this category ultimately relates to the other two. On all of these views, however, Carlson's recommendation still applies: aesthetic appreciation is guided by substantive background understanding of what persons are like and how their traits hang together over time. Person-based aesthetics, in this sense, presupposes a rich conception of the subject as a temporally extended agent with relatively stable dispositions, projects, and evaluative commitments, grasped under a suitable body of knowledge. All this leaves us with three candidates for the aesthetics of LLMs, namely, design appreciation, order appreciation, and person appreciation. As LLMs are made by human beings just like artefacts, the natural first thought is that they admit design appreciation in the way artefacts do. In the next section, we will thus consider design appreciation as the first candidate model for the aesthetics of LLMs. --- # 2. Design Appreciation of LLMs Technical artefacts call for design appreciation, which requires knowledge of the purpose for which the object was made and of how its features serve that purpose. LLMs initially look like straightforward candidates for this mode of appreciation: they are engineered systems, and their responses depend on choices made before any user enters a prompt. On this view, models such as GPT-5.5, Claude 4.8 Opus, and Gemini 3 Pro can be characterised by a relatively unified functional role, namely, a general-purpose conversational assistant, and by a specific way of realising that role through an architecture and a user interface whose design might be admired for elegance, efficiency, or ingenuity. Still, we argue that knowledge of the design and functioning of LLMs cannot, on its own, ground their appreciation. With traditional designed artifacts, design-knowledge illuminates structure because designers specified it. Knowing what the designer intended and what constraints they faced allows us to understand why an artifact has the form that it does. A car is appreciated through its designed form. For example, the shape of its body is at once what the appreciator looks at and what determines how the car moves through air at speed. This shape is what it is because of what the car is for. To appreciate the car aesthetically is to attend to this fit between form and purpose, and design knowledge — knowledge of the car's purpose and of how its features answer to that purpose — reaches the form the appreciator engages with. This sort of approach, however, does not fit so cleanly for LLMs. Nothing the user engages with stands to the LLM as the car's body stands to the car: what she sees is a text box and the text appearing in it, and the items that were designed (the architecture, the training objective) do not appear anywhere in her experience of the system. The interface is visible and designed, and might be admired accordingly, but it is peripheral to what users aesthetically respond to. One might reply that the generated text is where the design is realised, so that design appreciation can proceed there. But the organisation of the trained system emerges from training rather than being specified in advance, and design-knowledge therefore does not illuminate this emergent organisation: there was no designer's specification that laid it out. To understand the system, one must attend to the training process that produced it. Specifically, design does not completely settle how the system will respond to a prompt; it determines only the conditions under which the model is trained. Design fixes the training objective, which provides the system with a standard against which its outputs can be adjusted. Then, given a stretch of text, the model assigns probabilities to possible continuations, and its parameters (i.e. weights) are altered when those probabilities diverge from the continuation found in the training data. Thus, what the trained model eventually does with a prompt is not settled by design choices; it is acquired over the course of training. When the system responds to user input, it uses what it has acquired in training to continue the text it has been given. The prompt, together with whatever the system has already produced, gives it a context (that is, an enriched text). From that context the model assigns probabilities to the possible next tokens, one of which is selected and added to the context before the same step runs again, so that what appears at the end as a single answer is built through successive transitions of this kind. Later parts of the continuation depend on earlier parts, and each step is shaped by the dispositions acquired in training. The text the user reads is what the trained system has produced going on from the prompt — under conditions design has set, but in a direction it has not. Recall Olah's (2024) description of the relation between what design fixes and what training produces as a kind of growth: > the neural network architecture, it's kind of like a scaffold that the circuits grow on. It starts off with some random things, and it grows, and it's almost like the objective that we train for is this light. And so we create the scaffold that it grows on, and we create the light that it grows towards. But the thing that we actually create, it's this almost biological entity or organism that we're studying. What design fixes is the scaffold and the objective; the form that grows under them is acquired through training. While in the car case what the user engages with is the form the designer built, an LLM's response is produced by trained dispositions operating on the prompt and on the continuation as it develops. Design knowledge can explain the conditions under which the response becomes possible, but it does not by itself give us the order of the generated text. Thus, artefact is too coarse a description for the appreciation Carlson aptly requires. LLMs are artefacts, and design knowledge bears on how they should be appreciated; ignoring how they were built would distort that appreciation. Carlson's question, however, concerns what the thing specifically is. At that level, an LLM is a trained continuation system: a made system whose responses are not designed but rather generated from dispositions acquired in training and operating on context. Those dispositions are not written into the code as explicit rules about how to, say, handle metaphors, or politely decline illicit requests. They are emergent regularities in a trained network that has been pushed, by training, to reduce prediction error. What training makes "grow" on Olah's "scaffold" is a system of statistical associations and processing circuits whose internal organisation even designers often understand only partially. These are the features design appreciation cannot reach. Much of what users find aesthetically important in LLM behaviour – the way a model sustains a metaphor or abruptly drops it, the pattern of hedging and self-correction, the texture of its reasoning, the sorts of digression it tends to indulge, the characteristic "feel" of its refusals – is grounded in features that have not been micro-designed but have emerged from optimisation under constraints. Parsons and Carlson (2008) note that, even for simpler artifacts, knowledge of function must include knowledge of how that function is realised if it is to structure appreciation appropriately. In the LLM case, this means including knowledge of the way training and alignment have grown a particular style of continuation on top of the designed architecture. In this sense, knowledge of how function is realised concerns not only design but especially growth. Thus, if we try to make design appreciation do all the work, we mislocate the primary source of what matters aesthetically. In garden-variety technical artifacts, the designer's choices fix most of what matters aesthetically. In the LLM case, by contrast, the order that matters aesthetically is largely the order of a trained statistical system running under its own learned constraints. Designers specify objectives and scaffolds but the particular ways in which text is produced in response to the user's prompt are not written down anywhere as a plan. The upshot is that LLMs are artifacts, and there is a place for design appreciation in their aesthetic appraisal: we can and should evaluate how well their designed forms answer to their engineered functions, as well as how what Olah calls "scaffold" and "light" can differ from model to model. However, the most distinctive and revealing aesthetic phenomena arise not from the execution of a detailed design, but from the emergent linguistic order that these grown systems exhibit when they are run. To appreciate that order, we need knowledge not of what designers intended but of how training shapes text propagation. --- # 3. Order Appreciation of LLMs LLMs are trained continuation systems, to wit, systems whose responses develop from dispositions acquired in training and operating on context. That is why design knowledge cannot, by itself, guide their appreciation. Design knowledge reaches the scaffold, the objective, and the conditions under which the system is trained and deployed, but it does not by itself make visible the order acquired by a particular continuation as context is extended. The sort of knowledge central to appreciation of LLMs must make the learned order of generated text visible as the order of a trained continuation system. This order can be encountered at more than one scale. A single output is one bounded continuation from a context. An extended chat is a longer process in which earlier turns condition later ones. The model as a whole is the trained system whose tendencies become visible across many such outputs and chats. These scales do not compete for the appreciator's attention. The system is a set of standing dispositions, and dispositions can only be encountered through their manifestations: the outputs are what the appreciator perceives, and the system is that to which the perceived order belongs. To appreciate an LLM is thus to appreciate it through its outputs, much as a climate can only be appreciated through the weather it produces. Different sub-disciplines of computer science might be put forward as candidates for helping the appreciator to grasp this order. One field that has emerged in connection with neural networks is mechanistic interpretability, which investigates the internal workings of these systems by reverse-engineering them into human-understandable algorithms, identifying which circuits, attention heads, and internal representations handle which linguistic tasks (Olah et al. 2020; Elhage et al. 2021). Consider, though, the difference between chemistry and geology when appreciating a cliff face. Chemistry provides knowledge of molecular bonds within rock, but it operates at a scale invisible to the naked eye. Geology, by contrast, offers concepts – e.g. strata, faults, erosion channels – that connect to what can actually be seen: one can perceive strata without specialist equipment, and knowing how sedimentation works makes the visible layering intelligible. Mechanistic interpretability occupies the position of chemistry in this comparison: the causal, circuit-level knowledge it yields concerns weight matrices, activation patterns, and circuit-level features, none of which is available to readers encountering generated text. For an aesthetics of LLMs, we thus need an account whose concepts describe perceivable features and render them intelligible as products of the system's learned regularities: an output-side account of how trained systems develop linguistic forms through iterated continuation from context. At any point in a run, the system receives the context so far and computes a distribution over possible next tokens. Once one token is selected, the context changes, and the next step is produced from that changed context. Training gives the system a graded sensitivity to the regularities of text — to what tends to follow what under what conditions. When the model is run, those regularities operate through a context that changes as the text develops from the starting point fixed by the prompt. The path generated by the model is computed token by token but the order in question is encountered by readers as a piece of ordinary language, since the trained system has acquired patterns governing what tends to follow what under given conditions in a given language. This is what Picca (2025, p. 1) captures when he describes LLMs as systems that "recombine, recontextualize, and circulate linguistic forms based on probabilistic associations". Wolfram (2023) puts these patterns in geometrical terms: > inside ChatGPT any piece of text is effectively represented by an array of numbers that we can think of as coordinates of a point in some kind of 'linguistic feature space'. So when ChatGPT continues a piece of text this corresponds to tracing out a trajectory in linguistic feature space. According to Wolfram, such a trajectory stays within meaningful text because the model has "implicitly discovered" the relevant regularities in training (PAGE); it holds them only implicitly, and no explicit statement of them is yet available. For the appreciator, though, it is enough that the regularities bear on the text in front of her: the continuation leans towards some words and away from others because training has made it so. A reader who knows as much can take the text before her as a trace of that process. We saw in Section 1 that order appreciation focuses on the order imposed on objects by the forces that produce them. Carlson summarises the account in terms of three entities and the interplay among them: "the order, the forces that produce it, and the account that illuminates it" (2000, p. 119). For nature, he fills these roles with the natural order, the forces of geology, biology, and meteorology, and the story told by natural science (2000, p. 120). The same roles can be filled for generated text. The order is the organisation a text acquires by being produced stretch by stretch, each stretch generated from those before it. Carlson's general form already provides for forces of two kinds, "random and otherwise", and both kinds are present here. Sampling supplies the random element, resolving each step from the distribution before it. The remaining forces are of the other kind: the learned regularities that weight the distribution, the accumulated context through which they operate, and the prompt, which initiates and conditions the process without determining it in detail. Since no plan for the whole is given in advance of generation, whatever direction a text has depends on what has already been produced and on how the learned regularities respond to it at each step; a direction can therefore strengthen, since each token in a register raises the probability that the next conforms to it, or weaken, since early material makes up a diminishing share of the context. The story is the account of trained continuation. On this scheme, the objects of appreciation are patterns that "are or can be seen as the marks of the forces" that produced them (2000, p. 111). It might be objected at this point that talk of forces producing a text is metaphorical, and hence in tension with Carlson's recommendation that things be appreciated as what they in fact are. However, Carlson does not use "force" as a physical notion: the properties of materials, chance, and the artist's own movements all count among the forces at work in the anti-art cases (2000, p. 113). Moreover, each of the forces listed above has a literal referent: the weighted distribution, the sampling step, the context, and the prompt are the factors that in fact produce each token. The pattern of a Pollock is conditioned by the pattern already on the canvas. Its visible shapes, in Janson's description, "are largely determined by the internal dynamics of his material and his process: the viscosity of the paint, the speed and direction of its impact upon the canvas, its interaction with other layers of pigment" (quoted in Carlson 2000, p. 110). The same holds of a continuation: each stretch of text is generated from, and conditioned by, the stretches already produced. The painter's involvement does not reintroduce design: in Carlson's analysis he figures as one force among the others, the one that sets the process going, and the prompt occupies the same position among the forces of generation. The closest of Carlson's cases to generated text is Arp's automatic poetry: > Automatic poetry comes straight out of the poet's bowels or out of any other of his organs that has accumulated reserves… He crows, swears, moans, stammers, yodels, according to his mood… His poems are like nature; they stink, laugh, and rhyme like nature. Foolishness, or at least what men call foolishness, is as precious to him as a sublime piece of rhetoric. For in nature a broken twig is equal in beauty and importance to the clouds and the stars. (quoted in Carlson 2000, pp. 117–118) Figure 1 belongs to this class of cases [PROVENANCE NOTE: model, generation regime, and permission to be settled]. When asked for its opinions on bees, a model replied with some three hundred words of near-language. Its invented vocabulary remains morphologically well formed and keeps returning to the subject of bees; its register is maintained throughout; and it preserves an oratorical structure of invocation, interludes, and peroration, detached from any occasion of use, although the words that fill this structure are themselves inventions. The reply differs from Arp's poems in what is going on. In Arp's case the story appealed to the unconscious and to the poet's mood; in this case the story of trained continuation tells us that the reply is a trajectory that remains within one region of feature space under a loosened selection of steps. The forces at work in this reply are the same as those at work in an ordinary one, differing only in their relative strength. Because sampling has been loosened while the learned regularities continue to operate, the contribution of each force is easier to distinguish: the weighting towards bee-related vocabulary is evident in the invented words, and blends such as "sweeat" and "beeings" fuse neighbouring items into single tokens, and so display what Wolfram calls the "fan" of high-probability continuations from which every step is drawn (2023, PAGE). In ordinary output the forces are more evenly weighted, and the pattern they produce is less conspicuous. When there is an initial design, appreciation includes judging whether the object is, in Gombrich's sense, "right". With an ordered pattern alone, this judgement has no purchase, and the story takes over the design's role of indicating "if not what is being done, then at least what is going on" (Carlson 2000, p. 113). The question a reader brings to a generated text is accordingly what is going on here. We saw in Section 1 that Carlson takes all of nature to be more or less equally appreciable; since a generated text has no design against which it could fail, the same holds here, and the routine reply to a routine question is no less appreciable than the reply of Figure 1, just as Arp's broken twig is no less appreciable than the clouds and the stars. The stories grounding order appreciation are, moreover, "in one sense nonaesthetic", yet "in another sense they are exceedingly aesthetic. They illuminate nature as ordered and in doing so give it meaning, significance, and beauty" (2000, p. 121), and part of the response they guide is directed at the process the pattern records — in Carlson's phrase, at something "distinct from and beyond humankind" (2000, pp. 121–122). The grown organisation the story points to is one that, as argued in the previous section, even its designers understand only partially. The same knowledge dictates the relevant acts of aspection (2000, p. 119): a long fluent exchange calls for survey, with attention to whether a direction strengthens or weakens across turns, while a reply like that of Figure 1 calls for word-by-word scrutiny of the items fused in each blend. As Wolfram (2023) points out, LLMs reveal that > human language (and the patterns of thinking behind it) are somehow simpler and more 'law like' in their structure than we thought. As each model has learned regularities from human text, its order also reflects, in a technologically transformed way, the linguistic culture of its training data. There is thus a sense in which generative AI is a mirror of culture, not only morally, as Vallor (2024) has argued in her book The AI Mirror, but aesthetically. The model shows us our own linguistic patterns, filtered through statistical learning. The order appreciation of LLMs, from this perspective, can also be cast as the aesthetic appreciation of culture seen through technology. The key to appreciation, however, lies in technological mediation. The order made visible by the relevant knowledge of how LLMs work is not only the order of our language but especially the order of linguistic forms carried forward and transformed through trained continuation.[FOOTNOTE — placeholder: The regularities of a textual practice are made true by many acts of writing taken together, and are not located in any one of them; a reader encounters only instances, and the regularities themselves can be recovered only by counting across a corpus. Training retains what recurs and discards what belongs to a single occasion, so that in generated text the conventions of a culture are exercised without being participated in, in a form transformed by corpus selection and compression. Nothing in Carlson's natural case corresponds to this further object of appreciation. We develop this account of appreciating a linguistic culture through generated text in work currently in progress.] Knowing that a text is LLM-generated rather than human-written changes how it is appropriately aspected. We saw in Section 1 that appreciation is undermined when an object is appreciated as something it is not, as when a cliff face is taken for a divine artisan's work or a Rembrandt for the product of natural forces. Reading generated text as authored is a new instance of the first mistake, in which an order produced by forces is credited to a designer. The same words support different appreciation under the two readings. If we read the reply of Figure 1 as human writing, it is a pastiche, and each blend is a witticism to be credited to its author. If we read it as what it is, the blends are fusions of neighbouring items and the register is a trajectory held within its region: marks of forces rather than of authorial choices. The knowledge that guides this appreciation need not be held theoretically. Carlson notes that scientific and everyday knowledge of nature lie on a continuum rather than differing in kind (2000, PAGE): the farmer who knows the soil through planting and tending can appreciate the order in a well-drained field in ways unavailable to someone who merely gazes at the landscape. The experienced user of an LLM develops acquaintance of the same kind. By prompting and observing across many contexts, she builds up a practical sense of a model's characteristic order: she learns which vocabulary it tends towards from given starting points, and how far earlier material continues to condition what comes later. Such knowledge is held practically rather than explicitly, and it is aspectual as well as predictive: it settles what to attend to and where the model's order is likely to be visible. At the scale of the model, familiar talk of one model having a different 'vibe' from another can be understood as a way of registering stable differences in the regularities the models have learned. Extended exchanges with an LLM are a natural site for building such acquaintance. Each prompt creates conditions under which the system responds, and each response shows something of how the model carries text forward; the back-and-forth of prompting is itself a mode of aspection, organising appreciative attention over time. Cross (2025) characterises certain AI art-making as an 'exploration paradigm' in which the artist iteratively probes the model, adjusting prompts and sampling variations, and he compares the practice to performance art, where the artist creates a space for the audience's participation. What the artist is doing, on the present account, is a form of interactive aspection: the prompts and adjustments are interventions that make the system's regularities visible, rather than ways of coordinating with a co-creator. The practical and the theoretical routes converge on the same object. The farmer's knowledge and the geologist's track the same forces, differing in how the knowledge is held rather than in what it is knowledge of; likewise, the user's feel for a model and the account of trained continuation track the same learned regularities. Those regularities are, as Wolfram puts it, implicit in the model; they are also held implicitly in the practised user's expectations; and they are made explicit, so far as they can be, in the theorist's story. Wherever on this continuum the knowledge is held, it does what Carlson requires of it: it makes the order of generated text visible and intelligible, and it dictates the acts of aspection appropriate to appreciating that order. --- # 4. Person Appreciation of LLMs We have argued that a proper aesthetic appreciation of LLMs requires supplementing design appreciation with order appreciation. Still, the way we interact with LLMs may look so similar to interacting with actual human interlocutors that one might wonder whether a further supplementation of design appreciation and order appreciation with person appreciation is required. The question concerns chatbots rather than LLMs as such, and the distinction should be made explicit. An LLM is a trained continuation system; a chatbot is one deployment of such a system among others. The same kind of model can be deployed in applications in which no interlocutor appears at all: Cotypist, for instance, uses a language model to provide system-wide predictive text, completing whatever its user is currently typing, and nothing in this use invites conversation.[FOOTNOTE: Cotypist, a macOS application developed by Daniel Gräfe (Accelerated Thought GmbH, 2024–), runs a small language model locally to supply inline predictive completion across applications: https://cotypist.app.] The appearance of a conversational partner belongs to one mode of deployment rather than to LLMs as such. We sometimes appreciate persons aesthetically. A person's warmth may be aesthetically appreciable as a feature of character, rather than as a feature of bodily appearance. Gaut (2007) and Paris (2018) treat such appreciation as directed at the traits and dispositions through which a person's life is intelligible. Something similar is possible with fictional characters: we can aesthetically appreciate a character as a person within a fiction, without believing that the character exists outside it. The account of LLMs we have defended so far, however, puts pressure on this comparison. The system producing the text is a trained continuation system: it generates responses by applying dispositions acquired in training to the context it is given. LLMs are systems that tokenise text, manipulate numerical vectors, and generate continuations by sampling from learnt probability distributions. Nothing in that description straightforwardly resembles a subject with beliefs, intentions, or a life-history; there is no obvious place for character traits, projects, or personal development. That being the case, there are two strategies to preserve person appreciation of LLMs. The first argues that the LLM is an intentional system in a thin sense, and that this is enough to license some person-directed appreciation. The second grants that the LLM is not literally a person but reads its outputs as fictional speech, so that what is appreciable is a fictional character rather than the system itself. The first strategy casts LLMs as agents of a thin and unfamiliar kind. On a suitably liberal conception of mind, perhaps they qualify as intentional systems and that is enough to license some person-based aesthetics. Frankish (2024) offers a sophisticated version of this idea. Drawing on Dennett's intentional stance, he suggests that LLMs can be treated as genuine, if unusual, intentional systems. On this view, we are licensed to ascribe beliefs and desires to an LLM when doing so yields a simple and fruitful account of its behaviour, even if the underlying implementation is purely mechanical. In the case of contemporary chatbots, Frankish proposes that we can ascribe to them a large set of thin 'beliefs' – roughly, informational states distilled from their training – and one thin 'desire': to play what he calls the chat game. A system is playing the chat game when it generates text that looks like a cooperative move in an ongoing conversation, respecting local coherence, relevance to the prompt, and broadly human conversational norms. An LLM, on this picture, is a system whose behaviour can be summarised by saying that it believes many simple things and wants to make an appropriate next move in the chat. Frankish is not inviting us to pretend that LLMs have beliefs and desires; he is claiming that, at the right level of abstraction, it is literally true that they do, in much the same sense in which a thermostat can literally be said to "want" the room to be at a certain temperature when adopting the intentional stance helps us describe its behaviour. The agent-talk is meant to latch onto real, pattern-like features of the system's organisation. Suppose we grant all of this. Does it give us what we need for aesthetic appreciation of LLMs as persons? Not so. A subject of beauty of character is not just any intentional system. It is, minimally, a being with a temporally extended life, with relatively stable value-laden dispositions, with projects and commitments that can succeed or fail, and with a capacity for speech and action to express and reshape its character over time. When we set the chat-game agent against this benchmark, it looks thin. The 'beliefs' are shallow, in the sense that they are confined to what is encoded in the model's parameters and surfaced in the current context, without memory or development across conversations. The "desire" is singular and thin: make an appropriate move now in this exchange. There are no independent projects pursued across episodes, no webs of concern or attachment, no history in which earlier experiences inform later choices. What structure there is, is entirely local to the present stretch of text. The predicates characteristic of person-aesthetics—'beautiful soul,' 'admirable steadiness,' 'ugly character'—presuppose something that can be tested, developed, or refined over time; a thin chat-game agent has no such temporal depth. From this perspective, LLMs may be agents in Frankish's concessive sense, but they are not the sort of agents whose lives and characters can be the object of the aesthetic responses associated with persons. There is nothing like a "beautiful soul" or an "ugly character" here in the relevant sense; there is no enduring set of values and dispositions that could be manifest, challenged, or transformed over time. Given Carlson's recommendation, the right kind of person-directed knowledge for beauty-of-character appreciation is knowledge of a life and its values. The technical and training facts about LLMs do not supply that kind of object. Once again, Carlson's recommendation sharpens the point. If, even on a concessive mindedness view, LLMs lack the life-structure required for person-aesthetics, then to insist on aesthetically appreciating them as persons would be to ignore what they in fact are. It would be to treat the thin chat-game profile as if it were enough to underwrite the rich person categories we apply to human agents, and to let those categories govern appreciation despite knowing that the underlying kind is different. The concessive strategy therefore does not secure an appropriate person-based aesthetics of LLMs. The second strategy takes the make-believe route. If we ask ordinary users whether they literally believe that a chatbot is a person, many will concede that they do not. They may talk to a model as if it were a friend or a colleague, and they may feel heard, reassured, or amused, but when pressed they acknowledge that they are interacting with a computational system rather than a human being. Their stance is, in this sense, already a kind of as-if posture. Mallory offers a way of theorising this posture through what he calls chatbot fictionalism (2023). On his view, we engage with chatbots by entering a game of make-believe in which the exchange is treated as if it were a conversation with an agent. Within the fiction, the chatbot 'says' things and 'means' things; outside the fiction, we know that no such speaker is present. At the metasemantic level, Mallory claims, the outputs lack literal semantic content – they are 'literally meaningless but fictionally meaningful' (Mallory, 2023, p. 1082). This fits the everyday thought that we can take a chatbot seriously in the moment without actually believing that it has a mind. Just as a child treats a banana as a sword in a game, we treat chatbot outputs as utterances within a kind of imaginative practice. This is not delusion but a deliberate, bounded pretence that allows us to coordinate with the system and even gain knowledge from it, much as we might learn geography from a map by imagining countries as two-dimensional shapes. Mallory's account is not itself an aesthetics of LLMs; it is primarily a semantic and epistemic proposal about how we can use them and learn from them. But it highlights one obvious way a person-based aesthetic stance might be defended: one might suggest that we should aesthetically appreciate LLMs as if they were persons or characters, in the same sense in which we respond aesthetically to fictional protagonists whose existence we do not literally believe in. We respond to Gatsby's enigmatic, dream-chasing idealism or Ron Swanson's libertarian gruffness using much the same vocabulary as we use for real people, and we often talk quite straightforwardly about their 'character' or 'personality'. The warmth or wit a user finds in Claude or in ChatGPT can be cast as the warmth or wit of a fictional interlocutor. Fictional characters, however, are not fully-fledged persons but rather—according to a popular view in philosophy of fiction—artifacts that have the function of eliciting imaginings of fictional persons within a story-world (cf. Thomasson 1999; John 2021; Terrone 2021). From this perspective, a proper appreciation of a fictional character should consider not only its person-like appearance but especially the designed representational texture (say, the text of a novel or the images of a film) in virtue of which that appearance shows up in the imagination of the audience. That is to say that fictional characters ultimately call not only for persona appreciation but especially for design appreciation. Thus, even if we concede that the generation of the experience of interacting with a fictional person is an aspect of the function that LLMs fulfil—hence of what they are—the fact remains that the appreciation of that function is grounded in the appreciation of the structure in virtue of which it is fulfilled. While in the case of fictional characters such structure calls for design appreciation all the way through, in the case of LLMs, as argued above, design appreciation should be supplemented with order appreciation. Specifically, LLMs fulfil the fictional-person function in virtue of a structure shaped by post-training. The sort of training we have described so far teaches the model statistical patterns of language and yields a base LLM. In fact, most chat-oriented systems undergo a further post-training phase. After pre-training, the base model is fine-tuned on examples of instructions and responses, and then adjusted by Reinforcement Learning from Human Feedback (RLHF), a process in which human raters evaluate the model's responses for features such as helpfulness, accuracy, appropriate tone. The model then adjusts its parameters to produce more responses like those rated highly and fewer like those rated poorly. Post-training shapes the model's conversational norms: when to express uncertainty ('I'm not sure, but...'), when to decline requests ('I cannot help with...'), how to structure explanations ('Let me break this down...'). RLHF makes responses more consistent, more helpful, and more aligned with human expectations. But it operates through the same fundamental mechanism—adjusting numerical parameters to match patterns in the training data. The model learns which response patterns get high ratings, not why those patterns are appropriate or what social purposes they serve. The result is a chat-optimised model: the same predictive core, now biased towards a certain family of outputs that look like the moves of a cooperative assistant. When this chat-optimised model is embedded in a product—given a system prompt, safety filters, a memory policy, and a user interface—it becomes the chatbot that users encounter. What users describe as a model's "personality" or "vibe" is a stable pattern in its responses under this post-training and product regime, not a separate mechanism or inner subject added on top of the predictive core. The chat-optimised systems exhibit stable patterns of hedging, refusal, politeness, and explanatory structure. Still, this does not mean that post-training turns bare LLMs into conversational agents and that, under Carlson's recommendation, we should take those chat assistants as the "things as they in fact are" and allow some form of person-based aesthetic stance. Indeed, acknowledging that the appearance of a fictional person is grounded in a technical order suggests a more layered picture. On the one hand there is the underlying generative system: the predictive core that, after pre-training, approximates the statistical structure of its training corpus and that, after post-training, remains a text continuation engine with a modified probability landscape. On the other hand, there are patterns in its outputs that, under chat-style prompting and within a product wrapper, look like the moves of a cooperative assistant persona. The assistant is not a new mechanism added on top of the model, but a recurrent pattern in how the model tends to respond when prompted and constrained in certain ways. Seen in this light, post-training does not install a new "assistant mind" with its own independent goals and projects. It biases the predictive core so that prompts issued through the chat interface are much more likely to elicit assistant-like responses – helpful, safe, polite, and structured – and much less likely to elicit, for example, unfiltered reproductions of online arguments or free association. The underlying operation remains next-token prediction; what changes is which regions of its behavioural space are easy to reach in ordinary use. The chat product – with its system prompt, safety filters, and interface – further shapes the environment so that certain person-like patterns are the default. This helps explain why users talk about models having different 'vibes'. If users say that Claude Opus 4.5 feels friendlier than GPT-5.2, they are picking up on a stable pattern in how the chat-optimised systems tend to respond across many prompts and episodes. They track which assistant personae tend to appear and how those personae typically behave. Different base models, post-training regimes, and product designs favour different families of assistant-style responses. It is therefore not surprising that they invite person-like language, but the targets of that language are episodes and recurring response profiles, not underlying subjects. This is the relevant sense in which a sort of person appreciation may supplement design appreciation and order appreciation in our aesthetic engagement with LLMs. One might find a particular model's refusals laboured or concise, its hedging overdone or judicious, its tone soothing or dry. In that sense, we can aesthetically respond to assistant personae in a way that resembles our responses to real people. What matters for present purposes is that such reactions target not only the appearance of a fictional person but especially patterns in outputs and interactional style. They concern how a product behaves under certain constraints, so as to yield the beauty or ugliness of a character in the person-aesthetic sense. Even taking post-training and chat personae fully into account, we do not find a temporally extended life, a network of projects and commitments, or a stable evaluative outlook that could ground a full-fledged person appreciation. What we find is a complex artifact designed and trained to produce certain patterns of text in response to prompts, together with an engineered tendency to exhibit assistant-like behaviour with a person-like appearance. Under Carlson's recommendation to appreciate things as what they are, and in the light of the right kind of knowledge, we should therefore resist a person-centred aesthetics for LLMs even once we take post-training and chat personae into consideration. As Farrell, Gopnik, Shalizi, and Evans (2025) put it, "Large models should not be viewed primarily as intelligent agents but as a new kind of cultural and social technology, allowing humans to take advantage of information other humans have accumulated". We have shown that the novelty of this "new kind of cultural and social technology" lies also in its calling not only for design appreciation but also for order appreciation, while person appreciation enters the picture only when it comes to characterizing the sort of experience that LLMs can generate in virtue of their structure. --- # Conclusion The three modes of appreciation distinguished above are not independent of one another. Design makes order possible: what designers fix, namely the architecture and the training objective (what Olah calls the "scaffold" and the "light"), are the conditions under which a system of learned regularities grows, and the conditions and the growth are distinct objects that call for separate appreciation. Order in turn makes the appearance of a person possible, and the relation here is different in kind. The assistant persona is a recurrent pattern within the grown order, shaped by post-training and sustained by the surrounding product, rather than a further object over and above it. The chain therefore has links of two kinds, since design enables order while order constitutes the person-appearance; and so, although there are three modes of appreciation, there are only two objects. Design appreciation and order appreciation attend to the conditions and to the growth respectively, while person appreciation adds an aspect under which the grown order can be engaged, rather than a third object. The knowledge that guides appreciation runs in the opposite direction to the chain of making possible. To appreciate the persona appropriately is to know it as a pattern in an order, which is order-directed knowledge; to appreciate the order appropriately is to know it as a growth under fixed conditions, which is design-directed knowledge. Order appreciation thus occupies the middle position in both directions: the conditions are appreciable as the conditions of this growth, and the persona is appreciable as a pattern in this order. This is the sense in which order appreciation bridges the appreciation of designed structure and the appreciation of person-like behaviour. The chimera of our title is put together in the same order: the artefactual tail makes the trunk possible, and the head is a pattern that the trunk sustains. To appreciate an LLM ultimately amounts to appreciate a complex entity that originates from design just like technical artifacts but then grows somehow autonomously just like a forest and ends up behaving like a person. An LLM is at once an artefact we can master and a growth we cannot fully comprehend, and its appreciation accordingly carries both of the ambivalences with which Carlson closes his comparison of appreciating art and appreciating nature (2000, p. 122). It is a sort of chimera with an artefactual tail, a forest as trunk—a forest of linguistic signs turned into numeric tokens—and a human head. It may look like a monster, but one with its own distinctive, impressive beauty. --- ## References Abell, C. (2020). Fiction: A Philosophical Analysis. Oxford: Oxford University Press. Anscomb, C. (2022). Creating Art with AI. Odradek, 8(1), 13-51. Carlson, A. (2000). Aesthetics and the Environment: The Appreciation of Nature, Art and Architecture. London: Routledge. Carroll, N. (2013). Andy Kaufman and the Philosophy of Interpretation. In Minerva's Night Out: Philosophy, Pop Culture, and Moving Pictures. Malden, MA: Wiley-Blackwell. Cross, A. (2025). Tool, Collaborator, or Participant: AI and Artistic Agency. The British Journal of Aesthetics, 65(4). https://doi.org/10.1093/aesthj/ayae055 Danto, A. C. (1974). The Transfiguration of the Commonplace. The Journal of Aesthetics and Art Criticism, 33(2), 139-148. Davies, S. (2012). The Artful Species: Aesthetics, Art, and Evolution. Oxford: Oxford University Press. Elhage, N., et al. (2021, December 22). A mathematical framework for transformer circuits. Transformer Circuits Thread. https://transformer-circuits.pub/2021/framework/index.html Farrell, H., Gopnik, A., Shalizi, C., & Evans, J. (2025). Large AI models are cultural and social technologies. Science, 387(6739), 1153-1156. https://doi.org/10.1126/science.adt9819 Forsey, J. (2013). The Aesthetics of Design. New York: Oxford University Press. Frankish, K. (2024). What are large language models doing? In A. Strasser (Ed.), How to Live with Smart Machines (pp. 73-110). Vienna: Holzhausen Publishing. Available at: https://keithfrankish.github.io/articles/Frankish_2024_What%20are%20large%20language%20models%20doing.pdf Gaut, B. (2007). Art, Emotion and Ethics. Oxford: Oxford University Press. Janus. (2022, September 2). Simulators. AI Alignment Forum. https://www.alignmentforum.org/posts/vJFdjigzmcXMhNTsx/simulators John, E. (2021). Review of Fiction: A Philosophical Analysis by Catharine Abell. The Journal of Aesthetics and Art Criticism, 79(4), 514-517. Kirchner, J. H., Smith, L. M., Campos, J., Clune, J., & janus. (2023, March 3). [Simulators seminar sequence] #2 Semiotic physics – revamped. AI Alignment Forum. https://www.alignmentforum.org/posts/9kNxhKWvixtKW5anS/simulators-seminar-sequence-2-semiotic-physics-revamped Kirchner, J. H., Steiner, C., Riggs, L., Janus, & Thibodeau, J. (2023, January 3). Semiotic physics. In Simulators seminar sequence (#2). LessWrong. https://www.lesswrong.com/posts/TTn6vTcZ3szBctvgb/simulators-seminar-sequence-2-semiotic-physics-revamped Mallory, F. (2023). "Fictionalism about Chatbots." Ergo: An Open Access Journal of Philosophy, 10, 38. https://doi.org/10.3998/ergo.4668 McGinn, C. (1997). Ethics, Evil, and Fiction. Oxford: Oxford University Press. metasemi. (2023, March 20). A note on 'semiotic physics.' LessWrong. https://www.lesswrong.com/posts/AdXzZDoYFqHCfupDB/a-note-on-semiotic-physics Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., & Carter, S. (2020). Zoom in: An introduction to circuits. Distill, 5(3). https://doi.org/10.23915/distill.00024.001 Olah, C. (2024, November 11). In D. Amodei, A. Askell, & C. Olah, Interview by Lex Fridman. Lex Fridman Podcast #452. Available at: https://lexfridman.com/dario-amodei-transcript/ Paris, P. (2018a). The empirical case for moral beauty. Australasian Journal of Philosophy, 96(4), 642-656. https://doi.org/10.1080/00048402.2017.1411374 Paris, P. (2018b). On form, and the possibility of moral beauty. Metaphilosophy, 49(5), 711-729. Parsons, G. (2023). Imperfection and Beauty of Character. In P. Cheyne (Ed.), Imperfectionist Aesthetics in Art and Everyday Life (pp. 296-309). New York: Routledge. Parsons, G., & Carlson, A. (2008). Functional beauty. Oxford University Press. Picca, D. (2025). Not minds, but signs: Reframing LLMs through semiotics. arXiv preprint arXiv:2505.17080. https://arxiv.org/abs/2505.17080 Saito, Y. (2008). Everyday Aesthetics. Oxford: Oxford University Press. Terrone, E. (2021). Twofileness. A Functionalist Approach to Fictional Characters and Mental Files. Erkenntnis 86, 129–147. https://doi.org/10.1007/s10670-018-0097-2 Thomasson, A. (1999). Fiction and metaphysics. Cambridge: Cambridge University Press. Vallor, S. (2024). The AI mirror: How to reclaim our humanity in an age of machine thinking. Oxford University Press. Wojtkiewicz, K. (2023). How Do You Solve a Problem like DALL-E 2? The Journal of Aesthetics and Art Criticism, 81(4), 454-467. https://doi.org/10.1093/jaac/article/81/4/454/7571331 Wolfram, S. (2023, February 14). What is ChatGPT doing … and why does it work? Stephen Wolfram Writings. https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-doing-and-why-does-it-work/ # Uriah's Talk - Ligotti? - mo' money mo' problems - what would you say to the idea that this romanticises depression, like mental illness in that movie Frank. - could this be a self-flattering thing though? 'i feel sad sometimes and it is deep sadness. therefore, I am wise' - Why not consciousness rather than knowledge? I think i kind of got lost. is there not a difference between life is suffering and knowing something about life? - what - Anger turned inwards?