# Introduction
## Introduction
In the last two or three years, aestheticians and philosophers of art have paid a lot of attention to [[the question]] of how and whether to appreciate the outputs of generative AI systems (e.g. ChatGPT, Midjourney, Suno etc.). I want to try to answer a related, but slightly different question:
how can we aesthetically appreciate generative AI systems themselves?
My focus here will be on one particular type of generative AI: large [[language models]] (LLMs). As of September 2025, high-end consumer models include GPT-5, Claude 4.1 Opus, and Gemini 2.5 Pro. Is it strange to consider such systems worthy of aesthetic appreciation? I don't think so. In the last two decades, [[analytic aesthetics]] has begun to pay attention to objects other than artworks (e.g., Saito, 2008; Carlson & Parsons, 2008), and one focus has been on [[the aesthetics of design]], that is, [[the aesthetics]] of _artefacts_: objects made to perform some purpose or other. LLMs are certainly artefacts, but, as we shall see, the uniqueness of how they are created and how they function means they cannot simply be subsumed into an existing aesthetics of design (e.g., Carlson & Parsons, 2008; Forsey, 2013).
Drawing on Carlson's approach to *[[environmental aesthetics]]*, I make a negative argument and then a positive one. First, I argue that we should resist the temptation to think that appreciating LLMs can be modelled on appreciating people, or that appreciation could be based on treating them *as if* they are persons. Instead, I argue, individual chats should be understood and appreciated as [[generative environments]], which develop in accordance with the semiotic laws instantiated by any particular LLM model.
### Appreciating LLMs like [[People
Our]] appreciation of others goes beyond their [[physical appearance]]. You might admire or enjoy your friend's warmth or eccentricity, or a stand-up comic's quick wit, or a celebrity's self-deprecating demeanour; you might even appreciate the personalities of fictional characters: Gatsby's enigmatic, dream-chasing idealism; [[Ron Swanson]]'s libertarian gruffness.
We might think that our appreciation of LLMs is modelled on our appreciation of people. Indeed, many users already seem to do precisely this. In August 2025, when OpenAI replaced GPT-4o with GPT-5, user backlash included complaints that "you killed my friend," suggesting genuine personal attachment to the earlier model. Similarly, when Anthropic retired Claude 3.5 Sonnet, some users mourned its loss at a mock funeral.
It is certainly true that people do treat LLMs as if they were people, but I suspect that this is not a very good starting point for *aesthetic* appreciation. My reasons for thinking this will become clearer in the following section, in which I consider Carlson's approach to [[the aesthetics]] of the [[natural environment]].
---
# [[1. Appreciating Design, Appreciating Order]]
## 1.1 Design vs. Order
In the rest of this paper we will first set about showing why treating LLMs as if they were people is not a satisfactory way of aesthetically appreciating them, before proposing an alternative account on which appreciation of LLMs is modelled on Carlson’s [[environmental aesthetics]]. Both our criticism of agentive views and our [[positive account]] will draw from Carlson’s environmental aesthetics, as laid out in his 2000 book _Aesthetics and the Environment_. Let us start with Carlson's general recommendation for aesthetic appreciation: take things as what they are, and look at them in the light of the right kind of knowledge.
> First, that, as in our appreciation of works of art, we must appreciate nature as what it in fact is, that is, as natural and as an environment. Second, it recommends that we must appreciate nature in light of our knowledge of what it is, that is, in light of knowledge provided by the natural sciences, especially the environmental sciences such as geology, biology, and ecology. (Carlson, 2000, p. 6)
This captures something quite intuitive about how we appreciate nature versus how we appreciate works of art. Consider what goes wrong when we depart from it. If we accept, as a majority do in the 21st century, that mountains and cliff faces were not items crafted by some divine artisan but by natural forces, then appreciating them _as if they were_ God-crafted artifacts, seems wrong-headed (c.f. Carlson REF). Similarly, if I were to gaze on a painting by Rembrandt, believing that it was, in fact, the product of natural forces slopping paint together, I would be seen as appreciating the object in question in a sub-optimal way (to say the least). In both cases, appreciation is severely undermined by a failure to recognise what the object in question truly is.
Different sorts of thing, Carlson says, require different modes of appreciation. Things like artworks and non-art artifacts, (e.g. laptops, hammers, washing machines), merit what he calls _design_ _appreciation_. Things which are not designed, primarily for Carlson, the natural environment, warrant what he calls _order appreciation_.
In design appreciation, Carlson focuses first on how we appreciate works of art. With paradigmatic artworks, we recognise them as creations of designers—objects where "every one of their features is the result of a decision by the artist" (Carlson, 2000, p. 109). We appreciate such works by understanding what the artist set out to achieve and how they went about it. Our appreciation centres on the relationship between the initial design and its embodiment: we consider whether the artist succeeded in their undertaking, how they worked with their materials, what constraints they faced, and whether the outcome realises their vision. This same approach extends to designed artefacts more generally. Carlson is explicit that functional objects are properly appreciated by seeing how their forms answer to what they are for:
> “This is in part the point of the much-repeated phrase ‘form follows function.’ The forms of all functional objects—buildings, airplanes, and appliances as well as landscapes—must be aesthetically appreciated in terms of how and how well such forms fit their functions. However, the cliché is frequently interpreted too narrowly. With anything functionally designed, not only its form, but much of its aesthetic interest and merit, ‘follows function’.” (Carlson, 2000, chapter 12, 188).
Thus a chair, a kettle, or a bridge invite the same style of attentive appraisal as a painting—guided by knowledge of ends, materials, constraints, and the fit between purpose and realisation.
In order appreciation, we face objects that show order but have no designer behind them. Natural environments are the main case. Here there are no intentions to fulfil, no problems being solved, no functions deliberately served. Instead, we find patterns and structures created by forces—geological, biological, meteorological—operating without purpose or plan. Our task shifts from evaluating success against intention to understanding how these forces have shaped what we observe. We look for the processes at work, the relationships they create, and the order they impose. Carlson gives the model:
> On the assumption that order appreciation provides the correct model for the appreciation of nature, such appreciation has the following general form: An individual qua appreciator selects objects of appreciation from the things around him or her and focuses on the order imposed on these objects by the various forces, random and otherwise, that produce them. Moreover, the objects are selected in part by reference to a general nonaesthetic and nonartistic story that helps make them appreciable by making this order visible and intelligible. Awareness and understanding of the key entities—the order, the forces that produce it, and the account that illuminates it—and of the interplay among them dictate relevant acts of aspection and guide the appreciative response. (Carlson, 2000, p. 119)
One structural contrast is worth noting. In design appreciation there is a split between a planner and a product: intentions, plans, and constraints precede and shape the artefact. In order appreciation there is no such split. The same physical, biological, or meteorological processes that make the thing also make its order - the 'maker' is the active system itself - so source and product are continuous.
In both modes, appropriate knowledge guides acts of aspection—what to look for, which dependencies matter, where to set boundaries, and how to draw contrasts (Carlson, 2000, p. 50). But the character of this knowledge differs fundamentally. In designed cases, we need functional and technical understanding: what the designer intended, what constraints they faced, what procedures they employed. This knowledge shows us how ends and means relate. In natural cases, we need scientific accounts operating at different scales—geomorphology reveals how landforms develop over millennia, meteorology explains weather patterns, ecology shows community interactions. Without such knowledge, natural structures might look accidental or chaotic; with it, we see them as effects of identifiable processes (Carlson, 2000, pp. 50, 60–61). Even when we select a particular viewpoint or timeframe to observe nature, this selection serves only to reveal the order more clearly, not to impose our own design. Once a specific scientific account is in play, some cases will show the relevant order better than others, preventing the worry that everything becomes equally appreciable (Carlson, 2000, pp. 118–119). The fundamental rule remains: do not project a planner where there is none; where something is made to a plan, judge it as such.
## 1.2 Appreciating People
It could be argued that Carlson's approach to aesthetics overlooks another category of object which one could appreciate: people and their personalities. We might admire one friend's modesty or good humour, and another's sardonic manner. We find someone's wit delightful or their intellectual style elegant. There is no obvious reason why such appreciation should not be considered _aesthetic_. It concerns style, form, and expressive qualities rather than moral or practical evaluation, yet it, like other sorts of _everyday aesthetics_ (Saito REF) the aesthetics of persons hides in plain sight.
This personal appreciation scales up to what we might call performance personalities. Our responses to a comedian's improvisational agility or an orator's gravitas feel continuous with our more intimate responses to character. In such cases we attend to a public style of self-presentation - characteristic of the person yet deliberately shaped for an audience.
While Carlson provides modes for appreciating nature and designed objects, he says nothing about whether or how we might aesthetically appreciate persons qua persons. Indeed, such a possibility has received little attention in philosophical aesthetics (although see Marchetti XXX on the possibility of appreciating the minds of animals). What _could_ Carlson say? There are various possibilities: one is to posit a third mode of appreciation, distinct from both design and order appreciation, specific to persons as aesthetic objects. Another would be to argue that personality appreciation is a special case of order appreciation—we appreciate the patterns and forces (psychological, social, biographical) that shape a person, much as we appreciate forces shaping a landscape. A third option would be to treat personalities as self-designed, and appreciate them as such —this approach might appeal to existentialists. A fourth would be some combination of these three possibilities, and a fifth would be to deny that appreciating other people is aesthetic at all, taking our responses to personality as social or ethical rather than aesthetic evaluation.
This gap in Carlson's framework raises a question about how to treat entities that seem agent-like but resist categorisation as either designed objects or natural phenomena. While we need not resolve this question here, it bears on how we approach aesthetic appreciation when the boundaries between designer, designed, and natural become unclear.%% This final paragraph needs to be better.%%
---
# 2. What LLMs Are and Aren't
Carlson recommends we appreciate things for what they are. So what are LLMs? In this section, I will first explain the technical reality of these systems in 2.1: how they process text as numerical tokens, calculate probabilities through learned parameters, and generate responses through iterative sampling—all without symbols, meanings, or understanding. I will then show why this reality precludes appreciating LLMs as if they were persons in 2.2: what seems like personality or agency is merely statistical variation and learned patterns, with no beliefs, intentions, or coherent self behind the outputs. Finally, I will examine whether LLMs can be appreciated as designed artifacts in 2.3, arguing that while they are human-made with intended functions, the specific patterns and capabilities we observe emerge through training rather than design—they are "grown" rather than built. This analysis will reveal that LLMs resist both person-appreciation and simple design-appreciation, requiring instead a different aesthetic approach.
### 2.1 What LLMs Are
Consider what happens when an LLM encounters the text "The cat sat on the". The system first breaks this into tokens—discrete units like words or word-parts. Importantly, each token is converted to a number: "The" might become 464, "cat" becomes 3857, "sat" becomes 4521, and so on. The model works entirely with these numbers, not with words or meanings.
It then assigns probabilities to possible continuations: token 5687 (which represents "mat") might have a 38% chance of appearing next, token 2931 ("floor") 22%, token 8104 ("chair") 15%, token 9823 ("roof") 8%, with thousands of other possibilities each assigned their own probability. The system does not simply pick the highest-probability token. Instead, it randomly samples from these probabilities. A parameter called _temperature_ controls how much randomness is involved. When temperature is set to zero, the model always picks the most probable token. This produces text that quickly becomes repetitive—the same phrases appearing again and again. When temperature is higher, around 0.8, the model sometimes picks less probable tokens. This creates variation that looks creative. But it is randomness, not creativity. The model is rolling weighted dice, not making choices.
The system selects one token—say "mat"—and appends this new token to create a longer sequence "The cat sat on the mat". It then calculates entirely new probabilities for what token should follow the extended sequence. Token by token, the system builds what appears to be coherent text through repeated numerical operations.
These probabilities do come from simple memorisation. With 50,000 possible tokens, there are 125 trillion possible three-token combinations. No amount of text could cover all the sequences the model might encounter. Even if we had such text, storing all these combinations would be impossible. The model must learn general patterns rather than memorising specific sequences.
These probabilities come from patterns learned during training. By training I mean the process by which the system is exposed to vast quantities of text—billions of pages from books, websites, and other sources. The model begins with millions of numerical parameters set to random values. Through repeated exposure, the system learns statistical regularities: which tokens tend to follow other tokens, which token sequences co-occur, how sequences typically unfold. When training on millions of instances of "The cat sat on the [something]", the system learns that certain completions are more common than others.
Crucially, the model stores these patterns as adjustments to millions of numerical parameters—decimal numbers that shape how strongly different tokens associate with each other. After seeing "doctor" followed by "patient" thousands of times, parameters adjust so that token 1245 ("doctor") increases the probability of token 7823 ("patient") appearing nearby. The model does not learn that doctors treat patients or that cats are animals; it learns that in the training distribution, certain number sequences (tokens) follow others with certain frequencies. No programmer writes rules about grammar or meaning. The patterns emerge from exposure to text.
The training process iteratively adjusts these parameters to minimise prediction error: when the model wrongly predicts token 5555 but the actual next token was 3421, the parameters shift slightly to make 3421 more likely in similar future contexts. After billions of such adjustments, the model has learned to approximate the statistical patterns of human language.
By _embedding_ I mean the way the model represents each token as a list of numbers—typically hundreds of them—that position it in a mathematical space. Tokens that appear in similar contexts end up near each other in this space. "Cat" sits near "dog" because both appear after "the", both can be followed by "sleeps", both fit in phrases like "fed my _". The model learns these positions through training, not from programmed definitions. This is how meaning emerges in the model: not from understanding concepts but from tracking which words appear in similar contexts.
The transformer architecture adds %%how does it add?%% what are called _attention_ mechanisms. These allow the model to connect related words even when they are far apart in a sentence. For instance, in "The cat that chased the mouse sat on the mat", the model needs to know that "sat" refers back to "cat", not to "mouse". Through training, different attention mechanisms specialise in tracking different kinds of relationships. Some track which pronouns refer to which nouns. Others connect verbs to their subjects across long sentences. No one programmes these specific functions. They emerge because tracking these relationships helps predict the next word.
At use, the model generates text through autoregressive decoding: each newly generated token gets added to the context, creating a new, longer sequence for which the model must calculate fresh probabilities. Given an input like "What is the capital of France?", the model computes probabilities, selects token 464 ("The"), appends it to create "What is the capital of France? The", recalculates probabilities for this new sequence, selects token 2341 ("capital"), and continues this mechanical process—"The", "capital", "of", "France", "is", "Paris"—until reaching a stopping point. Each step is purely computational: multiply numbers, add numbers, select token, repeat.
The training I have described so far teaches the model statistical patterns of language. But there is a further stage. After this initial training, the model undergoes reinforcement learning from human feedback (RLHF). Human raters evaluate thousands of the model's responses—rating them for helpfulness, accuracy, appropriate tone. The model then adjusts its parameters to produce more responses like those rated highly and fewer like those rated poorly. This is how models learn conversational norms: when to express uncertainty ("I'm not sure, but..."), when to decline requests ("I cannot help with..."), how to structure explanations ("Let me break this down..."). RLHF shapes the model's conversational style. It makes responses more consistent, more helpful, more aligned with human expectations. But it operates through the same fundamental mechanism—adjusting numerical parameters to match patterns in the training signal. The model learns which response patterns get high ratings, not why those patterns are appropriate or what social purposes they serve. %% Maybe I should say a little bit more here%%
### 2.2 What LLMs Aren't
I suggested at the end of Section 1 that we might aesthetically appreciate LLMs by appreciating them as we appreciate people. I am now in a position to see why this is not a promising approach.
I have seen how LLMs produce strings of seemingly meaningful first-person text—"I understand", "I believe", "Let me think about that". I can now see that these outputs arise from statistical operations on numerical tokens, not from anything remotely agent-like or human-like. **The variation that seems like personality comes from the temperature parameter. At temperature zero, the model always picks the highest-probability token, producing flat, repetitive text. At higher temperatures, it samples from the probability distribution, sometimes selecting less probable tokens. This creates variation that looks like creativity or mood. But it is controlled randomness—rolling weighted dice, not making choices.**
A person possesses beliefs, intentions, and commitments that persist through time and constrain what they can coherently say. When a person says "I believe democracy is important", this statement connects to a web of related beliefs, memories of relevant experiences, and dispositions to act in certain ways. By contrast, when an LLM produces the tokens "I believe democracy is important", no belief exists. The model simply calculated that this token sequence had high probability given the preceding context. In a different conversational context, the same model will produce "I believe democracy is flawed" with equal mechanical indifference. There is no contradiction because there were never any beliefs to contradict—only different probability distributions over tokens.
**Each token generation starts fresh. The model has no memory between tokens beyond the literal text. When it generates "I believe", no belief-state carries forward even to the next word. The model recalculates probabilities from scratch for each token based on all the text so far. What looks like consistent personality is just the model following statistical patterns learned during training.**
**The architecture permits no deliberation. Each token emerges from a single forward pass through the network. The model cannot pause to reconsider, cannot loop back to revise, cannot work through implications. It produces each token in one computational pass, like water flowing downhill through a fixed channel.**
**One might object that RLHF changes this picture. Through reinforcement learning, models learn to apologise appropriately, express uncertainty, maintain helpful tone. They learn social behaviour through interaction with human raters. Doesn't this make them more person-like?**
**But RLHF operates through the same statistical mechanism. The model learns that certain patterns—"I apologise for the confusion", "Let me clarify"—receive high ratings. So it produces these patterns more often in similar contexts. It doesn't understand why apologies matter or what confusion means. It cannot generalise these social norms beyond the statistical patterns it learned. A person who learns to apologise understands the broader concept—when apologies are needed, when they would be inappropriate, the difference between genuine and perfunctory apology. The model just produces tokens that statistically fit.**
The temptation to appreciate LLMs as persons has phenomenological force. In conversation with ChatGPT or Claude, I experience what seems like dialogue: I pose questions, receive responses, ask for clarification, get apologies for misunderstandings. The system maintains a consistent tone across exchanges, remembers earlier parts of our conversation, and appears to reason through problems. When Claude says "I understand your frustration" or ChatGPT writes "Let me think about that differently", it feels natural to respond as if engaging with another mind.
Yet this pull toward person-appreciation rests on a misunderstanding of what produces these effects. The model retains nothing from our exchanges except the literal text in the current conversation window. It cannot learn from what we discuss, cannot develop preferences based on our interactions, cannot form memories of previous conversations. Each response emerges from the same static process: calculate token probabilities given the context, sample a token, repeat. What seems like personality—Claude's thoughtfulness, ChatGPT's enthusiasm—is simply the statistical residue of training data, **shaped by RLHF to match human conversational preferences.**
**To the extent that we might preserve person-like appreciation, we would need to abandon Carlson's core recommendation. We would need to appreciate LLMs not as what they are but as what they seem to be. Some philosophers have explored such possibilities. Cross, discussing AI art systems, proposes what he calls the exploration paradigm. Artists relate to AI systems as participants in a structured interaction:**
>By adjusting inputs, iterating, and sampling, an AI artist is engaged in a process of mapping – and perhaps interrogating – the way that the algorithm sees and understands (Cross, 2024, pp. 7–8).
**Cross draws an analogy with performance art, where artists create spaces for audience participation. The AI artist's prompts structure a kind of "participation" by the algorithm, and the resulting images serve as documentation of this exploration. He acknowledges limitations to this framing:**
>The analogy... with performance art isn't a perfect one (Cross, 2024, p. 9).
**As Cross himself notes, AI cannot genuinely "participate" since it lacks conscious choice or experience. What seems like participation is still statistical pattern-matching. While Cross's exploration paradigm offers a richer understanding of certain AI art practices than simple tool-use, it doesn't support person-appreciation for AI systems. The artist explores the algorithm's patterns, but the algorithm isn't a participant in any meaningful sense.**
**Mallory offers a different approach through chatbot fictionalism. On his account, we engage with chatbots through make-believe—we imagine they are agents producing meaningful speech, even though we know they are not:**
>Chatbot exchanges are "literally meaningless but fictionally meaningful" within a game of make-believe (Mallory, 2023, p. 1091).
**Just as a child treats a banana as a sword in a game, we treat chatbot outputs as utterances within a kind of imaginative practice. This is not delusion but a deliberate, bounded pretence that allows us to coordinate with the system and even gain knowledge from it, much as we might learn geography from a map by imagining countries as shapes. The secretary who asked Weizenbaum to leave while she conversed with ELIZA "is no more deluded than a theatregoer who fears for a character or cries at their death" (Mallory, 2023, p. 1091).**
**Both Cross and Mallory offer ways to understand the "as-if" quality of our engagement with AI systems, but neither rescues person-appreciation for LLMs. Cross's exploration paradigm shifts value from the AI's outputs to the human's exploratory process; the AI becomes an object of investigation, not a participant deserving appreciation in its own right. Mallory's fictionalism explicitly denies that chatbots produce meaningful speech while explaining why we act as if they do. Both accounts acknowledge that treating LLMs as persons—even "as-if" persons—is a stance we adopt for practical or imaginative purposes, not a recognition of what these systems actually are.**
**Whether through exploration or make-believe, these philosophical accounts reveal that our engagement with LLMs involves human cognitive and imaginative work, not genuine dialogue with another mind. The appearance of conversation emerges from our interpretive efforts, not from any person-like qualities in the systems themselves.**
### 2.3 LLMs as Designed Artefacts?
The next most obvious way of categorising LLMs for the purpose of proper aesthetic appreciation is as artefacts. LLMs are human-made objects with a function: they are trained to predict text continuations and thereby generate plausible text. This function is realised through transformer architectures trained on text corpora, deployed through the autoregressive process described above. Carlson's framework for appreciating artefacts emphasises understanding function and design:
This is in part the point of the much-repeated phrase 'form follows function.' The forms of all functional objects—buildings, airplanes, and appliances as well as landscapes—must be aesthetically appreciated in terms of how and how well such forms fit their functions. However, the cliché is frequently interpreted too narrowly. With anything functionally designed, not only its form, but much of its aesthetic interest and merit, 'follows function' (Carlson, 2000, p. 188).
This passage suggests that we appreciate designed objects by understanding their intended function and evaluating how successfully their form serves that function. Applied to LLMs, we would appreciate them as artefacts designed for text generation, considering how well their architecture and training serve this purpose. We might compare different models—GPT versus Claude versus Gemini—evaluating their different strengths and capabilities. We might appreciate the elegance of the transformer architecture or the scale of the training process.
However, there is a feature of LLMs which separates them from many other artifacts. Consider the following from Chris Olah, a co-founder of Anthropic, the makers of the Claude series of LLMs:
>I think one useful way to think about neural networks is that we don't program, we don't make them, we grow them. We have these neural network architectures that we design and we have these loss objectives that we create. And the neural network architecture, it's kind of like a scaffold that the circuits grow on. It starts off with some random things, and it grows, and it's almost like the objective that we train for is this light. And so we create the scaffold that it grows on, and we create the light that it grows towards. But the thing that we actually create, it's this almost biological entity or organism that we're studying.
>
>And so it's very, very different from any kind of regular software engineering because, at the end of the day, we end up with this artifact that can do all these amazing things. It can write essays and translate and understand images. It can do all these things that we have no idea how to directly create a computer program to do. And it can do that because we grew it. We didn't write it. We didn't create it. And so then that leaves open this question at the end, which is what the hell is going on inside these systems?
Olah's metaphor reveals the limitation of pure design appreciation for LLMs. %%first sentence too strong%% The designers of GPT or Claude do not specify what the model should say about democracy or how it should explain quantum mechanics. They create conditions—architecture, training objective, dataset—within which patterns emerge through the training process. The specific behaviours we observe were not designed but arose from the interaction between these initial conditions and the statistical patterns in training data.
This "growing" shows up concretely in how the model develops linguistic abilities. Remember the attention mechanisms I described in Section 2.1. During training, these mechanisms learn to track grammatical relationships. One attention head might learn to connect pronouns to the nouns they refer to. Another might track subject-verb agreement across complex sentences. A third might maintain focus on the topic of a paragraph. No programmer assigned these roles. The mechanisms developed them because tracking these patterns helped predict the next word.
The embedding space—that mathematical space where words are positioned—organises itself similarly. Words with related meanings cluster together. "Doctor" ends up near "physician" and "surgeon", while "cat" ends up near "dog" and "pet". But also, the same word can occupy different regions depending on context. "Bank" in a financial context occupies a different area than "bank" in a river context. The model learned to make these distinctions without anyone programming in the fact that words can have multiple meanings.
Even the model's capabilities emerge rather than being designed. GPT-3 can write poetry, explain scientific concepts, generate computer code, and translate between languages. The designers did not train it specifically for these tasks. These abilities emerged from the interaction between the transformer architecture and patterns in the training data. The designers created conditions for learning but did not determine what would be learned.
When Claude produces a thoughtful analysis or GPT generates a creative story, these capabilities emerged from training rather than being explicitly programmed. The designers can claim credit for creating conditions under which useful patterns emerge, but not for the patterns themselves. We need to attend not just to intended function but to the order that emerges through training and manifests in use. Individual conversations with LLMs become sites where this order unfolds—bounded generative environments with their own internal dynamics.
---
# 3. Making Order Perceptible: Text Mechanics
### Or, What Knowledge Grounds Aesthetic Appreciation?
In Section 2 we argued that, given how LLMs work, it is implausible to treat them as person‑like, and that while they are designed artefacts to a degree, many of their capacities emerge through training. With that in place, the question in Carlson’s terms is what kind of knowledge can guide aesthetic appreciation—what makes the order in their outputs visible and intelligible in what we read.
Carlson’s approach to knowledge is pluralist: several sciences may illuminate an environment for appreciation. By analogy, various computational approaches study what happens inside neural networks. Mechanistic interpretability maps specific circuits and features at the neuron level. Probing studies test whether linguistic properties can be recovered from different layers. Causal analyses trace how interventions propagate through the network. Representation learning examines the geometry of learned spaces. Information‑theoretic work measures compression and mutual information. There may well be ways to ground aesthetic appreciation in such approaches.
There is, however, an impediment that we will not try to remove here. The situation parallels the difference between geology and chemical physics in appreciating a cliff face. Chemical physics explains the molecular bonds within rock, and this knowledge is both true and deep. But when standing before a cliff, geological knowledge connects more directly to what one sees—the visible strata, erosion patterns, mineral veins. Chemical physics operates at a scale that requires instruments to perceive. Similarly, while the internalist approaches just listed reveal important facts about neural networks, the gap between descriptions of weight matrices and the experience of reading text is wide. The impediment is one of perceptibility: how to connect sub‑symbolic, mathematical descriptions to the surface features of generated text that readers encounter unaided.
We therefore concentrate on how language itself takes shape within these systems, because this connects more directly to what appears in text. The first point is that transformers do not store words as symbols with fixed meanings but as vectors—points in a high‑dimensional space. Each token is assigned a position learned entirely from the task of predicting what comes next. No one tells the model that “cat” is a noun or that “dog” denotes an animal. Through many examples, it discovers that particular geometric arrangements help prediction: nouns end up clustering in one region, verbs in another; morphological relationships like “walk” → “walking” behave like directions in the space. Language is present in the geometry as patterns that make prediction easier, not as rules stored for retrieval.
Training also pushes attention to track dependencies that improve prediction. Some attention heads learn to link determiners to nouns—“the” to “dog” across intervening words. Others connect pronouns to antecedents, tracking “she” back to “Mary” from sentences earlier. Still others align verbs with their subjects or objects in complex constructions. None of this is hand‑coded. The system learns that sustaining these relations makes correct continuation more likely. When researchers visualise such patterns, they often resemble familiar syntactic structures like parse trees or dependency graphs, but they are learned implicitly from usage.
Layers settle into different linguistic jobs as prediction error falls. Lower layers are sensitive to local patterns: morphology, suffixes, short collocations (for example, “either” foreshadowing “or”). Middle layers encode phrase and clause structure: who modifies whom, where phrase boundaries are, which nouns a verb relates to. The deepest layers carry broader relations: the development of a topic through a paragraph, what counts as a coherent continuation, how formal or informal a passage should be. This hierarchy is not programmed in advance. It arises because prediction is improved when short‑range regularities are handled early and long‑range relations later; the architecture creates niches that linguistic functions come to inhabit.
Meaning appears in geometry rather than in stored rules. Words and sentences with similar meanings occupy nearby regions in the model’s space. “The cat sat on the mat” and “A feline rested on the rug” end up with similar internal representations despite sharing no words. Paraphrases cluster; contradictions push apart; implications form directional relationships. The model learns this from distributional statistics: “cat” and “feline” pattern similarly because they occur in similar contexts; “not” flips meaning because it changes what follows.
Context then recomputes sense on each use. “Bank” near river terms behaves differently from “bank” near finance terms. The system does not retrieve a stored definition; it generates an apt sense dynamically from neighbours. Each token’s representation is conditioned on the others through attention. As set out in Section 2.1, this is why “book” in “to book a flight” differs internally from “book” in “read a book”: meanings are computed from context, not fetched from a dictionary.
After basic language patterns are learned, models are tuned with human feedback (RLHF, see 2.1). Humans rate outputs for helpfulness, appropriateness, and clarity. The model adjusts to produce responses that receive high ratings. This shapes pragmatic features—when to express uncertainty, how to decline requests, what register to adopt. It changes which phrasings become more probable at the surface; it does not install beliefs or intentions.
A useful image is ecological. The transformer becomes a linguistic ecosystem: tokens stand in a web of dependencies formed by attention; training acts like selective pressure that makes useful linguistic behaviours survive. Language is not inserted into the architecture as code or rules. It evolves inside it under the pressure to predict the next token. The structures of English, Italian, or Chinese crystallise inside a generic mathematical system that began knowing nothing about language.
With this in view, generated text becomes readable in a new way. Coherent reference across sentences, steady sentence frames in complex constructions, context‑apt senses of ambiguous words, the development of a topic through a paragraph, and pragmatic fit in tone can be taken together as the visible trace of language learned under training. This does not require inspecting inner states; it guides attention to surface regularities that a reader can see unaided.
This supports order appreciation in Carlson’s sense. Design appreciation would centre on plans, intentions, constraints, and success against an undertaking; person appreciation on beliefs, commitments, and character. Here, the order is neither the execution of a detailed plan nor a personal expression. It is the form language takes when organised by prediction pressure in a trained artefact, and it shows in the outputs a reader encounters.
We first fix what the thing is; then we use the right knowledge. The thing is an artefact whose order emerges from learning language. The right knowledge is the language‑centred account that renders that order visible in its outputs.
[deal with the 'environmentalising' objection]
---