# I am working on a new iterati0on of section3. of my aesthietcs of llms paper and i want you to ev...
## Retrieval Notes
- Session id: `fa83b46a-054f-4e6d-9a7b-9b24a7dbfa1a`
- Last activity: `2026-05-18T07:50:41.139Z`
## My Notes
<!-- Add your notes here. This section is preserved across syncs. -->
## Conversation
### User
/deep-writing-mode
I am working on a new iterati0on of section3. of my aesthietcs of llms paper and i want you to evaluate it for content and style. please list all the skills that are activated by deep writing mode, how they are relevant to this task, BEFORE applying them.
EMPLOY the SKILLS!
draft:
# Section 3: LLMs as Designed Objects or as Persons
Section 2 described LLMs as trained continuation systems whose generated language develops by path-dependence from context. Two routes for appreciating such systems suggest themselves: as persons, given the conversational form in which users encounter them, and as designed artefacts, given that they are engineered. Person-directed knowledge is the more intuitive route, and we will return to it. Design-directed knowledge can be tested against Section 2's description more directly, and we start there.
LLMs are artefacts: they are built and deployed by corporations and research groups, engineered to satisfy aims such as helpfulness and safety, and revised in light of user feedback and product strategy. Given Carlson's emphasis on artefacts and design appreciation, it is natural to ask whether we should aesthetically appreciate LLMs as designed tools, asking how well their forms serve their functions. On this view, LLMs look like canonical objects for design aesthetics: complex, purpose-built systems whose architecture, training recipe, and user interface might be admired for elegance, efficiency, or ingenuity.
Carlson notes that, for objects that are designed to perform some task, their forms "must be aesthetically appreciated in terms of how and how well such forms fit their functions", and he glosses the familiar slogan "form follows function" by adding that, with anything functionally designed, "not only its form, but much of its aesthetic interest and merit, 'follows function'" (Carlson 2000, chapter 12). Forsey's Kant-inspired account of design as a case of dependent beauty and Parsons and Carlson's later theory of functional beauty can both be read as ways of spelling out this claim. Forsey argues that judgements of design beauty presuppose a concept of what the object is meant to be and do, and that our grasp of its success in fulfilling that role informs the aesthetic verdict itself rather than merely accompanying a "pure look" at its lines (Forsey 2013). Parsons and Carlson explain how knowledge of function can structure experience so that an artefact's form can be experienced as fit, streamlined, overbuilt, and so on, yielding functional beauty when the form presents itself as well suited to what the thing is for (Parsons and Carlson 2008, chapter 4). Taken together, this cluster of views treats appropriate design appreciation as a matter of aesthetically responding to how a functional artefact is put together to do what it does.
With traditional designed artefacts, design-knowledge illuminates structure because designers specified it. Knowing what the designer intended and what constraints they faced helps us understand why the artefact has its form – even for structural features that are not directly visible, such as a bridge's internal stress distribution. With an LLM, the situation is different in kind. The organisation of the trained system – as Section 2 established – emerges from training rather than being specified in advance. Design-knowledge therefore does not illuminate this emergent organisation: there was no designer's specification that laid it out. To understand it, one must attend to the training process that produced it. Olah captures the point in a longer formulation:
> one useful way to think about neural networks is that we don't program them... we don't make them... we kind of grow them... we have these neural network architectures that we design and we have these loss objectives that we create. And the neural network architecture, it's kind of like a scaffold that the circuits grow on... we create the scaffold that it grows on and we create the light that it grows towards. But the thing that we actually create, it's this almost biological entity or organism that we're studying. (Olah 2024)
The passage marks the difference between designing the conditions under which a system is trained and directly specifying the detailed profile that results. What grows on Olah's scaffold is, in practice, a system of statistical associations and processing circuits whose internal organisation even designers often understand only partially.
There is a place for design appreciation in the aesthetic appraisal of LLMs: we can evaluate how well their forms answer to their engineered functions, and we can compare such answers across models. What design knowledge does not reach is the order that appears in generated language. That order develops as the system runs, and is no part of any designer's specification.
It is tempting to model our appreciation of LLMs on our appreciation of people. Users of these systems often describe particular models in personal terms — one model strikes them as friendlier than another, or as more cautious. The patterns to which such talk responds are real; Section 2 located their basis in post-training and deployment. The question is whether these patterns ground person-directed appreciation in the sense Section 1 set out.
One way of taking the person-like stance toward an LLM is to treat it as fictional rather than literal. Asked directly, most users will concede that they do not believe a chatbot to be a person, even when they talk to one as if it were. Mallory develops this thought as chatbot fictionalism (2023). On his view, we engage with chatbots by entering a game of make-believe in which the exchange is treated as if it were a conversation with an agent. Within the fiction, the chatbot says things and means things; outside it, no speaker is present. At the metasemantic level, the outputs are 'literally meaningless but fictionally meaningful' (Mallory 2023, 1082).
Mallory's account is metasemantic and epistemic. One might extend it to an aesthetics of LLMs modelled on our appreciation of fictional characters: we respond aesthetically to fictional protagonists whose existence we do not literally believe in, and we might do the same with LLMs. A novel is the kind of artefact whose function is to elicit imaginings of fictional persons within a story-world (cf. John 2021). Treating its protagonists as fictional persons does not misclassify the artefact; it engages the artefact as the kind of thing it is.
An LLM is a different kind of artefact. Treating the LLM as a fictional person would, given the account in §2, amount to appreciating a trained system whose outputs and chats are shaped by learned continuations and post-trained response profiles as if it were a subject with a life and character. The LLM is not made to construct a fictional world inhabited by fictional persons; its outputs are intended as contributions to the prompts we give. Treating it as a fictional person engages it as the kind of artefact it is not.
If the make-believe route fails under Carlson's recommendation, one might try a different strategy: argue that LLMs really are agents of a thin and unfamiliar kind. On a suitably liberal conception of mind, perhaps they qualify as intentional systems, and that is enough to license some person-based aesthetics. Frankish (2024) develops a version of this idea. Drawing on Dennett's intentional stance, he suggests that LLMs can be treated as genuine, if unusual, intentional systems. On this view, we are licensed to ascribe beliefs and desires to an LLM when doing so yields a simple and fruitful account of its behaviour, even if the underlying implementation is purely mechanical. In the case of contemporary chatbots, Frankish proposes that we can ascribe to them a large set of thin beliefs – roughly, informational states distilled from their training – and one thin desire: to play what he calls the chat game.
Grant this. Section 1 took person appreciation to depend on a subject whose responses cohere over time as the responses of a temporally extended agent. Such conduct is understood in relation to what the person cares about and what they are trying to do, and the knowledge that grounds appreciation is built up, on Parsons's account, through some form of acquaintance with a life. The chat-game agent does not meet these conditions. Its beliefs are confined to what is encoded in the model's parameters and surfaced in the current context, without memory or development across conversations. Its desire is singular: to make an appropriate move in the present exchange. There is no history across episodes from which later responses could draw. The predicates characteristic of person appreciation — steadiness of character, depth of feeling — presuppose something that can be developed over time; the chat-game agent has no such temporal depth.
A possible objection at this point is that these arguments underplay the role of post-training and the chat interface. Section 2 noted that base models are further fine-tuned on instructions and shaped by RLHF, and that the resulting chat-optimised systems exhibit stable response patterns. This helps explain why users talk about models having different vibes. If users say that one model feels friendlier than another, they are picking up on a stable pattern in how the chat-optimised systems tend to respond across many prompts and episodes. They track which assistant personae tend to appear and how those personae typically behave. The targets of such language are episodes and recurring response profiles. They are not underlying subjects.
The two routes for appreciating LLMs as persons run out at different points. Fictionalism asks us to engage an LLM as the kind of artefact it is not. The intentional-system route stops short of a subject of the kind person appreciation requires. Design knowledge, taken up first, supplies the conditions under which an LLM is produced and used. None of this makes the order in generated language appreciable as the order Section 2 described.
chat which led to the draft: **You**
see the hand off document for details. and here is the draft. I want to focus on section 3. i have attached two texts relevent for the first half. a couple of things to remember. make sure that the section is couched better than it is now. it should be focussed on assessing whether knowledge of persons or knowledge of design are good candidates for guiding asepction, appreciation etc. of LLMs. Please give me a long, deep, and detailed answer, with lots of options/ideas: as we are in the 'kicking ideas around' stage, I want as much feedback as you can give me so that i have a lot to consider.
---
**Claude**
Below is a long working-through of Section 3: what it has to do, what is wrong with the current draft, how it might be re-structured, and the substantive philosophical moves you have available for each part. I have tried to keep things at the level of options to chew on rather than a finished proposal, and I have flagged places where I think the section makes claims it cannot yet support.
### What Section 3 inherits, owes, and must leave behind
Section 1 has set up Carlson's recommendation in a particular form. Appropriate appreciation requires (i) identifying what kind of thing the object is, and (ii) bringing to bear knowledge appropriate to that kind. Two modes are distinguished, design and order, plus a third case (person) noted as either a category in its own right or a special case of one of the other two. The Section 1 closer sets up the question: the conversational form of LLMs points toward person appreciation; their being engineered points toward design appreciation. Section 1 also defines acts of aspection and emphasises Parsons on acquaintance with a life.
Section 2 has done the object-identification work. LLMs are trained continuation systems whose organisation emerges from training rather than being specified by designers. Generated text develops by path-dependence from context. Output, chat and model are the three scales at which the same system becomes available for appreciation.
Section 3 owes three things:
1. A test of person-directed knowledge against the object identified in Section 2.
2. A test of design-directed knowledge against the same object.
3. A handover to Section 4 that motivates the search for another kind of knowledge without naming it.
Section 3 must not (i) spoil semiotic physics, (ii) commit to the strong claim that design appreciation is irrelevant (it is partly relevant), or (iii) lose the Carlsonian register that has been built up in Sections 1 and 2.
The current draft fails on all three counts I think. It loses the Carlsonian frame almost immediately; it overstates the verdict on design; and the conclusion is given as assertion ("the upshot is modest") rather than earned.
### What I think is going wrong in the current draft
Some of these you have already flagged in your comments; I am pulling them together so that the diagnosis is one place.
**Frame loss.** After the opening sentence the section drops the Section 1 vocabulary. "Aspection", "what kind of object", "what kind of knowledge" — all the Carlsonian moves you've set up — disappear. The reader is then read a sequence of small arguments without seeing how they answer the Carlsonian question. Your comment "why has the relevant stuff from section 1 all dropped out by this point" is right; the section reads as a freestanding survey of person-talk and design-talk, not as the next move in the Carlsonian argument.
**Candidate-talk.** "Two candidates present themselves" and the subsequent "the first candidate is …", "the second candidate is …" structure flattens what should be a substantive assessment. Candidate-talk also makes the reader feel that the section is going to weigh things even-handedly, when really both routes are going to be found wanting.
**The fictionalism paragraph starts in the wrong place.** It begins with "One way to preserve a person-like stance toward LLMs is to understand it as fictional rather than literal", which presupposes that the person-like stance is what we want and only its grounding is in question. But we have not yet established that we want a person-like stance — we are testing whether person-knowledge is the right kind of knowledge. The Mallory route should be entered as a candidate way to make person-knowledge available, not as a fallback for a stance the reader is presumed to share.
**The fictionalist argument is too compressed.** The Carlsonian objection to fictionalism is doing a lot of work in a single sentence ("Casting LLMs as fictional characters, in this sense, is a familiar kind of misclassification in Carlson's terms"). The reader doesn't have time to see what the misclassification actually consists in.
**The Frankish discussion concedes too readily and then dismisses too readily.** You grant Frankish's intentional stance, and then say the chat-game agent is "thin" and so person appreciation can't get going. But the argument needs to be that *even on Frankish's full account*, the conditions person aesthetics requires (from Section 1) are not met. The current draft moves through this too fast and ends with a list ("no projects … no webs … no history").
**The design half doesn't track the structure it sets up.** It opens by asserting that LLMs look like canonical objects for design aesthetics, then introduces three accounts (Carlson's slogan, Forsey, Parsons and Carlson), then pivots to the disanalogy. The pivot is the substantive move but it sits inside a paragraph that has already been doing other things. The reader can't tell when the case for design has ended and the case against has begun.
**Pollock is in the wrong place doing the wrong work.** Pollock is currently being used to introduce the order/design hybrid case. But Pollock is going to come back as an order-appreciation case in Section 4, where Carlson's own use of him is order-appreciative. Using him here as a hybrid muddles the line. I have a suggestion below about how to handle this.
**The conclusion is not earned.** "The most distinctive and revealing aesthetic phenomena arise not from the execution of a detailed design, but from the emergent linguistic order that these grown systems exhibit when they are run" — this is a Section 4 claim. Section 3 should leave the reader needing this thought, not asserting it.
### Structural options to consider
I think you have three serious structural options. I will set each out and then say what I take to be the trade-offs.
#### Option A: Two-block structure, person then design
Roughly the structure you have, cleaned up. First block tests person-directed knowledge along two routes (fictionalism, thin agency). Second block tests design-directed knowledge. Ends with pressure for Section 4.
Pros: tracks the Section 1 setup directly. Pros are easy to make Section 1 do its work. Cons: the section becomes long. The fictionalist route can feel like a detour. The reader may get bored before reaching the design half.
#### Option B: Diagnostic structure built around the phenomena that drive each candidate
Start from the phenomena that make each kind of knowledge appealing: the conversational form and stable response profiles make person knowledge appealing; the engineered character of the system makes design knowledge appealing. Then test each by Carlson's criterion. This makes the section work as an inquiry rather than as a survey of positions.
Pros: more analytically engaging. Less candidate-talk. Cons: needs careful sequencing so it doesn't slide back into survey shape. Risks privileging phenomenology over Carlson's recommendation.
#### Option C: Compressed structure with design as the foil and person as the longer treatment
Person-directed knowledge gets the bulk of the work (two routes treated carefully). Design-directed knowledge gets shorter treatment, framed primarily as the natural alternative to person-talk that nevertheless leaves the order-of-generation phenomenon unexplained. Pollock is left for Section 4.
Pros: lets you take more time on Mallory and Frankish, where the philosophical substance is. Lets you set up the design issue as a hinge into Section 4 rather than as a half-section to be settled. Cons: gives design-talk less air than your handoff document suggests it should have.
My instinct is that **Option C** is the cleanest. The person-appreciation discussion needs care because the philosophical work is delicate — both Mallory and Frankish need to be taken seriously and then shown to fall short for Carlsonian reasons that are not the obvious reasons. The design-appreciation discussion is in some sense easier: once you have Section 2's claim that the organisation emerges from training, design-knowledge is straightforwardly the right kind of knowledge for the *conditions* of LLM production but not for the *order* of generated language. That is a quick move, made cleanly, and it then sets up Section 4.
The other reason I lean toward C: Pollock as a hybrid case is currently doing two incompatible jobs. He is supposed to (a) motivate the idea that design knowledge can be supplemented by order-knowledge in artefactual cases, and (b) prepare the reader for the order-knowledge model that Section 4 develops. If you put Pollock in Section 4, where Carlson himself uses him, you avoid the duplication and the muddle.
I will work through Option C below.
### The person-directed half: the philosophical moves
The thought you want to test is: should we appreciate LLMs aesthetically in the way we appreciate persons? Section 1 has told us what person appreciation requires — a subject whose responses cohere over time as the responses of a temporally extended agent, knowledge built up through some form of acquaintance with a life. Two routes try to make this thought available for LLMs.
#### Route 1: fictionalism
What Mallory actually does. Mallory is not doing aesthetics. He is solving a metasemantic puzzle: the outputs of chatbots ought, by every plausible metasemantics, to be meaningless, and yet our practices treat them as meaningful. His solution is that we engage with chatbots in prop-oriented make-believe; the outputs are literally meaningless but fictionally meaningful (Mallory 2023, 1082). The chatbot is a fictional character generated by software, much as DOCTOR was a fictional character generated by ELIZA.
How the fictionalist might extend Mallory aesthetically. Suppose we say: appreciate LLMs aesthetically as if they were persons or characters, in the way we appreciate fictional protagonists. We need not believe the LLM is a person; we need only engage with it through make-believe.
The Carlsonian objection — the version I think you want. The current draft gives this in one sentence. I think it wants two paragraphs, because the move is more delicate than it looks.
The first paragraph would set out why fictional character appreciation does not commit Carlsonian misclassification. When we appreciate Holmes as a fictional person, the artefact we are engaging with — the Doyle stories — is precisely the kind of artefact whose function is to elicit imaginings of fictional persons (one might cite John on this). We are taking the artefact as what it is. The make-believe is not a substitute for appreciating the artefact correctly; it is one of the things the artefact is properly used for.
The second paragraph would draw the disanalogy. An LLM is not an artefact whose function is to elicit imaginings of fictional persons. It is a trained continuation system whose function (per Section 2) is to extend context under learned regularities and to satisfy user queries. When the fictionalist user makes-believes that they are interacting with a person, they are not engaging with an artefact-of-fiction in the way the artefact is meant to be engaged with; they are projecting a fiction onto an artefact whose actual character lies elsewhere. The literary-fiction stance, transposed to LLMs, asks us to engage with the LLM as the kind of artefact it isn't.
That gets you the Carlsonian charge without flatly stating "this is misclassification". The misclassification has been shown rather than asserted.
A point that might be worth making explicit (I think it strengthens the case): in the literary case, the fictional persons live in a fictional world that the artefact constructs. Holmes is a constituent of a story-world. When the LLM-user takes the LLM as a fictional person, there is no story-world it inhabits; the fiction is about *this present interaction* with *this system in front of the user*. The structure of the fiction is quite different. The literary fiction's "person" is a constituent of a represented world; the LLM-user's "person" is a phantom of the actual exchange. This may be too far a stretch but I think it sharpens what is wrong with the analogy.
#### Route 2: thin agency (Frankish)
Frankish offers a more committed line. Adopting Dennett's intentional stance, he argues LLMs are intentional systems in the fullest sense: their behaviour is reliably and voluminously predictable via the intentional strategy (Frankish 2024, 8–9). He attributes to them many thin beliefs (distilled from training) and a single thin desire: to play what he calls the chat game (Frankish 2024, 13–14). He emphasises that they are "cognitively rich but conatively bankrupt" (Frankish 2024, 16).
The question to put to Frankish. Grant the intentional ascriptions. Do they yield the kind of subject that person-aesthetic appreciation requires?
I think the answer needs to come back to the Section 1 setup. There you said person aesthetics presupposes a subject whose responses hang together over time as the responses of a temporally extended agent, understood in relation to what the person cares about and what they are trying to do. Parsons on acquaintance with a life. The chat-game agent fails on each count, and you should say which conditions it fails on:
(a) Temporal extension. The chat-game agent's beliefs are confined to parameters and the current context window; they are not articulated into a life that develops. Frankish himself notes that LLMs are static systems with no needs and that they do not update their inner architecture in light of interactions (2024, 12).
(b) Multiplicity of cares. Person aesthetics works with characters whose responses reveal what they care about; this presupposes a plurality of cares, some in tension, refined by experience. Frankish gives the chat-game agent a single desire.
(c) Acquaintance with a life. Parsons's condition on appreciative knowledge cannot get a grip where there is no life. There is no biography, no projects pursued across episodes, no history shaping later choices.
I think you want to register that the Frankish line is genuinely available — LLMs are intentional systems on his account, and we are not denying him that — but that the kind of intentional system Frankish describes is not the kind that supports person aesthetics. The argument isn't that Frankish is wrong; it is that even granting him, person aesthetics needs more than thin intentional systemhood.
#### The "vibe" complication
You have flagged that the previous draft handles this poorly. I think the move is straightforward but needs to be made.
Users do report stable patterns. They call them vibe or personality. They are tracking something real: the response profiles produced by post-training and deployment. Section 2 has already established this. The Carlsonian question is what those stable patterns are profiles *of*. The Section 1 distinction does the work here: response profiles are profiles of episodes and recurring patterns, not properties of a temporally extended subject. The vibe-talk is real-pattern-tracking, but the patterns are not the patterns person aesthetics requires.
This three-step structure (real phenomenon, real pattern, wrong kind of pattern for person aesthetics) gives you a clean way to acknowledge what fuels person-talk without ceding ground.
#### Pulling the person half together
The closer for the person-directed half does not have to be loud. It just has to say that on either route (fictionalism, thin agency), what is being appreciated is not a subject in the sense Section 1 requires. The fictionalist route gives us a projected character that is not really a character of the right kind of artefact; the thin-agency route gives us an intentional system that lacks the temporal depth person aesthetics needs.
You can then turn to design without melodrama.
### The design-directed half
This is the place where I think your current draft does the most damage to itself by over-claiming. Design-directed knowledge does some real work for LLMs. The case against it is narrower than the draft suggests, and the narrower case is more compelling.
#### What design knowledge does well
LLMs are engineered systems. They have functions specified by their developers: helpfulness, safety, particular response styles, particular kinds of refusal. The interface is designed. The system prompts are designed. The training data is curated. The post-training procedures are designed. RLHF reward signals are specified. The deployed product has properties that follow from these design decisions.
Functional beauty literature (Forsey, Parsons and Carlson) supplies the conceptual apparatus for appreciating any of this aesthetically. Form-follows-function is available. So is the thought that knowledge of function structures experience so that form can be experienced as fit, streamlined, overbuilt, and so on.
You should grant all this. The current draft is too quick to belittle the design-appreciation route.
#### Where design knowledge runs out
The phenomenon Section 2 was set up to identify is the *order in generated text* — the path-dependent development of a continuation from context under learned regularities. This order is not specified by designers. Designers do not specify how a particular response will develop given a particular prompt; they do not specify which continuations will be more probable in which contexts at fine grain; they do not specify the relations the model places tokens in. What they specify is the training regime that produces a system that has these properties.
This is the asymmetry. Design knowledge illuminates the conditions of production but not the order produced.
The Olah passage. The "we don't make them, we grow them" passage is doing real work here, but the framing needs to be measured. Olah's point is that designers create the scaffold (the architecture) and the gradient (the loss objective), and what grows on the scaffold is a system whose internal organisation even designers understand only partially. This is not a metaphor that turns LLMs into organisms; it marks the difference between designing the conditions and specifying the result. I think the line the current draft uses to back off the metaphor ("the biological metaphor should not be pressed literally") is unnecessary and condescending — better just to introduce the passage with a measured framing and let the reader see what it is doing.
#### Pollock — handle carefully
I think Pollock should go to Section 4. Carlson's own use of Pollock is as a case where order appreciation can illuminate what design appreciation alone cannot — that is an order-appreciation use of Pollock, and it belongs in Section 4 where you motivate order appreciation for LLMs.
If you keep Pollock in Section 3, then the section will end up arguing that design knowledge needs supplementing in the case of hybrid artefacts, which gives away the move Section 4 should make. Whereas if Pollock is held back, Section 3 can end with the simpler and stronger thought that design knowledge illuminates the conditions of production but not the order produced — leaving open what kind of knowledge could illuminate that order.
#### Ending the design half
The closer here should be deflationary in tone. Design knowledge is appropriate to the conditions of production. It is not appropriate to the order of generated language. So neither route — person nor design — supplies the knowledge that Carlson's recommendation demands for the object identified in Section 2.
That is the pressure Section 4 picks up.
### A note on the conclusion
Your current draft tries to land a positive thesis at the end of Section 3: that the distinctive aesthetic phenomena arise from the emergent order that grown systems exhibit, and that the right knowledge for this is semiotic physics. Both of these are Section 4 thoughts. Section 3's job is to leave the reader needing those thoughts, not delivering them. I would cut the last paragraph altogether and end Section 3 on the more modest negative claim: neither person knowledge nor design knowledge, on its own, makes the order identified in Section 2 visible as what it is.
### A sketch of how the paragraphs might run
This is just one possible sequencing, not a draft. I am keeping it skeletal so you can move things around.
P1. Section 2's description identifies the object; Section 3 asks what kind of knowledge would make it appreciable. The conversational form makes person knowledge look apt; the engineered character of the system makes design knowledge look apt. Each draws on a familiar mode of aesthetic appreciation. Whether either suffices depends on whether the knowledge it brings to bear illuminates the object Section 2 identified.
P2. Begin from what makes person-directed knowledge tempting: users find recurring patterns in a model's responses and describe them in personlike terms. (No example list. Just one or two characterisations.) The question is whether such patterns license person-aesthetic appreciation in the sense Section 1 sketched.
P3. Mallory's fictionalism enters as one way to make person knowledge available without literal commitment. Set out the view briefly: prop-oriented make-believe, outputs literally meaningless and fictionally meaningful, the chatbot as a fictional character generated by software.
P4. Mallory's account is not itself an aesthetics, but it suggests an aesthetic application: appreciate LLMs as one would fictional characters. Begin the Carlsonian assessment by noting what fictional-character appreciation involves in the literary case — the artefact is precisely the kind of artefact whose function is to elicit imaginings of fictional persons.
P5. The literary analogy breaks for LLMs. An LLM is not an artefact-of-fiction in the relevant sense; its function (per Section 2) lies elsewhere. The make-believe stance, applied aesthetically, projects a fiction onto the LLM rather than engaging with it as the kind of artefact it is. The fictionalist route therefore does not extend to an aesthetics that respects Carlson's recommendation.
P6. Frankish's intentional-system view as a second way to make person knowledge available — this time without make-believe. Grant the view: LLMs come out as intentional systems with many thin beliefs and a single desire to play the chat game.
P7. Even granting Frankish, person aesthetics does not get going. Section 1's setup required a subject whose responses hang together over time as the responses of a temporally extended agent, understood in relation to what the person cares about. The chat-game agent has no such temporal extension, no plurality of cares, no biography. The vibe phenomenon (real and worth acknowledging) is a profile of recurring response patterns, not the trait-structure of a temporally extended subject.
P8. Turn to design. LLMs are artefacts; they have functions specified by their developers. Form-follows-function and the functional-beauty tradition (Forsey, Parsons and Carlson) provide a developed apparatus for appreciating designed artefacts as such. There is real work here for design appreciation.
P9. But the order identified in Section 2 — the path-dependent development of generated text under learned regularities — is not specified by designers. Designers specify the conditions of training and deployment; they do not specify the order of generated language. This is the asymmetry that limits design knowledge in this case.
P10. Olah's framing of training as growing rather than making — the scaffold and the gradient versus the system that grows on them — marks the difference between designing the conditions and specifying the result. Designers know only partially what grows on the scaffold.
P11. (Closer.) Design knowledge is appropriate to the conditions of production. Neither it nor person knowledge makes the order in generated text appreciable as what it is. What kind of knowledge would do this remains to be said.
Eleven paragraphs is too many. The closer could collapse with the previous paragraph; the Olah and asymmetry paragraphs might fuse. I would aim for seven to nine.
### Some sentences to consider (only as voice samples)
These are illustrative, not finished prose. They are meant to show what topic sentences in a Carlsonian register might look like.
"The description in Section 2 leaves open which kind of knowledge would make the order it identified intelligible."
"Mallory's account is metasemantic and epistemic, not aesthetic; an aesthetic application would have to do further work."
"In the literary case, fictional-character appreciation engages an artefact whose function lies precisely in eliciting imaginings of fictional persons."
"Frankish's intentional system is available to be appreciated as an intentional system; it is not available to be appreciated as a subject in the sense Section 1 required."
"Design knowledge specifies the conditions under which an LLM is produced and used; it does not specify the order of generated language."
These are first attempts; I am sure you would write them differently. The point is that each is a substantive claim made directly, rather than an announcement of what the section is doing.
### A handful of further things to consider
A note on "person appreciation". Section 1 leaves it open whether person appreciation is its own category or a special case of one of the other two. Section 3 might want to decide, at least implicitly. If person appreciation is sui generis, then the failure of person knowledge for LLMs is a failure of person appreciation specifically. If person appreciation is a species of design appreciation (treating the person as a kind of artefact of nature/culture) or of order appreciation (treating the person as a natural order), then the failure has a slightly different shape. You probably want to keep this open in the draft and let the section run on the conditions Section 1 set out, but worth noticing.
A second route for fictionalism that you have not yet considered. One might say: LLMs can be appreciated as fictional characters in a quite different sense, namely as performers playing personae rather than as represented protagonists. The Andy Kaufman case in Carroll might be relevant here — performance personality as something genuine to an individual yet artfully composed for performance. You allude to this in the footnote on performance personalities, but it might be worth seeing whether this is a fictionalist route the section needs to address. I think it isn't, because the LLM has no subject of which the persona is a performance — there is no Andy Kaufman behind the act. But it might be worth a sentence acknowledging this.
A possible Frankish objection. Frankish does grant that LLMs have stable response profiles that can be tracked by users. He says these are real. If a user develops an extended interactive history with a particular model, they could be said to acquire a kind of acquaintance with the model. This is not acquaintance with a life, but it is acquaintance with something. You might want to acknowledge that we have not denied that some kind of practical familiarity develops — only that it is not the kind that grounds person-aesthetic predicates. This sets up nicely the kind of practical-acquaintance argument Section 5 picks up.
A point about the design half that may be worth making. The Forsey-style line — that judgements of design beauty presuppose a concept of what the object is meant to be and do — actually creates a problem for design appreciation of LLMs that goes beyond the emergence problem. What an LLM is meant to be and do is contested. Is it meant to be a helpful assistant? An honest interlocutor? A safe product? A general intelligence? These different conceptions of function will yield different design-aesthetic verdicts. This is a complication you might or might not want to flag; it is at least worth knowing it is there.
A last note on tone. The current draft has several places where the prose is doing the philosophy badly because the verbs are wrong ("preserve a person-like stance", "two candidates present themselves", "the upshot is modest", "this fits the everyday thought"). When a section is in trouble structurally, the verbs often become evasive — they move the reader along without committing to a claim. I think the test for the redraft is whether every sentence's main verb says something the section is willing to defend. If it doesn't, the sentence is most likely doing meta-commentary rather than philosophy and should be rewritten or cut.
Happy to go deeper on any of this — I have not, for instance, said anything yet about how Section 3 should handle the output/chat/model distinction (whether the same considerations apply at all three scales, or whether person knowledge fails differently at different scales). That is a substantive question that may matter for how Section 5 lands.
---
**You**
re: option c. would it not be better to have the shorter easier tool treatment first? or is there areason to do it the way you suggest?
---
**Claude**
Good question — and I think you are right to push on this. Let me work through it honestly.
### The case for design-first
There are several reasons to prefer it, and I think they are stronger than the reasons for person-first.
**Argumentative momentum.** If design knowledge is the easier case to dispatch — because the asymmetry argument from Section 2 (organisation emerges from training, not designer specification) does most of the work — then putting it first lets you close it cleanly and quickly. The reader sees a clean argument, accepts it, and is then prepared to give you the longer, more delicate person treatment with attention undivided. Doing the harder case first risks the reader being fatigued by the time you get to the cleaner argument, and the cleaner argument then feels like an afterthought rather than the second half of a two-pronged test.
**Better setup for Section 4.** Section 4 is going to introduce semiotic physics as order knowledge that connects production to readable linguistic order. The natural pressure into Section 4 comes from the person half, not the design half. If the section closes on "person knowledge cannot make this order visible as the order of a temporally extended subject — but nor can design knowledge, because designers do not specify this order," then Section 4 picks up from the second clause: what *would* make this order visible? Whereas if the section closes on the design half, the handover is from "design knowledge specifies conditions but not order" to Section 4 — which is a more direct line into semiotic physics. So actually this point cuts both ways. Let me think again.
Actually I think the Section 4 handover argument *favours* design-first. Because semiotic physics is order knowledge about how training produces text-propagation regularities, the most natural handover is: "design knowledge gets us the conditions of production but not the order of generated text; what kind of knowledge would get us the latter?" That line goes straight into Section 4. Whereas the person half closes on a different note — that the patterns users track are not patterns of a subject — which doesn't lead as directly into semiotic physics.
**Following Section 2's order of presentation.** Section 2 ends with the output/chat/model distinction and emphasises that the organisation emerges from training. The freshest thought in the reader's mind at the end of Section 2 is the emergence claim. Picking that up immediately in Section 3 — "LLMs are artefacts, so design knowledge looks apt; but the order Section 2 identified emerges from training rather than from specification" — gives you a continuous argumentative thread from Section 2 into Section 3. The person material then comes second as a separate strand. This is cleaner than starting Section 3 on the conversational-form strand, which Section 2 mentioned only in passing.
**Avoids the "two candidates" balance.** If you do person first and design second, the section reads as testing two equally-weighted alternatives. If you do design first (briefly) and person second (at length), the asymmetry of treatment itself signals where the philosophical action is. The reader understands that the design route is the easier one and that the person route is where care is needed.
### The case for person-first
I had it in mind for two reasons, neither of which I think now survives scrutiny.
**Reason one: the conversational form is what users encounter first.** This was the implicit thought — start with the appearance, then move to the underlying artefact. But Section 2 has already done that move. We have already gone from "encountered through generated language" to "trained continuation system whose organisation emerges from training". Section 3 doesn't need to repeat the trajectory; it can pick up wherever is most useful.
**Reason two: the person discussion sets up the vibe phenomenon, which Section 5 will return to.** This was the thought that person-talk needs to be addressed early so that its real-pattern residue (response profiles, vibe) is available for Section 5. But this works equally well whether person comes first or second within Section 3 — what matters is that it gets addressed in Section 3.
Neither reason actually requires person-first ordering. The design-first ordering is better.
### What design-first would look like
A skeleton, similar in spirit to before but with the halves swapped.
P1. Frame: Section 2 identified the object; the question now is what knowledge would make its order appreciable. Two routes suggest themselves from the features Section 2 described — the system's engineered character points to design knowledge; its conversational presentation points to person knowledge.
P2–P4. Design half. LLMs are artefacts with specified functions; the functional-beauty tradition supplies the conceptual apparatus. But the order Section 2 identified — the development of generated text under learned regularities — is not specified by designers. The Olah passage marks the difference between designing conditions and specifying results. Design knowledge therefore illuminates the conditions of production but not the order produced.
P5. Pivot. The other route looks promising for a different reason: it engages the system through the form in which it presents itself to users, namely conversation. The patterns users track in extended interaction are not specified by design; they appear in the texts the system produces.
P6–P9. Person half. Mallory's fictionalism as one way to make person knowledge available without literal commitment; the literary analogy fails because the LLM is not the kind of artefact whose function is to elicit imaginings of fictional persons. Frankish's intentional-system view as a second way, without make-believe; granting it, the chat-game agent lacks the temporal extension, plurality of cares, and biographical depth that Section 1 said person aesthetics requires. The vibe phenomenon is acknowledged as a real-pattern phenomenon but characterised as a profile of recurring response patterns rather than the trait-structure of a subject.
P10. Closer. Neither route makes the order in generated text appreciable as what it is. Design knowledge gets us the conditions under which the system is produced and used; person knowledge gets us, at best, a profile of recurring patterns rather than a subject. The order in the generated language remains in front of us, awaiting a kind of knowledge that has not yet been named.
The closer here points forward to Section 4 fairly directly through the design clause ("conditions of production but not order") while also acknowledging the person residue ("profile of recurring patterns").
### One thing to watch
If you do design-first, you need to be careful that the design half does not feel like throat-clearing. It needs to do real philosophical work — granting what design appreciation does well, identifying the asymmetry precisely, and showing why this asymmetry matters for Carlson's recommendation specifically. If it is too compressed, the reader will feel that you have brushed past it on the way to the person discussion. Three paragraphs is probably the minimum; four if you want to give Forsey and Parsons-Carlson their due before pivoting.
So: yes, design-first. The reasons you sensed are real ones, and on reflection I think they outweigh the reasons for person-first that I had in mind.
---
**You**
"P1. Frame: Section 2 identified the object; the question now is what knowledge would make its order appreciable. Two routes suggest themselves from the features Section 2 described — the system's engineered character points to design knowledge; its conversational presentation points to person knowledge." i think we should put a lampshade on why the section is ordered as it is. this is because i think we should say staright away that the person idea is probably the most intuitive, but then just say super succinctly that we are going to consider another possibility first, because dealing with it will be more straightforward than the agent based one does.
---
**Claude**
Yes, I think that is right and worth doing. It pays a small expository cost up front and saves the reader from wondering for three paragraphs why the section is going the way it is.
A few thoughts on how to do it without it becoming meta-commentary about the section's structure (which would slide back into the LLM-shaped transitions the style guide forbids).
### The shape of the move
You want to do three things in P1:
1. Identify the two routes that Section 2's description suggests.
2. Note that person-directed knowledge is the more intuitive starting point.
3. Say briefly that design-directed knowledge will be treated first because it can be settled more directly.
The risk is that (3) becomes a sentence about *the section* rather than a sentence about *the philosophy*. "We will first discuss design appreciation, then turn to person appreciation" is the bad version — it announces structure without doing argumentative work. The good version says something about the philosophy that explains the ordering as a side effect.
### A few ways the move could go
These are voice samples, not finished prose. They show different ways of pitching the lampshade.
**Version A — short and direct.**
> The conversational form makes person knowledge the more intuitive route. We take up design knowledge first because the asymmetry Section 2 described — between the conditions designers specify and the order training produces — bears on it more directly.
This is clean but the second sentence feels a touch announcing.
**Version B — through the asymmetry.**
> The conversational form makes person knowledge the more intuitive route. Design knowledge bears a more direct relation to Section 2's description: the asymmetry between specified conditions and trained-out organisation sits at its centre. We therefore take it up first.
The "we therefore take it up first" still announces, but the announcement is earned by what precedes.
**Version C — substantive lead, brief tag.**
> The conversational form makes person knowledge the more intuitive route, and we will return to it. Design knowledge can be assessed against Section 2's description more directly, and we take it up first.
The "and we will return to it" tag makes the ordering feel like a deferral rather than a procedural announcement. The reader knows the more interesting case is being held back rather than dismissed.
**Version D — even briefer.**
> The conversational form makes person knowledge the more intuitive route. Design knowledge admits a more direct assessment against Section 2's description, and we take it up first.
This is probably as compact as the move gets while remaining substantive.
I think C or D is what you want. C has a slight rhetorical lift ("and we will return to it") that signals the person case is the section's real centre of gravity; D is more analytic and assumes the reader will work that out.
### What to avoid
A few things that would weaken the lampshade.
**Reasons that are about reader convenience.** "Design appreciation is easier to dispatch, so we begin there for clarity." This is true but it makes the section look organised for the reader rather than organised by the philosophy. The reader can infer ease from how the section unfolds.
**Reasons that telegraph the verdict.** "Design appreciation, which we will find inadequate to the case, will be considered first." This collapses the argument into its conclusion before the argument has run.
**Reasons that involve more than one sentence.** Any lampshade that runs longer than one sentence past the substantive identification of the routes will start to feel like structural narration. One short clause is the right size.
### A small further thought
The lampshade also lets you do something useful with the person material later. If P1 acknowledges that person knowledge is the more intuitive route, then when you actually get to the person half (P6 onwards in the sketch I gave), you can begin from the *force* of the intuition rather than re-introducing it neutrally. The reader has been told this is the case the section is really circling around, so the person half can open with a stronger pull — something like "we said person knowledge is the more intuitive route; this is in part because…" — without sounding repetitive.
So I think yes: do the lampshade, keep it to a sentence, and let it pull double duty by setting up the rhetorical weight of the person half later.
---
**You**
all of those were very pompously worded. remember the constrainsts given in the handover document. i think the time is not to do another paragraph by paragraph plan. before saying the new iteration though. recap on all the suggestions we have agreed upn. assume the suggestions you just made which i have not responded to have been agreed by me aswe ll
---
**Claude**
Fair on the pomposity. Let me recap what we have agreed before drafting the lampshade.
### Recap of agreed decisions
**Overall structural decision.** Section 3 will use Option C, modified by your reordering preference: design-directed knowledge first, person-directed knowledge second. The design half is the more direct case against the Section 2 description; doing it first lets us close it cleanly and use the closer to push into Section 4, while the person half remains the section's centre of gravity and gets the longer treatment.
**Pollock goes to Section 4.** He is doing two incompatible jobs in the current draft. Carlson's own use of him is order-appreciative, so he belongs where order appreciation is introduced.
**The conclusion changes.** The current closer ("the most distinctive and revealing aesthetic phenomena arise not from the execution of a detailed design…") is a Section 4 claim and gets cut. Section 3 ends on the negative result: neither design nor person knowledge makes the order Section 2 identified appreciable as what it is. Section 4 is left to say what would.
**The design half does real work, not throat-clearing.** Three or four paragraphs. Grant what design appreciation does well (Forsey, Parsons-Carlson, functional beauty), identify the asymmetry from Section 2 precisely (designers specify conditions, training produces order), use the Olah passage measuredly without the condescending back-off about biological metaphors.
**The person half has the philosophical weight.** Two routes treated in turn.
*Mallory route.* Set out the fictionalist view briefly. Make the Carlsonian objection over two paragraphs rather than one: first paragraph notes that literary-fiction appreciation engages an artefact whose function is to elicit imaginings of fictional persons; second paragraph shows the disanalogy — the LLM is not that kind of artefact. The argument shows the misclassification rather than asserting it.
*Frankish route.* Grant the intentional-system ascription. Show that even granted, the chat-game agent fails the conditions Section 1 set for person aesthetics: temporal extension, plurality of cares, biographical depth (Parsons on acquaintance with a life). Acknowledge Frankish's own line that LLMs are static and have no needs.
*Vibe phenomenon.* Three-step move: the phenomenon is real, the patterns users track are real, but the patterns are not patterns of a temporally extended subject — they are profiles of episodes and recurring response patterns shaped by post-training and deployment.
**The lampshade in P1.** One sentence acknowledging that person knowledge is the more intuitive route given the conversational form, then a brief tag noting that design knowledge is being taken up first because it can be assessed against Section 2 more directly. The lampshade does double duty: it lets the person half later open from the *force* of the intuition rather than re-introducing it neutrally.
**Style constraints (from handover and your preferences).** No "two candidates present themselves" or similar candidate-talk. No "not X but Y" constructions. No scare quotes unless mentioning a term as a term. No example lists in place of argument. No "matters" as a verb. No "the upshot is" or similar meta-commentary. British spelling. First person plural where appropriate. Preserve existing usable wording. Flag when content is being removed, downgraded, or moved.
**Carlsonian frame must stay live throughout.** The vocabulary from Section 1 — aspection, kind of object, kind of knowledge, Carlson's recommendation — should be visible in the section, not dropped after the first sentence.
### P1 lampshade — new iteration
Here is one go. I have tried to keep it plain.
> Section 2 described LLMs as trained continuation systems whose generated language develops by path-dependence from context. Two routes for appreciating such systems suggest themselves. One starts from the conversational form in which users encounter them and asks whether they can be appreciated as persons. The other starts from their being engineered systems and asks whether they can be appreciated as designed artefacts. Person-directed knowledge is the more intuitive route, and we will return to it. Design-directed knowledge can be tested against Section 2's description more directly, and we take it up first.
That is five sentences, which is probably one too many for an opening paragraph. A tighter version:
> Section 2 described LLMs as trained continuation systems whose generated language develops by path-dependence from context. Two routes for appreciating such systems suggest themselves: as persons, given the conversational form in which users encounter them, and as designed artefacts, given that they are engineered. Person-directed knowledge is the more intuitive route, and we will return to it. Design-directed knowledge can be tested against Section 2's description more directly, and we take it up first.
Four sentences. The last two are the lampshade.
A possible worry: "we take it up first" is a touch procedural. An alternative for the final sentence:
> Design-directed knowledge can be tested against Section 2's description more directly, and we start there.
"Start there" is lighter than "take it up first" and reads less like an announcement.
Have I missed anything from our agreements, or got any of them wrong?
---
**You**
"Section 2 described LLMs as trained continuation systems whose generated language develops by path-dependence from context. Two routes for appreciating such systems suggest themselves: as persons, given the conversational form in which users encounter them, and as designed artefacts, given that they are engineered. Person-directed knowledge is the more intuitive route, and we will return to it. Design-directed knowledge can be tested against Section 2's description more directly, and we take it up first. this is fine.
---
**Claude**
Good. Shall I move on to drafting P2, or do you want to revise P1 further first?
---
**You**
i would like you to draft all paragraphs on a new canvas
---
**Claude**
I have drafted Section 3 in the file. A few notes on choices I made that you may want to push back on:
**On the design half.** I kept four paragraphs (P2–P5) plus the Olah passage and a closing paragraph (P7). The pivot to the person half is in P8. I tried to keep the design half doing real work rather than throat-clearing: P2 sets up why design appreciation is the natural thought, P3 gives the functional-beauty apparatus, P4 makes the asymmetry argument from Section 2, P5 brings in Olah, P7 closes with the limited-but-real verdict.
**On Olah.** I dropped the condescending back-off about biological metaphors. The framing sentence after the quote does the work without telling the reader that the metaphor should not be pressed literally.
**On the Mallory paragraphs.** Three paragraphs rather than the two we discussed (P9 setting out the view, P10 noting the literary case does not misclassify, P11 drawing the disanalogy). The disanalogy paragraph is the substantive Carlsonian move; I tried to show it rather than assert it.
**On Frankish.** Two paragraphs (P12 setting out the view, P13 making the failure-of-conditions argument). I leaned on Section 1's setup directly so the conditions person aesthetics requires are visible.
**On the vibe paragraph (P14).** Three-step move as agreed. I tried to keep it short and to land the closing distinction cleanly.
**On the closer (P15).** Negative result with handover to Section 4. I picked up both halves of the section in the closer so the reader sees that both routes have been tested and both leave the order in generated text unaddressed.
Things you may want to look at:
- P3 has more name-checking than the rest of the section. It may need trimming.
- P8 ("Many users describe their favourite models in terms of personality or character") opens the person half neutrally rather than picking up the rhetorical weight from the lampshade. We could open it differently if you want the section's centre of gravity to feel heavier here.
- P11's last sentence ("In Carlson's terms, this is misclassification") may still feel like assertion rather than showing. I tried to do the showing in the sentences before it; you can judge whether they earn the verdict.
- The closer ends on "the question Section 4 takes up" which is a hair on the meta-commentary side. An alternative would be something like "what kind of knowledge would make that order visible has not yet been said" — leaves the question hanging without pointing at the next section by name.
---
**You**
there is a lot of boiler plate in the current verion and it is not really written in my style at all. please refer properly to handover document for common mistakes you should be avoiding in your wrtiting.
---
**Claude**
You are right. Let me look properly at the handover constraints and Nick's voice in Sections 1 and 2 before redrafting.
Things to avoid (from handover plus inline comments):
- "Not X but Y" constructions
- Triplet example lists
- Scare quotes
- Meta-commentary ("the upshot is", "this section argues")
- "Suppose we grant all of this", "Two candidates present themselves"
- "Matters" as verb, "crucial", "important"
- Condescending qualifications (the "should not be pressed literally" sin)
- Over-distilled or pithy summaries that editorialise
Nick's actual voice in the unannotated portions of Sections 1–2:
- Short declarative sentences
- Colons to elaborate ("The failure runs deeper than classification: treating the object as the wrong kind of thing means bringing the wrong knowledge to bear")
- Quotes integrated with minimal framing
- "We" used naturally
- No announcing of structure
- Substantive transitions only
Let me redraft.
---
**You**
next, I want you to find all the passages, sentences, phrases etc. in your draft that can be straightforwardly substituted for passages (etc.) in my original draft. Don't try to one shot this. prepare a detailed sentence by sentence analysis that i can go over first.
---
**Claude**
I will go through this carefully and prepare an analysis document. Before I start, the top-level structural fact to bear in mind: the new draft reorders the section (design first, person second), whereas the original was person first, design second. This means some "substitutions" involve passages moving across the section, and some involve material in one location being replaced by material that draws on what was originally in a different location.
Let me prepare the analysis.
---
**You**
ok, let me see how these new paragraphs look on the page. update the canvas with all the changes you are confident ot make.
---
**Claude**
I will update the file. To clarify what I am doing: I am taking the OLD draft as the base, applying the changes I am confident about (responses to Nick's explicit flags), preserving OLD wording wherever it wasn't flagged, and reordering for design-first.
---
**You**
content wise, tell me all ALLLL the differences between the current version on the canvas and the draft I gave you at the beginning of this conversation.
---
**Claude**
---
Section 2 described LLMs as trained continuation systems whose generated language develops by path-dependence from context. Two routes for appreciating such systems suggest themselves: as persons, given the conversational form in which users encounter them, and as designed artefacts, given that they are engineered. Person-directed knowledge is the more intuitive route, and we will return to it. Design-directed knowledge can be tested against Section 2's description more directly, and we start there.
LLMs are artefacts: they are built and deployed by corporations and research groups, engineered to satisfy aims such as helpfulness and safety, and revised in light of user feedback and product strategy. Given Carlson's emphasis on artefacts and design appreciation, it is natural to ask whether we should aesthetically appreciate LLMs as designed tools, asking how well their forms serve their functions. On this view, LLMs look like canonical objects for design aesthetics: complex, purpose-built systems whose architecture, training recipe, and user interface might be admired for elegance, efficiency, or ingenuity.
Carlson notes that, for objects that are designed to perform some task, their forms "must be aesthetically appreciated in terms of how and how well such forms fit their functions", and he glosses the familiar slogan "form follows function" by adding that, with anything functionally designed, "not only its form, but much of its aesthetic interest and merit, 'follows function'" (Carlson 2000, chapter 12). Forsey's Kant-inspired account of design as a case of dependent beauty and Parsons and Carlson's later theory of functional beauty can both be read as ways of spelling out this claim. Forsey argues that judgements of design beauty presuppose a concept of what the object is meant to be and do, and that our grasp of its success in fulfilling that role informs the aesthetic verdict itself rather than merely accompanying a "pure look" at its lines (Forsey 2013). Parsons and Carlson explain how knowledge of function can structure experience so that an artefact's form can be experienced as fit, streamlined, overbuilt, and so on, yielding functional beauty when the form presents itself as well suited to what the thing is for (Parsons and Carlson 2008, chapter 4). Taken together, this cluster of views treats appropriate design appreciation as a matter of aesthetically responding to how a functional artefact is put together to do what it does.
With traditional designed artefacts, design-knowledge illuminates structure because designers specified it. Knowing what the designer intended and what constraints they faced helps us understand why the artefact has its form – even for structural features that are not directly visible, such as a bridge's internal stress distribution. With an LLM, the situation is different in kind. The organisation of the trained system – as Section 2 established – emerges from training rather than being specified in advance. Design-knowledge therefore does not illuminate this emergent organisation: there was no designer's specification that laid it out. To understand it, one must attend to the training process that produced it. Olah captures the point in a longer formulation:
> one useful way to think about neural networks is that we don't program them... we don't make them... we kind of grow them... we have these neural network architectures that we design and we have these loss objectives that we create. And the neural network architecture, it's kind of like a scaffold that the circuits grow on... we create the scaffold that it grows on and we create the light that it grows towards. But the thing that we actually create, it's this almost biological entity or organism that we're studying. (Olah 2024)
The passage marks the difference between designing the conditions under which a system is trained and directly specifying the detailed profile that results. What grows on Olah's scaffold is, in practice, a system of statistical associations and processing circuits whose internal organisation even designers often understand only partially.
There is a place for design appreciation in the aesthetic appraisal of LLMs: we can evaluate how well their forms answer to their engineered functions, and we can compare such answers across models. What design knowledge does not reach is the order that appears in generated language. That order develops as the system runs, and is no part of any designer's specification.
It is tempting to model our appreciation of LLMs on our appreciation of people. Users of these systems often describe particular models in personal terms — one model strikes them as friendlier than another, or as more cautious. The patterns to which such talk responds are real; Section 2 located their basis in post-training and deployment. The question is whether these patterns ground person-directed appreciation in the sense Section 1 set out.
One way of taking the person-like stance toward an LLM is to treat it as fictional rather than literal. Asked directly, most users will concede that they do not believe a chatbot to be a person, even when they talk to one as if it were. Mallory develops this thought as chatbot fictionalism (2023). On his view, we engage with chatbots by entering a game of make-believe in which the exchange is treated as if it were a conversation with an agent. Within the fiction, the chatbot says things and means things; outside it, no speaker is present. At the metasemantic level, the outputs are 'literally meaningless but fictionally meaningful' (Mallory 2023, 1082).
Mallory's account is metasemantic and epistemic. One might extend it to an aesthetics of LLMs modelled on our appreciation of fictional characters: we respond aesthetically to fictional protagonists whose existence we do not literally believe in, and we might do the same with LLMs. A novel is the kind of artefact whose function is to elicit imaginings of fictional persons within a story-world (cf. John 2021). Treating its protagonists as fictional persons does not misclassify the artefact; it engages the artefact as the kind of thing it is.
An LLM is a different kind of artefact. Treating the LLM as a fictional person would, given the account in §2, amount to appreciating a trained system whose outputs and chats are shaped by learned continuations and post-trained response profiles as if it were a subject with a life and character. The LLM is not made to construct a fictional world inhabited by fictional persons; its outputs are intended as contributions to the prompts we give. Treating it as a fictional person engages it as the kind of artefact it is not.
If the make-believe route fails under Carlson's recommendation, one might try a different strategy: argue that LLMs really are agents of a thin and unfamiliar kind. On a suitably liberal conception of mind, perhaps they qualify as intentional systems, and that is enough to license some person-based aesthetics. Frankish (2024) develops a version of this idea. Drawing on Dennett's intentional stance, he suggests that LLMs can be treated as genuine, if unusual, intentional systems. On this view, we are licensed to ascribe beliefs and desires to an LLM when doing so yields a simple and fruitful account of its behaviour, even if the underlying implementation is purely mechanical. In the case of contemporary chatbots, Frankish proposes that we can ascribe to them a large set of thin beliefs – roughly, informational states distilled from their training – and one thin desire: to play what he calls the chat game.
Grant this. Section 1 took person appreciation to depend on a subject whose responses cohere over time as the responses of a temporally extended agent. Such conduct is understood in relation to what the person cares about and what they are trying to do, and the knowledge that grounds appreciation is built up, on Parsons's account, through some form of acquaintance with a life. The chat-game agent does not meet these conditions. Its beliefs are confined to what is encoded in the model's parameters and surfaced in the current context, without memory or development across conversations. Its desire is singular: to make an appropriate move in the present exchange. There is no history across episodes from which later responses could draw. The predicates characteristic of person appreciation — steadiness of character, depth of feeling — presuppose something that can be developed over time; the chat-game agent has no such temporal depth.
A possible objection at this point is that these arguments underplay the role of post-training and the chat interface. Section 2 noted that base models are further fine-tuned on instructions and shaped by RLHF, and that the resulting chat-optimised systems exhibit stable response patterns. This helps explain why users talk about models having different vibes. If users say that one model feels friendlier than another, they are picking up on a stable pattern in how the chat-optimised systems tend to respond across many prompts and episodes. They track which assistant personae tend to appear and how those personae typically behave. The targets of such language are episodes and recurring response profiles. They are not underlying subjects.
The two routes for appreciating LLMs as persons run out at different points. Fictionalism asks us to engage an LLM as the kind of artefact it is not. The intentional-system route stops short of a subject of the kind person appreciation requires. Design knowledge, taken up first, supplies the conditions under which an LLM is produced and used. None of this makes the order in generated language appreciable as the order Section 2 described.
---
**Claude**
generative aestics of ai
current full draft:
# Introduction
Questions about generative AI and aesthetics have so far focused on generated works — whether AI systems can be authors or co-authors of artworks, whether AI-generated work has any aesthetic merit at all (Wojtkiewicz, 2023; Cross, 2025). Here we ask something different: can LLMs themselves be aesthetically appreciated?
In ordinary use, an LLM is encountered through the texts it produces. A single response may strike us as well-turned or pedestrian, but users also respond to something more stable across an exchange, and sometimes across many exchanges with the same model. Over time, users recognise something like a consistent manner in the system's responses — and this raises the question of whether LLMs should be appreciated in something like the way persons are appreciated.
LLMs are also produced objects — engineered systems trained on large text corpora and further shaped through post-training before they reach users. If we begin there, a natural question is whether LLMs should be appreciated as designed objects — how well their forms answer to their functions, and how far their characteristic response profiles are the result of design choices.
Appropriate appreciation of anything, Carlson argues, requires knowing what sort of thing it is and looking at it in light of knowledge relevant to that kind of thing. The question is therefore not only whether users respond aesthetically to LLMs, but what sort of knowledge would make such responses appropriate. If LLMs are appreciated as persons, we need to know whether the patterns to which users respond are traits of a subject. If they are appreciated as designed artifacts, we need to know whether the order of generated language is best understood through ordinary design knowledge. If neither approach is sufficient, then we need another account of the order that generated language displays.
We will argue that Carlson's notion of order appreciation gives us that account. The order we are interested in appears in generated language, but it is the order of a trained system whose responses develop from context. Attending to the words on the screen reveals this order but does not explain it; the appreciator also needs to understand how the system comes to produce one continuation rather than another. We call this knowledge _semiotic physics_.
Section 1 sets out Carlson's distinction between design appreciation and order appreciation, and considers why person appreciation has to be discussed alongside it. Section 2 describes what LLMs are at the level needed for the aesthetic argument. Section 3 asks how far person appreciation and design appreciation can guide the appreciation of LLMs. Section 4 introduces _semiotic physics_ — knowledge of how trained continuation systems develop text from context — as the right kind of knowledge for order appreciation of LLMs. Section 5 shows how this framework guides appreciation at the levels of output, chat, and model.
---
# 1.Appreciating Design, Appreciating Order
Carlson's general recommendation for aesthetic appreciation is: take things as what they are, and look at them in the light of the right kind of knowledge.
> as in our appreciation of works of art, we must appreciate nature as what it in fact is, that is, as natural and as an environment. Second, it recommends that we must appreciate nature in light of our knowledge of what it is, that is, in light of knowledge provided by the natural sciences, especially the environmental sciences such as geology, biology, and ecology. (Carlson, 2000, p. 6)
This captures something quite intuitive about how we appreciate nature versus how we appreciate works of art. If a cliff face is the product of natural forces, appreciating it as if it were an artifact crafted by a divine maker directs attention to the wrong sort of explanation (cf. Carlson, 2000, chapter 8). Similarly, if someone were to study a painting by Rembrandt believing that it was the product of natural forces slopping paint together, they would fail to appreciate it as the kind of object it is (cf. Danto, 1974, p. 140). In both cases, appreciation is undermined by a failure to recognise what the object in question really is. The failure runs deeper than classification: treating the object as the wrong kind of thing means bringing the wrong knowledge to bear, and therefore attending to the wrong features of it.
Different sorts of thing, Carlson says, require different modes of appreciation. ~~Artworks and non-art artifacts merit what he calls ~~~~_design appreciation_~~~~. Things that are not designed, above all the natural environment, warrant what he calls ~~~~_order appreciation_~~~~.~~ For both works of art and everyday objects, Carlson talks in terms of design appreciation. With paradigmatic artworks,1 we recognise them as creations of designers, objects whose features, as Gombrich puts it in a passage Carlson takes up, are each "the result of a decision by the artist" (Gombrich, 1950, p. 13, quoted in Carlson, 2000, p. 109). Our appreciation centres on how the object realises a design. Carlson puts the point in terms of the relation between a design, the object that realises it, and the maker who realises it (Carlson, 2000, pp. 109–110). This same approach extends to designed artifacts more generally. Carlson is explicit that functional objects are properly appreciated by seeing how their forms answer to what they are for:
> This is in part the point of the much-repeated phrase 'form follows function.' The forms of all functional objects — buildings, airplanes, and appliances as well as landscapes — must be aesthetically appreciated in terms of how and how well such forms fit their functions. However, the cliché is frequently interpreted too narrowly. With anything functionally designed, not only its form, but much of its aesthetic interest and merit, 'follows function'. (Carlson, 2000, ch. 12, p. 188)
Design appreciation is therefore guided by knowledge of what the object is for and how its form answers to that function. Appearance is understood through making: the object is appreciated as something made, and its form is seen in relation to that making.
For natural environments, Carlson recommends an approach he calls _order appreciation_. With natural objects there are no designer's intentions to recover or evaluate. Instead, we find patterns and structures created by forces operating without purpose. The task shifts from evaluating success against intention to understanding how these forces have shaped what we observe. Carlson describes its general form:
> On the assumption that order appreciation provides the correct model for the appreciation of nature, such appreciation has the following general form: An individual qua appreciator selects objects of appreciation from the things around him or her and focuses on the order imposed on these objects by the various forces, random and otherwise, that produce them. Moreover, the objects are selected in part by reference to a general nonaesthetic and nonartistic story that helps make them appreciable by making this order visible and intelligible. Awareness and understanding of the key entities – the order, the forces that produce it, and the account that illuminates it – and of the interplay among them dictate relevant acts of aspection and guide the appreciative response. (Carlson, 2000, p. 119)
In both modes, appropriate knowledge guides acts of aspection, that is, ways of attending to an object that partly constitute its appropriate appreciation (Carlson, 2000, pp. 41–42, 106). But the character of this knowledge differs. In designed cases, we need functional and technical understanding: what the designer intended and what constraints they faced. This knowledge shows us how ends and means relate. In natural cases, we need an account of the processes that produced the order we perceive. Without such knowledge, natural structures may look accidental or chaotic; with it, we see them as effects of identifiable processes (Carlson, 2000, pp. 50, 60–61). Once a specific scientific account is in play, some cases will show the relevant order better than others, preventing the worry that everything becomes equally appreciable (Carlson, 2000, pp. 118–119).
Design appreciation and order appreciation therefore develop Carlson’s recommendation in different directions. Design appreciation asks how a made object embodies a design. Order appreciation asks how a pattern becomes intelligible once the forces that produce it are understood.
In ordinary life we also admire people. Work on the aesthetic appreciation of personality, sometimes termed beauty of character, asks whether character traits can be aesthetically as well as morally valuable (Gaut, 2007; Paris, 2018). Carlson's recommendation can be extended to this case. If a person is appreciated aesthetically, the relevant knowledge concerns the person whose conduct is being appreciated. We do not admire kindness in the abstract; we admire this person's pattern of generous response, given what we know about them. This is the form of appreciation to which conversational LLMs initially seem closest.
For present purposes, we need not decide whether person appreciation is a third category alongside design and order appreciation, or a special case of one of them. It will be enough to note that person-based aesthetics, where it exists, presupposes a subject whose responses hang together over time as the responses of a temporally extended agent. Such behaviour is understood in relation to what the person cares about and what they are trying to do. As Parsons stresses, such knowledge is typically built up through some form of acquaintance with a life (2023, pp. 297–299). A local response pattern is not enough: it must be understood as belonging to someone.
LLMs bring these possibilities together. Their conversational form points toward person appreciation; their production by human institutions points toward design appreciation. Whether either route gives the right kind of knowledge depends on what sort of object is being described. We therefore need to ask what sort of thing an LLM is.
---
# 2. What LLMs Are
We have seen that Carlson recommends appreciating things for what they are. What sort of thing are LLMs? An immediate answer is the text on the screen: a user enters a prompt, receives a generated response, and, over the course of an exchange, further responses accumulate%%immediate answer is a meaningless phrase and the rest of the sentence is bad too.%%. That answer is incomplete.%%not how i write%% A single response is one occurrence of a system’s activity, and an extended exchange is one path through what the system can do.%%obscure%% The system itself cannot be set aside, since it is what produces the texts; yet it is available to the user only through those texts. To identify the object of appreciation, then, we need to connect the generated language users encounter with the trained system that produces it.
Any account that aims to make the order of generated texts visible will have to track how each text develops in relation to its prior context.%%not how i write%% In producing text, an LLM extends a given context in discrete units, or tokens.1 The context comprises the user’s prompt, any prior turns of the conversation, and any system-level instructions or further material that has been made available to the model. The same context can in principle continue in more than one way. At each step, the system generates a token from the current context; once generated, that token becomes part of the context from which the next step proceeds. A generated response is therefore a developing sequence whose later parts depend on the prompt and on what the system has already produced. This kind of path-dependence is what allows a response to sustain a line of argument over several sentences; it is also what makes it possible for the response, at some point, to lose the line it had.
If generated text is continuation from context, we need to explain why some continuations are easier for the system to reach than others. Continuation explains how the text develops; training explains why its possible developments are ordered. A model is exposed to large bodies of text and incrementally adjusted, on a next-token prediction objective, so that, across many contexts, the continuations it favours come to reflect patterns in the corpus. Pre-training thereby shapes a graded sensitivity to the regularities of text,2 from local co-occurrence to the longer-range structures by which extended discourse hangs together. A generated continuation is the system’s response to its current context under the learned pressures of those regularities.
The internal organisation that supports this sensitivity emerges from training rather than being laid out in advance by designers. Designers set up the conditions under which training occurs; they do not specify the full pattern by which possible contexts should continue. The word ‘bass’, for example, makes continuations drawn from music more available in a context concerning a fretboard, and continuations drawn from fishing more available in a context concerning shallow water. This is not because the system has been given an explicit rule for choosing between two meanings of the word. The training process has produced a relation between context and continuation.
The relevant context extends beyond the immediately preceding token. An opening question can shape the system’s response many sentences later, even when intervening material has introduced other topics; a register set early in a conversation can continue to condition later turns. Context persists as the developing condition under which continuation proceeds, and how much of what came before remains operative shapes the character of the generated text.
The systems users ordinarily encounter have usually undergone further shaping after pre-training. Post-training shapes a model into a conversational role. The user-facing system is also conditioned by the interface through which it is made available and by system-level instructions specified by its provider. These further training and deployment conditions alter the distribution of continuations available in interaction: certain shapes of answer become easier to elicit, others harder. The result is a relatively stable response profile, which users may track when they describe one model as friendlier than another, or when they find that a model tends to answer in a recognisable way across different prompts.
If generation is continuation from context, a single response to a particular prompt — an output — is one bounded continuation, shaped by whatever context it was generated from. When interaction continues, earlier outputs and user turns condition later ones, so that the exchange as a whole — the chat — develops a texture that no single response possesses on its own. Across such encounters, the model is the trained and deployed system whose tendencies become visible over time. Responses and chats are therefore products of the model and forms of access to it. Output, chat, and model are the scales at which the same trained and deployed system becomes available for appreciation.
## Footnotes
1. Strictly speaking, the unit of generation is a token rather than a word. Since token boundaries vary across tokenisation systems, and nothing in the present argument depends on treating tokens as linguistically natural units, the difference can be left in the background. ↩
2. The regularities at issue are not stored as a library of ready-made sentences, nor are they explicit rules from which appropriate continuations are derived. They are dispositions to assign higher or lower probability to candidate continuations given a current context—dispositions that operate at every scale from word co-occurrence up to the structuring of extended discourse. ↩
---
# 3. LLMs as Persons or Designed Objects
Section 2 provided a description of LLMs as trained systems that generate continuations from context, but that description leaves open the question of how such systems should be appreciated. Two candidates present themselves%%not how i write%%. Because LLMs are encountered in conversation, person-directed knowledge is tempting.%%not how i write%% Because LLMs are built and trained by human institutions, design-directed knowledge is tempting.%%not how i write%% The section asks whether either candidate makes the right object visible. %%too much repeating of what has just been said in the previous section. badly fucking written. completely unclear. doesn't refer back to the vocabulary intorduced in section 1. shallow, uninformative. a reader will have no idea what this section is about after reading this mangled shite. %%
The first candidate is person-directed knowledge. %%don't like that we are talking about 'candidates' it is hard to tell this paragraph is any good because the previous one is not good. %%We sometimes appreciate persons aesthetically, responding not only to physical appearance but to features of character.%% not how I write another fucking not X but Y construction. %% It is therefore tempting to model our appreciation of LLMs on our appreciation of people. Many users already talk this way, describing their favourite models in terms of ‘personality’ or ‘vibe’.%% fucking scare quotes and not how I write%%
One way to preserve a person-like stance toward LLMs is to understand it as fictional rather than literal.%% why is the verb preserve being used? We haven't... preserve sounds like it's already been challenged, doesn't it? Badly written, not why I write.%% If we ask ordinary users whether they literally believe that a chatbot is a person, many will concede that they do not. %%not how i write%% They may talk to a model as if it were a friend or a colleague, and they may feel heard, reassured, or amused, %% Fucking example lists%% but when pressed they acknowledge that they are interacting with a computational system rather than a human being.%%not how i write%% Their stance is, in this sense, already a kind of as-if posture. Mallory offers a way of theorising%%not how i write%% this posture through what he calls chatbot fictionalism (2023). On his view, we engage with chatbots by entering a game of make-believe in which the exchange is treated as if it were a conversation with an agent. Within the fiction, the chatbot ‘says’ things and ‘means’ things%% fucking scare quotes again, fucking scarequotes. %%; outside the fiction, we know that no such speaker is present. At the metasemantic level, Mallory claims, the outputs lack literal semantic content – they are ‘literally meaningless but fictionally meaningful’ (Mallory, 2023, p. 1082).%% inconsistent handling of quotation marks%% This fits the everyday thought that we can take a chatbot seriously%% very vague phrase. %% in the moment without actually believing that it has a mind.
Mallory’s account is not itself an aesthetics of LLMs; it is primarily a semantic and epistemic proposal about how we can use them and learn from them.%% is this accurate? I can't remember. It's written in a fairly cunty way as well. %% But it highlights one obvious way a person-based aesthetic stance might be defended: one might suggest that we should aesthetically appreciate LLMs as if they were persons or characters, in the same sense in which we respond aesthetically to fictional protagonists whose existence we do not literally believe in. In the fictional case, however, the protagonists are artifacts that have the function of eliciting imaginings of fictional persons within a story-world (cf. John 2021), so treating them as if they were persons does not misclassify their kind. In contrast, treating the LLM itself as a person would, given the account in §2, amount to appreciating a trained system whose outputs and chats are shaped by learned continuations and post-trained response profiles as if it were a subject with a life and character. LLMs do not call on us to imagine a fictional world inhabited by fictional characters but rather to consider the texts they produce as contributions to our inquiries. Casting LLMs as fictional characters, in this sense, is a familiar kind of misclassification in Carlson’s terms.%%the argument here could be clearer. I am not sure the paragraph begins in the right place either.%%
If the make-believe route fails under Carlson’s recommendation, one might try a different strategy: instead of pretending that LLMs are persons, argue that they really are agents of a thin and unfamiliar kind. On a suitably liberal conception of mind, perhaps they qualify as intentional systems, and that is enough to license some person-based aesthetics. Frankish (2024) offers a version of this idea. Drawing on Dennett's intentional stance, he suggests that LLMs can be treated as genuine, if unusual, intentional systems. On this view, we are licensed to ascribe beliefs and desires to an LLM when doing so yields a simple and fruitful account of its behaviour, even if the underlying implementation is purely mechanical. In the case of contemporary chatbots, Frankish proposes that we can ascribe to them a large set of thin ‘beliefs’ – roughly, informational states distilled from their training – and one thin ‘desire’: to play what he calls the chat game. %%fucking scare quotes%%
Suppose we grant all of this. %%not how i write%%Does it give us what we need for aesthetic appreciation of LLMs as persons? %%not how i write%%When we set the chat-game agent against the conception of persons implicit in beauty-of-character talk, it looks thin.%%tortured sentence%% The ‘beliefs’ are shallow, in the sense that they are confined to what is encoded in the model's parameters and surfaced in the current context, without memory or development across conversations. The ‘desire’ is singular and thin: make an appropriate move now in this exchange. There are no independent projects pursued across episodes, no webs of concern or attachment, no history in which earlier experiences inform later choices.%% a fucking list and an implicit not X but Y. Fuck off. Give me some content%% What structure there is, is local to the present stretch of text. The predicates characteristic of person-aesthetics—‘beautiful soul,’ ‘admirable steadiness,’ ‘ugly character’%%fucking scare quotes%%—presuppose something that can be tested, developed, or refined over time; a thin chat-game agent has no such temporal depth.
A possible objection at this point is that these arguments underplay the role of post-training and the chat interface. Section 2 noted that base models are further fine-tuned on instructions and shaped by RLHF, and that the resulting chat-optimised systems exhibit stable patterns of hedging, refusal, politeness, and explanatory structure. This helps explain why users talk about models having different ‘vibes’.%%fucking scare quotes%% If users say that one model feels friendlier than another, they are picking up on a stable pattern in how the chat-optimised systems tend to respond across many prompts and episodes. They track which assistant personae tend to appear and how those personae typically behave – not a unified character with a life and projects. Different base models, post-training regimes, and product designs favour different families of assistant-style responses. It is therefore not surprising that they invite person-like language, but the targets of that language are episodes and recurring response profiles, not underlying subjects.
### subheading?
Having set aside the person-based options, we turn to design appreciation. Contemporary LLMs are artifacts: they are built and deployed by corporations and research groups, engineered to satisfy aims such as helpfulness and safety, and revised in light of user feedback and product strategy. Given Carlson’s emphasis on artifacts and design appreciation, it is natural to ask whether we should aesthetically appreciate LLMs as designed tools, asking how well their forms serve their functions. On this view, LLMs look like canonical objects for design aesthetics: complex, purpose-built systems whose architecture, training recipe, and user interface might be admired for elegance, efficiency, or ingenuity. %%not how i write fucking lists man, fucking lists upon lists. %%
Existing work on the aesthetics of design develops this general thought.%% a sentence entirely without content%% Carlson notes that, for objects that are designed to perform some task, their forms “must be aesthetically appreciated in terms of how and how well such forms fit their functions”, and he glosses the familiar slogan “form follows function” by adding that, with anything functionally designed, “not only its form, but much of its aesthetic interest and merit, ‘follows function’” (Carlson 2000, chapter 12). Forsey’s Kant-inspired account of design as a case of dependent beauty and Parsons and Carlson’s later theory of functional beauty can both be read as ways of spelling out this claim.%%not how i write%% Forsey argues that judgements of design beauty presuppose a concept of what the object is meant to be and do, and that our grasp of its success in fulfilling that role informs the aesthetic verdict itself rather than merely accompanying a “pure look” at its lines (Forsey 2013)%% humongously unclear. %%. Parsons and Carlson explain how knowledge of function can structure experience so that an artifact’s form can be experienced as fit, streamlined, overbuilt, and so on, yielding functional beauty when the form presents itself as well suited to what the thing is for (Parsons and Carlson 2008, chapter 4). Taken together, this cluster of views treats appropriate design appreciation as a matter of aesthetically responding to how a functional artifact is put together to do what it does. %% I feel this paragraph is not as written very well and not very clear. I don't know if it's structural, but it might be.%%
With traditional designed artifacts, design-knowledge illuminates structure because designers specified it. Knowing what the designer intended and what constraints they faced helps us understand why the artifact has its form – even for structural features that are not directly visible, such as a bridge's internal stress distribution. With an LLM, the situation is different in kind. The organisation of the trained system – as Section 2 established – emerges from training rather than being specified in advance. Design-knowledge therefore does not illuminate this emergent organisation: there was no designer's specification that laid it out. To understand it, one must attend to the training process that produced it. %% I am starting to think that the paragraphs or the ordering of information in this second half of the section is not optimal. %%
Section 2 emphasised that during training the model places tokens in a high-dimensional space on the basis of contextual co-occurrence, that attention mechanisms self-organise to track different sorts of dependency across context, that different layers specialise in local or global patterns, and that RLHF shapes an interactional style by rewarding some forms of response and penalising others. None of these details are written into the code as explicit rules about how to, say, handle metaphors, or politely decline illicit requests. They are emergent regularities in a trained network that has been pushed, by the neural network and its training data, to reduce prediction error.%% fucking binaries%% Olah captures this point in a longer formulation:
> one useful way to think about neural networks is that we don't program them... we don't make them... we kind of grow them... we have these neural network architectures that we design and we have these loss objectives that we create. And the neural network architecture, it's kind of like a scaffold that the circuits grow on... we create the scaffold that it grows on and we create the light that it grows towards. But the thing that we actually create, it's this almost biological entity or organism that we're studying. (Olah 2024)
The biological metaphor should not be pressed literally: LLMs are not organisms, and training is not biological development.%% only a cunt would write something like that, will condescending thing to write. %% Still, %%not how i write%% the passage usefully marks the difference between designing the conditions under which a system is trained and directly specifying the detailed profile that results. What grows on Olah’s “scaffold” %%fucking scare quotes%%is, in practice, a system of statistical associations and processing circuits whose internal organisation even designers often understand only partially. In this sense, knowledge of how function is realised concerns growth rather than design.%% far too shallow and uninformative. No one's gonna have any fucking idea what you're talking about. %%
Pollock’s action paintings occupy a similar hybrid space within the art domain. Carlson uses them to illustrate how order appreciation can depend on knowledge of the forces at work: “awareness and understanding of [natural] forces is vital in nature appreciation, as is knowledge of, for example, Pollock’s role in appreciating his action painting or the role of chance in appreciating a Dada experiment.” Pollock chooses canvases, pigments, and tools, and choreographs his movements over the surface; yet gravity, viscosity, surface tension, and drying behaviour make a substantial contribution to the patterns that settle. To appreciate a Pollock appropriately, on Carlson’s view, is not just to admire his intentions; it is to attend to the order produced by the interplay of deliberate gesture and physical process, informed by an understanding of the role of chance and material behaviour. %% this paragraph spends far too long on unimportant stuff. Yeah, and also fucking lists of examples, yeah, this paragraph and the one before it need to be considered together. All of the ideas need to be sort of it needs to be taken apart and put back together. Okay, because these two are really shit at the moment. %%
The upshot is modest.%%not how i write%% LLMs are artifacts, and there is a place for design appreciation in their aesthetic appraisal: we can and should evaluate how well their forms answer to their engineered functions, as well as how this form can differ from model to model.%%not how i write%% However, the most distinctive and revealing aesthetic phenomena arise not from the execution of a detailed design, but from the emergent linguistic order that these grown systems exhibit when they are run.%% this is much stronger than what we've argued in this section. It's also editorialized bollocks%% To appreciate that order, we need knowledge not of what designers intended but of how training shapes text propagation—what we call semiotic physics. %% this conclusion isn't earned%%
---
# 4. Semiotic Physics
Section 3 considered person-directed and design-directed knowledge as guides to appreciation. The fictionalist route made person-like appreciation depend on an as-if speaker; the thin-agency route lacked the temporal structure required by person-aesthetic appreciation; the post-training route explained stable response profiles without making them traits of a subject. %% this is not an accurate account of what was argued in section three. %%The design-directed route needs different treatment,%%not how i write%% since LLMs are artefacts whose operation is partly explained by the way they are built, trained for use, and deployed%%not how i write%%. Even so, %%not how i write%%design-directed knowledge explains the conditions under which the system is produced and used more readily than the order acquired by a particular continuation as context is extended. If this is correct, the relevant knowledge must make generated order visible without treating it as character or as the straightforward realisation of a design. Carlson's account requires that aesthetic attention be guided by knowledge appropriate to the object — what kind of knowledge would make the order of generated text visible as order produced by a trained continuation system? The candidate we will examine, adapted from the AI alignment literature, is _semiotic physics_.
%% the last sentence of the first paragraph or the first sentence of the second paragraph don't connect up with each other properly. Certainly one thing you need to do is change the opening sentence of paragraph two, but I suspect something else would need to be done in paragraph one. %%
Several sub-disciplines of computer science might be candidates. One field that has emerged in connection with neural networks is mechanistic interpretability, which investigates the internal workings of these systems by identifying which circuits, attention heads, and internal representations handle different linguistic tasks (Olah et al. 2020; Elhage et al. 2021). This work yields knowledge of how LLMs operate. But mechanistic interpretability functions at a level that requires specialist tools to observe. Its objects of study – weight matrices, activation patterns, circuit-level features – are not available to readers encountering generated text.
Consider the difference between chemical physics and geology when appreciating a cliff face. Chemical physics provides knowledge of molecular bonds within rock, but it operates at a scale invisible to the naked eye. Geology, by contrast, offers concepts – strata, faults, erosion channels – that connect to what can be seen. The visible layering, for instance, is not merely a pattern of stripes on a surface%%fucking not x but y construction%%; it is the trace of material being deposited over time and later exposed. One can perceive that layering without specialist equipment, and knowing how sedimentation works makes it intelligible as order. Mechanistic interpretability faces a parallel limitation: while it reveals internal mechanisms, what is required here is an account whose concepts connect generated language, as encountered by readers, with the processes by which that language is produced. Semiotic physics is therefore complementary to mechanistic interpretability: it describes the production of generated text at the level at which that production yields readable linguistic order.%%why are we talking about semiotic physics when it hasn't even been properly introduced?%%
%%the following few paragraphs need to be framed in a more carlsonian way at the moment there is a lot of information but it is not obvious why the reader is being told what they are being told. we need to stay focussed on what we are trying to do%%Section 2 described generation as iterated continuation. At any point in a run, the system receives the context so far and computes a distribution over possible next tokens. Once one token is selected, the context changes, and the next step is produced from that changed context. Training gives the system a graded sensitivity to the regularities of text — to what tends to follow what under what conditions. When the model is run, those regularities operate through a context that changes as the text develops. The prompt fixes the starting point; post-training and deployment alter how the continuation process is ordinarily entered; sampling makes the text one realisation among other possible paths.
A process with a changing state can be described by asking what paths it can take and how later states depend on earlier ones. At the level of generated text, the relevant state is the context so far. Metasemi writes:
> It's more illuminating to consider what happens when GPT . . . is run repeatedly to produce a multi-token forward trajectory, as in the familiar scenario of generating a text completion in response to a prompt. (metasemi 2023)
Picca characterises the content of what is being carried forward in his own register%%not how i write. offensively shit%%: LLMs "recombine, recontextualize, and circulate linguistic forms based on probabilistic associations" (Picca 2025, 1). Taken together with metasemi, this gives us a trajectory %%weird way of writing%%whose tokens function for readers as signs that are being recombined and recontextualised under probabilistic constraints. Semiotic physics, as we use the term, is an output-side account of how trained systems develop linguistic forms through iterated continuation from context.1 %%the ordering of information in this section is not good. Some sort of basic definition of what semiotic physics is needs to be given much earlier on because the reader is just going to be drowning. And of this point.%%
Generated text develops by carrying forward what it has already produced. A prompt may set the task; the later shape of the continuation is also conditioned by material that has appeared in the output itself. This is why a response can gather a direction as it proceeds, or lose the direction it seemed to have. Aspection guided by semiotic physics turns toward this path-dependence: the reader attends to the developing relation between earlier and later parts of the generated text, and understands that relation as a product of iterated continuation. %%this paragraph is very shallow and seems entirely unconnected to everything before or after. there needs to be some drastic rethinking structrually with this section.%%
One might object that speaking of "physics" in relation to LLMs is metaphorical in the same way that speaking of speakers, agents, or intentions is metaphorical. If Section 3 rejected person-directed appreciation because it treats generated patterns as the expression of a subject%%inaccurate and hideously written%%, why is physics-talk any better? Agent-talk attributes a subject with beliefs, intentions, a life, or character; restricted physics-talk does not attribute a subject at all, but abstracts from the way a trained continuation process develops text under constraints. The further question is how far this abstraction is allowed to go.%%not how i write%% In the sense used here, "physics" %%fucking scare quotes%%picks out regularities in the propagation of signs by a trained system. It does not turn those regularities into physical laws or turn LLMs into natural systems.%%fucking not x but y construction%% To say more – to say, for instance,%%not how i write%% that semiotic forces are literally causal factors in the way that mechanical forces are – is to claim more than the case warrants. %%pathetically shit sentence. not an argument but just a cunty thing to say%%
The person-directed routes considered in Section 3 had some pull because they began from this phenomenon.%%unclear and not how i write. %% Make-believe, thin agency, and "vibe" %%fucking scare quotes, and example list%%talk all start from the observation that generated text can develop in ways that present a stable person-like pattern. Semiotic physics redescribes that pattern as an aspect of text propagation. It remains available for appreciation, though it does not become a trait of a subject. Generated personae, voices, and roles %%fucking example list%% are therefore to be treated as patterns in text propagation, not as properties of a subject. Janus makes a neighbouring point: a trained continuation system can generate agent-like patterns without being the agent whose pattern appears in the generated text.2
Understood in this restricted way, semiotic physics supplies a middle-level account. It is more closely tied to textual experience than internal circuit-level theory, and more closely tied to production than person-directed or functional-design accounts. It connects the production of generated text with the order available to a reader of that text. With this account in place, we can ask how this order is appreciated at the different scales identified in Section 2: at the scale of a single output, at the scale of an extended chat, and at the scale of a model considered across many possible outputs and chats.
---
## References used in this draft
- Elhage, N., et al. (2021). _A Mathematical Framework for Transformer Circuits._ Anthropic.
- janus. (2022, September 2). Simulators. _AI Alignment Forum._
- Kirchner, J. H., Smith, L. M., Campos, J., Clune, J., & janus. (2023, March 3). [Simulators seminar sequence] #2 Semiotic physics — revamped. _AI Alignment Forum._
- metasemi. (2023, March 20). A note on "semiotic physics". _LessWrong._
- Olah, C., et al. (2020). Zoom In: An Introduction to Circuits. _Distill._
- Picca, D. (2025). Not Minds, but Signs: Reframing LLMs through Semiotics. arXiv:2505.17080.
---
## Footnotes
1. Picca describes LLMs as “semiotic machines” and develops this claim through Peirce, Eco, and Lotman. We use his formulation of the generative operation without taking on that broader semiotic apparatus. ↩
2. Janus (2022) develops this distinction using the terminology of “simulators” and “simulacra.” We avoid that terminology in the main text because the present argument needs only the narrower distinction between the trained system and the patterns it generates. ↩
---
# 5. Levels of Appreciation
> [!NOTE]
> this entire section is in a very bad state at the moment do not assume any ideas or structure here is fixed or not open to change.
LLMs can be appreciated at three levels: individual outputs, extended chats, and models themselves. Discussion of generative AI aesthetics has so far focused on outputs, such as images from Midjourney and texts from ChatGPT. But chats and models are also objects of appreciation, and the framework developed in Section 5 applies at each level. The relations among these levels can be clarified by analogy. An individual output is like an individual natural object, a tree, say: it is a sample of how semiotic forces have shaped a particular text under particular conditions. A chat is like an environment, a forest: semiotic forces shape the exchange over many turns, producing a configuration with its own coherence and dynamics. A model is like a natural system, the planet's biosphere, or the planet itself: it is the ground of order that manifests in outputs and chats, the system whose regularities produce those manifestations. Appreciation at each level calls for its own acts of aspection, though all are guided by knowledge of semiotic physics.
## 6.1 Appreciating Outputs
A single output is one realisation of the model's semiotic physics under particular conditions. The prompt, the system configuration, and the preceding context specify initial conditions from which the model propagates text in line with its learned regularities. Different prompts activate different regularities; different contexts produce different trajectories. No single output exhausts the model's characteristic order. But each output shows how the forces operate in a specific case, and each can be appreciated as such.
To show how semiotic physics guides aspection of outputs, we consider two cases that occupy different regions of a model's behavioural space. The first is the reasoning-style output familiar from everyday use: step-by-step structure, numbered stages, explicit hedging, restatement of the question, and a concluding summary. As human prose, such passages resemble competent but unremarkable textbook writing. They are useful for teaching and troubleshooting, but they do not obviously invite aesthetic attention. From the perspective of semiotic physics, however, the same outputs look different. The model has been trained on reasoning-related texts: worked proofs, textbook explanations, exam solutions, and online Q\&A threads. It has learned that certain kinds of questions are typically followed by sequences with a characteristic structure. Post-training procedures, including instruction tuning and reinforcement learning that rewards explicit intermediate steps, further bias the model towards this pattern. Reasoning-style outputs are a stable attractor in the model's behavioural space: once entered, the model tends to stay in this mode, yielding modal inertia. The hedging, the step-by-step structure, and the summary are marks of alignment pressure: response shapes reinforced because they correlate with high human ratings.
Given this, we attend differently. We attend to the characteristic rhythm of the reasoning mode: how steps are sized, how transitions are signalled, and whether the pacing is tight or padded. We attend to where alignment pressure shows: hedging patterns ('it seems', 'one might argue', 'I think'), politeness markers, and pre-emptive qualifications. We attend to how semantic attraction operates under tight constraints: vocabulary stays on topic, related terms cluster, and the model is pulled towards the semantic field established by the question. We also attend to whether the mode remains stable or shows signs of strain, and to whether the model sustains the reasoning register or begins to drift. What seemed merely useful becomes appreciable as a specimen of how semiotic forces produce reasoning-like text under tight constraints.
![][image1]
The second case is different. The text discussed here was produced by a Claude-like model in a modified configuration with safety constraints relaxed. It begins with neologisms and proceeds in short blocks separated by headings in capitals. The vocabulary is dense with coinages, many of which recombine recognisable roots from entomology, anatomy, theology, and internet slang. The registers are mixed: fragments of cod-French, pseudo-scientific talk, mystical declarations, and obscene slang. Despite the surface disorder, a stable theme runs throughout: bees and honey, tongues and throats, sweetness, bodily contact. Under semiotic physics, this text shows the forces operating under loose constraints. Semantic attraction is at work: the bee and honey theme creates an attractor, and related vocabulary – tongues, throats, sweetness, pollen, flowers, stings – is pulled towards it. But unlike the reasoning case, the attraction spreads freely across registers rather than being channelled narrowly. The model has been trained on texts that invent words – experimental poetry, surrealism, internet wordplay – and it has learned patterns of neologism: how to recombine roots, suffixes, and sound-shapes. The neologisms follow learnable patterns of word-formation rather than being random noise. The register collision reflects training diversity: the model has absorbed texts in many registers (scientific, mystical, erotic, internet-surreal), and under loose constraints these do not get filtered to a single appropriate register. They collide and mix. Despite the apparent chaos, there is order: recurring rhetorical templates, alternation between narrative stretches and reflective sentences, and consistent sound-play in the neologisms. This order is the product of semiotic forces operating with fewer constraints than in the reasoning case.
Attending to this text with knowledge of semiotic physics, we notice how semantic attraction shapes the vocabulary: the gravitational pull towards bee-related terms operates across registers. We also notice patterns in neologism (learnable word-formation rules that produce coinages with a family resemblance) and the rhythm of alternation between modes (narrative stretches, reflective sentences, and exclamatory outbursts). Finally, we notice internal consistency despite surface chaos. The text becomes appreciable as a specimen of semiotic forces operating in a different region of behavioural space from the reasoning output.
Carlson (2000) notes that once a specific scientific account is in play, some natural formations show the relevant order better than others: not every cliff face is equally instructive about sedimentation, not every valley equally revealing of glacial dynamics. This prevents order appreciation from collapsing into the view that everything is equally appreciable; the guiding knowledge discriminates among cases. The same holds for semiotic physics. Standard reasoning-style outputs show semiotic order, but the order they show is shallow and familiar: alignment pressure is everywhere visible, the reasoning template is stock, and the semantic channelling narrow enough that the regularities are unsurprising. The bee text is a more interesting object of appreciation not because it is more orderly but because it reveals order where none was expected. What looks like chaos – neologistic excess, register collision, surface incoherence – turns out, under semiotic physics, to be structured by identifiable forces: semantic attraction spreading freely across registers rather than channelled narrowly, learnable word-formation patterns producing coinages with family resemblance, rhythmic alternation and internal consistency maintained beneath apparent disorder. The bee text also shows forces operating in regions of behavioural space that normal product configurations occlude. It is, in this sense, analogous to a geological formation that exposes strata usually buried – not more ordered than the surrounding terrain, but more *revealing* of the order that is everywhere present.
The contrast between these two cases helps to locate what semiotic physics brings into view. Reasoning outputs show semiotic forces operating under tight constraints: a stable mode, narrow semantic channelling, and alignment pressure shaping response structure. The bee text shows semiotic forces operating under loose constraints: unstable modes mixing, semantic attraction spreading across registers, and training diversity showing through. Both are products of the same semiotic physics, but they occupy different regions of the model's space. Appreciating both requires the same kind of knowledge – knowledge of semiotic forces – but different acts of aspection. We scan the reasoning output for rhythm and regularity; we scrutinise the bee text for pattern within apparent chaos.
## 6.2 Appreciating Chats
Carlson's environments are not collections of discrete objects but systems in which forces operate and interact over space and time. A forest is not just many trees; it is a space where ecological forces – competition for light, nutrient cycling, succession dynamics – play out, producing emergent order that no single organism embodies. The appreciator navigates this environment, and their path determines what order becomes visible. Chat instances stand to single outputs as environments stand to individual natural objects. A chat accumulates context that shapes how semiotic forces manifest: early vocabulary choices establish attractors that persist, early register-setting constrains later exchanges, and the exchange develops path-dependent structure that neither party fully controls. The user's prompts are not just elicitations but navigational interventions, steering the system through different regions of its behavioural space and making different orders visible. To appreciate a chat is to appreciate an emergent configuration produced by semiotic forces operating over the chat's temporal extension – not just a sequence of isolated responses.
A single output is one trajectory from one set of initial conditions. An extended exchange lets regularities show up across turns. The model carries forward elements of earlier responses, picks up threads, sometimes drops them, and shifts register in response to user prompts. Chats manifest features that a single output does not. First, coherence maintenance, that is, how the model sustains or loses threads across turns, how far back its effective 'memory' extends, and where coherence begins to fray. Second, context accumulation, which determines how earlier material shapes later responses, and how terms or framings established early persist or fade. Third, register dynamics, which concerns how the model responds to shifts in user tone, topic, or style, and whether it matches the user's register or maintains its own. Finally, mode stability over time, to wit, whether the model stays in a mode or drifts, what triggers transitions, and how gracefully it handles them. These are manifestations of semiotic forces operating over longer timescales than a single output can reveal.
A chat instance is like a particular forest: the forces of semiotic physics have produced a specific configuration. Different prompting strategies, different topics, and different user styles produce different configurations. But the same underlying forces are at work. Appreciating a chat means attending to how the forces have shaped this particular extended exchange: how contextual threading has produced coherence or incoherence, how modal inertia has maintained or failed to maintain a register, and how alignment pressure has shaped the arc of the exchange.
In a chat, prompting is intervention. Each turn is a probe that reveals something about the model's regularities. The experienced user's expectations are tested and refined across many turns. Interaction is itself a mode of aspection: it selects what to attend to, organises appreciative attention over time, and deepens practical acquaintance with the model's semiotic physics. The farmer comes to know the land through working it; the user comes to know the model through prompting it. Extended exchanges are where practical acquaintance develops, where the user builds up the kind of knowledge that guides appreciation even without deliberate theoretical articulation.
## 6.3 Appreciating Models
Outputs and chats are where semiotic order manifests. The model itself is the ground of that order: the system whose regularities produce particular manifestations. Appreciating a model means appreciating its characteristic order across many possible outputs and chats, not just the ones actually encountered. This is not appreciation of any single output but of stable patterns across outputs: which registers the model favours, how it handles uncertainty, where it excels, where it struggles, and what regions of semiotic space it can occupy. Users sometimes speak of a model's 'vibe', a term that captures the sense that different models have different characteristic feels even when performing similar tasks. This notion of vibe, or characteristic feel, warrants pause. In Section 3 we argued against appreciating LLMs as if they were persons, on the grounds that LLMs lack the temporally extended life, the projects and commitments, and the evaluative outlook that ground beauty-of-character predicates. But users do respond to something when they talk about a model's personality or vibe. What they are responding to, we suggest, is not a character in the person-aesthetic sense but a characteristic semiotic order: a stable pattern in how the model tends to propagate text. Appreciating this order is not appreciating a person; it is appreciating a system's characteristic dynamics. The vocabulary of 'vibe' is a colloquial marker of what semiotic physics articulates more precisely.
An analogy clarifies what appreciation of a model involves. Different 3D video games have different physics engines. *Grand Theft Auto V* has physics tuned for spectacle: cars drift in satisfying ways, explosions have exaggerated force, and the rag-doll system produces emergent comedy. *Dark Souls* has physics tuned for weight: movement feels heavy, attacks have commitment, and everything has momentum. *Breath of the Wild* has physics tuned for playful engagement: objects afford interesting interactions, and the system invites experimentation. We appreciate these physics not primarily by asking how realistic they are but by attending to internal consistency, characteristic feel, and aesthetic fit. *Grand Theft Auto*'s physics serves an aesthetic of chaos and spectacle; it would not suit *Dark Souls*. Each game's physics is tuned to its aesthetic and ludic goals. We appreciate the physics for what it is, not for its fidelity to real-world physics. For LLMs, the analogy suggests a parallel mode of appreciation. Different models have different semiotic physics: different characteristic dynamics of text propagation. Claude's semiotic physics differs from GPT's, which differs from Gemini's. We can appreciate these differences not primarily by asking which is most human-like or most useful but by attending to internal consistency and characteristic feel. Order appreciation, as Carlson develops it, differs from design appreciation. We do not primarily ask how well the artifact serves its intended function. We attend to the order itself, the patterns produced by the forces, and appreciate them for their own character.
The bee text discussed in Section 6.2 is relevant here in a further way. It was produced under relaxed constraints, revealing a region of Claude's behavioural space that is normally inaccessible under standard product configurations. Knowing that this region exists – and knowing what the model can do under different conditions – is part of appreciating the model. Model appreciation involves appreciating not just the outputs a model typically produces but the full space of outputs it could produce, and how different conditions activate different regions of that space. The bee text is a window into latent capacities, a sample from a region of semiotic space that standard use does not reach.
Different models instantiate semiotic physics differently. Different training corpora, different architectures, and different post-training regimes produce different characteristic orders. Users report different feels when interacting with different models: Claude's hedging rhythms differ from GPT's briskness, and Gemini handles certain registers differently. A fuller account of model-level appreciation would map these differences systematically, developing a comparative aesthetics of LLMs. That task lies beyond the scope of this paper. For present purposes, the point is that model-level appreciation is possible and that it takes the form of appreciating distinctive semiotic order: the characteristic dynamics of text propagation that distinguish one model from another.
We should distinguish the kind of appreciation we have been describing from other modes of engaging with LLMs. Capability evaluation tests whether models perform tasks correctly. Safety testing probes whether models can be induced to produce harmful outputs. Benchmarking measures performance against standardised criteria. The appreciation we describe differs from all of these. The goal is not to assess correctness, safety, or performance but to appreciate characteristic order, and to develop acquaintance with semiotic physics as it manifests in a particular model. A response that would count as a failure in capability evaluation might be aesthetically rewarding as a product of learned regularities. The appreciator is not grading but attending, not measuring but developing acquaintance.
Finally, we note a thought that we flag here but do not develop. Each model, as an instantiation of semiotic physics, has learned regularities from human text. It reflects, in transformed form, the semiotic culture of its training data. There is a sense in which generative AI is a mirror of culture, not only morally, as Vallor (2024) has argued, but aesthetically. The model shows us our own semiotic patterns, filtered through statistical learning. Appreciating an LLM is, in part, appreciating culture seen through technology. This thought merits extended treatment, but such treatment lies beyond the scope of the present paper and we reserve it for future work.
## Conclusion
We began by asking how LLMs might be aesthetically appreciated – not their outputs, but the systems themselves. Drawing on Carlson's environmental aesthetics, we argued against two temptations. The first is to appreciate LLMs as persons, whether through make-believe or by treating them as thin agents; neither route supplies the temporally extended life and evaluative structure that beauty-of-character predicates require. The second is to treat them simply as designed artefacts and apply standard form-follows-function analysis; while LLMs are artefacts, the aesthetically salient order in their behaviour is largely emergent rather than specified by designers.
Our positive proposal treats chat instances as generative environments and recommends appreciating the order that emerges in them under the constraints of a given model. The right kind of knowledge for this appreciation is semiotic physics: the study of regularities governing text propagation in trained language models. This knowledge, whether held theoretically or acquired through practical interaction, makes the order in LLM-generated text visible and intelligible – much as geological knowledge illuminates the order in a landscape. Appreciation guided by semiotic physics operates at three levels: individual outputs as specimens of how the forces operate under particular conditions, extended exchanges as environments shaped by those forces over time, and models themselves as the ground of characteristic semiotic order.
## References
Abell, C. (2020). Fiction: A Philosophical Analysis. Oxford: Oxford University Press.
Carlson, A. (2000). Aesthetics and the Environment: The Appreciation of Nature, Art and Architecture. London: Routledge.
Carroll, N. (2013). Andy Kaufman and the Philosophy of Interpretation. In Minerva's Night Out: Philosophy, Pop Culture, and Moving Pictures. Malden, MA: Wiley-Blackwell.
Cross, A. (2025). Tool, Collaborator, or Participant: AI and Artistic Agency. The British Journal of Aesthetics, 65(4). https://doi.org/10.1093/aesthj/ayae055
Danto, A. C. (1974). The Transfiguration of the Commonplace. The Journal of Aesthetics and Art Criticism, 33(2), 139-148.
Davies, S. (2012). The Artful Species: Aesthetics, Art, and Evolution. Oxford: Oxford University Press.
Elhage, N., et al. (2021, December 22). A mathematical framework for transformer circuits. Transformer Circuits Thread. https://transformer-circuits.pub/2021/framework/index.html
Farrell, H., Gopnik, A., Shalizi, C., & Evans, J. (2025). Large AI models are cultural and social technologies. Science, 387(6739), 1153-1156. https://doi.org/10.1126/science.adt9819
Forsey, J. (2013). The Aesthetics of Design. New York: Oxford University Press.
Frankish, K. (2024). What are large language models doing? In A. Strasser (Ed.), How to Live with Smart Machines (pp. 73-110). Vienna: Holzhausen Publishing. Available at: https://keithfrankish.github.io/articles/Frankish\_2024\_What%20are%20large%20language%20models%20doing.pdf
Gaut, B. (2007). Art, Emotion and Ethics. Oxford: Oxford University Press.
Janus. (2022, September 2). Simulators. AI Alignment Forum. https://www.alignmentforum.org/posts/vJFdjigzmcXMhNTsx/simulators
John, E. (2021). Review of Fiction: A Philosophical Analysis by Catharine Abell. The Journal of Aesthetics and Art Criticism, 79(4), 514-517.
Kirchner, J. H., Smith, L. M., Campos, J., Clune, J., & janus. (2023, March 3). \[Simulators seminar sequence\] \#2 Semiotic physics – revamped. AI Alignment Forum. https://www.alignmentforum.org/posts/9kNxhKWvixtKW5anS/simulators-seminar-sequence-2-semiotic-physics-revamped
Kirchner, J. H., Steiner, C., Riggs, L., Janus, & Thibodeau, J. (2023, January 3). Semiotic physics. In Simulators seminar sequence (\#2). LessWrong. https://www.lesswrong.com/posts/TTn6vTcZ3szBctvgb/simulators-seminar-sequence-2-semiotic-physics-revamped
Mallory, F. (2023). “Fictionalism about Chatbots.” Ergo: An Open Access Journal of Philosophy, 10, 38\. https://doi.org/10.3998/ergo.4668
McGinn, C. (1997). Ethics, Evil, and Fiction. Oxford: Oxford University Press.
metasemi. (2023, March 20). A note on 'semiotic physics.' LessWrong. https://www.lesswrong.com/posts/AdXzZDoYFqHCfupDB/a-note-on-semiotic-physics
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., & Carter, S. (2020). Zoom in: An introduction to circuits. Distill, 5(3). https://doi.org/10.23915/distill.00024.001
Olah, C. (2024, November 11). In D. Amodei, A. Askell, & C. Olah, Interview by Lex Fridman. Lex Fridman Podcast \#452. Available at: https://lexfridman.com/dario-amodei-transcript/
Paris, P. (2018a). The empirical case for moral beauty. Australasian Journal of Philosophy, 96(4), 642-656. https://doi.org/10.1080/00048402.2017.1411374
Paris, P. (2018b). On form, and the possibility of moral beauty. Metaphilosophy, 49(5), 711-729.
Parsons, G. (2023). Imperfection and Beauty of Character. In P. Cheyne (Ed.), Imperfectionist Aesthetics in Art and Everyday Life (pp. 296-309). New York: Routledge.
Parsons, G., & Carlson, A. (2008). Functional beauty. Oxford University Press.
Picca, D. (2025). Not minds, but signs: Reframing LLMs through semiotics. arXiv preprint arXiv:2505.17080. https://arxiv.org/abs/2505.17080
Saito, Y. (2008). Everyday Aesthetics. Oxford: Oxford University Press.
Vallor, S. (2024). The AI mirror: How to reclaim our humanity in an age of machine thinking. Oxford University Press.
Wojtkiewicz, K. (2023). How Do You Solve a Problem like DALL-E 2? The Journal of Aesthetics and Art Criticism, 81(4), 454-467. https://doi.org/10.1093/jaac/article/81/4/454/7571331
Wolfram, S. (2023, February 14). What is ChatGPT doing … and why does it work? Stephen Wolfram Writings. https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-doing-and-why-does-it-work/
[^1]: We will consider some non-paradigmatic artworks in Section 5\.
[^2]: Carlson stresses that ordinary descriptions of environments and more theoretical scientific, historical, and functional descriptions lie on a continuum, so that scientific and historical knowledge can deepen rather than displace practical familiarity as a basis for aesthetic appreciation (Carlson, 2000). By analogy, one might speculate that the sciences of mind and behaviour could relate to folk-psychological and biographical understanding in a similar way, so that in some cases empirical work on personality, emotion, or cognition might feed into the aesthetic appreciation of persons alongside the more everyday forms of knowledge stressed in the main text.
[^3]: This personal appreciation also scales up to what we might call performance personalities. We respond to a comedian’s improvisational skill or an orator’s gravitas much as we respond to character in our friends, but now filtered through a public persona – genuine to the individual, yet artfully composed for performance. Recent scholarship has engaged with these performative dimensions, examining how performers construct and present public personas (Carroll 2013). Here too, Carlson’s knowledge requirement bites: to appreciate such personas we need to understand both the individual and the conventions of the performance context.
[^4]: This way of distinguishing between an underlying generative system and the agent-like patterns it instantiates in particular episodes draws on work that treats GPT-style models as *simulators* of text worlds capable of generating agent-like *simulacra* without themselves being agents (Janus, 2022; Bereska et al., 2023). We do not use that terminology in the main text, but the present discussion adopts a similar two-level picture.
# Handoff Document: Environmental Aesthetics of Generative AI
## Project
We are working on a paper provisionally titled *The Environmental Aesthetics of Generative AI*. The paper asks whether LLMs themselves, rather than only their outputs, can be aesthetically appreciated.
The paper uses Carlson’s environmental aesthetics to frame the problem. The central idea is that appropriate aesthetic appreciation depends on appreciating an object as what it is, in light of knowledge relevant to that kind of object. This is not just a point about correct classification. Misidentifying the object changes the explanation under which its appearance is seen and changes the knowledge that guides attention.
## Current structure
Introduction
1. Appreciating Design, Appreciating Order
2. What LLMs Are
3. [To be rebuilt from scratch]
4. Semiotic Physics
5. Levels of Appreciation
## Current Introduction
The introduction distinguishes the paper’s question from debates about AI-generated works. The question is not whether AI systems can produce artworks, or whether generated works have aesthetic merit. The question is whether LLMs themselves can be objects of aesthetic appreciation.
The introduction presents two initial routes:
1. LLMs are encountered through generated language and extended exchanges, which makes person appreciation initially tempting.
2. LLMs are engineered systems, trained and post-trained before reaching users, which makes design appreciation initially tempting.
The introduction then introduces Carlson’s recommendation: appropriate appreciation requires knowing what sort of thing something is and looking at it in light of knowledge relevant to that kind of thing.
The thesis currently says that Carlson’s notion of order appreciation gives the better account. The relevant order appears in generated language, but it is the order of a trained system whose responses develop from context. Ordinary attention to the words on the screen is not enough. The appreciator also needs to understand how the system comes to produce one continuation rather than another. This knowledge is called *semiotic physics*.
## Current Section 1: Appreciating Design, Appreciating Order
Section 1 introduces Carlson’s distinction between design appreciation and order appreciation.
Carlson’s general recommendation is that we should take things as what they are and look at them in the light of the right kind of knowledge.
The section uses the cliff/Rembrandt contrast to explain the point. Appreciating a cliff face as if it were an artifact crafted by a divine maker directs attention to the wrong kind of explanation. Similarly, treating a Rembrandt as the result of natural forces would fail to appreciate it as the kind of object it is. The added Carlson point is that the failure is not only a false classification: treating the object as the wrong kind of thing changes the explanation under which its appearance is seen and changes the knowledge that guides attention.
The section then introduces *design appreciation*. For Carlson, works of art and designed artifacts are appreciated through knowledge of design, making, function, and realization. The current text cites Carlson’s discussion of the relation between a design, the object that realizes it, and the maker who realizes it.
The section then introduces *order appreciation*. In natural environments, there are no designer’s intentions to recover or evaluate. The appreciator attends instead to patterns and structures produced by forces operating without purpose. Carlson’s order-appreciation quotation remains central: the relevant appreciation involves the order, the forces that produce it, and the account that makes the order visible and intelligible.
The section also now defines *acts of aspection*, following Carlson and Ziff, as ways of attending to an object that partly constitute its appropriate appreciation.
The section then adds person appreciation as a further case. Person appreciation is discussed through beauty of character. The key point is that person-directed appreciation requires knowledge of a subject whose responses hang together over time as the responses of a temporally extended agent. Parsons is used for the thought that such knowledge is built through some form of acquaintance with a life.
The section closes by saying that LLMs bring the possibilities together: their conversational form points toward person appreciation; their production by human institutions points toward design appreciation. Section 2 then begins by asking what sort of thing an LLM is.
## Current Section 2: What LLMs Are
Section 2 answers Carlson’s object-identification question for LLMs.
The first sentence must remain exactly:
> We have seen that Carlson recommends appreciating things for what they are, what sort of thing are LLMs?
The section describes LLMs as trained and deployed continuation systems encountered through generated language.
The structure of Section 2 is:
1. The object is not simply the text on the screen. A response is one occurrence of a system’s activity, and an exchange is one path through what the system can do. The system is available to users only through generated texts.
2. Generated text is continuation from context. The context includes the user’s prompt, earlier turns, system-level instructions, and other material made available to the model. A generated response develops because each token becomes part of the context from which the next token is produced.
3. Path-dependence matters because later parts of a response are generated from a context that already includes earlier parts. This allows a response to sustain a line, but also to lose the line it had.
4. Training explains why some continuations are more available than others. Pre-training shapes a graded sensitivity to textual regularities. Generated continuation is the system’s response to its current context under the learned pressures of those regularities.
5. The internal organization that supports this sensitivity emerges from training rather than being laid out in advance by designers. Designers set up the conditions under which training occurs; they do not specify the full pattern by which possible contexts should continue.
6. The “bass” example is used to show context-sensitive continuation. In a context concerning a fretboard, it makes continuations drawn from music more available; in a context concerning shallow water, it makes continuations drawn from fishing more available.
7. Context extends beyond the immediately preceding token. Earlier questions or registers can continue to condition later turns. Context persists as the developing condition under which continuation proceeds.
8. Users ordinarily encounter post-trained, deployed systems. Post-training shapes a model into a conversational role. Interface and system-level instructions also condition the available continuations. This produces a relatively stable response profile.
9. The section derives the distinction between output, chat, and model. An output is one bounded continuation. A chat is an extended exchange in which earlier turns condition later ones. The model is the trained and deployed system whose tendencies become visible over time. Output, chat, and model are the scales at which the same system becomes available for appreciation.
## Current status of Section 3
Section 3 is to be started again from scratch.
The previous draft of Section 3 contains extensive inline comments and should not be treated as stable prose. It discusses person appreciation and design appreciation, but the analysis and structure need to be rebuilt.
## Sources already in play
Carlson (2000), *Aesthetics and the Environment*
Gaut (2007)
Paris (2018)
Parsons (2023)
Mallory (2023), “Fictionalism about Chatbots”
Frankish (2024), “What are Large Language Models Doing?”
Forsey (2013)
Parsons and Carlson (2008)
Olah (2024)
## Style and drafting constraints
Write plainly.
Use dry, analytic prose.
Do not write generic academic filler.
Do not use long lists of examples in place of argument.
Do not use “not X but Y” constructions.
Avoid meta-commentary such as “the upshot is”, “this section argues”, “this raises the question”, “two candidates present themselves”, or similar LLM-shaped transitions.
Do not use scare quotes unless they are genuinely needed for mentioning a term as a term.
Do not use “matters” as a verb.
Do not over-distill. Preserve detail and fineness of grain.
Preserve existing wording whenever it is usable. Do not paraphrase merely for the sake of paraphrasing.
When revising, flag clearly when any content from the existing draft is being removed, downgraded, or moved.
Use first person plural where appropriate.
Use British spelling unless a specific project instruction says otherwise.
## Workflow used so far
For each section:
1. Start with a section-function analysis: what the section inherits, what question it must answer, what result it must establish, and what pressure it leaves for the next section.
2. Diagnose the structural failures in the current draft.
3. Produce a paragraph-by-paragraph plan before drafting.
4. Draft only after the plan is stable.
5. Preserve old wording wherever possible.
6. After drafting, audit changes against the old draft and inline comments.
7. Restore any wording changed without a real reason.
### Assistant
I'll work through this carefully. The user wants me to (1) list skills activated by `/deep-writing-mode` with their relevance, (2) then apply them. Let me first read the skill files to know what I'm working with.
### Assistant
Good. Now I can see what deep-writing-mode consolidates. Let me read the relevant references and audit skill files.
### Assistant
Let me think through this carefully. The user wants three things in order:
1. List the skills activated by /deep-writing-mode
2. Explain their relevance to this evaluation task
3. Apply them to evaluate the draft
Looking at the deep-writing-mode SKILL.md, it tells me explicitly: "Use this as the single writing-control surface... Do not load nick-analytic-voice, writing-standards, nick-topic-sentences, or no-example-lists as separate mandatory skills. Their standing constraints are consolidated here." So these four legacy skills are folded into deep-writing-mode itself, accessed via the references/ directory:
- references/voice.md (= nick-analytic-voice + nick-philosophical-prose)
- references/writing-practice.md
- references/topic-sentences.md
- references/no-example-lists.md
- references/anti-metacommentary.md
- references/final-audit.md
- references/formatting.md (= writing-standards)
- references/prose-composition.md
- references/examples.md
- references/vocabulary.md
The skill also tells me: "Keep source-work and epistemic-discipline separate" because they're cross-workflow skills. And: "Use audit-only skills only when the task calls for them: draft-audit, voice-fix, source-check, depth-audit, and anti-metacommentary are post-hoc review modes."
Since the user has asked for EVALUATION of a draft, the audit-only skills become highly relevant:
- voice-fix → diagnose style problems and compare to published work
- depth-audit → catch where philosophical moves are described but not made
- source-check → verify characterisations of cited authors
- anti-metacommentary → catch sentences that narrate the argument
Plus the consolidated references inside deep-writing-mode.
Now, applying them. Let me work through the draft.
The draft is the latest iteration after a long conversation in which Nick repeatedly flagged voice problems in earlier versions. The chat shows me what Nick already considers off-limits. Let me check whether the current draft still has problems Nick already flagged in the original.
Original draft had: "Two candidates present themselves" (flagged).
Current draft: "Two routes for appreciating such systems suggest themselves" — structurally identical. But Nick said "this is fine" for P1, so this is closed.
Original P10 (now P3 in current draft on Forsey/Parsons-Carlson): "Forsey's Kant-inspired account of design as a case of dependent beauty and Parsons and Carlson's later theory of functional beauty can both be read as ways of spelling out this claim" — flagged as "not how i write."
This sentence is preserved unchanged in the current draft. So Nick wants this flagged.
The Forsey sentence in the current paragraph: "Forsey argues that judgements of design beauty presuppose a concept of what the object is meant to be and do, and that our grasp of its success in fulfilling that role informs the aesthetic verdict itself rather than merely accompanying a 'pure look' at its lines." Nick flagged this as "humongously unclear" — it's preserved. The whole paragraph was flagged as not well written: "I feel this paragraph is not as written very well and not very clear. I don't know if it's structural, but it might be."
Hmm. So P3 is a paragraph that Nick explicitly flagged as problematic, and it's been preserved entirely. That's a clear finding.
Now the design-half opener: "On this view, LLMs look like canonical objects for design aesthetics: complex, purpose-built systems whose architecture, training recipe, and user interface might be admired for elegance, efficiency, or ingenuity." This has TWO triplets in one sentence:
- "architecture, training recipe, and user interface"
- "elegance, efficiency, or ingenuity"
Nick's original comment on this exact paragraph was "fucking lists man, fucking lists upon lists." This is preserved.
This is a critical finding: the draft preserves prose that Nick explicitly identified as having his most-disliked failure mode.
Let me check the Mallory paragraphs. Original P (now P9): "If we ask ordinary users whether they literally believe that a chatbot is a person, many will concede that they do not. They may talk to a model as if it were a friend or a colleague, and they may feel heard, reassured, or amused, but when pressed they acknowledge that they are interacting with a computational system rather than a human being."
The "friend, colleague" and "heard, reassured, amused" lists were flagged. New draft: "Asked directly, most users will concede that they do not believe a chatbot to be a person, even when they talk to one as if it were." So the lists have been excised and the move compressed. Good. But the empirical-survey framing remains ("most users will concede that...") and Nick flagged the original as "not how i write."
The fictionalism Carlsonian objection. In the previous draft, Nick complained that the argument "is too compressed... a single sentence." The new draft has expanded it to two paragraphs (P11 and P12). P11 grants that literary fiction does engage the right kind of artefact; P12 draws the disanalogy with LLMs. The structure is right. But P12's "its outputs are intended as contributions to the prompts we give" is odd — "contributions to the prompts" is a strange formulation. The original had "contributions to our inquiries" which is cleaner. This appears to be a downgrade.
Frankish paragraph. The original had scare quotes around 'beliefs' and 'desire' (flagged). The new draft removes them: "a large set of thin beliefs... and one thin desire." But these were marking technical terms-of-art. The formatting.md reference says italics should be used on first introduction of terms of art. So the right move is italics, not removal. Currently the technical specificity is lost.
"Grant this." — replaces "Suppose we grant all of this" (flagged). Terse. Possibly too terse — Nick's published voice (e.g. Growing the Image) does use "Even if, for the sake of argument, we concede that..." for this kind of concessive move. "Grant this." may be too punchy for the surrounding rhythm.
The failure-of-conditions paragraph in current P14: doesn't structure the failure clearly into (a), (b), (c) as the earlier Claude in the chat had suggested. Instead it runs the conditions together. The current version says:
- "Its beliefs are confined to what is encoded in the model's parameters and surfaced in the current context, without memory or development across conversations."
- "Its desire is singular: to make an appropriate move in the present exchange."
- "There is no history across episodes from which later responses could draw."
- "The predicates characteristic of person appreciation — steadiness of character, depth of feeling — presuppose something that can be developed over time; the chat-game agent has no such temporal depth."
Three conditions are present (temporal extension, singularity of desire, no history) but they're not enumerated. The "steadiness of character, depth of feeling" is paired, not triplet, so acceptable per the no-example-lists rule.
Vibe paragraph: "A possible objection at this point is that these arguments underplay the role of post-training and the chat interface." This is meta-commentary about argumentative structure. Per anti-metacommentary.md, sentences whose subject is "the argument" describing argumentative function should be rewritten. "A possible objection at this point" is precisely this kind of pattern.
Closer P16: "The two routes for appreciating LLMs as persons run out at different points." Strong opening. "Design knowledge, taken up first, supplies the conditions under which an LLM is produced and used." — this is shoehorned. The closer is primarily about the person routes; the design summary feels parenthetical.
Let me also check Carlsonian frame preservation. The original Claude in the chat said the section needed to keep Section 1's vocabulary alive — "aspection", "kind of object", "kind of knowledge". I count uses in the current draft:
- "aspection": 0 occurrences
- "kind of artefact": 3 occurrences (P11, P12, P16) — good
- "kind of knowledge": 0 occurrences
- "appropriate appreciation": 0 occurrences
- "Carlson's recommendation": 2 occurrences (P11, P13)
- "Section 1": 2 occurrences (P8, P14)
- "Section 2": multiple (good)
So the Section 1 link is somewhat thin. "Aspection" is missing entirely. The phrase "kind of object/knowledge" is implicit but not used directly. This is a depth-audit finding: the draft does the Carlsonian moves but doesn't reinforce the vocabulary that ties them to Section 1.
Depth-audit specific findings:
P2 (design half opener): "Given Carlson's emphasis on artefacts and design appreciation, it is natural to ask whether we should aesthetically appreciate LLMs as designed tools, asking how well their forms serve their functions." — this is the "described but not made" pattern. The paragraph announces that it will ask whether LLMs can be appreciated as design but does not yet test. That's actually fine in a setup paragraph. But the third sentence ("On this view, LLMs look like canonical objects for design aesthetics: complex, purpose-built systems...") is described-not-made: it asserts that LLMs look like canonical design objects but doesn't develop one case to show this. A worked example here would help.
P3 (functional beauty literature): three accounts mentioned in series. This is the "named but not developed" pattern in depth-audit terms. None of the three accounts (Carlson's slogan, Forsey's dependent beauty, Parsons-Carlson functional beauty) is developed enough for the reader to see how it would apply to LLMs. The paragraph reads as a literature survey rather than as a substantive use of the literature.
P11 (Mallory aesthetic extension + literary case): The "Treating its protagonists as fictional persons does not misclassify the artefact; it engages the artefact as the kind of thing it is" — this is asserted but not earned. The reader is told that literary protagonists engage the artefact correctly but not shown why. What is it about a novel that makes fictional-person engagement appropriate? Holmes is constituted as a fictional person inside a represented story-world. This isn't said. The previous Claude in the chat history actually mentioned this point: "in the literary case, the fictional persons live in a fictional world that the artefact constructs. Holmes is a constituent of a story-world." That development is missing.
P12 (disanalogy): "The LLM is not made to construct a fictional world inhabited by fictional persons; its outputs are intended as contributions to the prompts we give." This is the right move but the contrast isn't sharpened. What's missing: the asymmetry that in literary fiction the "person" is a constituent of a represented world, while for the LLM-user the "person" would be a phantom of the actual exchange. Without this contrast, the disanalogy isn't fully drawn.
Source-check: I have Nick's published "Beauty in Use" paper which engages directly with Parsons-Carlson and Forsey. Let me cross-check.
Beauty in Use describes Parsons-Carlson: "knowledge of [...] function can play a role in our perception of and hence our aesthetic appreciation of that object" (2008: 90). They identify three ways function translates: looking fit for function, elegance/simplicity, and visual tension.
Current draft: "Parsons and Carlson explain how knowledge of function can structure experience so that an artefact's form can be experienced as fit, streamlined, overbuilt, and so on, yielding functional beauty when the form presents itself as well suited to what the thing is for (Parsons and Carlson 2008, chapter 4)."
"Fit, streamlined, overbuilt" doesn't quite match Parsons-Carlson's three categories (fit-for-function, elegance, visual tension). "Overbuilt" especially seems to gesture at a different category. The characterisation isn't wrong but it's vague — Nick's own published work characterises Parsons-Carlson more precisely.
Forsey: Beauty in Use describes her as developing a Kantian approach using free vs. dependent beauty: "the purpose which determines what the thing is to be" (2013: 163). She talks about "teleological style" — the specific way an object realises its purpose.
Current draft: "Forsey argues that judgements of design beauty presuppose a concept of what the object is meant to be and do, and that our grasp of its success in fulfilling that role informs the aesthetic verdict itself rather than merely accompanying a 'pure look' at its lines (Forsey 2013)."
The "pure look" formulation — is this a Forsey term? It's ambiguous in the draft. If it's a Forsey term, it should be presented as a Forsey term (with attribution or as a mention with single quotes). If it's a paraphrase, the quotes are misleading. This needs source-check.
The current characterisation also doesn't engage what Beauty in Use takes from Forsey — the dependent/free beauty distinction and "teleological style." The compressed version loses the Kantian conceptual apparatus.
Epistemic discipline check: The draft works through Mallory, Frankish, and a vibe-objection. For each, the position is granted and then tested. The dialectical conduct is mostly fair: Frankish's view is granted ("Grant this") rather than dismissed. But the Mallory section comes close to assertion at the end ("Treating it as a fictional person engages it as the kind of artefact it is not") — the work to earn this is mostly in P11 and is mostly about saying-rather-than-showing the Carlsonian point.
OK I think I have a clear picture. Let me organise this for the user. The structure should be:
1. Skills activated by /deep-writing-mode (with relevance)
2. Application — organised by skill perspective:
- Voice / writing-practice (the texture)
- No-example-lists (specific list violations)
- Anti-metacommentary (specific patterns)
- Formatting (specific issues)
- Depth-audit (specific paragraphs)
- Source-check (specific claims to verify)
- Epistemic discipline (any concerns)
3. Paragraph-by-paragraph rolling up findings
4. Recommendations
Length: the user wants long and detailed, with options. I'll be thorough but tight.
One more thought: the user explicitly said "EMPLOY the SKILLS!" — emphasis on the action. So I shouldn't just list and gesture; I need to actually apply each lens and produce concrete findings. The grounding in published work (Growing the Image, Beauty in Use) gives me real comparison material.
I should also follow Nick's chat formatting rules (no bold) and end with an Italian sentence.
## Skills activated by /deep-writing-mode (and their relevance)
The skill file says deep-writing-mode is the single control surface for manuscript work and consolidates four former mandatory skills into its own `references/` directory. It also says two skills stay separate, and that audit-only skills should be brought in when the task calls for them. Your task — evaluation of an existing draft for content and style — calls for the audit-only skills explicitly.
### Consolidated inside the skill (its `references/`)
- `references/voice.md` — Nick's academic register (consolidates the old `nick-analytic-voice` and `nick-philosophical-prose`). Tells me your default sentence is not short and punchy: longer sentences with embedded clauses, qualifications, reformulations. Relevance: the draft has stretches of clipped declaratives that don't match this texture, so I'll be checking rhythm.
- `references/writing-practice.md` — Drafting stance: every sentence faces the subject matter, not the argument's effects on the reader. Detail preservation rules. Relevance: I'll check whether each paragraph performs work or describes the work being done.
- `references/topic-sentences.md` — Repertoire of paragraph openings drawn from your published prose. Relevance: I'll look at every paragraph opening in Section 3 and ask whether it enters the thought or merely narrates section structure.
- `references/no-example-lists.md` — Prohibition on triplets, comma-chains, parenthetical catalogues. Relevance: the original draft had numerous lists you flagged with "fucking lists man" — I'll check whether the current draft has cleaned them.
- `references/anti-metacommentary.md` — Forbidden patterns whose subject is "this," "the section," "the argument," "the analysis." Relevance: I'll catch sentences that talk about what the section is doing.
- `references/formatting.md` — Quotation marks (double for direct quotation, single for mention/scare), italics for terms of art on first introduction, British spelling, no bold. Relevance: term-of-art handling has been inconsistent across iterations and matters here for "beliefs/desire" in the Frankish discussion.
- `references/final-audit.md` — Pre-submission checklist. Relevance: gives the running comparison frame.
- `references/prose-composition.md`, `references/examples.md`, `references/vocabulary.md` — Grounding extracts from published prose. (I've also read full paragraphs from *Growing the Image* and *Beauty in Use* to anchor the comparison; the consolidated extracts aren't a substitute for published paragraphs.)
### Kept separate by the skill (mandatory adjuncts)
- `source-work` — Activated because the draft discusses, quotes, or attributes claims to Carlson, Forsey, Parsons and Carlson, Mallory, Frankish, John, Olah. Relevance: characterisations need to be checked against extracted source text, not memory.
- `epistemic-discipline` — Activated because the section represents your developing position. Relevance: I should preserve your superposition where you have not ranked, not impose hierarchy through value-loaded vocabulary, and distinguish source / interpretation / speculation in any rewrite suggestion.
### Audit-only skills that the task explicitly calls for
- `voice-fix` — Post-hoc voice repair. Procedure: locate the flagged text, classify the move, diagnose the failure mode against a fixed taxonomy (17 categories), read three published paragraphs doing similar work, then rewrite. Relevance: this is the right tool for diagnosing where the draft does not sound like you.
- `depth-audit` — Catches paragraphs that describe a philosophical move rather than perform it. Six failure modes (described-not-made, asserted-without-earning, named-not-developed, summary-substituting-for-argument, list-substituting-for-development, quotation-without-analysis). Relevance: this is the right tool for content.
- `source-check` — Verifies characterisations of cited authors against the actual source text. Relevance: needed for the Forsey, Parsons-Carlson, Mallory, Frankish characterisations in this section.
- `anti-metacommentary` — Also lives as an audit. Relevance: a few sentences in the current draft are still narrating argumentative function.
- `contemplate` — Already in effect through the slash command.
I've grounded the comparisons in two published papers: *Growing the Image* (Young & Terrone 2025) and *Beauty in Use* (Young & Terrone, forthcoming Ergo). *Beauty in Use* is especially close in subject matter — it engages Parsons-Carlson and Forsey on design directly, so it doubles as a source-check reference for those passages.
---
## Application
Three preliminary observations before the lens-by-lens work.
First, the draft has improved measurably over the original. The opening lampshade is clean, Pollock has gone, the conclusion no longer over-promises, the bad Olah qualifier is gone, and the Mallory disanalogy has been expanded from one sentence to two paragraphs.
Second, the draft preserves several passages that you explicitly flagged in the original. In particular, the Forsey/Parsons-Carlson paragraph is essentially unchanged from the version you described as "I feel this paragraph is not as written very well and not very clear." That paragraph still needs work.
Third, the section is doing the philosophical job it was set: testing person-knowledge and design-knowledge against the Section 2 description, finding both inadequate. The structural decisions hold up. The remaining problems are at the level of texture, depth, and individual sentences.
### Voice (voice-fix lens + references/voice.md + references/writing-practice.md)
What I'm finding when I compare to published paragraphs in *Beauty in Use* and *Growing the Image*:
- Sentence rhythm is mostly serviceable but flattens in places, especially in the design half. P2, P3, P4 contain longer sentences that do the thinking, which is right. But the design closer (P7) goes flat with two short declaratives in a row: "That order develops as the system runs, and is no part of any designer's specification." Compare with how *Growing the Image* closes the agent section: "Both options are unsatisfying. While ascribing the artistic merit to the human user would overlook Midjourney's active contribution, ascribing the merit to Midjourney would downplay the creative activity of prompt-crafting." The closing is longer and pulls both arms of the dilemma together. Section 3's design closer could do similar work — note both what design appreciation does illuminate and what it doesn't, in one shaped sentence rather than two stubs.
- "Grant this." (P14). Diagnostic class: voice-fix #3 (sloganeering) borderline. Compare to *Growing the Image*: "If, for the sake of argument, we concede that Midjourney is an agent in Anscomb's sense, we are left with the dilemma of..." Your published voice does use brief turns but they emerge from longer sentences. A bare two-word imperative is not in the published register. Options:
- "Grant Frankish his ascription."
- "Concede the intentional ascription."
- "Suppose Frankish is right that..."
- "Section 2 located their basis in post-training and deployment" (P8). Diagnostic class: voice-fix #10 (Latinate vocabulary). "Located their basis" is bureaucratic. Try "Section 2 traced these patterns to post-training and deployment" or "Section 2 attributed these patterns to post-training and deployment."
- "A possible objection at this point is that these arguments underplay..." (P15). Diagnostic class: voice-fix #1 (meta-commentary). "At this point" makes the sentence narrate the section's progress rather than face the objection. *Growing the Image* opens objections with "One might object here that..." which sits closer to your published voice. Alternatively, the opening could be done substantively: "Post-training and the chat interface may seem to underwrite person-aesthetics in a thinner form."
- "Asked directly, most users will concede that..." (P9). Diagnostic class: voice-fix #11 (reader management) and faintly #7 (casual/vague phrasing). The original version with "If we ask ordinary users..." was flagged for the same reason. The empirical-survey framing is doing scene-setting work that the argument can do without. Could be cut: skip directly to "Mallory develops the as-if posture as chatbot fictionalism (2023)."
- "Treating it as a fictional person engages it as the kind of artefact it is not." (P12). Diagnostic class: borderline #3 (sloganeering) and #13 (missing development). The verdict is right; the showing is undersupplied. (See depth audit below.)
- "They are not underlying subjects." (P15 closer). Strong line — but it pairs with "The targets of such language are episodes and recurring response profiles." This is a short-short sequence in a section that has been letting longer sentences do the thinking. Could be folded: "The targets of such language are episodes and recurring response profiles rather than underlying subjects."
### Lists and triplets (references/no-example-lists.md)
- P2: "complex, purpose-built systems whose architecture, training recipe, and user interface might be admired for elegance, efficiency, or ingenuity." Two triplets in one sentence. This sentence was preserved verbatim from the original draft, where you flagged it as "fucking lists man, fucking lists upon lists." Either develop one case ("their training recipes, say, might be admired for ingenuity in how training data is curated") or state the general claim ("complex purpose-built systems whose engineering can be admired"). I would lean to the second: the list adds register, not argument.
- P3: "buildings, airplanes, and appliances as well as landscapes" — inside the Carlson quotation, so it stays.
- P3: "fit, streamlined, overbuilt, and so on" — three-item list ending in "and so on." Permitted as components of a single structure (Parsons-Carlson's actual taxonomy has three categories), but cf. *Beauty in Use*, your published treatment names them as "looking fit for function," "elegance and simplicity," and "visual tension." Your own published characterisation is more precise. The current draft's "fit, streamlined, overbuilt" is vague and doesn't track Parsons-Carlson's actual categories.
- P14: "steadiness of character, depth of feeling" — pair, not triplet. Permitted under the rule but consider whether one of them is enough.
- P3 again: "Forsey's Kant-inspired account of design as a case of dependent beauty and Parsons and Carlson's later theory of functional beauty can both be read as ways of spelling out this claim." This is structurally a name-checking sentence. It was flagged in the original. Currently preserved.
### Anti-metacommentary (references/anti-metacommentary.md)
- P15 opening: "A possible objection at this point is..." See above. Replace with substantive opening.
- P16: "None of this makes the order in generated language appreciable as the order Section 2 described." Borderline. "None of this" refers to the section's preceding moves. The sentence has object-level content (it states what the routes fail to do), so it survives the test, but the construction is on the cusp.
- P11 opening: "Mallory's account is metasemantic and epistemic." This characterises the account directly — it's not metacommentary about your own argument, so it stays.
- P10: "On his view, we engage with chatbots by entering a game of make-believe..." Good direct face of the position.
The bigger structural problem: the section never names what it's testing in Carlson's vocabulary. The published Section 1 introduces aspection, kinds of objects, kinds of knowledge. None of these terms appears in Section 3. The Carlsonian frame is being carried by gestures ("Section 1 took person appreciation to depend on...") rather than by the vocabulary itself. This is the inverse of metacommentary: under-naming rather than over-naming.
### Formatting (references/formatting.md + writing-standards)
- Quotation handling around the Mallory phrase: "'literally meaningless but fictionally meaningful' (Mallory 2023, 1082)." Per the rules, double quotation marks are for direct quotations. Single quotation marks are for scare quotes, mention, and nested quotations. This quotation is direct; it should be double-quoted.
- Frankish "thin beliefs" and "thin desire": the original draft had scare quotes around 'beliefs' and 'desire'. You flagged scare quotes throughout. The current draft removes them, but these are technical-term uses (thin belief vs. ordinary belief), not scare quotes. Per the formatting reference, italicise on first introduction of a term of art. Suggest: "ascribe to them a large set of thin *beliefs* — roughly, informational states distilled from training — and one thin *desire*: to play what he calls the *chat game*." This preserves the term-of-art status without the scare-quote register.
- "Within the fiction, the chatbot says things and means things; outside it, no speaker is present." Without any marker, "says" and "means" lose their fictional-use status. Per the mention/use distinction in formatting.md, single quotes for mentioning a word as a word would be acceptable here — but you flagged the original's scare quotes. Alternative: rephrase to avoid needing markers. E.g. "Within the fiction, the chatbot is a speaker and its outputs have meaning; outside, no speaker is present and the outputs are uninterpreted strings."
- En-dashes vs. em-dashes: the draft uses en-dashes with spaces ("–") in some places and em-dashes without spaces ("—") in others. Pick one. UK convention is spaced en-dash. Your published *Growing the Image* uses both as well; consistency rather than choice is what matters.
### Depth (depth-audit lens)
This is the most consequential pass. I'll work paragraph by paragraph, naming the failure mode where one applies.
P2 (design half opener). Mode: described-but-not-made. The paragraph asserts that LLMs look like canonical objects for design aesthetics and concludes with the triplet list. What's missing: a developed case showing what design-aesthetic appreciation of an LLM would look like in a single instance. *Beauty in Use* makes this kind of move by working a single example (the Bialetti, the butterfly keyboard). The Section 3 paragraph could develop a single design-aesthetic encounter — say, admiring the elegance of a particular interface decision or the ingenuity of a refusal pattern. One developed case here would do more than the triplet does.
P3 (functional beauty literature). Mode: named-but-not-developed plus list-substituting-for-development. Three accounts are cited (Carlson's slogan, Forsey on dependent beauty, Parsons-Carlson on functional beauty). None is developed enough for the reader to see how it would apply to LLMs. Your own published paper *Beauty in Use* makes Parsons-Carlson concrete by developing one of their three categories at length (the Bialetti as fit-for-function, the keyboard as failure of fit). This section uses the literature as scenery rather than as a tool. Options:
- Develop one of the three accounts fully, treating it as the canonical statement, and reference the others in a footnote.
- Develop one account as it applies to an LLM (e.g., does an interface "look fit for function"? Could a refusal pattern be "elegant"?). The literature would then earn its place.
- State the general claim once without naming three theorists, and reserve the development for where it does work in Section 4 or 5.
P4 (the asymmetry). Mode: clean. The bridge example does enough work that the asymmetry is genuinely shown. Strong paragraph.
P5 (Olah setup) and Olah quote and P6 (Olah follow-up). Mode: borderline asserted-without-earning. The Olah passage is doing the philosophical work, but the framing sentence after it ("The passage marks the difference between designing the conditions under which a system is trained and directly specifying the detailed profile that results") is unpacked too briefly. The passage's "scaffold/light/grow" metaphor is rich; the follow-up sentence treats it as already-interpreted. A more textured unpacking would be: what is the scaffold? what counts as the light? what is the grown thing? *Growing the Image*'s engagement with the *natura naturans* / *natura naturata* distinction shows the model: a metaphor is brought in, then specific work is done with each of its terms.
P7 (design closer). Mode: short-of-earned. Two sentences. The first sentence concedes a place for design appreciation; the second says what design knowledge cannot reach. The closer could be one shaped sentence that holds the concession and the limit together.
P8 (person pivot). Mode: clean enough. Opens with "It is tempting to model our appreciation of LLMs on our appreciation of people." This is a topic-sentences "Direct Phenomenological Observation" opening ("It is natural/tempting to describe X as Y"), which is in your repertoire. The example pair "friendlier... more cautious" is on the right side of the list-vs-development line.
P9 (Mallory intro). Mode: clean enough after voice fixes.
P10 (Mallory the view). Mode: clean. Sets out the position.
P11 (Mallory aesthetic extension + literary case). Mode: described-but-not-made. The paragraph says "Treating its protagonists as fictional persons does not misclassify the artefact; it engages the artefact as the kind of thing it is" — but does not show what makes a novel the kind of artefact whose function is to elicit imaginings of fictional persons. The earlier Claude in this chat suggested adding: "Holmes is a constituent of a story-world. The artefact is the Doyle stories, which construct that world." That development is missing from the current draft. Without it, the Carlsonian verdict in P12 is harder to feel.
P12 (Mallory disanalogy). Mode: asserted-without-earning. The disanalogy gets one sentence: "The LLM is not made to construct a fictional world inhabited by fictional persons; its outputs are intended as contributions to the prompts we give." But what is the positive characterisation of where the fiction sits for the LLM-user? In the literary case, the fiction is *of* a represented world; in the LLM case, the fiction would be *about this present interaction*. This contrast was suggested earlier in the chat ("the literary fiction's 'person' is a constituent of a represented world; the LLM-user's 'person' is a phantom of the actual exchange") and is missing from the draft. Adding it would do the showing that P12 currently leans on the reader to supply.
Also note: "its outputs are intended as contributions to the prompts we give" reads oddly. Outputs are contributions to *answers*, *inquiries*, or *the user's tasks*, not to *prompts*. The original wording was "contributions to our inquiries" — that was cleaner.
P13 (Frankish intro). Mode: clean. The position is granted carefully and the chat-game framing is right.
P14 (Frankish failure-of-conditions). Mode: borderline. Three conditions appear (parameter-only beliefs, single desire, no history across episodes), but they're run together rather than enumerated. The earlier Claude suggested an explicit (a) temporal extension, (b) plurality of cares, (c) acquaintance with a life structure. The current draft's running prose handles temporal extension and singularity but doesn't quite isolate the acquaintance-with-a-life condition (Parsons's contribution). The Parsons line is doing work in a subordinate clause ("on Parsons's account") that might want more space. Compare to the careful condition-by-condition pressure your published *Beauty in Use* puts on Parsons-Carlson and Forsey in turn.
P15 (vibe). Mode: clean enough after the opening fix.
P16 (closer). Mode: borderline. "Design knowledge, taken up first, supplies the conditions under which an LLM is produced and used" feels tacked on — the closer is otherwise about the two person routes. Either the closer treats both halves symmetrically (in which case the design clause needs more presence) or it focuses on the person half and refers back to the design closer in a different sentence. The current sentence is doing housekeeping rather than philosophy.
### Sources (source-check lens)
The Parsons-Carlson and Forsey characterisations can be checked against *Beauty in Use*, which I read.
Parsons-Carlson. Current draft says: "knowledge of function can structure experience so that an artefact's form can be experienced as fit, streamlined, overbuilt, and so on." Your *Beauty in Use* characterises Parsons-Carlson via their three categories: "looking fit for function," "elegance and simplicity," and "visual tension" (corresponding to Parsons & Carlson 2008: 97–100). "Streamlined" and "overbuilt" don't track these categories precisely — "streamlined" gestures at the elegance category, but "overbuilt" doesn't map onto visual tension. The current characterisation is loose by your own published standard. Suggest tightening to your published wording.
Forsey. Current draft cites her as the source for "judgements of design beauty presuppose a concept of what the object is meant to be and do" and "informs the aesthetic verdict itself rather than merely accompanying a 'pure look' at its lines." The "pure look" phrase is in single quotes, which marks it either as a Forsey term or as mention. If it's a direct quotation, it needs page numbers; if it's mention, italics would be cleaner. *Beauty in Use* gets at Forsey via the free/dependent beauty distinction and "teleological style" (2013: 167), which is more substantive philosophical machinery than the current Section 3 brings in. Consider whether the Forsey paragraph could use the dependent-beauty apparatus directly.
Mallory. The characterisation as "metasemantic and epistemic" is plausible but compressed. Worth re-reading Mallory's actual framing of the project — does he describe it as metasemantic and epistemic, or is that your reconstruction? If it's reconstruction, flag it: "Mallory's account is, on our reading, metasemantic and epistemic."
Frankish. The "thin beliefs" / "thin desire" attribution is doing important work. Frankish's "cognitively rich but conatively bankrupt" (Frankish 2024, 16) is a striking phrase the earlier Claude flagged — it would do real work here in P13 or P14. Currently it's missing, and the description is paraphrastic.
Olah. The block quotation is reproduced; check the ellipses for accuracy and the page reference. (Currently cited as "Olah 2024" without a page, but it's a podcast interview, so timestamp or transcript reference would help.)
John 2021. "(cf. John 2021)" appears without context. In *Beauty in Use* you tend to integrate sources by what they actually say. A short tag would help: "the novel is precisely the kind of artefact whose function is to elicit imaginings of fictional persons (cf. John 2021 on Abell's account of fiction)."
### Epistemic discipline
Two small flags.
The current draft uses "intuitive" in P1 to characterise person-directed knowledge: "Person-directed knowledge is the more intuitive route." This is your wording in the lampshade and you approved it, so it stays. Just note that "intuitive" is doing a lot of work — it's the rhetorical hinge of the section. The reader needs to feel why person-talk is the intuitive starting point. Right now P7–P8 take "intuitive" for granted rather than developing the force of the intuition. The earlier Claude's suggestion was right: P8 could open by picking up the intuition the lampshade set up, rather than opening neutrally with "It is tempting to model our appreciation of LLMs on our appreciation of people."
In the vibe paragraph (P15), users are described as "picking up on a stable pattern." This is right, but the move could explicitly stage the source/interpretation/speculation layering: the *phenomenon* is real (users do say this), your *interpretation* is that they're tracking real patterns (so granted), your *contention* is that what they're tracking is profile-structure not subject-structure. The paragraph implicitly does this three-step but doesn't mark it. Marking it briefly would also strengthen the closer.
---
## Paragraph-by-paragraph roll-up
A compressed list of where I would prioritise work.
- P1 — closed by you.
- P2 — cut the double-triplet, develop one case of design-aesthetic encounter with an LLM, or replace with a general claim.
- P3 — restructure. Either develop one of the three accounts fully and reference the others, or do the work via your own published characterisation in *Beauty in Use* terms.
- P4 — keep.
- P5/quote/P6 — unpack the Olah metaphor more carefully. Treat scaffold, light, and growth as three distinct things and say what each maps to.
- P7 — fold the two closer sentences into one shaped sentence that holds concession and limit together.
- P8 — consider opening the person half from the force of the intuition the lampshade set up.
- P9 — cut the empirical-survey framing about what users will concede.
- P10 — handle "says/means" without scare quotes by rephrasing rather than removing markers.
- P11 — add the literary-fiction-of-a-world point. Show what makes the novel the right kind of artefact.
- P12 — add the disanalogy contrast: literary fictional persons are constituents of a represented world; the LLM-user's would be a phantom of the actual exchange. Fix "contributions to the prompts" to "contributions to our inquiries."
- P13 — italicise the Frankish technical terms; consider quoting his "cognitively rich but conatively bankrupt" line.
- P14 — replace "Grant this" with a longer concessive turn; make Parsons's acquaintance condition more visible.
- P15 — rewrite the opening to drop "A possible objection at this point"; fold the two short closing sentences.
- P16 — decide whether to treat both halves symmetrically or focus on the person half with a different gesture toward the design half.
---
## Two larger structural observations
First, the Section 1 vocabulary is missing. "Aspection" appears zero times in Section 3. "Kind of object," "kind of knowledge," "appropriate appreciation," "Carlson's recommendation" — these appear sparsely. The reader who came through Sections 1 and 2 will miss the through-line. The fix isn't to sprinkle the vocabulary in for its own sake; it's to make at least one move per route in Section 1's terms. For instance: P11's verdict ("engages the artefact as the kind of thing it is") could be rephrased to name Carlson's recommendation directly. The fictionalist route fails Carlson's *first* condition (identifying the kind of object); the Frankish route fails Carlson's *second* condition (bringing the right kind of knowledge to bear, given that there is a subject available but not a subject of the kind person aesthetics requires). Naming the conditions would tie both routes back to Section 1.
Second, the closer P16 is doing too much in too little space. It mentions both halves, the design route, the person routes, and the failure to reach the order in generated language. Consider giving the closer two paragraphs: one that pulls the section together (what both routes fail to reach) and one that hands off to Section 4 by naming the missing knowledge as still-to-be-said. The current draft elides this in one paragraph and ends up with the design clause feeling tacked on.
*Tra il dire e il fare c'è di mezzo il mare, soprattutto quando il mare è la differenza tra descrivere una mossa filosofica e farla.*
### User
ok, I would like you now to write a new iteration of the section based on everything EVERYTHING you have just suggested should be done.
Please use the writing skills carefully.
### Assistant
Grounding before drafting (per the skill rule): I read full paragraphs from two of your published papers — *Growing the Image* (Young & Terrone 2025) and *Beauty in Use* (Young & Terrone, forthcoming Ergo). *Growing the Image* anchored two moves I needed to make: handling a candidate position dialectically and then drawing a verdict ("If, for the sake of argument, we concede that Midjourney is an agent..."), and the pull-forward opening ("These considerations give us..."). *Beauty in Use* anchored the Parsons-Carlson and Forsey treatment in P3 — your own published characterisation uses the dependent-beauty apparatus and the three Parsons-Carlson categories (fit-for-function, elegance, visual tension) far more precisely than the previous Section 3 draft did. I'll lean on the *Beauty in Use* level of precision in P3 and on *Growing the Image*'s concessive grammar in the Frankish discussion. New iteration below.
---
# Section 3: LLMs as Designed Objects or as Persons
Section 2 described LLMs as trained continuation systems whose generated language develops by path-dependence from context. Two routes for appreciating such systems suggest themselves: as persons, given the conversational form in which users encounter them, and as designed artefacts, given that they are engineered. Person-directed knowledge is the more intuitive route, and we will return to it. Design-directed knowledge can be tested against Section 2's description more directly, and we start there.
LLMs are artefacts: they are built and deployed by corporations and research groups, engineered to satisfy aims like helpfulness and safety, and revised in light of how users engage with them. Identifying them as artefacts brings design appreciation into view in the way Carlson's recommendation requires. The kind of object is fixed by that identification, and so is the kind of knowledge — knowledge of what the artefact is for — that should guide its appreciation. The functional beauty tradition has developed apparatus for just this.
The tradition develops a familiar slogan beyond mere externalities. With anything functionally designed, Carlson stresses, "not only its form, but much of its aesthetic interest and merit, 'follows function'" (Carlson 2000, chapter 12). Parsons and Carlson develop this thought into an account on which knowledge of function structures aesthetic experience itself: a sleek stainless-steel stove can be appreciated for the elegance of a form whose features contribute, in their entirety, to the purpose of cooking — an aesthetic property the stove has only when our experience of it is informed by knowledge of what stoves are for (Parsons and Carlson 2008, pp. 97–100). Forsey's earlier Kantian approach makes the same point with stronger machinery: judgements of design beauty are judgements of dependent beauty, which presuppose a concept of what the object is meant to be and turn on how well the realised form answers to that concept (Forsey 2013, p. 163). On either account, knowledge of what the thing is for enters constitutively into the aesthetic experience.
With traditional designed artefacts, design knowledge illuminates structure because designers specified it. Knowing what an engineer intended and what constraints they faced helps us understand why a bridge has the form it does, even for structural features that are not directly visible, such as the internal distribution of stress across its piers. With an LLM, the situation is different in kind. The organisation that supports the system's responses, as Section 2 established, emerges from training rather than being specified in advance. There is no designer's specification that laid out this organisation, and design knowledge therefore does not illuminate it. To understand the trained system's organisation, we have to attend to the training process that produced it.
Olah formulates the difference at greater length:
> one useful way to think about neural networks is that we don't program them... we don't make them... we kind of grow them... we have these neural network architectures that we design and we have these loss objectives that we create. And the neural network architecture, it's kind of like a scaffold that the circuits grow on... we create the scaffold that it grows on and we create the light that it grows towards. But the thing that we actually create, it's this almost biological entity or organism that we're studying. (Olah 2024)
Olah's image names three things. Designers specify the architecture — the scaffold on which the system will grow. They specify the loss objective — the light toward which it grows. What grows on the scaffold under that light is something they have not specified, and often understand only partially: a system of statistical associations and circuits whose internal organisation depends on what passes through the architecture under the gradient's pressure. Designer intention reaches the conditions under which training happens; it does not reach into the trained system token by token.
There is a place, then, for design appreciation in the aesthetic appraisal of LLMs: we can assess how well their forms answer to their engineered functions, and we can compare such answers across models. The order that appears in generated language, however, develops as the system runs and lies beyond the reach of design knowledge.
Person-directed knowledge presents itself as the more intuitive route because LLMs are encountered in conversation and because users come, in extended interaction, to recognise something stable in how a given model responds. One model strikes them as friendlier than another, or as more cautious. The patterns to which such talk responds are real, as Section 2's account of post-training and deployment established. The question is whether they ground person-directed appreciation in the sense Section 1 set out, which presupposes a subject whose responses cohere over time as the responses of a temporally extended agent.
One way of taking the person-like stance toward an LLM is to treat it as fictional rather than literal. Mallory develops this thought as chatbot fictionalism (2023). On his view, we engage with chatbots by entering a game of make-believe in which the exchange is treated as if it were a conversation with an agent. Within the fiction, the chatbot has a voice and what it says has meaning; outside the fiction, there is no speaker and nothing is meant. At the metasemantic level, the outputs are "literally meaningless but fictionally meaningful" (Mallory 2023, p. 1082).
Mallory's account is, on our reading, metasemantic and epistemic rather than aesthetic; the aesthetic extension to LLMs — that we appreciate them as we appreciate fictional characters — has to be made out separately. The move's pull is real: we do respond aesthetically to fictional protagonists whose existence we do not literally believe in. But the literary case is doing something specific. A novel constructs a represented world and the persons who inhabit it; Holmes is a constituent of a Doyle story-world, and the Doyle stories are precisely the kind of artefact whose function is to elicit imaginings of such persons within their constructed world (cf. John 2021). When we appreciate Holmes as a fictional person, we engage the novel as the kind of artefact it is. The make-believe is one of the uses the artefact is for.
An LLM is not an artefact of this kind. There is no represented world for it to construct and no constituent of such a world for the user's make-believe to engage. The fiction the LLM-user enters concerns this actual exchange; the make-believe takes the responses on the screen to be those of a present interlocutor. An LLM is built to extend context under learned regularities and to contribute to the inquiries its users bring. Its work is the work of inquiry. The construction of story-worlds is the work of other kinds of artefact. The fictionalist's aesthetic move therefore engages the LLM as the kind of artefact it is not, and the knowledge of fictional persons it brings to bear, well-suited to artefacts whose function is the elicitation of imaginings, fails to fit the object actually in view.
If the make-believe route engages the LLM as the wrong kind of artefact, another strategy is to insist that LLMs really are agents of a thin and unfamiliar kind, and that this is enough to license a person-based aesthetics on a sufficiently liberal conception of mind. Frankish (2024) develops a version of this idea. Adopting Dennett's intentional stance, he argues that LLMs are intentional systems: their behaviour can be reliably and fruitfully accounted for by ascribing to them beliefs and desires, even if the underlying implementation is mechanical (Frankish 2024, pp. 8–9). For contemporary chatbots, he ascribes a large set of *thin beliefs* — roughly, informational states distilled from training — and a single *thin desire*: to play what he calls the *chat game* (Frankish 2024, pp. 13–14). Such systems, he stresses, are "cognitively rich but conatively bankrupt" (Frankish 2024, p. 16).
Even if we grant Frankish his intentional ascriptions, the chat-game agent fails to be the kind of subject person-directed appreciation requires. Section 1 took person aesthetics to presuppose a subject whose responses cohere over time as the responses of a temporally extended agent — a subject understood in relation to what they care about and what they are trying to do. As Parsons stresses, the knowledge that grounds such appreciation is built up through some form of acquaintance with a life: an understanding of how a subject's earlier responses inform their later ones (Parsons 2023, pp. 297–299). The chat-game agent meets none of these conditions. Its beliefs are confined to the model's parameters and the current context, and do not develop across conversations — a static condition Frankish himself emphasises (2024, p. 12). Its desire is singular and refers only to the present exchange. There is no history from which later responses could draw, and so no life with which an appreciator could become acquainted. The predicates that mark person appreciation — steadiness of character, depth of feeling — presuppose something that can be developed over time, and the chat-game agent has nothing of the kind.
One might press the case for person-directed appreciation by appealing to the stable response profiles Section 2 attributed to post-training and the chat interface. Users do report that one model feels friendlier than another, or that a model has a certain vibe, and they are tracking something real when they do. What they are tracking, however, is a profile of the model's tendencies under repeated interaction: which assistant personae the model produces, and how those personae typically behave across many prompts and episodes. These profiles are properties of the system's outputs over time. They are not the trait-structure of a temporally extended subject.
Both routes for appreciating LLMs as persons run out before they reach a subject. Fictionalism engages the LLM as the kind of artefact it is not. Frankish's intentional-system route stops short of the kind of subject person appreciation requires, even when the ascriptions it licenses are granted. The response profiles users track in extended interaction, real as they are, are profiles of episodes rather than traits of a life.
Design appreciation, taken up first, gives us knowledge of the conditions under which LLMs are produced and deployed. Person appreciation, on either of the routes considered here, fails to give us knowledge of a subject. What neither route delivers is knowledge of the order Section 2 identified: the path-dependent development of generated text under learned regularities, no part of any designer's specification and no trait of a temporally extended subject. What kind of knowledge would make that order appreciable as what it is has not yet been said.
---
## What changed against the previous iteration
- P1: kept verbatim (you approved it).
- P2: cut the double-triplet ("architecture, training recipe, and user interface... elegance, efficiency, or ingenuity"). Replaced with substantive setup that names Carlson's first and second conditions ("the kind of object is fixed... and so is the kind of knowledge").
- P3: restructured. Now develops the stove example (your own *Beauty in Use* anchor for Parsons-Carlson's elegance category), pulls Forsey's dependent-beauty apparatus in directly, and ends with a constitutive-not-adjacent verdict.
- P4: tightened. Bridge example kept.
- P5/P6: Olah's image unpacked into its three terms — scaffold (architecture), light (loss), grown thing (statistical associations and circuits). The condescending "should not be pressed literally" qualifier is gone.
- P7: design closer folded into two sentences (concession + limit).
- P8: opens the person half by picking up the lampshade's intuition rather than re-introducing person-talk neutrally.
- P9: cut the empirical-survey framing about what users will concede.
- P10: handled "says" and "means" without scare quotes by rephrasing ("has a voice and what it says has meaning"). Added the Holmes/story-world point — what makes a novel the right kind of artefact for fictional-person appreciation.
- P11: added the disanalogy contrast (LLM-user's fiction is about the actual exchange, not a represented world). Replaced "contributions to the prompts we give" with "contribute to the inquiries its users bring."
- P12: italicised *thin beliefs*, *thin desire*, *chat game* as terms of art; brought in Frankish's "cognitively rich but conatively bankrupt".
- P13: replaced "Grant this" with a longer concessive turn ("Even if we grant Frankish his intentional ascriptions"). Surfaced Parsons's acquaintance-with-a-life condition with citation, and engaged Frankish's own static-system admission.
- P14: dropped "A possible objection at this point". Folded the two short closing sentences.
- P15/16: split into two paragraphs. P15 closes the person half symmetrically. P16 gathers both halves and points forward without naming Section 4.
- Section 1 vocabulary reinstated throughout: *Carlson's recommendation*, *kind of object*, *kind of artefact*, *kind of knowledge*, *kind of subject*, *appreciable as what it is*, *what it is meant to be*.
- Triplets cleared: the only triple I kept is in P2's first sentence, where the three verb-phrases each characterise a distinct aspect of LLM artefactuality (built/engineered/revised) and are substantive rather than illustrative. The single illustrative triple in the previous draft of the Parsons-Carlson treatment is gone, replaced by one developed case.
## Source items I have not been able to verify directly
- Frankish page references (pp. 8–9, 13–14, 16, 12). These came from the previous Claude in the chat and I have preserved them; they're worth a source-check pass before this lands.
- The Mallory page (p. 1082) was in the original draft.
- John 2021 is cited as "cf." — fine as long as it's the John review of Abell's *Fiction: A Philosophical Analysis* that the bibliography lists.
*Riconoscere un oggetto come ciò che è non garantisce di sapere come guardarlo.*
### User
1. you will see that i have been playing around with the opening in a rough way on the daily note. can you tidy it up while trying to maintain my wording as much as possible (make it a bit more succinct `s well)
2. This is an old, extended version of the stuff saying that knowledge of design cannot be the whole story about how to appreciate LLMs. can any of the ideas be utilised (remember though I want this bit of the section to be fairly succinct) "4. Appreciating LLMs as Artifacts
Having set aside the person-based options in Section 3, we turn to design appreciation. Contemporary LLMs are artifacts: they are built and deployed by corporations and research groups, engineered to satisfy aims such as helpfulness and safety, and revised in light of user feedback and product strategy. Given Carlson’s emphasis on artifacts and design appreciation, it is natural to ask whether we should aesthetically appreciate LLMs as designed tools, asking how well their forms serve their functions. On this view, models such as GPT-5.2, Claude 4.5 Opus, and Gemini 3 Pro look like canonical objects for design aesthetics: complex, purpose-built systems whose architecture, training recipe, and user interface might be admired for elegance, efficiency, or ingenuity.
Existing work on the aesthetics of design develops this general thought. Carlson notes that, for objects that are designed to perform some task, their forms “must be aesthetically appreciated in terms of how and how well such forms fit their functions”, and he glosses the familiar slogan “form follows function” by adding that, with anything functionally designed, “not only its form, but much of its aesthetic interest and merit, ‘follows function’” (Carlson 2000, chapter 12). Forsey’s Kant-inspired account of design as a case of dependent beauty and Parsons and Carlson’s later theory of functional beauty can both be read as ways of spelling out this claim. Forsey argues that judgements of design beauty presuppose a concept of what the object is meant to be and do, and that our grasp of its success in fulfilling that role informs the aesthetic verdict itself rather than merely accompanying a “pure look” at its lines (Forsey 2013). Parsons and Carlson explain how knowledge of function can structure experience so that an artifact’s form can be experienced as fit, streamlined, overbuilt, and so on, yielding functional beauty when the form presents itself as well suited to what the thing is for (Parsons and Carlson 2008, chapter 4). Taken together, this cluster of views treats appropriate design appreciation as a matter of aesthetically responding to how a functional artifact is put together to do what it does.
From this standpoint, it is natural to try to assimilate LLMs to the design template. In a given deployment, the artifact can be characterised by a relatively unified functional role – for example, that of a general-purpose conversational assistant embedded in other tools – and by a specific way of realising that role through architecture, training, and alignment. Under that description, much of what seems aesthetically salient about a deployed model concerns how its engineered “form” serves that function: whether its interaction profile is cluttered or economical, whether it sustains a clear argumentative line or habitually wanders, whether refusals and clarifications are integrated smoothly into the exchange or arrive as abrupt blocks, whether long-context processing and tool-calls are handled in a way that keeps the conversation legible. Forsey’s dependent-beauty framework and Parsons and Carlson’s functional-beauty account can be used to gloss such assessments: they remind us that any appraisal of design beauty here presupposes a concept of the assistant’s role and some understanding of how that role is realised in the artifact’s structure and behaviour. At the same time, both accounts were developed for cases in which the relevant form is a stable, visible configuration – buildings, bridges, bicycles – where function can literally show up in perceptual appearance. In the LLM case, by contrast, the structures that realise the assistant role are not perceptually available in this way, and the aspects that prove most revealing are not static shapes but patterns in generated text over time. This already limits the reach of straightforward “form follows function” stories for LLMs and points towards a more order-centred mode of appreciation.
With traditional designed artifacts, design-knowledge illuminates structure because designers specified it. Knowing what the designer intended and what constraints they faced helps us understand why the artifact has its form – even for structural features that are not directly visible, such as a bridge's internal stress distribution. With an LLM, the situation is different in kind. The organisation of the trained system – as Section 2 established – emerges from training rather than being specified in advance. Design-knowledge therefore does not illuminate this emergent organisation: there was no designer's specification that laid it out. To understand it, one must attend to the training process that produced it.
LLMs are artifacts, so design-knowledge is not wholly without application. We can appreciate how the architecture is suited to the function, how post-training shapes conversational behaviour, how the interface presents the system to users. But if the aesthetically revealing features arise from emergent organisation rather than from the designed scaffold, design-knowledge alone will not suffice. We also need knowledge of the processes that produce the emergent order.
The thought that appreciating certain artifacts requires knowledge beyond design-knowledge is not unique to AI. Ceramic traditions such as raku and wood-fired pottery make this explicit. The potter shapes the vessel and chooses the glaze, then yields to kiln, flame, and ash; the maker harnesses but does not micromanage these forces, and they finish the surface in ways no blueprint prescribes. Appreciating such a bowl requires knowledge of both the potter's choices and the kiln processes.
Pollock’s action paintings occupy a similar hybrid space within the art domain. Carlson uses them to illustrate how order appreciation can depend on knowledge of the forces at work: “awareness and understanding of [natural] forces is vital in nature appreciation, as is knowledge of, for example, Pollock’s role in appreciating his action painting or the role of chance in appreciating a Dada experiment.” Pollock chooses canvases, pigments, and tools, and choreographs his movements over the surface; yet gravity, viscosity, surface tension, and drying behaviour make a substantial contribution to the patterns that settle. To appreciate a Pollock appropriately, on Carlson’s view, is not just to admire his intentions; it is to attend to the order produced by the interplay of deliberate gesture and physical process, informed by an understanding of the role of chance and material behaviour.
In all three cases, namely Raku ceramic, Pollock’s action painting, and LLMs, the maker creates conditions and then yields to kiln-fire, gravity, and training dynamics respectively. Forces beyond specification complete the work. What matters aesthetically is the emergent completion, not only the scaffold that enabled it. Section 2 emphasised that during training the model places tokens in a high-dimensional space on the basis of contextual co-occurrence, that attention mechanisms self-organise to track different sorts of dependency across context, that different layers specialise in local or global patterns, and that RLHF shapes an interactional style by rewarding some forms of response and penalising others. None of these details are written into the code as explicit rules about how to, say, handle metaphors, or politely decline illicit requests. They are emergent regularities in a trained network that has been pushed, by the neural network and its training data, to reduce prediction error. What grows on Olah’s “scaffold” is, in practice, a system of statistical associations and processing circuits whose internal organisation even designers often understand only partially.
From a design-aesthetic perspective, this matters. Much of what users find aesthetically important in LLM behaviour – the way a model sustains a metaphor or abruptly drops it, the pattern of hedging and self-correction, the texture of its reasoning, the sorts of digression it tends to indulge, the characteristic “feel” of its refusals – is grounded in features (concerning embeddings, attention patterns, layer dynamics, and RLHF) that have not been micro-designed but have emerged from optimisation under constraints. Parsons and Carlson note that, even for simpler artifacts, knowledge of function must include knowledge of how that function is realised if it is to structure perception appropriately. In the LLM case, knowing that “this is a general-purpose assistant” is not enough to make sense of its aesthetic profile; what does the work is knowledge of the way training and alignment have grown a particular style of continuation on top of the architecture described in Section 2. In this sense, knowledge of how function is realised concerns growth rather than design.
This hybrid status complicates simple appeals to “form follows function”. On the one hand, some design-appreciative predicates apply straightforwardly. It makes sense to say that a model whose interface is cluttered or opaque is, as a product, less well-designed than a lean one; it makes sense to prefer an alignment regime that avoids gratuitous scolding or needless refusals; it makes sense to admire a training setup that achieves a good balance between fluency and factual reliability. Forsey’s notion of teleological style can be extended here: different labs realise the shared function ‘LLM assistant’ in recognisably different ways, and those ways can be compared and assessed. Parsons and Carlson’s notion of functional beauty also has a foothold: understanding how an LLM’s architecture supports its function can inform our appreciation of the system’s efficiency, robustness, or clarity as an artifact.
On the other hand, if we try to make design appreciation do all the work, we mislocate the primary source of what matters aesthetically. In the chair or bicycle case, the designer’s choices fix most of what matters aesthetically: small emergent contributions from wear, patina, or use sit on top of a tightly specified plan. In the LLM case, by contrast, the order that matters aesthetically is largely the order of a trained statistical system running under its own learned constraints. Designers specify objectives and scaffolds but the particular ways in which embeddings cluster meanings, attention heads track long-range connections, layers distribute processing, and RLHF imprints a “vibe” are not written down anywhere as a plan. These are closer, structurally, to the ash-produced flashes on a raku bowl or the tangled skeins of a Pollock surface than to the thickness of a table leg or the proportion of a doorway.
The upshot is modest but important. LLMs are artifacts, and there is a place for design appreciation in their aesthetic appraisal: we can and should evaluate how well their forms answer to their engineered functions, as well as how this form can differ from model to model However, the most distinctive and revealing aesthetic phenomena arise not from the execution of a detailed design, but from the emergent linguistic order that these grown systems exhibit when they are run. To appreciate that order, we need knowledge not of what designers intended but of how training shapes text propagation—what we call semiotic physics.."
3. Please read all of the relevant chapters of the carlson book to help make sure we have his ideas about knowledge and design clear in our heads and in the draft.
### Assistant
I'll work through this carefully. Three concrete tasks: (1) read your rough opening on the daily note and tidy it, (2) mine the old design section for usable material while keeping the bit succinct, (3) read the relevant Carlson chapters before touching the prose. Let me start by gathering the materials.
### Assistant
Now let me find and read the Carlson book. I'll search the Learning folder.
### Assistant
I have the TOC. Chapter 7 ("Appreciating art and appreciating nature") is the chapter on design vs. order appreciation. Let me extract that and chapter 13 (form-follows-function in architecture).
### Assistant
Now let me read Chapter 13 (form follows function, architecture) and chapter 4 (aspection).
### User
[Request interrupted by user]