# created ```dataview LIST WITHOUT ID file.link FROM -"windsurf" WHERE file.cday = date(this.file.name) AND !startswith(file.folder, "windsurf") SORT file.cday ASC ``` # modified ```dataview LIST WITHOUT ID file.link FROM -"windsurf" WHERE file.mday = date(this.file.name) AND !startswith(file.folder, "windsurf") SORT file.mday ASC ``` --- # [[diary and thoughts]] #thought #diary --- # Notes ## Introduction Part of the reason paintings, novels, or films merit [[aesthetic appreciation]] is that they are [[the result]] of their makers’ efforts. A painter, writer, or director may spend months or years honing a work through sustained attention and revision; viewers or readers often experience these works precisely _as_ the products of such capacities. The 1069 pages and 388 endnotes [[David Foster Wallace]]’s _Infinite Jest_ make [[the author]]'s meticulous attention to detail apparent. For the last few years, however, generative AI has become adept at producing images, stories, and movies on demand. It does so without effort, attention, or skill, and extremely quickly indeed. It would seem, then, that at least one of the reasons why human-made works merit appreciation does not seem as though it can be applied to AI-generated works. An AI image might be pleasing to one's eye, or an AI song pleasing to one's ear, but we can doubt that this sort of  'eye candy' and 'ear candy' merits [[aesthetic appreciation]] for the same reason that actual candy does not merit [[aesthetic appreciation]]. A _Snickers [[Limited Edition Extreme]] Caramel and Nuts_ chocolate bar will taste delightful to someone with a sweet tooth, but it does not seem to merit [[aesthetic appreciation]] in the same way that _Infinite Jest_ does, as it was not was not produced, nor does it exhibit, care, attention, effort etc.  Is this all AI art can amount to, aesthetically? An efficient means of producing aesthetically worthless digital candy? Here, we argue 'no'. Generative AI does merit [[aesthetic appreciation]], but not for the reasons that traditional artworks do. Rather, an aesthetics of AI should be modelled on _environmental aesthetics_. Drawing on Carlson's work this topic, as well as Janus' conception of LLMs as simulators, we argue that individual conversations with an LLM can be thought of as instances of temporally evolving [[generative environments]], and therefore merit [[aesthetic appreciation]] in something like the way that the [[natural environment]] does. - [more here about the [[second half]] of the paper] ## 1. [[Appreciating Nature]] as a Generative Environment Carlson’s _Natural Environmental Model for environmental aesthetics rests on two ideas:_ > First, that, as in our appreciation of works of art, we must appreciate nature as what it in fact is, that is, as natural and as an environment. Second, it recommends that we must appreciate nature in light of our knowledge of what it is, that is, in light of knowledge provided by the natural sciences, especially the environmental sciences such as geology, biology, and ecology. (2000 p. 6) Regarding his first point—that we must appreciate nature as what it in fact is—Carlson argues that the environment is a system of interconnected elements shaped by various processes and forces, and not simply a collection of objects or scenes (ibid. p. 44). Environments thus “come about ‘naturally,’ [in that] they change, grow, and develop by means of natural processes.” From this perspective, aesthetic appreciation involves recognising _what_ one is encountering—a natural environment in which components are interrelated—and understanding it in light of scientific or other relevant knowledge that illuminates its composition and development. Just as an informed grasp of artistic traditions can enrich one’s appreciation of an artwork, Carlson argues that familiarity with geology, biology, or ecology can guide our attention to patterns and processes that might otherwise remain unnoticed. For instance, one might begin by noticing only the colours or shapes of a coastal cliff’s sedimentary layers. However, discovering that these layers formed over thousands of years of deposition and compaction reveals changes how one regards it aesthetically. If we see a forest or reef as subject to diverse forces and processes, an appropriate aesthetic engagement will centre on how those forces have shaped what we observe. Although Carlson does not the word 'generativity,' his first recommendation—that we appreciate nature as both natural and as an environment—fits easily with this term. To see an environment as natural is to recognise that its features arise from autonomous causal processes rather than from design. This distinction can be articulated using Spinoza’s concepts of _natura naturans_ and _natura naturata_. For Spinoza, _natura naturans_ refers to nature as an active, self-creating system—substance and its attributes, or the immanent causal laws that govern all things. It is nature in its dynamic, productive aspect. In contrast, _natura naturata_ refers to the products of this activity: the collection of individual modes, or the particular things and events that constitute the universe. Appreciating nature as 'natural', in this sense, is to apprehend its phenomena (_natura naturata_) as the determinate outcomes of its underlying generative processes (_natura naturans_). To see it as an environment, then, is to attend to the unity of these products within the single system from which they arise. Taken together, Carlson’s recommendation asks us to appreciate both process and product as inseparable aspects of one generative whole. Carlson’s second recommendation—that aesthetic judgement be informed by the natural sciences—strengthens this reading. Geology, biology, and ecology investigate the forces that generate the very phenomena we perceive; scientific knowledge therefore discloses an environment’s generative history and continuing activity. Appreciating nature “in light of this knowledge” is, in effect, appreciating its generativity—the order that emerges from undirected yet law-governed processes. The various natural sciences—geology, biology, ecology, physics—are disciplines that study the processes and forces that generate natural phenomena. Appreciating nature 'in light of this knowledge' is therefore an appreciation of its generative character. A geologist appreciates costal cliff by understanding the generative forces that produced it: "geological uplift and marine erosion". Their perception of the cliff is of a generated product and a segment of "nature's ongoing processes". This principle applies across the natural sciences. A biologist appreciates a forest as an ecosystem generated by processes of growth, competition, and decay. A physicist appreciates a rainbow as a phenomenon generated by the refraction and dispersion of light through water droplets. In each case, scientific knowledge reveals the generative process, which in turn informs the aesthetic appreciation of the generated product. Note that this pluralism in understanding the environment should not be mistaken for an 'anything goes' approach. A framework is only admissible if it provides a correct account of the generative processes in question. Phrenology or vitalism, for instance, were unilluminating because they posited false causal connections (between cranial features and character, between living matter and a vital force), thereby failing to correctly identify either the generative forces or their resultant phenomena.[^1] Carlson provides a general formulation for this mode of appreciation, which he terms _order appreciation_: > On the assumption that order appreciation provides the correct model for the appreciation of nature, such appreciation has the following general _form_: An individual qua appreciator selects objects of appreciation from the things around him or her and focuses on the order imposed on these objects by the various forces, random and otherwise, that produce them. Moreover, the objects are selected in part by reference to a general nonaesthetic and nonartistic story that helps make them appreciable by making this order visible and intelligible. Awareness and understanding of the key entities—the order, the forces that produce it, and the account that illuminates it—and of the interplay among them dictate relevant acts of aspection and guide the appreciative response.(ibid. p. 119) This scientific understanding allows the observer to see "unity in what might otherwise appear as disparate features", because the cliff's shape, the waves, and the local plant life are all understood as products of the same interconnected generative system. The "organic unity" that Carlson identifies is a unity of generation. As he observes: > natural objects possess [...] an organic unity with their environments of creation: such objects are a part of and have developed out of the elements of their environments by means of the forces at work within those environments. Thus the environments of creation are aesthetically relevant to natural objects. (ibid. p. 44) The organic unity that Carlson identifies is therefore a unity of generation. The aesthetic character of the cliff, for example, is clarified by understanding it as a generated product of its environment. Geological knowledge re-frames the cliff from a set of surface features into a record of deposition and compaction over millennia, shifting the focus of appreciation from appearance to generativity. The visible strata and the tectonic forces revealed by geology thus exemplify the relationship between _natura naturata_ and the underlying _natura naturans_. This principle extends across the sciences: biology reveals processes of growth and decay, while physics examines energy flows. Each discipline offers a complementary lens on a single generative environment, allowing for an appreciation that attends both to the unity of the whole and the plurality of its orders. Carlson’s model, therefore, directs us to evaluate nature as the outcome of non-designed processes, where scientific knowledge serves to clarify the generative order already present. %%the paragraph above could be tightened up a bit, it seems somewhat redundant.%% ## 2. LLMs as Generative Environments One might think that environmental aesthetics is a poor fit for generative AI for a straightforward reason: such systems are not natural but man-made. Indeed, generative AI systems might be understood as, in a sense, *doubly* man-made. They are human-created artifacts, which are themselves created through ingestion of vast quantities of other human-created artifacts (texts, images, audio). Such a worry can be assuaged by noting two things. First, Carlson is happy to extend his account to _man-made_ environments: > environments typically are not the products of designers and typically have no design. Rather they come about “naturally,” they change, grow, and develop by means of natural processes. Or they come about by means of human agency, but even then only rarely are they the result of a designer embodying a design. In short, the paradigm of the environmental object of appreciation is unruly in yet another way: neither its nature nor its meaning are determined by a designer and a design. (Carlson p. xiii) What we think Carlson is getting at here is that while environments grow, change, and develop by means of natural processes, man is able to initiate, or guide, or curtail these natural processes. This leads to a reasonably intuitive distinction between wholly natural environments (a forest, a swamp), man made environments (a specially planted timber forest, a garden), and man-made _places_ (a department store, a gym). A gym is not understood as an environment because it did not come about, or change, or grow, through natural processes. It is as 'artificial' as a vaccine or a tennis racket. > (footnote: it is of course quite hard to say exactly what counts as a natural force or not. It could be argued that the will and efforts of humans are also natural. We will not go into this question here, but rely on this as a rough and ready distinction). Second, AI engineers themselves talk in these terms. Consider the following from Chris Olah, one of the co-founders of Anthropic: > I think one useful way to think about neural networks is that we don’t program and we don’t make them. We kind of, we grow them…we have these neural network architectures that we design and we have these loss objectives that we create. And the neural network architecture, it’s kind of like a scaffold that the circuits grow on, it starts off with […] random things and it grows...And so we create the scaffold that it grows on and we create the, you know, the light that it grows towards. But the thing that we actually create, it’s this almost biological, you know, entity or organism that we’re studying. The outcome in each case is a system shaped by undirected forces, forming a unified whole. LLMs, understood this way as grown generative systems, thus align with environmental models, and their proper appreciation begins with examining their core function as next-token predictors. In this section, we adopt the two-part method for appropriate aesthetic appreciation outlined in §1. First, in order to appreciate the LLM _as what it is_, we must understand its actual operational nature; that is, the specific computational architecture and statistical processes that govern its function. This will be the focus of 2.1, which also serves as a primer for non-specialists on how LLMs are trained and operate. In 2.2 we turn to the question of illumination, and argue that Janus's simulator theory (2022) offers a promising light in which to understand, and thereby appreciate, LLMs. ### 2.1 What LLMs in fact are Large Language Models (LLMs) are neural networks trained to predict text. At the most basic level, these systems learn patterns in language by processing vast amounts of textual data through a training method called self-supervised learning. This approach differs fundamentally from traditional programming, where explicit rules are coded; instead, the model discovers patterns through exposure to examples. The training process works by presenting the model with sequences of text broken into smaller units called tokens—these can be words, parts of words, or punctuation marks. For any given sequence, the model learns to predict what token should come next. For example, given the sequence "The cat sat on the", the model might learn that "mat", "floor", or "chair" are likely continuations, while "purple" or "quickly" are less probable. This training happens through millions or billions of examples, with the model adjusting its internal parameters each time its prediction differs from the actual next token in the training data. The architecture consists of layers of artificial neurons with weighted connections between them. During training, these weights are adjusted through a process called backpropagation, gradually encoding patterns that allow increasingly accurate predictions. The training objective is straightforward: minimize the difference between predicted and actual next tokens across the entire dataset. This creates a system that has effectively compressed the statistical patterns of language into its parameters. When deployed for text generation, these models operate through what is called autoregressive generation. Starting with an initial prompt, the model predicts the next token, adds that token to the sequence, then predicts the subsequent token based on the extended sequence, and continues this process iteratively. Each prediction takes into account all previous tokens in the sequence, allowing the model to maintain context and coherence across longer passages. A crucial aspect of this generation process is its probabilistic nature. Rather than producing a single definitive next token, the model outputs a probability distribution over all possible tokens in its vocabulary. For instance, after "The weather today is", the model might assign high probabilities to words like "sunny", "cloudy", or "beautiful", and lower probabilities to less likely continuations. The actual token selected is sampled from this distribution, introducing controlled randomness that allows the same prompt to generate different outputs on different runs. This sampling process creates what can be thought of as branching paths of possible text continuations. At each step in generation, the model could take any of thousands of different directions, with the probability of each path determined by the patterns learned during training. The temperature parameter used during sampling controls this randomness—higher temperatures lead to more diverse and unpredictable outputs, while lower temperatures make the model more likely to choose high-probability tokens. The model itself contains no explicit database of facts, no symbolic reasoning system, and no programmed understanding of grammar or meaning. Instead, it consists entirely of numerical parameters—billions or trillions of them in modern LLMs—that encode statistical regularities extracted from the training data. These parameters implicitly capture patterns at multiple levels: letter combinations, word formations, grammatical structures, stylistic choices, factual associations, and even logical relationships, all emerging from the single objective of predicting the next token. It is important to note that while we often speak of these models "understanding" or "knowing" things, what they actually possess is a sophisticated capacity to recognize and reproduce patterns. When a model correctly completes "The capital of France is Paris", it is not accessing a stored fact but rather reproducing a pattern that appeared frequently in its training data. This pattern-matching extends to more complex behaviors—the model can appear to reason, create, or converse because these activities, when expressed in text, follow learnable patterns. The computational process during text generation involves the prompt passing through the network's layers, with each layer transforming the representation based on its learned parameters. Modern LLMs use an architecture called transformers, which employ attention mechanisms allowing the model to dynamically focus on different parts of the input when making predictions. This enables the processing of long-range dependencies in text, where the meaning of a word might depend on context from much earlier in the passage. This technical architecture—a neural network trained to predict tokens that can be run recursively to generate sequences—constitutes the fundamental mechanism of LLMs. Through this simple process of iterative next-token prediction, combined with vast scale in both model size and training data, these systems develop the ability to generate coherent, contextually appropriate text across a remarkable range of domains and styles. ### 2.2 The Simulator Framework as Scientific Understanding Having established the basic mechanism of LLMs, we can now apply Carlson's framework for environmental appreciation. Recall that Carlson's second recommendation is that aesthetic judgment be informed by the natural sciences, which **"investigate the forces that generate the very phenomena we perceive."** In the context of LLMs, this requires identifying an appropriate scientific framework—what Carlson calls a "general nonaesthetic and nonartistic story that helps make them appreciable by making this order visible and intelligible." Not all proposed frameworks succeed in this role. Just as phrenology and vitalism failed as frameworks because they posited false causal connections... thereby failing to correctly identify either the generative forces or their resultant phenomena, we must evaluate different conceptual frameworks for understanding LLMs based on whether they accurately capture the generative processes at work. Janus systematically examines and rejects several common framings. The agent framework fails because "GPT does not consistently act according to any particular objective (except the tautological prediction objective)." The oracle framework is inadequate because "GPT does not consistently try to say true/correct things... Spouting falsehoods in some circumstances is incentivized by GPT's outer objective." The tool framework similarly falls short because "unlike specialized Tool AIs that optimize for a particular optimand, GPT wasn't optimized to do anything specific at all." These frameworks fail precisely because they misidentify the generative process. As Janus argues: "I think that implicit type-confusion is common in discourse about GPT. 'GPT', the neural network, the policy that was optimized, is the easier object to point to and say definite things about. But when we talk about 'GPT's' capabilities, impacts, or alignment, we're usually actually concerned about the behaviors of an algorithm which calls GPT in an autoregressive loop repeatedly writing to some prompt-state." This brings us to Janus's positive proposal: understanding LLMs as simulators. "The natural thing to do with a predictor that inputs a sequence and outputs a probability distribution over the next token is to sample a token from those likelihoods, then add it to the sequence and recurse, indefinitely yielding a simulated future. Predictive sequence models in the generative modality are simulators of a learned distribution." The simulator framework succeeds where others fail because it correctly identifies the generative process. As Janus explains: "The outer objective of self-supervised learning is Bayes-optimal conditional inference over the prior of the training distribution, which I call the simulation objective, because a conditional model can be used to simulate rollouts which probabilistically obey its learned distribution by iteratively sampling from its posterior (predictions) and updating the condition (prompt)." This framework reveals the deep structure of how LLMs generate text. The semiotic physics note elaborates: "It's in this analogical sense that a simulator like GPT implements a 'physics' whose 'elementary particles' are linguistic tokens. When we experience the generated output text as meaningful, the tokens it's composed of are serving as semiotic signs. Thus we can refer to the simulator's physics-analogue as semiotic physics." The concept of semiotic physics provides a particularly rich understanding of LLM behavior. As the note explains: "Like real-world physics, the simulator's 'physics' leads to emergent phenomena of immediate significance to human beings. In real-world physics, these emergent phenomena include stars and snails; in semiotic physics, they're the stories the simulators tell and the simulacra that populate them." Janus emphasizes the distinction between simulator and simulacra: "GPT is to a piece of text output by GPT as quantum physics is to a person taking a test, or as transition rules of Conway's Game of Life are to glider. The simulator is a time-invariant law which unconditionally governs the evolution of all simulacra." This distinction is crucial for understanding how diverse behaviors emerge from a single underlying system. The simulator framework also explains puzzling aspects of LLM behavior that other frameworks struggle with. For instance, regarding agency: "GPT-driven agents are ephemeral – they can spontaneously disappear if the scene in the text changes and be replaced by different spontaneously generated agents. They can exist in parallel, e.g. in a story with multiple agentic characters in the same scene." Through the lens of semiotic physics, we can appreciate LLMs as implementing generative laws analogous to physical laws. Janus notes: "Models trained with the strict simulation objective are directly incentivized to reverse-engineer the (semantic) physics of the training distribution, and consequently, to propagate simulations whose dynamical evolution is indistinguishable from that of training samples." This understanding fulfills Carlson's criterion for order appreciation. Just as understanding geological processes allows us to aesthetically appreciate a cliff face as "a record of deposition and compaction over millennia," understanding LLMs as simulators implementing semiotic physics allows us to appreciate the generated text as the product of learned generative laws operating on initial conditions (prompts). We can thus establish the simulator framework, with its concept of semiotic physics, as suitably fruitful for understanding and appreciating LLMs. It provides what Carlson requires: a correct account of the generative processes that reveals the "organic unity" between the generated phenomena and their underlying system. The framework makes "visible and intelligible" the order by which token sequences emerge from the interplay of learned patterns and stochastic sampling. In embracing this framework, we are certainly not saying that this is the only light in which to understand, know and appreciate LLMs. Other approaches, such as mechanistic interpretability—which seeks to understand the internal representations and computations within neural networks—may prove equally fruitful for different purposes. The plurality of scientific approaches enriches rather than diminishes our appreciation, just as biology, geology, and physics offer complementary perspectives on natural environments. The simulator framework, grounded in the concept of semiotic physics, thus provides the scientific understanding necessary for aesthetic appreciation in Carlson's sense. It correctly identifies the generative processes at work, distinguishes between the generating system and generated phenomena, and reveals the lawlike regularities that govern the production of linguistic trajectories. In the next section, we will get into the nitty-gritty of exactly how this works, examining specific examples of how prompts function as initial conditions and how the interplay of deterministic computation and stochastic sampling generates the rich variety of textual phenomena we observe.