Here is the plan. --- - It seems obvious that LLMs are a type of tool. - We use Claude, or ChatGPT, or whatever, to do this or that. - Erik Hoel gives that ordinary thought a sharper form in a recent essay, "Bits In, Bits Out." - His framing is Bergson's: we are *Homo faber*, tool-makers; LLMs are the latest in a line stretching back 3.3 million years to stone axes. - If LLMs were a genuine source of surplus intelligence, Hoel argues, that surplus should have shown up first in writing — the domain closest to what these systems are made of. - > "Yet, beyond mass-producing stilted emails and stilted social media posts and stilted essays, the impact of LLMs on writing itself has not really been to improve or accelerate good writing overall. We are not in a glut of good writing. We are in a dearth of it." - I find something right in this — the public effect of these systems on prose has often looked like a flood of stilted language, and it would be absurd to pretend otherwise. - What I am less sure of is whether "tool" is the right description of the thing that has produced this effect. - Hoel's opening helps explain why the category of tool feels natural in the first place. - He describes finding a stone axe on a beach in Cornwall as a child — recognising it immediately for what it was. - There is something to this: paradigmatic tools are recognisable as tools. - But notice that in many of these cases, recognisability travels together with a rough grasp of what the thing is for. - Hoel's stone axe is not just recognisable as a tool; it is recognisable as a tool for cutting. - A hammer is not just recognisable as a thing-you-swing; it is recognisable as a thing for driving nails. - If that is what ordinary tool-recognition looks like — recognition accompanied by at least a rough functional grasp — then what happens when we try to place ChatGPT under the same description? - One reason ordinary tools are usually easy to classify is that they tend to have a more or less specifiable proper function. - By this I do not mean just anything an object can be used for, but the use that distinguishes what it is *for* from the uses to which it can merely be put. - A hammer can hold down loose papers, and a heavy book can prop open a door, but these are accidental functions rather than proper ones. - Nor does the point collapse the moment we turn to multi-functional objects: a Swiss Army knife has several proper functions, not none. - That is why the familiar examples still work — a hammer is for hammering, a vacuum cleaner is for vacuuming, Google is for searching. - Once we ask that question of ChatGPT, the obvious answers all seem partial, unstable, or wrong in different ways. - Perhaps its function is to predict the next token; but that is a description of mechanism, not of use — nobody opens ChatGPT in order to predict tokens, any more than we describe the function of the heart as contracting rhythmically. - Perhaps it is for chatting; Keith Frankish has suggested that LLMs can be understood, from the intentional stance, as wanting to play the chat game — but that fits badly with code generation, translation, or philosophical use. - Hoel's suggestion — writing — catches something real, but not quite enough; I want to come back to it. - "Helping with tasks" is so general as to individuate nothing — it is as though one described the function of a hammer as "helping with projects." - The difficulty is not that these answers are all wrong, but that none of them settles the question the way "for hammering" settles it for a hammer. - That does not show that LLMs are not tools; but it does suggest that they are a quite different type of tool, or not quite a type of tool at all. - I am not claiming to have proved that LLMs cannot possibly be tools — the force of the previous step is classificatory hesitation, not metaphysical victory. - I am claiming that the tool description is less straightforward than it first sounds. - The issue is not just that LLMs can be used for many things — so can Swiss Army knives, and nobody finds that puzzling. - The issue is instability at the level of ordinary functional description: none of the obvious answers to "what is it for?" quite works, and the answers that come closest tend to describe mechanism or to be so general as to describe everything. - And that uncertainty about proper function quickly becomes uncertainty about evaluation. - If we do not know clearly what a thing is for, it is much harder to say in advance what would count as a good test of it. - We can test a hammer by asking whether it drives nails; we can test a search engine by asking whether it returns relevant results; in each case the proper function gives us the standard. - When there is no settled proper function, test-selection becomes a substantive question rather than something we can take for granted. - I am not saying that no test of LLMs is possible — that would be absurd. - I am saying that when Hoel selects writing as the privileged proving ground, that choice requires more argument than his framing initially suggests. - But there is a straightforward reason to start with writing. - These systems operate through language, are trained on language, and produce language; so it is not at all arbitrary to think that writing is where their character should show first and most clearly. - Hoel puts the point vividly: - > "for an LLM, words are its womb, its mother, its literal atoms — yet their impact on writing as a whole has been mostly to generate mountains of slop" - He is right about the constitutive connection: if there is surplus intelligence in these systems, text is the place it ought to appear. - The question is not why one would look at writing — that question has a good answer. - The question is what exactly we are measuring when we look there. - The difficulty is that "writing" is too coarse a heading for the very different things people do with text. - A finished essay, a published book, a social media post — these are texts as artefacts, meant to stand on their own and be judged as products. - But when someone throws a half-formed idea at an LLM and uses the response to sharpen what they actually think, the text is not the product — it is the working surface. - Hoel's evidence — the Amazon data, the social media slop, the unchanged quality of the best books — largely concerns text as artefact. - That is a reasonable thing to measure, but it is not the only thing that goes on under the name of "writing." - [OPTION: whether to add a sentence here listing other text-uses (exploration, testing, feedback) or whether the artefact/working-surface distinction is enough on its own — I lean toward the latter, since listing risks triplet-padding, but you may want more concreteness] - What Hoel mostly measures is text as artefact; the practice I want to describe treats text as a working surface. - A finished blog post or book is meant to be read and evaluated on its own merits. - A prompt-response-revision loop may not produce anything worth publishing — and that is not a failure, because publishing was not the point. - The point was to test a distinction, or to discover what you actually think by responding to what the system sends back. - "Has writing improved?" presupposes that the relevant success condition is improvement in the quality of end-products. - If much of the interesting practice is not aimed at end-products in the first place, then the question is too blunt to capture what is going on. - The difference becomes clearer once we compare an LLM not to a pen but to a system that sends material back altered. - A pen makes marks on a surface; it extends my ability to inscribe; it does not return a proposal, a misreading, or a line of continuation I had not seen. - A pen does not prompt me back. - When I use an LLM, I write something — or half-write something, or throw a distinction at the system in rough form — and what comes back is language already reshaped by the model's learned patterns: sometimes flatter than what I wanted, sometimes unexpectedly connective, sometimes productively irritating. - I then have to decide what to reject, what to sharpen, and what pressure to reapply. - We are not only prompting the system; the return prompts us back. - [OPTION: the existing draft has a fuller version of this — "sometimes wrongly confident" is a fourth item I cut to avoid a quadruplet list. You may want it back, since each names something genuinely different about what the system returns.] - Given that recursive structure, Hoel's evidence is not refuted — but it is being measured at the wrong level of description. - I am not denying the slop, and I am not denying that public prose has often worsened. - What I doubt is that this settles the character of the system. - A lot of bad prose may show what happens when returned text is treated as a product to be published rather than as material to be resisted or worked through. - The same technology supports two quite different modes of use: one that floods the zone with generic artefacts, and one in which returned text functions as part of a thinking process. - Hoel's evidence tells us something about the first mode; it is less clear that it tells us much about the second. - The category I want to borrow from *Growing the Image* is medium. - I do not mean merely that the system lies between a user and an outcome — in that weak sense almost anything could count as a medium. - I mean something closer to what Wollheim and Thomson-Jones describe: a structured field of resources and practices whose characteristic resistances and possibilities become visible only in the work itself. - Plenty of tools give the user less than perfect control; that alone does not make something a medium. - What makes the difference is that the system does not simply wait to be directed toward a fixed end — it pushes back, filters, smooths, distorts, and sends something back that has to be dealt with in the working itself. - A tool is ordinarily understood by what it is for; a medium is understood by the characteristic way it makes work proceed. - [OPTION: this is where the "I want to leave the agent question aside" move could go, if you want it — something like: "I am not interested, in this essay, in asking whether LLMs are minds or collaborators in the full interpersonal sense; the weaker claim is enough." I have left it out since the Codex plan does not give it a paragraph, but it may be worth a sentence or two here as a qualification.] - The comparison that has seemed most useful to me here is Frippertronics. - Robert Fripp and Brian Eno set up two Revox reel-to-reel tape machines so that what Fripp played into one returned from the other a few seconds later, and then returned again after another pass through the loop. - The loop does not merely preserve what went in — it sends it back changed; delay, thinning, layering, and accumulation are not accidents external to the practice but the very conditions under which it proceeds. - What Fripp is doing is not well captured by saying that he is merely using recording devices as tools — of course the tape machines are tools in one ordinary sense, but the interesting fact is the mode of engagement. - He plays into the system, hears the delayed return, and adjusts subsequent playing in light of that return; the loop becomes part of the process by which the next intention is formed. - The comparison should not be pushed in the wrong way: I am not claiming that LLMs produce the textual equivalent of a new musical genre, but that in both cases the interesting fact concerns the method of making — what the return from the system does to the next move of thought. - Hoel is right to insist on slop, and any account that tried to wave it away would deserve to fail. - One of the most useful things in his essay is the sense that the public effect of these systems has often been a thinning of language rather than an enrichment of it. - There is a musical analogue: in Alvin Lucier's *I Am Sitting in a Room*, a spoken passage is played back into a room and re-recorded again and again, so that the room's own resonances gradually take over — the content is not developed or deepened but washed into the characteristic signature of the system. - Something similar happens with LLM use when there is too little resistance: text passed through the model without discrimination gets drawn toward the model's easier habits — familiar transitions, familiar emphases, familiar shapes of explanation. - Slop is what happens when that tendency is allowed to dominate the exchange. - But this does not straightforwardly show that LLMs are tools; it may instead show what happens when a medium with strong generic tendencies is used with too little resistance from the person inside the loop. - Hoel asks, in effect, whether writing has improved, and his answer is no. - From this he concludes that LLMs are tools: bits in, bits out. - I think he asks the right empirical question at the wrong level of description. - The novelty here does not lie in the possibility of producing a new kind of sentence; it lies in a mode of making — a recursive practice in which language is sent into a system, returned in altered form, and then either resisted or pursued by the person who receives it. - Hoel is right that there has been no textual singularity, and right that what has happened to public writing tells us something. - What I doubt is that it tells us that LLMs are tools in the familiar sense his opening picture encourages. - They may be a quite different type of tool, or not quite a type of tool at all — and that is worth taking seriously before we conclude that the current technology is merely another stone in *Homo faber*'s long line of rocks. --- A few notes on choices I made and things worth kicking around: On the opening block quote: I used the "stilted emails... dearth of good writing" passage rather than the "bits in, bits out" conclusion because it captures Hoel's *evidence*, not just his slogan. The slogan then arrives in paragraph 15 as a payoff. But you could swap them — put "You put more bits in, you get better bits out. Fine. That's a tool" in paragraph 1 instead, and keep the longer slop passage for paragraph 14. Depends on whether you want the reader to encounter Hoel's argument or his conclusion first. On the stone axe (paragraph 2): I have kept it brief — one sentence of description, then the move to recognisability and function. Your editorial comment on the bullet draft says "We definitely don't need to go into his Neolithic Axe anecdote," but the Codex plan uses it. I think one sentence is the right balance. If you want less, you could cut directly to "Paradigmatic tools are recognisable as tools" without mentioning the anecdote at all. On paragraphs 8 and 9 (artefact vs. working surface): These are the least grounded in existing draft material. The current Codex Revision draft does not make this distinction explicitly — it goes from "writing is too blunt" straight to the pen comparison. Paragraphs 8-9 are filling the gap the Codex plan identified. I am not fully sure the sentences I have here are earning their keep. The risk is that "text as working surface" sounds like jargon, even though I have tried to cash it out concretely. You may want to rethink how these two paragraphs work, or compress them into one. On medium placement (paragraph 12): The Codex plan delays it until the reader has been shown why "tool" and "writing" are both too blunt. I agree with this — in the current draft, medium arrives too early and has to do too much heavy lifting before the reader has seen the problem clearly. But delaying it means paragraphs 8-11 go without the positive category. That is four paragraphs of "here is why the existing description does not work" before you say what works better. If that feels too long, one option is to drop a hint earlier — something like "I will come back to what better description might be available" — but I have not done this because it is meta-commentary (talking about what you are going to do rather than doing it). Another option: cut paragraphs 8 and 9 to one paragraph, so there are only three paragraphs between Hoel's writing (para 7) and medium (para 12). On Frippertronics placement (paragraph 13): The Codex plan puts it here so it arrives as clarification of something already shown, not as an imported analogy. I followed this, but it means the essay's most vivid material comes very late. You might want to pull it forward to paragraph 11 or 12 and introduce medium after Frippertronics rather than before. The argument would then be: show the recursive structure (pen, para 10) → make it vivid (Frippertronics, para 11) → name the category (medium, para 12). That has the advantage of letting Frippertronics *motivate* medium rather than *illustrate* it. Worth considering. On the agent-question move: The current draft has a paragraph that says "I want to leave the agent question aside. I am not interested in whether LLMs are minds or collaborators." The Codex plan does not give this its own paragraph. I have marked where it could go (paragraph 12, as a qualification). I think it is worth keeping as one or two sentences — it prevents a misreading where people think you are arguing LLMs are collaborators or minds. But it does not need a whole paragraph. On the compression/Chiang idea: This appears in the brainstorm note from the Sunday session — LLMs as a "compressed representation of the training corpus" and the loop as "compression–decompression–compression." I have not included it because the Codex plan does not use it and because you flagged "compression loop" as jargon to be removed. But the underlying idea — that what the model does to your text is analogous to lossy compression — is genuinely interesting and might earn a sentence somewhere in paragraph 10 or 13. If you want it, it could go after the Frippertronics description: "What passes through the loop is not mere text in a thin sense; it is language carrying thought, and what comes back has been reorganised by the model's learned patterns — some specificity lost, some connections surfaced that had not been salient before." That sentence is already in the current draft and it captures the compression idea without using the word. *Ogni strumento ha una funzione che lo rende riconoscibile — tranne quelli che ci costringono a inventare nuove parole per ciò che fanno.*