# If you look in the downloads folder, you'll see that the most recent file in it
## Retrieval Notes
- Session id: `7864c5e6-703f-4855-9af5-ba5b0a7e3512`
- Last activity: `2026-03-11T21:45:14.594Z`
## My Notes
<!-- Add your notes here. This section is preserved across syncs. -->
## Conversation
### User
If you look in the downloads folder, you'll see that the most recent file in it is a JSON of a conversation. Use that to get up to speed as to what we are are doing and the text we are working with at the moment is at this address. Don't edit it but just use it as context.
### Assistant
Those JSON files look like Obsidian config files, not conversation exports. Let me look more broadly at the most recent files in Downloads.
### Assistant
Let me try a different approach — Glob might work for external paths.
### Assistant
Those are all from an old vault backup. Let me look at the root level of Downloads for recent files.
### Assistant
Found it — the most recent file is a JSONL session export from today. Let me read it.
### Assistant
Large file — let me use a subagent to extract the conversation content.
### Assistant
OK, I'm up to speed. Here's where things stand.
The session you exported was a long one — compacted three times — that started with deciding between the Generating Philosophy paper and the Substack essay, and ended up making real progress on the essay. The draft I've just read ([[Writing/substack/LLMs Are Not Tools.md]]) reflects the essay's current state, which incorporates work from that session but does not yet include the session's most significant development: the substance idea.
## What the current draft does
It runs from Hoel's "mere tools" position through the proper function problem (LLMs lack one), through the artefact/working-surface distinction, through the pen comparison and Frippertronics, to slop-as-Lucier, and closes by questioning whether "bits in, bits out" settles things. The category it reaches for is "medium," borrowed from Wollheim/Thomson-Jones via *Growing the Image*.
## What the session developed beyond the draft
The real breakthrough was the turn from "medium" as the positive category to "substance" as the prior one:
- LLMs are not tools (no proper function) but substances — stuff with latent capacities, like iron or oil. Substances don't have functions; they have properties from which uses emerge.
- Hoel's own oil analogy gets turned against him: when we discovered oil, we didn't know what we had. Kerosene came first, internal combustion later. We may be in the "kerosene phase" with LLMs.
- The substance/medium relationship gives you two moves for the price of one: substance is what the stuff *is*; medium is what it becomes when taken up into practice. Oil is a substance; oil painting is a medium.
- Aristotelian potentiality (dynamis) fits naturally: prompt engineering, chain-of-thought, etc. are discoveries about the substance's latent capacities, not instructions to a tool.
- The epistemic conclusion sharpens: Hoel assumes we already know what LLMs are for. We don't. "Bits in, bits out" describes the current phase, not the nature of the thing.
## The CEV structure from the session
The ideal essay that emerged was roughly: (1) Hoel's position → (2) concede slop → (3) proper function fails → (4) turn: substance, not tool → (5) oil analogy reframed → (6) potentiality gesture → (7) substance becomes medium in practice → (8) recursive structure (pen doesn't prompt back) → (9) Frippertronics → (10) artefact vs. working surface → (11) slop as Lucier → (12) conclusion: we invented a substance, evaluation is premature.
## What's in the draft's editorial notes
The notes at the bottom (lines 131–148) discuss several open choices — block quote selection, whether to cut the stone axe, how to handle paragraphs 8–9 (artefact/working surface), when medium arrives, where Frippertronics goes, whether to include the agent-question qualification, and the compression/Chiang idea. These were written by the assistant during the session.
So the draft is solid from paragraphs 1–15 but hasn't yet been restructured to include the substance turn (steps 4–7 of the CEV). That's the gap between where the draft is and where the thinking got to.
What do you want to do from here?
*La sostanza precede la forma — prima si scopre la materia, poi si impara cosa farne.*
### User
i carried on the chat you just looked at with codex. this is what he said. what do you thnk?
CHAT: Yes. I have read the newest file, 797a7eb4-f119-4e0d-9298-0047aa41a7f3 (1).jsonl, and the main thing I take from it is that your thinking has shifted quite a lot.
The present live thought is no longer just "LLMs are not straightforward tools" or even "LLMs might be better thought of as media". It is that the essay may need to be built around the stronger claim that LLMs are substances: stuff with capacities, not implements with proper functions. On that picture, the proper-function argument is still useful, but it no longer just destabilises Hoel's classification. It opens the way to a different category altogether. We do not ask what iron is for; we ask what can be done with iron. Likewise, we may not yet know what LLMs are for, because we are still discovering what can be elicited from them.
That also means Hoel's oil analogy now cuts the other way. In the conversation, the strongest new move was the "kerosene phase" thought: when oil was first discovered, its transformative uses were not yet known, and early uptake did not reveal the full range of its capacities. So the right reply is not a hypey "just wait, it will get better". It is a methodological point: Hoel's argument assumes that the history is already settled, that we already know what this thing is for, and that public writing therefore gives us the right measure. The substance framing says that this may be too quick.
I also noticed that the conversation preserved a lot from the earlier tool/medium draft rather than throwing it away. The things marked for retention were:
opening with Hoel's "mere tool" position
the proper-function problem
the fact that we do use these systems, but that does not yet settle what sort of thing they are
the distinction between text as artefact and text as working surface
the pen comparison
the recursive structure where the return prompts us back
the medium material
Frippertronics
Lucier/slop
But their role changes. On the new picture, "substance" becomes the main positive category, and "medium" names what happens when that substance is taken up into practice. That seems to be the cleanest way of putting the relation between the two. The LLM is the substance; recursive text-exchange is the medium. That is a much more determinate brief than the one I was working from before.
So, in short, I do think I now see your current thinking more clearly: the essay's live centre of gravity is the substance thesis, with Hoel's tool/oil framing being answered by saying that we may still be in the kerosene phase of a newly invented stuff whose latent capacities are not yet fully understood. I will take that as the current source of truth unless you redirect it. Qui il punto non è più soltanto che “tool” sia una categoria troppo rozza, ma che forse stiamo ancora imparando che cosa sia questa nuova sostanza e come lavorarla.
Okay, so first of all, other than the very early stages of the the current version of this. Do you know which document you need to be looking at now? I'll copy the file, hold on. Anyway, I am happy for you to change or restructure things quite considerably um after the very first block quote. Okay, so the so basically I have an opening paragraph and then I have a block quote. You can do whatever you want afterwards with the block quote. You can reuse some of my stuff, you can um add your own stuff but um yeah so try and think about a new structure for what I want to do now okay so based yeah with uh the idea that it's a substance and that it's a substance which potentially at least we have not yet discovered all of its latent capacities. Can you think of the CEV of the an essay that would be recalibrated like that? And can you give me a very detailed structure please in a bullet point framework? you must invoke the following skills BEFORE DOING ANYTHING
* Skill contemplate
* Skill nick-analytic-voice
* Skill nick-philosophical-prose
* Skill twork
* Skill source-work
* Skill epistemic-discipline
* Skill writing-standards
I am invoking `contemplate`, `nick-analytic-voice`, `nick-philosophical-prose`, `twork`, `source-work`, `epistemic-discipline`, and `writing-standards` for this pass.
First, the document question. I do not know with certainty which file you mean beyond the current Substack draft lineage, because you have not pasted the file yet. My best inference is that the working document is still the current draft in LLMs Are Not Tools.md or its immediate successor, but I am treating that as an inference, not as fact. What I do know is the structural constraint you have now given me: the opening paragraph and the first block quote stay, and after that the piece can be rebuilt quite freely.
That changes the task in a useful way. I no longer need to preserve the old middle. I need to think about the CEV of the essay if it is recalibrated around the substance thought. And that means I should stop thinking in terms of "how do I retrofit the old tool/medium structure?" and instead ask what the strongest version of the essay would look like if the new centre of gravity were granted from the outset.
The first thing to notice is that the substance thought does not merely add a new metaphor. It changes the logic of the response to Hoel. On the older structure, the proper-function discussion mainly served to show that "tool" was too quick or too blunt. On the new structure, that discussion becomes a diagnostic step. It shows that the reason tool-talk breaks down is that we are not dealing, or not only dealing, with an implement whose identity is fixed by proper function. We are dealing with something more like a stuff, a material, a substance with capacities, tendencies, and potentialities. That is a more radical thought. But it is also, I think, a much more generative one.
The second thing to notice is that the substance thesis does not eliminate the medium material. It reorganises it. "Substance" and "medium" are not rival labels at the same level. A substance is what the thing is like considered as a material with capacities. A medium is what it becomes when those capacities are taken up into a practice. Oil is a substance; oil paint is a medium. Silicon is a substance; semiconducting architectures are a medium of computation. That is why the newer line from the JSON conversation felt stronger than the earlier tool/medium wavering. It gives you a clean relation between the categories instead of leaving them to compete.
The third thing is that Hoel has to remain central. The substance essay is still a response to him, and indeed one of its strongest virtues is that it answers his own oil analogy with a better reading of oil. He uses oil as the model of a transformative discovery: if LLMs were a true source of intelligence, they ought to have transformed writing by now. But the reply is that discovering a substance does not mean immediately knowing what all its powers are or which practices will best actualise them. There is a lag between discovery and mature use. There are provisional uses. There are crude uses. There are low-grade uses that nevertheless reveal something real. That is where the kerosene-phase thought does real work. It is not just a flourish. It gives historical content to the methodological point.
The fourth thing is that the essay needs to be careful not to sound like hype. This is the obvious danger of the substance framing. "We have not yet discovered all of its latent capacities" can sound like an AI booster saying "just wait". That would kill the piece. So the structure has to make it clear that the point is epistemic and methodological before it is predictive. The claim is not that the buried greatness will inevitably emerge. The claim is that Hoel's evaluative confidence presupposes a settled understanding of what the thing is for, and that this presupposition is precisely what the substance framing denies. That is much stronger and much more defensible than a vague promissory note.
The fifth thing is that the slop material becomes more interesting under the new structure, not less. Under the old structure, slop risked looking like a problem to be explained away. Under the substance structure, slop can be redescribed as one manifestation of the substance's characteristic tendencies when it is taken up under weak conditions. That is: if this stuff has powerful smoothing, averaging, and genericising tendencies, then of course one of the first large-scale social consequences of its use will be mountains of dead prose. That is not evidence against the substance framing. It is part of the evidence for how this substance behaves when it is used lazily or at scale. So the essay can concede Hoel's observation more fully without giving up the deeper point.
Now, what would the best structure look like given your fixed opening paragraph and block quote? I think the essay wants to pivot quite quickly, after the quote, into the thought that Hoel's framing only looks settled because "tool" is one of those words that shuts down inquiry too fast. Then it should move into the proper-function problem, not as an end in itself but as the route into the stronger substance claim. Once that move has been made, the oil analogy can be turned. Then the essay can slow down and explain how a substance becomes a medium in practice, which is where the pen comparison, recursive prompting, artefact versus working surface, and Frippertronics material can all be retained but made to answer a different question.
I also think the structure wants one clear paragraph that says, in effect: substances do not have proper functions, but they do have properties and tendencies. That is where the essay can say what kind of thing an LLM is supposed to be on this picture. Pattern-completion, contextual pliability, smoothing, associative reach, responsiveness to iterative prompting, and so on. Not as a list for its own sake, but as the beginning of a material profile. Otherwise "substance" remains too airy.
There is also a choice about whether to include Aristotle or Locke. I think Aristotle's potentiality language can appear briefly if it is kept functional rather than scholastic. Locke's "something I know not what" is attractive, but I suspect it may take the essay sideways unless it is used in one deft sentence. For CEV purposes, I would keep Aristotle as a light touch and leave Locke out unless you later find a way of making him pull real weight.
The software-development analogy should stay, but not as a rival main example. It belongs after the recursive-use material, where it can clarify the methodological lesson. There, it shows that one way to mismeasure a new substance or medium is to look only at the public outputs and not at what kinds of activity and iteration it enables for competent users. That gives the writing case some support without letting the essay become "actually the real case is coding".
So the shape I now want is this: fixed opening and quote; immediate pressure on "tool"; proper function; turn to substance; answer to oil via kerosene phase; substance-to-medium transition; recursive practice; artefact versus working surface; software as methodological analogy; slop as characteristic tendency; return to Hoel with a narrower conclusion. That feels like the CEV structure.
The remaining question is how much detail belongs in the scaffold. Since you want the detailed framework, I think the right thing is to give you paragraph blocks with explicit argumentative jobs, candidate claims, examples, and the relation between paragraphs. Not polished prose, because you asked for structure; but not thin one-line labels either. The result should be something you could actually draft from without having to reinvent the argument.
Final Answer
I do not know for certain which file you mean until you paste it. My best inference is that the working document is the current Substack draft line, probably LLMs Are Not Tools.md, but I am treating that as an inference.
Given your constraint, I am treating the opening paragraph and the first block quote as fixed, and rebuilding everything after that around the substance thesis.
What follows is the CEV structure as I currently see it.
Bullet framework
Fixed opening
Keep your present opening paragraph.
Keep the first block quote from Hoel.
Treat that block quote as the last moment in which the reader is allowed to feel that the category of tool is straightforward.
Paragraph 1: the pivot after the quote
Start by granting that Hoel has identified something real: people do reach for the category of tool almost automatically here.
Then turn immediately to the pressure point: "tool" sounds clarifying, but in this case it may actually be one of those words that makes inquiry stop too early.
The paragraph should not yet announce "substance". It should make the reader feel that Hoel's classification is more settled than the object deserves.
The closing sentence should open the proper-function question: if this really is a tool, what sort of tool is it supposed to be?
Paragraph 2: proper function enters
Introduce the proper-function thought calmly and without drama.
The point here is not that all tools are single-purpose. It is that even fairly versatile tools usually admit a more or less stable answer to the question what they are for.
Use two or three quick cases: hammer, Google, Swiss Army knife.
Make the contrast exact: multi-functionality is not the problem; the problem is instability at the level of ordinary use-description.
Paragraph 3: the obvious candidates fail
Run the standard answers one by one.
"It predicts tokens" gives a mechanism, not a use.
"It chats" is too thin.
"It helps with tasks" is vacuous.
"It writes" catches something true, but not enough to settle what sort of thing we are dealing with.
The tone should be exploratory rather than triumphant. The point is not that you have disproved toolhood, but that the expected answer does not arrive.
Paragraph 4: the first verdict
This is where your line belongs.
Say that the lesson is not yet that LLMs are not tools in any sense whatsoever.
Say, rather, that they look like "a quite different type of tool, or not quite a type of tool at all".
Then add one more step which the earlier versions sometimes missed: if the category begins to wobble here, the standards of evaluation begin to wobble with it.
Paragraph 5: the methodological consequence
Make the point explicit that uncertainty about proper function becomes uncertainty about testing.
Do not say that no test is possible.
Say that once we no longer know straightforwardly what the thing is for, we no longer know straightforwardly which effects should count as the decisive measure of it.
This sets up the return to Hoel's writing argument, but now under pressure.
Paragraph 6: Hoel's case in its strongest form
Re-state his writing case carefully and charitably.
Use the block-quoted line about words being its "womb", "mother", and "literal atoms" if it is not already the fixed quote; if it is already the fixed quote, echo it without repeating it.
Spell out the logic: if these systems really were a new source of intelligence rather than just a family of tools, then writing ought to have shown it by now, because writing is the domain in which their native material is most directly at issue.
End with Hoel's empirical observation: what we have instead is scale, efficiency, editing help, feedback, and a great deal of slop.
Paragraph 7: the decisive turn
This is where the essay should stop trying to refine the category of tool and replace it.
The reason the proper-function question keeps failing, you now say, is that we are asking the wrong kind of question.
LLMs are not best understood as implements with functions, but as substances with capacities.
That is the conceptual leap on which the recalibrated essay stands or falls.
It needs to be stated plainly.
Paragraph 8: what "substance" means
Slow down here. This paragraph has to do more than coin a label.
Say that substances are not ordinarily understood in terms of proper function.
Iron is not for anything in the way a hammer is for hammering; it has properties, and from those properties uses emerge.
Oil was not discovered with a full list of applications attached to it.
The point is not chemistry as such. The point is category: stuff with capacities, not implement with purpose.
Then begin to sketch the analogue: LLMs have characteristic powers and tendencies from which practices and uses are still emerging.
Paragraph 9: the material profile
Give the substance thesis some concrete content.
Say what kind of capacities you take the LLM substance to have: pattern-completion, contextual pliability, associative reach, responsiveness to iterative prompting, and a strong tendency towards smoothing and genericity.
Do not make this a dead list. Frame it as a profile of behaviour.
The reason this paragraph matters is that without it "substance" remains metaphorical, whereas the essay needs it to feel explanatory.
Paragraph 10: turn Hoel's oil analogy
This should be one of the essay's strongest moments.
Hoel says that if LLMs were a true source of intelligence, discovering them should be like discovering oil.
The reply is that discovering oil did not mean immediately knowing what oil was for.
For a long time oil was lamp fuel. The larger transformations came later, once more of the substance's capacities had been drawn out and stabilised in practice.
That is the "kerosene phase" thought.
The conclusion should be methodological, not predictive: Hoel's argument assumes that the history is already settled, and that assumption is exactly what the substance picture denies.
Paragraph 11: guard against hype
I think this paragraph is necessary, because otherwise the previous one can sound like "just wait, the singularity is still coming".
Say directly that this is not a promissory argument and not a piece of tech evangelism.
You are not predicting that the hidden greatness of LLMs will inevitably unfold.
You are saying something more limited: we may still be in too early a phase of discovery and practice for Hoel's preferred test to bear the weight he wants to put on it.
This keeps the essay from sounding adolescent.
Paragraph 12: from substance to medium
Introduce the relation between the two categories.
A substance becomes a medium when its capacities are taken up into a practice.
Oil is a substance; oil painting is a medium.
The LLM is a substance; recursive text-exchange is the medium that has begun to form around it.
This move preserves the best material from the earlier drafts without making "medium" compete with "substance" as a rival thesis.
Paragraph 13: the pen comparison
Now bring back the pen comparison, but under the new framing.
A pen extends inscription; it does not return altered material.
An LLM does.
That difference matters because it marks the difference between using a tool to record a thought and working in a medium that pushes back on thought.
This paragraph should be simple and concrete.
Paragraph 14: recursive practice from the inside
Expand the line that the return prompts us back.
The system's output can flatten a distinction, over-generalise, misread, connect two things you had not linked, or hand you a banal summary that forces you to restate the point more sharply.
In each case, what matters is not only the output but what the output does to the next move.
This is where the essay needs to sound as though it knows what the practice actually feels like.
Paragraph 15: artefact and working surface
Now make the distinction that earlier versions were orbiting.
Text can be a finished artefact, but it can also be the working surface of thought.
Hoel's evidence concerns, by and large, artefacts: published prose, public writing, the visible outputs of civilisation.
Much interesting LLM use concerns text as a working surface: drafting, testing, reframing, probing, reformulating.
This does not show that Hoel is looking at the wrong phenomenon. It shows that he may be measuring only one use of text under a heading that hides several.
Paragraph 16: software-development analogy
Bring in software here, not earlier.
The point is not that coding is the real case and writing is not.
The point is methodological: in software, the interesting question is often not whether code overall has become better in some public average sense, but what people can now do, how they iterate, how they move through unfamiliar terrain, and what kinds of exploratory work become possible.
This helps cash out the evaluative point without displacing the writing case.
Paragraph 17: Frippertronics
Use Frippertronics here as the analogy that makes the recursive structure vivid.
The point is not product-equivalence but method.
What returns from the loop is one's own material sent back altered, and the next act is shaped in light of that return.
That is why the comparison belongs late in the essay: only once the recursive structure has been described does the analogy stop looking ornamental.
Paragraph 18: slop reinterpreted
This paragraph should concede a great deal to Hoel.
Slop is real. It is not a superficial embarrassment that your view has to explain away.
On the substance picture, slop can be treated as one characteristic manifestation of the material's genericising tendencies when it is used with too little resistance, too little skill, or under conditions of scale.
In other words, the mountains of dead prose are not evidence that nothing is there; they are evidence about how this substance behaves in one of its earliest mass-use forms.
Paragraph 19: what Hoel still gets right
Before concluding, explicitly say what you take from Hoel.
He is right that there has been no text singularity.
He is right that public prose matters.
He is right that efficiency and scale have so far been more visible than any civilisation-wide increase in quality.
This paragraph matters because it makes the conclusion feel responsive rather than oppositional.
Paragraph 20: final conclusion
The closing claim should now be narrower and stronger than "Hoel is wrong".
He has described the phenomenon too quickly, because he has treated a newly invented substance as though it were already a settled implement with a function.
We may not yet know what this thing is for, because that is not yet the right question.
We may still be learning what can be drawn from it, what practices stabilise around it, and what its latent capacities amount to when they are not reduced to kerosene-level use.
That is why writing, though still the right place to look, cannot straightforwardly function as the final measure of what an LLM is.
Shorter map of the whole shape
Fixed opening and Hoel quote.
Pressure on "tool".
Proper-function problem.
First verdict: "a quite different type of tool, or not quite a type of tool at all".
Testing problem.
Hoel's writing argument at full strength.
Turn to substance.
Substance explained through properties and capacities.
Oil analogy reversed through the kerosene-phase thought.
Guard against hype.
Substance becomes medium in practice.
Pen, recursive prompting, artefact versus working surface.
Software as methodological analogue.
Frippertronics as vivid structural analogy.
Slop as characteristic tendency.
Return to Hoel with a narrower conclusion.
What I think is structurally best
If I had to pick the points that now do the heaviest work, they would be these:
The move from proper function to substance.
The reversal of Hoel's oil analogy through the kerosene-phase thought.
The claim that medium is what the substance becomes in practice.
The reinterpretation of slop as a characteristic early-use tendency rather than a simple reductio.
Those four moves seem to me to be the ones that make this version a genuinely new essay rather than a revised draft of the older one.
If you paste the actual file, I can take this structure and map it directly onto the document from the first block quote onward. Se questa nuova linea funziona, il saggio non dirà più soltanto che “tool” è troppo poco, ma che stiamo ancora imparando a trattare una sostanza di cui non conosciamo bene le potenze.
### Assistant
I think the logic is right and the length is wrong.
The substance→medium relation is the cleanest thing in the proposal. It solves a problem the earlier drafts had, where "medium" was doing double duty as both the name for what the thing is and the name for what happens when you work with it. Splitting that into two — substance is what it is, medium is what it becomes in practice — gives each concept a job. And the kerosene-phase reversal of Hoel's oil analogy is genuinely the essay's strongest single move. Codex is right to flag it.
But twenty paragraphs is a lot for a Substack essay, and the structure front-loads too much negative argument. Six paragraphs of "tool doesn't work" before the substance turn arrives at paragraph 7 — that's a long time to wait for the positive thought. The reader who clicks on "LLMs Are Not Tools" wants to know what they are instead, and this structure makes them sit through proper function, candidates-that-fail, first verdict, methodological consequence, and Hoel restated before offering an answer. Your current draft moves faster than that.
Some specific things I'd push on:
Paragraphs 2, 3, and 4 could probably be two paragraphs, not three. The proper-function problem and the candidates-that-fail are really one move: here's the test, it doesn't pass. The "first verdict" paragraph (4) restates what the reader already feels. Your current draft handles this in two paragraphs (lines 24–46) and it's tighter for it.
The "guard against hype" paragraph (11) coming right after the kerosene paragraph (10) makes the essay sound defensive in exactly the place where it should sound confident. Better to build the anti-hype qualification into the kerosene paragraph itself — one or two sentences that make the point epistemic rather than predictive, then move on. A whole paragraph that says "I am not saying what you think I'm saying" breaks momentum.
Paragraphs 13 (pen), 14 (recursive practice), and 15 (artefact/working surface) are three paragraphs making closely related points. In the current draft these are already somewhat compressed. I'd keep the pen comparison and the recursive structure in one paragraph, and let artefact/working surface be its own. Two, not three.
Frippertronics at paragraph 17 is too late. By the time you've done pen, recursion, artefact/surface, AND software analogy, you're deep into the essay's second half and the reader is ready for the turn toward slop and conclusion. Frippertronics is the essay's most vivid material and it should arrive earlier — right after the substance-to-medium transition, where it can make the abstract relation concrete. The current draft already has it at that position and it works.
The software-development analogy (paragraph 16) is the weakest standalone paragraph in the proposal. The methodological point it makes — that looking at public outputs can miss what competent users are doing — is real, but it's a point the artefact/working-surface distinction already makes for writing. Adding a whole paragraph about coding risks making the essay look like it's searching for better evidence rather than trusting its own.
So if I were compressing Codex's structure, the shape I'd want is closer to:
1. Fixed opening and Hoel quote
2. Pressure on "tool" + proper function (one paragraph, not two)
3. Candidates that fail + first verdict (one paragraph)
4. Testing problem (brief — three or four sentences, not a full paragraph)
5. Hoel's writing case at full strength
6. Turn to substance (the decisive move)
7. What substance means + material profile (one paragraph that does both)
8. Oil reversed: kerosene phase (with anti-hype built in)
9. Substance becomes medium
10. Frippertronics (vivid illustration of the medium claim)
11. Pen comparison + recursive practice
12. Artefact vs. working surface
13. Slop as characteristic tendency (Lucier)
14. What Hoel gets right + conclusion
That's fourteen paragraphs. It preserves everything that matters in Codex's twenty but doesn't make the reader wait six paragraphs for the positive thought, doesn't break momentum with a standalone anti-hype paragraph, and gets Frippertronics in early enough to do structural work rather than arriving as a late ornament.
One thing I'd add that Codex doesn't mention: where exactly does the "we *use* them" line go? Your current draft has it at line 24 as a transitional thought between the opening and proper function. I think it's worth keeping as the bridge into the proper-function discussion — it's a concession that earns trust before the argument pressures the category.
The deeper question is whether the material-profile content (Codex's paragraph 9 — pattern-completion, contextual pliability, smoothing, associative reach) can be made to feel like observation rather than list. If it reads as a catalogue of LLM properties it will feel like a technical aside. If it reads as noticing what the stuff does when you work with it, it earns its place. That's a prose problem more than a structural one, but the structure needs to set it up right — which means it should come after the substance turn, not separated from it by the oil analogy.
*Venti paragrafi sono una mappa onesta del territorio, ma il saggio deve camminare più leggero di così.*
### User
i'm thinking of maybe cutting this all back and removing the patronics medium stuff for the time being and just going to the substance. Cause then I can say, well, what if we've accidentally synthesized a new substance and have not yet managed to extract or even find its full potential? make sure that you explain your analysis, justification, reasons etc etc. for your answer BEFORE giving me your final answer. Please give me a long, deep, and detailed answer, with lots of options/ideas: as we are in the 'kicking ideas around' stage, I want as much feedback as you can give me so that i have a lot to consider.
### Assistant
Let me sit with this for a while before jumping anywhere.
The proposal is: cut the Frippertronics, cut the medium material (Wollheim, Thomson-Jones, *Growing the Image*), and run the essay on the substance thesis alone. The pitch would be something like: we have accidentally synthesized a new substance and have not yet managed to extract or even find its full potential. That's a striking sentence. Let me think about whether the essay can actually stand on it.
First, what does "medium" do in the current draft that "substance" would need to absorb or replace?
In the current draft, "medium" does several things. It names the recursive structure — the fact that the system pushes back, sends material back altered, and makes the next move of thought a response to that return. It gives a frame for Frippertronics. It provides the contrast with "tool" — a tool is understood by what it's for, a medium is understood by the characteristic way it makes work proceed. And it borrows authority from aesthetic theory (Wollheim, Thomson-Jones) to make the claim feel grounded rather than metaphorical.
Now, if you cut "medium," what happens to each of those?
The recursive structure doesn't depend on the word "medium." You can describe the fact that the system prompts you back, that the return reshapes the next intention, without ever invoking the category. That's just phenomenology of use. So that survives.
Frippertronics... hmm. Frippertronics is vivid and people remember it. But it's an illustration of the medium claim — the tape loop as a structured field of resources and resistances. Without the medium frame, Frippertronics becomes an analogy in search of a category. You could still use it as a loose comparison ("the way I work with an LLM reminds me of how Fripp worked with the tape loop"), but it would lose the theoretical anchor that makes it more than a nice parallel. So cutting medium probably means cutting or severely trimming Frippertronics.
The Lucier/slop point... that one is interesting. In the current structure, Lucier illustrates what happens when the medium's tendencies dominate. On a substance-only picture, slop would be redescribed as what happens when the substance's characteristic tendencies — smoothing, genericity, averaging — are allowed to run without resistance. That actually works. You don't need "medium" to say that. You just need the idea that the substance has tendencies which, under conditions of lazy or unresisted use, produce characteristic degradation. That's a perfectly natural thing to say about a substance. Think about alcohol — a substance whose characteristic effects, under conditions of unresisted consumption, produce very predictable degradation. You don't need to invoke "medium" to say that.
OK so far the score is: recursive structure survives, Frippertronics probably goes or shrinks to a passing mention, slop-as-tendency survives, and the *Growing the Image* theoretical apparatus goes. What does the essay gain?
It gains compression and clarity. Right now the essay has to do a complicated two-step: first, LLMs aren't tools; second, they're media. The reader has to learn one new category. On the substance-only version, the reader still has to learn one new category, but it's a simpler one. Everyone already knows what a substance is in a pre-theoretical sense. You don't need Wollheim. You don't need "structured field of resources and practices whose characteristic resistances and possibilities become visible only in the work itself." That's a mouthful. "Substance" is plain English.
But wait. Let me push back on myself. Does "substance" do enough? When I say "iron is a substance," everyone nods. When I say "LLMs are a substance," does the reader know what to do with that? The worry is that "substance" is too empty — too much of a placeholder. "We've synthesized a new substance" sounds evocative but what does it *commit you to*?
Let me think about what the substance claim actually asserts. It asserts at least these things:
- LLMs have properties and capacities that exist prior to and independently of any particular use
- Those capacities are discovered through exploration, not determined by design intention alone
- The right questions to ask are "what can be done with this?" and "what does this do?", not "what is this for?"
- Evaluation is premature when we are still in the discovery phase
That's actually a lot. And it's all philosophically determinate. The first point is an ontological claim about the kind of thing we're dealing with. The second is an epistemological claim about how we come to know it. The third is a claim about the right form of inquiry. The fourth is a claim about the conditions under which Hoel-style assessment can be trusted.
Hmm, but there's a subtlety I want to track. When we say iron is a substance, part of what makes that intelligible is that iron has mind-independent properties. Its tensile strength, melting point, conductivity — these are features of the stuff itself. LLMs are artifacts. Their "properties" were designed in, even if not fully understood by the designers. Does this matter?
I don't think it undermines the substance framing, but it introduces a wrinkle. Synthetic substances are real substances. Nylon, polyethylene, carbon fibre — all synthesized, all have properties that were discovered through use, all had applications that were not foreseen at the moment of synthesis. Carbon fibre was developed in the 1960s but its use in aircraft, sporting goods, prosthetics, etc. was not anticipated in full. The "kerosene phase" thought works for synthetic substances too. Actually it might work even better for them, because the discoverer of a synthetic substance typically knows what they were trying to make but is often surprised by what they actually got.
That's interesting. "What if we've accidentally synthesized a new substance" actually gets a lot of mileage from the "accidentally" — it implies that the designers were aiming at something (a chatbot, a text predictor, an assistant) but what they produced has properties and capacities that exceed or differ from the design brief. That's a genuinely interesting claim and it has historical precedent. Vulcanized rubber, Teflon, penicillin — substances whose properties were discovered by accident or exceeded what was anticipated.
OK, so the substance-only essay could look quite different from the substance-plus-medium essay. Let me think about what its shape would be.
The opening is fixed: Hoel's "mere tool" position and the block quote. Then you need to pressure "tool" through proper function. Then the turn: not a tool but a substance. Then: what that means (capacities, tendencies, discovery-through-use). Then: Hoel's oil analogy reversed through the kerosene phase. Then: what the substance actually does when you work with it (this is where some of the recursive-practice material survives, but framed as *exploring a substance's capacities* rather than *working in a medium*). Then: slop as characteristic tendency of the substance under weak conditions. Then: conclusion — evaluation is premature because we are still discovering what we've made.
That's... eight or nine moves. Much tighter than twenty. And the essay would read faster, hit harder, and stay closer to one thought all the way through.
But let me think about what's lost. What genuinely depends on the medium concept and cannot be recovered from the substance concept alone?
The recursive structure can be framed as a feature of the substance — it's one of its characteristic capacities that when you send language into it, it returns language altered in specific ways. You don't need "medium" for that.
The pen comparison can survive as a way of marking the difference between a tool (which extends action toward a fixed end) and a substance (which has its own behaviour that you have to deal with). A pen doesn't push back because it's a tool. The LLM pushes back because it's a substance with tendencies.
The artefact/working-surface distinction... this one I want to think about more carefully. On the medium picture, the distinction is between the product of the medium and the process of working in it. On the substance picture, the distinction is between... hmm... between using the substance to produce something and exploring or working with the substance. That still works but it's slightly less natural. When we say "oil painting is a medium," the artefact/process distinction is built into the concept. When we say "oil is a substance," the distinction is between what you make from it and what you learn about it through working with it. That's close enough but it's a different emphasis.
Actually, I wonder whether the artefact/working-surface distinction even needs to be in the substance-only version. What's its argumentative job? It answers the question "if writing hasn't improved, does that refute the substance claim?" And the answer is: "improved writing" looks at text as artefact, but much interesting work with LLMs treats text as a working surface. You don't need "medium" to make that point. You can say: much of what competent users do with this substance does not result in a published artefact; it results in a changed way of thinking, a sharpened distinction, a discovered connection. The substance is being explored, not just consumed. That's a substance-native way of making the same point.
OK, so I think the artefact/working-surface point can survive the cut. But I want to ask: does cutting the medium material make the essay *too* narrow? One virtue of the tool/medium contrast is that it gives the reader a pair of concepts, and pairs are satisfying. "It's not a tool, it's a medium" has a rhetorical snap to it. "It's not a tool, it's a substance" has... a different kind of snap. Maybe a weirder one. Let me sit with that.
"We have not invented a new tool. We have synthesized a new substance."
Actually that's pretty good. It has a different kind of force — more material, more concrete, more surprising. "Medium" is an art-theory word that some readers won't have a strong grip on. "Substance" is a word everyone understands, and the surprise comes from applying it to software. The rhetorical gain might actually be larger, not smaller.
Now, what about the Aristotelian potentiality stuff? In the JSON conversation, there was a thought about dynamis — latent capacities that get actualised through use. On a substance-only essay, this fits very naturally. Substances have potentialities. You don't need to go full Aristotle, but a sentence or two about latent capacities being drawn out through practice could do real work. It would explain why prompt engineering, chain-of-thought, few-shot learning, etc. feel like discoveries rather than instructions. You're not telling a tool what to do; you're finding out what a substance can do. That's a nice way of putting it.
Let me think about the Frippertronics question specifically. Nick said "removing the Frippertronics medium stuff." Is Frippertronics necessarily tied to medium? Let me think...
The current draft uses Frippertronics as an example of working in a medium — the tape loop as a structured field with characteristic delays, layering, and accumulation. Without "medium," Frippertronics would have to be an example of... working with a substance? That's a stretch. Tape delay is not a substance in any natural sense. It's a process, a system, a setup. The reason it felt right as a medium example is precisely that "medium" is a process/practice concept, and Frippertronics is a practice. If you're working in the substance register, Frippertronics doesn't have an obvious home.
You could use it as a *loose analogy* — "the recursive structure reminds me of..." — but then it's ornamental rather than structural. And the essay is trying to cut ornament. So I think if you go substance-only, Frippertronics probably goes or becomes at most a parenthetical gesture. That's a loss — it's vivid and memorable — but it's a loss that buys you focus.
The Lucier comparison, on the other hand, fits substance better than medium. Lucier's room is a physical system with characteristic resonant frequencies — something very like a substance's tendencies. When speech is passed through the room repeatedly, the room's own character takes over. When text is passed through an LLM without resistance, the model's own tendencies take over. That's a substance-level point, not a medium-level point. So Lucier stays, and it might actually be strengthened by the reframing.
Let me think about another option: what if you don't cut medium entirely but demote it to a brief aside? Something like: "If we wanted to push the thought further, we might say that the substance becomes a medium when it is taken up into a practice — but that is a thought for another essay." That would acknowledge the direction of travel without making the current essay do the work of establishing both categories. It would also leave the door open for a sequel. Is that too cute? Maybe. But it's worth considering.
Or — different option — what if "medium" appears not as a theoretical category but as a plain-English observation? "In practice, people have begun to use this substance as a medium of thought — a working surface on which ideas are tested, resisted, and reshaped." That uses "medium" descriptively rather than theoretically. It doesn't invoke Wollheim or Thomson-Jones. It doesn't need Frippertronics to cash it out. It just says: this is what people do with the stuff. That might be the lightest possible version.
Hmm, let me think about what the essay's argumentative opponent would say to the substance thesis. The strongest objection is probably: "This is just the AI-optimist move dressed up in philosophical language. 'We haven't found all the uses yet' is what every startup says about every product. What makes this more than hype?"
The answer has to be: the substance framing is not predictive but epistemic. It's not saying "great things are coming." It's saying "we don't yet have the settled understanding of this thing that would be needed to make Hoel's evaluative argument work." That's a defensive claim, not an offensive one. You're not arguing that LLMs will transform civilisation. You're arguing that Hoel's confidence is premature. That distinction is what saves the essay from being boosterism.
But can the essay make that distinction clearly enough in the substance-only version? I think it can, and it might actually be easier without the medium apparatus. Here's why: "medium" invites the reader to think about what the practice looks like, what the experience is, what kinds of work become possible. That's intrinsically forward-looking and optimistic-sounding. "Substance" invites the reader to think about what the stuff is, what its properties are, what we don't yet know about it. That's more naturally epistemic and exploratory. So the substance framing may actually be better at maintaining the defensive posture the essay needs.
Now, what about the pen comparison? The current draft says: "A pen makes marks on a surface; it extends my ability to inscribe; it does not return a proposal, a misreading, or a line of continuation I had not seen. A pen does not prompt me back." This is one of the draft's best moments. Does it survive the cut?
I think it does, but its job changes. On the medium picture, the pen contrast says: unlike a pen (a tool), an LLM is a medium that pushes back. On the substance picture, the pen contrast says: unlike a pen (a tool that extends agency), an LLM is a substance that has its own tendencies — when you engage with it, those tendencies interact with your intentions, and what emerges is shaped by both. The pen comparison marks the difference between transparent instrumentality and the kind of thing that has its own behaviour. That's a substance-level point. So the pen stays.
Actually, thinking about it more — the pen contrast might be even sharper on the substance reading. A pen is so inert, so purely instrumental, that calling it a tool is trivially correct. An LLM is so reactive, so tendency-laden, so capable of surprising you, that calling it a tool misses the most interesting fact about it. The interesting fact is that it *has its own behaviour*. That's what makes it substance-like rather than tool-like.
Let me think about another angle. One thing the medium concept does well is explain *practice* — how people actually work with LLMs in a disciplined way. Prompt engineering, iterative refinement, the back-and-forth of drafting — all of this is naturally described in medium language. On a substance-only picture, how do you describe practice?
You could say: working with a substance involves learning its properties, learning how to elicit specific capacities, learning where it resists, learning where it smooths. That's not "practice" in the medium-theoretic sense, but it's practice in a craft sense — the way a potter learns clay, or a glassblower learns glass. Actually, "learning the material" is a very natural way to describe what skilled LLM users do. They learn what kinds of prompts elicit what kinds of response, where the model is strong and where it's weak, how to resist its tendencies and how to exploit them. That's material knowledge. And it's learnable, which is what makes the substance framing non-trivially different from saying "it's just random."
OK, I want to think about one more thing before I start converging. What is the essay's final sentence, or final image, on the substance-only version?
On the current draft, the ending is: "They may be a quite different type of tool, or not quite a type of tool at all — and that is worth taking seriously before we conclude that the current technology is merely another stone in *Homo faber*'s long line of rocks."
On Codex's version, the ending is something about not yet knowing what the thing is for because we are still learning what can be drawn from it.
On the substance-only version, the ending might be: "We have synthesized a new substance. We may not yet know what it is for, because 'what is it for' may not yet be the right question. We are still learning what it can do."
Or more concretely: you could return to Hoel's oil. "Hoel says discovering LLMs should have been like discovering oil. I think he may be more right than he realises — but we are still in the kerosene phase, and kerosene was not what oil turned out to be for."
That's a strong closing line. It turns Hoel's own image against his conclusion. And it doesn't predict anything — it just says: the story isn't over yet, and measuring kerosene-era products doesn't tell you what the substance is.
Hmm, but actually, I want to be careful. Is the "kerosene phase" line historically accurate? Let me think... Early petroleum use included lamp oil, lubricant, and patent medicines. The internal combustion engine came later. Kerosene was indeed the first large-scale refined product. Gasoline was initially a byproduct — sometimes discarded. So yes, there was a genuine phase in which the most visible use of petroleum was illumination, and the transformative uses came decades later. The analogy holds. (I'm noting that I haven't checked a source here — this is from general knowledge, not from a text I've extracted. If this were going into a draft, I'd want to verify the history.)
Let me also think about whether cutting Frippertronics and medium means cutting too much of what makes the essay *interesting to read*. There's a risk that a substance-only essay, while tighter and clearer, might read as more abstract — more philosophy, less vivid example. The current draft has Fripp, Eno, Lucier, tape loops, pen and paper. A substance-only version would have... oil, kerosene, iron, pen, Lucier. That's still a range of concrete reference. And the "accidentally synthesized a new substance" line is itself vivid. But the musical examples had a charm to them.
One option: keep Lucier (it works with substance) and replace Frippertronics with a different concrete example that fits the substance frame better. What kind of example? Maybe something about the history of a synthetic material — Bakelite, nylon, silicon. The early history of silicon is actually quite good here: silicon was known for decades before anyone figured out how to make semiconductors. The first uses were mundane (silicones for lubricant and sealant). The transformative use (transistors, integrated circuits) came much later and required discoveries about the material's properties that were not obvious from its early applications. That's a kerosene-phase story about a synthetic substance.
But... that might make the essay too tech-history and not enough about what it's actually like to use an LLM. Lucier gives you that experiential dimension. The pen comparison gives you some. But if Frippertronics goes, you lose the most evocative picture of what recursive engagement with a system feels like.
Another option: forget the named examples and just describe the experience directly. "I write a sentence. The system sends back three paragraphs. Most of it is flatter than what I wanted. But buried in the second paragraph is a connection I hadn't made — not because the system understood my thought, but because its associative patterns ran a different path through the same territory. I delete the three paragraphs and write a better sentence." That's vivid without needing Fripp or Eno. And it's directly about LLM use, not about music.
Actually, I think that might be the move. Replace Frippertronics not with another analogy but with a direct phenomenology of use that is explicitly framed as exploring a substance's capacities. "This is what it's like to work with this substance. You send language into it and something comes back — shaped by the substance's own tendencies, not by your intention. You then have to decide what to do with what came back."
That's more immediate, less borrowed, and it avoids the risk of the analogy overshadowing the thing it's supposed to illuminate.
Let me think about one final structural question. If you cut medium and go substance-only, does the essay still have enough to say? Or does it become a one-move essay — "it's a substance, not a tool" — and then struggle to fill out the middle?
I think there are at least five distinct things to say under the substance heading:
1. What the proper-function failure reveals (not just "tool doesn't fit" but "the failure points toward a different category")
2. What "substance" means and what it commits you to (capacities, tendencies, discovery-through-use)
3. How Hoel's oil analogy gets reversed (kerosene phase)
4. What this substance actually does when you work with it (tendencies, pushback, the pen contrast)
5. Why slop is a feature of the substance, not a refutation of it (Lucier, characteristic degradation)
Five moves is enough for a Substack essay. It's not padded. Each one does different work. And together they build to the conclusion that Hoel's evaluative confidence requires a settled understanding of the thing being evaluated — which we don't yet have.
Let me also think about whether the essay needs the artefact/working-surface distinction if medium is gone. I think it does, because the distinction answers Hoel's evidence. He points to public writing (artefacts) and says quality hasn't improved. The reply is: much of what this substance enables is not aimed at artefacts. Cutting that distinction would leave Hoel's evidence unanswered. So it stays. But it can be brief — a few sentences within the "what this substance does" paragraph, not a standalone paragraph.
One more thought. There's something appealing about the *simplicity* of the substance-only essay. The current draft and Codex's proposal are both architecturally complex — multiple categories, multiple analogies, multiple theoretical sources. A substance-only essay would have one category, one reversal (oil/kerosene), and one extended contrast (tool vs. substance). That simplicity is a virtue in an essay for a general audience. The reader walks away with one thought: "Huh, maybe we should think of LLMs not as tools with functions but as substances with capacities, and maybe we're still early in understanding what those capacities are." That's a thought you can hold in your head. "LLMs are media, not tools, and medium means a structured field of resources and practices whose characteristic resistances..." — that's a thought that starts to slip away before you finish saying it.
So my overall assessment is: yes, cutting to substance-only is a genuine option, and it might produce a tighter, more readable, and more surprising essay than either the current draft or Codex's twenty-paragraph version. The losses (Frippertronics, the *Growing the Image* theoretical backing, the medium concept as a named category) are real but can be compensated. The gains (clarity, compression, one-thought legibility, stronger rhetorical contrast with Hoel) are significant.
The question is whether you feel that the medium material is work you want to save for later or work you want to let go of. If it's the former — save it for a longer piece, an academic paper, a sequel — then the substance-only Substack essay becomes a way of staking the territory without using all the ammunition. If it's the latter, then you're simplifying the thought itself, not just the presentation. Both are legitimate choices but they feel different.
## Analysis
I've been sitting with this and I think cutting to substance-only is a strong move, and probably the right one for this essay. Let me lay out why, what you'd lose, what you'd gain, and several options for how it could work. Since you want a lot to kick around, I'm going to be thorough.
### Why substance-only works
The substance claim does enough philosophical work on its own. It gives you:
- An ontological point: LLMs are not implements with proper functions; they are stuff with capacities and tendencies.
- An epistemological point: those capacities are discovered through exploration, not determined in advance by design intention.
- A methodological point: the right question is not "what is this for?" but "what can be done with this?" — and that question is not yet settled.
- A direct reply to Hoel: his evaluative confidence presupposes that we already know what the thing is, and the substance framing denies exactly that.
Those four things are enough to carry a Substack essay. You don't need the medium apparatus to make any of them.
The medium concept was doing two things in the earlier drafts: naming the recursive practice (the return prompts you back), and providing theoretical grounding from aesthetic theory. But the recursive practice can be described as a feature of the substance — it's one of its characteristic capacities that when you send language into it, it sends language back altered. You don't need Wollheim or Thomson-Jones or a "structured field of resources and practices" to say that. And the theoretical grounding, while genuinely interesting, is a lot of machinery for a piece that wants to be readable by people who don't know what a Wollheimian medium is.
### What you'd lose
Frippertronics. This is the real loss. The tape-loop analogy is vivid, memorable, and it makes the recursive structure feel concrete. But Frippertronics is an illustration of working in a medium, not of working with a substance. Tape delay is a process, a setup, a system — not a stuff. Without the medium concept, Frippertronics doesn't have a natural home. You could use it as a passing comparison ("the recursive structure has something in common with...") but it would be ornamental rather than structural. And the essay is trying to cut ornament.
The *Growing the Image* theoretical backing. This is less of a loss than it sounds, because (a) most Substack readers haven't read the book, and (b) the substance framing has its own theoretical resources — Aristotelian potentiality, the philosophy of materials, the history of synthetic substances — that are arguably more accessible.
The "medium" concept itself. But here's the thing: you're not losing it permanently. You're saving it. If the substance essay works, you could write a follow-up that takes the next step: "OK, so LLMs are a substance — what happens when that substance is taken up into a practice? What kind of medium does it become?" That's a different essay, and one that could go deeper precisely because the groundwork was already laid.
### What you'd gain
One-thought legibility. The reader walks away with a single claim they can hold in their head: "Maybe LLMs aren't tools with functions; maybe they're substances with capacities, and we're still early in understanding what those capacities are." That's a thought you can remember, repeat, argue with. "LLMs are media, not tools, where medium means a structured field of resources and practices..." — that starts to slip away before you finish the sentence.
Compression. The substance-only essay could be twelve or thirteen paragraphs. Much less architecturally complex than Codex's twenty.
Rhetorical surprise. "We have not invented a new tool. We have synthesized a new substance." That's a weird and interesting sentence. It makes the reader stop. "It's a medium, not a tool" is interesting but less surprising — it stays within the register of art theory. "It's a substance" pulls the reader into a different register entirely — chemistry, materials science, craft — and that defamiliarization is a strength.
Stronger Hoel reversal. On the substance picture, Hoel's oil analogy doesn't just get answered — it gets weaponized. He uses oil as the model of a transformative discovery. The reply is: yes, exactly, and when we first discovered oil we used it for lamp fuel. The kerosene phase. That reversal is sharper and more self-contained without the medium machinery.
Better anti-hype posture. This is I think underappreciated. The medium framing invites the reader to imagine what LLM practice looks like at its best — recursive engagement, thoughtful resistance, creative exploration. That's intrinsically forward-looking and sounds optimistic. The substance framing invites the reader to acknowledge what we don't yet know about this stuff. That's more naturally epistemic and cautious. It maintains the defensive posture the essay needs: I'm not predicting greatness, I'm saying we can't evaluate yet.
### Options for how the substance-only essay could work
Here are several structural options. I'm grouping these as alternatives, not as a sequence:
**Option A: The tight version (9–10 paragraphs)**
1. Fixed opening + Hoel quote
2. Pressure on "tool" through proper function (one paragraph)
3. Candidates that fail + first verdict (one paragraph)
4. The turn: not a tool, a substance
5. What "substance" means — capacities, tendencies, discovery-through-use, with a brief material profile
6. Hoel's oil reversed: kerosene phase (with anti-hype built in)
7. What this substance does when you engage with it — pen contrast, recursive return, artefact vs. working surface woven together
8. Slop as characteristic tendency (Lucier)
9. What Hoel gets right + conclusion
This is very tight. Each paragraph does one thing. No filler. The danger is that it might feel rushed — the substance turn arrives at paragraph 4 and then you have five paragraphs to develop it. But for a Substack essay that might be exactly right. Get to the point and then develop it.
**Option B: The phenomenological version (11–12 paragraphs)**
Same as A, but after the oil/kerosene paragraph, insert a paragraph of direct phenomenology: what it's actually like to work with this substance. Not Frippertronics, not an analogy borrowed from music, but your own experience described in your own terms.
Something along the lines of: you write a half-formed thought, the system sends back something shaped by its own tendencies, you reject most of it but one move gives you something to push against, and the next thought is sharper for having had that resistance. Framed explicitly as: this is what it's like to explore a substance's capacities.
This option replaces the borrowed vividness of Frippertronics with the direct vividness of your own practice. It also makes the essay more personal, which suits Substack.
**Option C: The synthetic-substances version (12–13 paragraphs)**
This version leans into the history of synthetic materials. After the oil reversal, add a paragraph about Bakelite, nylon, or silicon — substances that were synthesized with one purpose in mind and turned out to have capacities that were discovered only later, through exploration and use. The point is that the history of technology is full of cases where a new material was mismeasured by its early applications. This gives the kerosene-phase thought more support and makes the substance claim feel less metaphorical, more historically grounded.
The risk is that this makes the essay a bit tech-history-heavy. But one paragraph of concrete historical example might be worth it if it keeps "substance" from floating free.
**Option D: The minimal-medium version**
Don't cut medium entirely; demote it to one sentence. The substance paragraphs do the heavy lifting, and at the point where you'd normally introduce medium, you instead say something like: "When a substance is taken up into a practice, it becomes a medium — but that is a thought I want to develop elsewhere. For now, the point is about the substance itself."
This acknowledges the direction of travel, signals that there's more to say, and leaves the door open for a sequel. It costs you one sentence and gains you an intellectual promissory note that readers might find intriguing.
**Option E: The substance-plus-phenomenology version (my speculation about what might work best)**
This one tries to replace the work that medium and Frippertronics were doing without using either concept. After the substance turn + oil reversal + material profile, you'd have a paragraph that describes the recursive practice directly but frames it entirely in substance language.
"Working with this substance is not like using a hammer, where the intention precedes the action and the tool merely executes. It is more like working with a reactive material — something that has its own tendencies, its own characteristic forms of pushback. You send language in; language comes back, but shaped by the substance's learned patterns. Sometimes flatter, sometimes unexpectedly connective, sometimes productively wrong. The next move is always a response to what came back, not just a continuation of what you intended. A pen does not do this. A search engine does not do this. This substance does."
That paragraph does the work of the pen comparison, the recursive-return point, and the medium claim, all in substance language, without needing Fripp or Wollheim.
### Things worth thinking about
The "accidentally" in "accidentally synthesized." This is doing interesting work. It implies that the designers were trying to make one thing (a chatbot, a text predictor) and what they got has properties they didn't fully anticipate. That resonates with actual LLM history — the "emergent abilities" discussion, the way capabilities appeared at scale that weren't designed in. Worth keeping.
The material profile question. Codex's paragraph 9 listed LLM properties (pattern-completion, contextual pliability, etc.). On the substance-only version, some version of this is needed — you have to say what properties the substance has, not just that it has properties. But it has to read as observation about working with the stuff, not as a technical spec. The difference between "this substance tends to smooth out specificity, to find generic paths through a topic, to connect things associatively rather than logically, and to respond dramatically to how it's asked" and "LLMs exhibit pattern completion, contextual sensitivity, and token-level prediction." The first is material knowledge, the kind of thing a craftsperson would say. The second is mechanism.
Whether you need the artefact/working-surface distinction at all. On the substance-only version, the distinction can be compressed into a few sentences within the phenomenology-of-use paragraph: "Most of what I've described does not result in a published text. The point is not to produce an artefact but to explore the substance — to test a distinction, discover what I think, find out what comes back when I push in a particular direction." That's enough to answer Hoel's evidence without needing a standalone paragraph.
Whether the Aristotelian potentiality language earns its keep. I think one sentence does — something about latent capacities (dynamis) being actualised through use. It gives the substance claim philosophical depth without turning the essay into a seminar. More than one sentence risks scholasticism.
Whether there's a danger of the essay sounding like it's *about* LLMs but *for* philosophers. The substance language is accessible. The Aristotle can be kept light. The oil/kerosene reversal is concrete. But you'd want to watch the tone — make sure it reads as an argument about a live topic, not as a conceptual exercise. The "accidentally synthesized" framing helps with this because it sounds like someone describing something that really happened, not someone deploying a category from the history of philosophy.
### What I'd want to flag honestly
I'm not sure whether "substance" is the final word or a staging post. It does more than "medium" alone, and it does it more simply. But there might be a version of the thought that's even simpler — something like: "We don't yet know what kind of thing this is, and Hoel's argument requires that we do." That's the epistemic claim stripped of all metaphor. The substance language gives it concreteness and explanatory power. But if it starts to feel like you're replacing one classification (tool) with another (substance) and the reader wonders whether you've just renamed the problem, you might need a sentence that says: the point is not that "substance" is the final answer, but that it captures the *kind of uncertainty* we're in — the uncertainty of someone confronted with a new material, not the uncertainty of someone holding a broken tool.
I'm speculating there — I don't know whether readers would feel that worry. But it's the strongest objection I can imagine to the substance-only version, and it's worth having an answer to it in reserve.
*Una sostanza nuova non si giudica dai primi usi che se ne fanno — si giudica, se mai, dopo che qualcuno ha imparato a lavorarla.*
### User
"That's interesting. "What if we've accidentally synthesized a new substance" actually gets a lot of mileage from the "accidentally" — it implies that the designers were aiming at something (a chatbot, a text predictor, an assistant) but what they produced has properties and capacities that exceed or differ from the design brief."the accidently of it it all isn't important. what i mean is that substances (natural or synthetic) can have latent properties that we do not know how to invoke immediately. we cannot say the same for tools –there are no such things as natural tools and the ones that we create do not have latent, hidden at first, properties.
### Assistant
OK, let me sit with this correction carefully because it changes the shape of the argument in a way I need to track.
Nick is saying: the "accidentally" isn't the point. The point is about the difference between substances and tools with respect to latent properties. Let me make sure I understand what he's actually claiming.
Substances — whether natural or synthetic — can have properties that are not immediately apparent and that we discover over time through exploration, experimentation, and practice. Iron has properties we didn't know about when we first encountered it. Nylon has properties DuPont didn't fully anticipate. Oil had latent capacities that were drawn out over decades. The point is not that these substances were *accidentally* created or *accidentally* discovered. The point is that *substances as a category* are the kind of thing that can have latent, initially hidden properties. That's just what substances are like.
Tools, by contrast, don't work this way. And Nick is making a sharper claim than I initially registered: there are no natural tools. Tools are created. They are made for purposes. And because they are made for purposes, they don't have latent, hidden properties in the way substances do. A hammer doesn't have undiscovered capacities waiting to be drawn out. You can use a hammer in unexpected ways (as a paperweight, as a weapon), but those aren't latent properties of the hammer — they're just alternative applications of properties that are already fully apparent. The hammer's hardness, weight, shape — these are all on the surface. There's nothing hidden.
That's actually a much cleaner distinction than what I was working with. Let me think about whether it holds up...
Is it true that tools don't have latent properties? Let me think of counterexamples...
A computer? Well, Nick might say a computer is precisely the kind of thing that challenges the tool category — which is partly the point of the essay. A computer does have capacities that exceed its original design brief. But then maybe that's evidence that a computer is not purely a tool, or not only a tool.
A Swiss Army knife? No — its capacities are all designed in and immediately apparent. You can see the blade, the corkscrew, the file. There's nothing hidden.
What about a simple lever? Could you discover new things about how a lever works? You could discover new applications of leverage, but the lever's properties — rigidity, length, fulcrum point — are all manifest. There's no hidden property of the lever waiting to be drawn out.
What about a material that was initially used as a tool but turned out to be more? Like... glass? Glass was first used as a decorative material (beads, vessels). Then lenses were discovered. Then optical instruments. Then fiber optics. But wait — glass is a substance, not a tool. A glass bead is a tool/artefact; glass itself is a substance. And it's the substance that has the latent properties (optical transparency, ability to be drawn into fibres, refractive index), not the artefact.
So the distinction Nick is drawing is between:
- Substances: have latent properties that are discovered over time; this is constitutive of what it means to be a substance
- Tools: are made for purposes; their properties are determined by design and are immediately apparent; they don't have latent, initially hidden properties
And the claim about LLMs is: they behave more like substances than tools in this respect. They have properties and capacities that are being discovered through exploration — chain-of-thought, few-shot learning, in-context learning, emergent abilities at scale, the way they respond to iterative prompting. These weren't designed in (or at least, weren't anticipated as part of the design brief). They're latent properties of the stuff that have been drawn out through use and experimentation.
Now, this is interesting because it removes the need for "accidentally" to do any work. The argument doesn't depend on whether the designers were surprised. It depends on the ontological distinction between substances and tools. Substances have latent properties; tools don't. LLMs have latent properties; therefore they're more like substances than tools.
Actually, wait. Let me push on this a bit more. Is it really true that the latent-property distinction maps perfectly onto substance vs. tool? Or is it more precisely about natural kinds vs. artefacts?
Hmm, but Nick explicitly said "natural or synthetic" — so he's not restricting substance to natural kinds. Nylon is a synthetic substance and it has latent properties. So the distinction is not natural vs. artificial. It's substance vs. tool. What makes something a substance is that it's a *stuff* with properties, and those properties can include ones that aren't immediately apparent. What makes something a tool is that it's made for a purpose, and its relevant properties are determined by and apparent from its design.
Now I want to think about what this does to the essay structure.
If the point is not about accidental discovery but about the ontological difference between substances and tools with respect to latent properties, then the argument gets sharper. The essay doesn't need to tell a story about surprise or unintended consequences. It needs to make a philosophical distinction about kinds of things and how we come to know their properties.
The kerosene-phase thought still works, but its function changes slightly. It's not illustrating that oil's uses were *accidentally* discovered. It's illustrating that substances, as a category, reveal their properties over time through exploration. Oil was always capable of powering an internal combustion engine. That capacity was always there, latent in the substance. It just took time and experimentation to find it and exploit it. The kerosene phase is not about accident — it's about the temporal structure of how we come to know what a substance can do.
This is cleaner. And it makes the anti-Hoel argument more precise. Hoel's mistake is not that he's failed to notice an accident. His mistake is that he's treating LLMs as tools — things whose properties are determined by design and immediately apparent — when they might be substances — things whose properties include latent capacities that are still being discovered. If they're tools, then what they do now is a fair basis for evaluation. If they're substances, then what they do now might be kerosene — a genuine but limited expression of latent capacities that have not yet been fully drawn out.
Now, what does this mean for the "no natural tools" observation? That's a genuinely interesting philosophical point and I want to think about whether it belongs in the essay.
"There are no natural tools." What does this mean exactly? It means that tools are always artefacts — always made by someone for some purpose. A rock is not a tool until someone picks it up and uses it as one. And even then, the rock-as-tool is understood in terms of the purpose to which it's being put, not in terms of the rock's intrinsic properties. The rock has intrinsic properties (hardness, weight, shape) — those are substance-level properties. But the *tool* "hammer-rock" is defined by the use, not by the stuff.
This is actually a really nice point. It draws the substance/tool distinction at the right joint. Substances have intrinsic properties. Tools have functions. And the essay is claiming that LLMs have intrinsic properties (latent capacities) rather than (or in addition to) functions. So they should be investigated like substances, not evaluated like tools.
Does this mean the essay should include a brief statement of the "no natural tools" point? I think it could work as a one-sentence clarification: "Substances, unlike tools, are not brought into existence for a purpose — even synthetic ones. They have properties, and from those properties uses emerge. Tools are defined by the other direction of fit: the purpose precedes the properties."
Hmm, but is the "no natural tools" claim strictly correct? What about an ant using a leaf as a tool? Or a chimpanzee using a stick? These are natural objects used as tools. But Nick's point would be that the *toolhood* is imposed by the user — the stick becomes a tool when the chimp uses it to extract termites. The stick's substance properties (length, flexibility, diameter) were always there. The toolhood is added. So even "natural tools" are really natural substances pressed into service as tools. The substance pre-exists the tool.
That's actually an even deeper point than the one I initially heard. It's not just that tools and substances are different categories. It's that substance is the more basic category. Everything that's a tool is also a substance (it's made of stuff), but not everything that's a substance is a tool (lots of stuff has no function). And when we encounter a new substance, the right first move is to investigate its properties, not to assign it a function.
OK, and now the LLM connection becomes very clear. When Hoel treats LLMs as tools and asks "are they good at their job?", he's skipping the prior question: "what kind of stuff is this, and what are its properties?" That prior question hasn't been settled yet. So evaluating LLMs by their current prominent use (writing assistance) is like evaluating oil by its performance as lamp fuel. It's not wrong — oil does work as lamp fuel — but it's measuring a substance by an early, partial, and potentially non-central application.
Now I want to think about what happens to the essay's middle section if we go with this framing. The earlier sessions had a lot of material about recursive practice, the pen comparison, artefact vs. working surface. Does that stuff survive the shift from "medium" to "substance with latent properties"?
The pen comparison survives easily. A pen is a paradigmatic tool. It has a clear function (making marks), fully apparent properties (point, ink, grip), and no latent capacities. An LLM is not like that. The contrast between pen-as-tool and LLM-as-substance is sharp and useful.
Recursive practice... this can be reframed. Instead of saying "the medium pushes back," you can say "one of the substance's latent properties is that when you engage with it through language, it returns language shaped by its own tendencies, and this return reshapes what you do next." That's a description of a capacity, not a description of a medium. It's something the substance *does*, not something a practice *is*. The emphasis shifts from practice to property.
Artefact vs. working surface... this maps naturally onto the distinction between evaluating a substance by its products and investigating a substance by working with it. Hoel evaluates by products (published writing). But much interesting LLM engagement is investigative — people working with the substance to find out what it can do, not to produce polished artefacts. This doesn't need "medium" at all.
Slop/Lucier... also maps cleanly. The substance has characteristic tendencies (smoothing, averaging, genericising). When those tendencies are allowed to run without resistance, the result is slop. That's a substance-level observation. Lucier's piece illustrates what happens when a physical system's characteristic resonant properties dominate: speech degrades into something determined by the room, not by the speaker. Same logic: LLM-generated text degrades into something determined by the model's tendencies, not by any human intention. That's slop. And it's a prediction of the substance framing, not a problem for it.
So actually, I think almost everything survives the cut except Frippertronics and the *Growing the Image* theoretical apparatus. And the essay becomes leaner and more focused. The question is whether it becomes *too* lean — whether there's enough material for a satisfying essay.
Let me sketch a possible structure:
1. Fixed opening + Hoel quote
2. Pressure on "tool" — proper function
3. Candidates that fail — what is an LLM for?
4. Turn: substances and tools are different kinds of things with respect to latent properties
5. No natural tools / tools defined by purpose / substances defined by properties and capacities
6. LLMs have latent properties: chain-of-thought, in-context learning, emergent abilities, etc.
7. Oil reversed: kerosene phase — not about accident, about the temporal structure of discovering what a substance can do
8. What this substance does when you engage with it — pen contrast, recursive return
9. Artefact vs. working surface — Hoel measures products, but much interesting engagement is investigative
10. Slop as characteristic tendency (Lucier)
11. Conclusion: we have synthesized a new substance and are still learning what it can do; Hoel's evaluation presupposes a settled understanding we don't yet have
That's eleven moves. Still tight. And every move does distinct work.
But I want to think about something else. The "no latent properties for tools" claim — is this doing *philosophical* work or is it doing *rhetorical* work? If it's doing philosophical work, it needs to be stated carefully and defended against obvious objections. If it's doing rhetorical work, it just needs to be suggestive enough to shift the reader's perspective.
One objection: couldn't you say that a computer has latent properties? When you buy a laptop, it has hardware capabilities you haven't yet discovered or used. Running novel software on it reveals new capabilities. But this is really a substance-level observation about the hardware and software, not a tool-level observation about the computer-as-implement. The computer-as-tool (email machine, word processor) has a clear function. The computer-as-substance (silicon, circuits, architecture) has latent capacities. So even here the distinction holds: what's latent is the substance, not the tool.
Another objection: couldn't you say that a hammer reveals unexpected properties when you swing it in a new context? But those aren't *latent* properties — they're the same manifest properties (mass, hardness) doing manifest things in a different setting. Latent properties are properties that you didn't know the thing had. A hammer doesn't turn out to have magnetic properties you hadn't noticed. Oil turns out to be combustible in engines in a way that nobody initially anticipated.
I think the distinction is philosophically sound. Substances have latent properties in the sense that investigation can reveal properties that were not previously known or anticipated. Tools have manifest functions and their relevant properties are determined by design. The distinction tracks something real about how we come to know things.
Now, let me think about a wrinkle that Nick's correction introduces. If the "accidentally" isn't doing work, then the essay can't lean on stories about unintended consequences or surprising discoveries. What it can lean on is the structural point about substances having the *kind of properties* that can be latent. That's a claim about kinds, not about particular historical episodes.
This is actually philosophically stronger but rhetorically harder. Historical stories are vivid (kerosene! Bakelite! penicillin!). Category distinctions are abstract. The essay needs to find a way to make the category distinction vivid without relying on accident narratives.
One way: the kerosene-phase story still works, because it's not about accident — it's about temporal lag in discovering what a substance can do. Oil wasn't accidentally useful as engine fuel; it always had that capacity. We just hadn't learned how to invoke it yet. The word "accidentally" isn't needed. "We hadn't yet learned how to invoke its latent capacities" is the right way to put it.
Another way: describe specific latent properties of LLMs that have been discovered through exploration. Chain-of-thought prompting is the best example. Nobody designed GPT-3 to be better at reasoning when prompted to think step by step. That was a latent property of the substance that was discovered through experimentation. Same with few-shot in-context learning. Same with the way these models respond to persona prompts, to structured formatting, to iterative refinement. These are all cases where the substance turned out to have capacities that weren't designed in and weren't initially apparent. They were drawn out through use.
That's concrete, LLM-specific, and doesn't require the word "accidentally." It just requires the observation that these capacities were latent and were discovered through exploration rather than deduced from a design brief.
Hmm, but I want to be careful about how much I claim here. The researchers who discovered chain-of-thought prompting didn't think of themselves as discovering a latent property of a substance. They thought of themselves as finding a better prompt. And many AI researchers would describe these as emergent properties of the architecture, not latent properties of a substance. Nick's essay is proposing a reframing — seeing these discoveries *as if* they were discoveries about a substance. That's a philosophical move, not a straightforward empirical observation. The essay needs to be clear that it's offering a category and then showing that the category fits, not reporting a fact that everyone already agrees on.
This connects to the epistemic discipline requirement. The substance claim is Nick's interpretation, not something Hoel says or something the AI research community says. It needs to be presented as such. "I want to suggest that..." or "What if we think of..." — not "LLMs are substances" as though it were a settled fact.
Wait, but Nick has committed to this framing. He's past the tentative stage on the substance claim itself — what he's working out is the essay structure, not whether the claim is right. So the essay can be assertive about the substance framing while still being epistemically careful about what follows from it.
Let me think about one more thing. The claim "there are no natural tools" is interesting because it gestures at the deeper philosophical point about the direction of fit between properties and purposes. For substances, properties come first and purposes emerge from them. For tools, purposes come first and properties are selected/designed for them. LLMs were designed (purpose first), but their properties exceed the design brief in ways that are discovered through exploration (properties emerging independently of purpose). So LLMs are in a weird intermediate position — artefacts that behave like substances. Made for a purpose, but with capacities that transcend the purpose.
That's actually a really interesting philosophical observation. It might be worth a sentence in the essay: "LLMs were designed, as tools are; but what they turned out to be has exceeded the design brief, as substances do. They are artefacts that behave like discovered stuff."
Or is that too clever? Maybe. But it captures something genuinely interesting about the ontological status of LLMs: they sit at the boundary between the made and the found. They were made, but what we're now doing with them is more like finding out what a newly encountered substance can do than like using a tool we built.
OK, I think I've explored this enough. Let me move toward a synthesis.
Nick's correction sharpens the argument in several ways:
1. It removes the dependence on "accident" or "surprise" and makes the claim structural: substances have latent properties as a category; tools don't.
2. It deepens the substance/tool distinction by noting that there are no natural tools — tools are always artefacts defined by purpose, while substances (natural or synthetic) are defined by properties that can include latent ones.
3. It makes the anti-Hoel argument more precise: Hoel evaluates LLMs as tools (by their performance at a function), but if they're substances, the right approach is investigation (discovering what their properties are), not evaluation (judging how well they fulfil a function).
4. It keeps the kerosene-phase thought but reframes it: not about accidental discovery, but about the temporal structure of discovering what a substance can do. Oil always had the capacity to power engines; we just hadn't learned how to invoke it yet.
5. It suggests that the essay doesn't need "medium" because the substance/tool distinction already does enough work, and the recursive practice, pen comparison, artefact/surface distinction, and slop/Lucier material can all be reframed in substance terms.
The one thing I want to flag is a potential vulnerability. If someone pushes back and says "but LLMs were designed, so they're artefacts, so they're tools," the essay needs an answer. And the answer is: yes, they were designed, but their properties include latent ones that were not anticipated by the design. They're artefacts that have substance-like characteristics. The design gave them some properties, but exploration has revealed others. That makes them different from hammers, Swiss Army knives, and other tools whose properties are fully determined by design.
This is where the essay might want to briefly mention that synthetic substances are a real category — nylon, carbon fibre, various alloys — where the stuff is manufactured but its full range of properties is discovered through subsequent investigation and use. LLMs might belong in this category. Not a natural substance, but not a mere tool either.
You're right, and the correction sharpens the argument considerably. Let me lay out what I think changes and what it opens up.
## What the correction does
The point isn't about surprise or unintended consequences. It's about a structural difference between two kinds of things. Substances — whether you dig them up or synthesize them — can have latent properties. Properties that are there in the stuff but that you don't know how to invoke, haven't yet discovered, haven't yet found the right conditions to draw out. That's just what substances are like. It's constitutive of the category.
Tools don't work this way. And the reason they don't is revealing: there are no natural tools. Every tool is an artefact. Every tool is made for a purpose. The purpose precedes the properties — or rather, the properties are selected and shaped by the design brief. A hammer doesn't have undiscovered capacities lurking in its head. You can use a hammer for purposes other than hammering, but those aren't latent properties — they're the same manifest properties (mass, hardness, surface area) doing manifest things in a different setting. There's nothing hidden.
The direction of fit runs the other way for substances. Properties come first; purposes emerge from them. Nobody looked at a lump of iron and said "this will make good bridges" — they discovered iron's tensile strength, its malleability, its ability to alloy with carbon, and from those discoveries, uses emerged. The properties were always there, latent in the stuff. The uses came later, sometimes much later, as people learned how to invoke what was already present.
## What this does to the LLM argument
If LLMs are tools, then their relevant properties are determined by their design and are in principle apparent from the design brief. What they do now is a fair basis for evaluation, because tools don't have hidden depths. You assess a tool by asking whether it does its job well.
If LLMs are substances, then their properties include latent ones — capacities that were not anticipated by the design brief and that are being discovered through exploration and use. Chain-of-thought prompting is a good example. Nobody designed GPT-3 to reason better when asked to think step by step. That was a latent property of the stuff — something that was there in the model's learned patterns and that was drawn out through experimentation. Same with in-context learning, few-shot generalization, the way these models respond to persona prompts, to structured formatting, to iterative refinement. These are discoveries about what the substance can do, not instructions to a tool about what it should do.
That distinction — discovery vs. instruction — is what makes the substance framing philosophically productive, not the question of whether anyone was surprised.
## How the kerosene-phase thought changes
The kerosene thought still works, but it's no longer a story about accidental discovery. It's a story about the temporal structure of coming to know what a substance can do. Oil was always capable of powering an internal combustion engine. That capacity was always latent in the substance. It wasn't that people stumbled upon it by accident — it's that it took time, experimentation, and the development of new technologies and practices before that particular capacity could be invoked. The kerosene phase wasn't a mistake or a failure of imagination. It was an early, genuine, but limited expression of capacities that had not yet been fully drawn out.
The analogy to LLMs: what we currently do with them — chatbots, writing assistance, code completion, summarization — may be kerosene. Real uses of real properties, but not yet the full expression of what this substance can do. And the point is not predictive (I'm not promising that GPT-for-engines is coming). The point is methodological: Hoel's evaluative confidence presupposes that we already know what the substance's properties are and which of them matter. The substance framing says that this question is still open.
## Options for what this opens up in the essay
Here are several directions this could go. I'm presenting them as parallel options, not ranked:
The artefacts-that-behave-like-substances thought. LLMs were designed — they're artefacts, like tools. But what they turned out to be has exceeded the design brief in ways that are discovered rather than specified. They're made things that behave like found things. That's a genuinely unusual ontological position: manufactured substances. Nylon, carbon fibre, and various alloys are precedents — stuff that was synthesized for one purpose but turned out to have properties that were discovered through subsequent investigation. The essay could note this without making it a whole section — just a sentence that acknowledges the wrinkle and moves on.
The material-profile paragraph. If substances have latent properties, then part of the essay's job is to say *which* properties this substance has — both the ones we've already found and the ones that might be latent. Pattern-completion, contextual pliability, associative reach, strong smoothing and genericising tendencies, responsiveness to framing and persona, the capacity to generate structurally coherent text at scale. Those are known properties. The latent-properties claim says there may be others. The essay doesn't need to speculate about what they are — it just needs to establish that the substance is the kind of thing where further properties could exist.
The evaluation question. If a substance has latent properties, then evaluating it by its current prominent uses is like evaluating oil by kerosene. It's not wrong — oil does work as lamp fuel, and LLMs do help with emails — but it's measuring a substance by one early application. The right question is not "how good is this tool at its job?" but "what can this stuff do, and have we found all of it yet?" That reframing is what makes Hoel's "bits in, bits out" look premature without looking wrong.
The slop-as-tendency observation. This gets stronger under the latent-properties framing. If the substance has characteristic tendencies — smoothing, averaging, genericising — then of course large-scale unresisted use will produce mountains of generic output. That's not evidence that there's nothing there. It's evidence about what happens when a substance's tendencies are allowed to run without skilled resistance. Lucier's room is the analogy: the room has characteristic resonant frequencies (substance-level properties), and when speech is passed through it repeatedly without intervention, those frequencies take over. The result reveals the substance's tendencies, not the substance's limits.
The "no hidden properties for tools" claim as a standalone philosophical point. This could be a paragraph in its own right or just a few sentences within the substance-turn paragraph. The thought is: tools are epistemically transparent in a way substances are not. When you hold a hammer, you know what it can do. Its capacities are apparent from its design. Substances aren't like this. You can hold a piece of silicon and have no idea that it can be used to make a transistor. That hidden-properties feature is what makes substances epistemically interesting, and it's what makes premature evaluation of a substance dangerous in a way that premature evaluation of a tool is not. You can't be wrong about what a hammer is for, but you can be wrong — or more precisely, premature — about what a substance is capable of.
Whether to mention the designed-for/emerged-from distinction explicitly. The essay could have a brief moment where it notes that there are two directions of fit: for tools, function precedes properties (we make the hammer to hammer); for substances, properties precede function (we discover what iron can do, and then we use it for that). LLMs were made with a function in mind (predict the next token, chat, assist), but what they can do exceeds what they were made for. That excess is what the substance framing captures. It's not an accident — it's what happens when you manufacture something complex enough that its properties outrun your design brief.
Whether to keep the pen comparison. I think yes. The pen is a paradigmatically transparent tool. Its properties are apparent. It has no latent capacities. It extends your ability to inscribe; it doesn't push back, surprise, or reveal hidden powers. The contrast with LLMs — which do push back, do surprise, and do reveal properties through use — is sharpened by the latent-properties framing. The pen is the essay's clearest example of what a tool is, and its poverty as a model for LLMs is part of the argument.
Whether to include the Aristotelian language at all. "Latent capacities" is already doing the work that "dynamis" would do, and it does it in English. One option: use the Aristotelian language in a light parenthetical — "(what Aristotle would have called *dynamis* — potentiality waiting to be actualised)" — to give the philosophically literate reader a foothold without asking the general reader to learn Greek. Another option: skip it entirely and let "latent capacities" speak for itself. My sense is that the lighter version works better in a Substack essay than either the full Aristotelian treatment or total omission.
Whether the essay needs a positive vision of what LLM substance-use looks like. Under the old structure, Frippertronics and the recursive-practice material provided this. Under the substance-only structure, you'd need something else — either a direct phenomenological description of working with the substance, or a few concrete examples of latent capacities being drawn out. The chain-of-thought prompting example could serve here: "Nobody told the model to reason step by step. Someone tried it, and it turned out the substance could do that. That is what discovery looks like — not designing a new function, but finding a latent capacity in an existing material."
Whether the essay needs the artefact/working-surface distinction. I think it does, because it answers Hoel's evidence. He points to published writing (artefacts) and says quality hasn't improved. The reply: much of what competent users do with this substance is not aimed at producing artefacts. It's aimed at exploring the substance — testing a distinction, finding out what comes back when you push in a particular direction. Evaluating a substance by its manufactured products misses the exploratory engagement that reveals its properties. This can be brief — a few sentences, not a paragraph — but it needs to be there.
## A possible compressed structure
I'm offering this as one way the moves might sequence, not as the only way:
1. Fixed opening + Hoel quote
2. Pressure on "tool" — if this is a tool, what is it for? (Proper function + candidates that fail, compressed)
3. First verdict — "a quite different type of tool, or not quite a type of tool at all"
4. The turn: substances have latent properties; tools don't. No natural tools. Properties precede purposes for substances; purposes precede properties for tools.
5. LLMs have latent properties — chain-of-thought, in-context learning, etc. as discoveries about a substance, not instructions to a tool
6. Hoel's oil reversed — kerosene phase, not as accident but as the temporal structure of discovering what a substance can do. Anti-hype built in: this is epistemic, not predictive.
7. What this substance does when you engage with it — pen contrast, recursive return, the fact that artefact-level evaluation misses exploratory engagement
8. Slop as characteristic tendency — Lucier — the substance's tendencies dominating under conditions of unresisted use
9. Conclusion — we have synthesized a new substance whose latent capacities are still being discovered; Hoel's "bits in, bits out" may describe the kerosene phase, not the nature of the thing
Nine paragraphs after the fixed opening. Tight. Each one does distinct work. Nothing ornamental. The medium concept doesn't appear, the Frippertronics don't appear, but nothing that those were doing is left undone — it's just done in substance language instead.
The remaining uncertainty I'd flag: the latent-properties claim about LLMs is strong and defensible, but it's an empirical claim as much as a philosophical one. Chain-of-thought prompting, in-context learning, emergent abilities — the AI research community is divided on whether these represent genuine "latent" capacities of the architecture or artefacts of training data, scale, and evaluation methodology. The substance framing doesn't need that debate to be settled — it just needs to be true that we are still discovering what these systems can do, which is observationally obvious. But the essay should be aware that "latent properties" is a philosophical interpretation of the observations, not a neutral description of them.
*Le proprietà latenti non si dichiarano — si scoprono lavorando la materia, e il lavoro non è ancora finito.*
### User
"Tools don't work this way. And the reason they don't is revealing: there are no natural tools. Every tool is an artefact. Every tool is made for a purpose. The purpose precedes the properties — or rather, the properties are selected and shaped by the design brief. A hammer doesn't have undiscovered capacities lurking in its head. You can use a hammer for purposes other than hammering, but those aren't latent properties — they're the same manifest properties (mass, hardness, surface area) doing manifest things in a different setting. There's nothing hidden." this was not the point i was making. i was just saying that tools don't have discoverable but at first unknown properties for two reasons,: 1. they are not natural objects. 2. it doesn't ake sense to say that designed/manufacured objects have hidden properites
### Assistant
OK, let me slow down and actually listen to what Nick is saying rather than elaborating my own version of it.
He's making a simpler, cleaner point than what I attributed to him. Let me try to state it precisely.
The claim is: tools don't have discoverable-but-initially-unknown properties. And there are two reasons for this:
1. Tools are not natural objects.
2. It doesn't make sense to say that designed/manufactured objects have hidden properties.
So the point is not about "direction of fit" or "purpose precedes properties" or any of the philosophical machinery I was building. It's a more straightforward observation.
Let me think about what these two reasons actually amount to.
Reason 1: Tools are not natural objects. Natural objects — iron ore, oil, silicon, wood — exist independently of human purposes. Because they exist independently of us, they can have properties we haven't yet encountered or understood. That's just what it means for something to exist independently: its nature is not exhausted by what we currently know about it. There can always be more to discover. Natural objects have a kind of depth that comes from their independence from human design.
Tools, by contrast, are made. They're artefacts. They don't exist independently of human purposes — they exist because someone made them for a reason. So reason 1 is really saying: only things that exist independently of us can have genuinely hidden properties, because only independent things have a nature that exceeds our current knowledge of them.
Reason 2: It doesn't make sense to say that designed/manufactured objects have hidden properties. This is the stronger claim and I think it's the one that does the real work. If you design and manufacture something, you determine its properties. You choose the materials, the shape, the mechanism. The properties of the finished object are a consequence of your design decisions. So in what sense could there be properties you don't know about? You made the thing. Its properties are a product of your decisions.
Now, someone might object: "But what about unintended side effects of design? What about emergent properties of complex systems?" And Nick's response (if I'm reading him right now) would be: those might exist, but they're not the point. The point is that the *concept* of a tool doesn't allow for hidden properties in the way the *concept* of a substance does. It's a conceptual point, not an empirical one.
Wait, but actually — and I need to be careful here — Nick is making a point about what makes sense to say, not about what's empirically true. "It doesn't make sense to say that designed/manufactured objects have hidden properties." That's a claim about the logic of the concepts, not about what happens in practice. In practice, complex manufactured systems sometimes do exhibit unexpected behaviours. But calling those "hidden properties" would be a category mistake, because "hidden properties" is the kind of thing that applies to natural substances, not to artefacts.
Hmm, but does this fully hold? Let me think of cases...
A pharmaceutical drug is designed and manufactured. But it can have side effects that weren't anticipated — properties that emerge from the interaction between the molecule and biological systems. Are those "hidden properties" of the drug? In ordinary language, we might say yes — "the drug turned out to have properties we didn't know about." But Nick's point might be that the drug-as-substance (the molecule) has those properties, not the drug-as-tool (the medication designed to treat a condition). The drug-as-tool was designed to lower blood pressure. The molecule it's made of turned out to also affect liver function. The hidden property belongs to the substance (the molecule), not to the tool (the medication).
That's actually a very clean way of putting it, and it might be what Nick is driving at. When we discover unexpected properties of a manufactured thing, we're really discovering properties of the substance it's made of, not properties of the tool it was designed to be. The tool is defined by its design brief. The substance it's made of has its own nature.
So the argument for LLMs would be: LLMs were designed as tools (chatbots, text predictors, assistants). But they turned out to have properties — chain-of-thought reasoning, in-context learning, etc. — that belong not to the tool (the chatbot) but to the substance (whatever the trained model actually is as a material with capacities). The interesting thing about LLMs is that the substance seems to have more going on than the tool-level design anticipated. And that's a reason to investigate them as substances rather than evaluate them as tools.
But wait — I'm elaborating again. Let me come back to what Nick actually said.
He said two things:
1. Tools are not natural objects.
2. It doesn't make sense to say designed/manufactured objects have hidden properties.
And his correction was that I was over-complicating it with "direction of fit" and "purpose precedes properties." He just wants the simpler observation stated clearly.
Let me try to state it as simply as I can:
Substances can have latent properties because they exist independently of our purposes — there can always be more to discover about them. Tools can't have latent properties because they are designed and manufactured — their properties are determined by the design. There's nothing there to discover because we put everything there that's there.
That's the two-reason version stated as simply as I can.
And the LLM argument would then be: LLMs seem to have latent properties (capacities discovered through use that weren't part of the design specification). If that's right, then they're more like substances than tools. And if they're more like substances, then Hoel's evaluative approach (judging them by how well they perform their designed function) is the wrong approach.
Is there anything more I should think about here?
I think the interesting question is whether Nick wants this two-reason argument stated in the essay as such, or whether it's just background for the substance claim. In the essay, it might be enough to say: "Substances can have latent properties; tools can't" and let the reasons be implicit. Or it might be worth one sentence: "Tools are made — their properties are determined by design, and there's nothing there to discover that wasn't put there. Substances are different: they exist with properties that can exceed what we currently know about them."
Actually, I realize I should also think about whether Nick's claim is that this applies to ALL manufactured objects or specifically to tools. Because there's a difference. A manufactured object that's simple enough (a hammer, a screwdriver) clearly has no hidden properties. But a manufactured object that's complex enough (a computer, a chemical compound, a neural network) might have properties that weren't anticipated by the designers.
But I think Nick's point is that *qua tool*, the object doesn't have hidden properties. The hammer *qua hammer* has no hidden properties. If the hammer turns out to have some unexpected physical property (maybe the metal alloy resonates at a frequency that repels mosquitoes — silly example), that's a property of the substance it's made of, not of the hammer-as-tool. The tool concept doesn't have room for hidden properties. The substance concept does.
And that's why the question "is an LLM a tool or a substance?" matters. If it's a tool, hidden properties don't make sense — what it does is what it was designed to do. If it's a substance, hidden properties are exactly what you'd expect — and the question becomes what else it can do that we haven't found yet.
OK, I think I understand the correction now. Let me respond simply, acknowledge that I over-complicated it, and state the point as Nick intended it.
Right. I was overbuilding. The point is simpler than what I made of it.
Two reasons tools don't have latent properties:
1. They're not natural objects. There's no independent nature to have more to discover about. A lump of iron exists independently of us — its properties are whatever they are regardless of what we know. A hammer doesn't exist independently of us. We made it.
2. Designed and manufactured objects don't have hidden properties because we determined what's there. We chose the shape, the weight, the material. The properties are consequences of design decisions. There's nothing to discover because we put everything there.
Substances are different not because of "direction of fit" or any fancy philosophical machinery — they're different because they can have properties we haven't yet found. Oil had combustion properties before anyone knew how to build an engine. Silicon had semiconducting properties before anyone knew what a transistor was. The properties were there, in the stuff, waiting to be found. That's what latent means — present but not yet drawn out.
The LLM argument then becomes: these systems appear to have latent properties — capacities that were not part of the design specification and that have been discovered through exploration. Chain-of-thought prompting, in-context learning, the way they respond to iterative refinement. Nobody designed those in. They were found, the way you find properties in a substance. And if that's what's going on, then evaluating LLMs as tools — asking "how well do they do their job?" — misapplies a framework that assumes there's nothing there beyond what was designed. Whereas the substance framing says: there might be more, and we're still finding out.
I was dressing up something clean in unnecessary philosophy of action. The observation itself is sharper without it.
*Ciò che è stato progettato non nasconde nulla — ciò che esiste indipendentemente da noi, invece, può sempre riservare qualcosa che non avevamo ancora trovato.*
### User
you got cut off, please continue from where you left off.
### Assistant
I think my response did land complete but was too thin — I acknowledged the correction without working through what it actually changes. Let me do that now.
So what does the cleaner version of the distinction do to the essay?
The two reasons are: (1) tools aren't natural objects, (2) it doesn't make sense to say designed/manufactured objects have hidden properties. Substances can have latent properties because they exist independently of us — there can be more to them than what we currently know. Tools can't, because we made them, and what we made is what's there.
The first thing I notice is that this gives the essay a much simpler pivot point. The earlier structure was: proper function fails → therefore "tool" is the wrong category → turn to substance. But with Nick's cleaner distinction, the argument can run differently. It's not just that proper function fails. It's that *even if* you could assign LLMs a proper function, the tool category would still be wrong, because LLMs exhibit something tools can't exhibit: discoverable-but-initially-unknown properties. The proper-function failure and the latent-properties observation are two independent reasons to reject "tool," not stages in a single argument.
Hmm, is that right? Let me think...
Actually, I think they might be connected rather than independent. If something has latent properties — properties you haven't yet discovered — then of course you can't assign it a stable proper function yet. The proper-function failure might be a *symptom* of the deeper fact that we're dealing with a substance. You can't say what it's for because it's not the kind of thing that's "for" anything — it's the kind of thing that has properties, some of which you haven't found yet. That's a nicer logical structure: proper function fails *because* this is a substance, not a tool.
But wait, does the essay need that causal connection, or is it enough to just observe both things? I think the causal connection is stronger and more interesting, but the essay could go either way. The simpler version just says: here's a thing that lacks proper function AND has discoverable latent properties — both of those are marks of substance, not of tool. The more ambitious version says: it lacks proper function *precisely because* it's a substance, and substances aren't the kind of thing that has proper functions.
Let me think about what this means for the kerosene-phase thought. Under Nick's cleaner version, the kerosene point isn't about anyone being surprised or anyone accidentally discovering anything. It's about the temporal structure of how we come to know what a substance can do. Oil's combustion properties were always there. They were latent. They were discovered when conditions, practices, and technologies developed enough to draw them out. The kerosene phase was a period during which only some of oil's properties had been invoked. Others were latent — present in the substance, but not yet drawn out.
And the LLM analogy: what we currently do with LLMs — chatbots, writing assistance, code completion — may represent a kerosene-phase engagement with a substance whose full range of properties has not yet been drawn out. Not because anyone was surprised, not because the designers made a mistake, but because *that's what substances are like*. They take time to know. Their properties are drawn out through exploration, experiment, and practice.
Now, what about the objection that LLMs are manufactured? Nick acknowledged this — he said "natural or synthetic." But his two reasons seem to apply to manufactured objects generally. If it doesn't make sense to say designed/manufactured objects have hidden properties, then doesn't that apply to LLMs too, since LLMs are designed and manufactured?
I think the answer Nick would give is: LLMs were designed, yes, but what was designed was the architecture and training procedure. The resulting trained model — the thing that actually sits there and generates text — has properties that were not part of the design specification. Those properties emerged from the training process in ways that the designers did not anticipate and cannot fully predict. So the trained model is more like a synthetic substance (produced by a process, having properties that the process gave it but that weren't fully anticipated) than like a tool (designed for a purpose, with properties determined by the design).
This is actually a crucial move. The designers designed the training process. The training process produced the model. But the model's properties aren't specified by the design — they're produced by the process and discovered after the fact. It's like how a chemist designs a synthesis procedure but doesn't fully predict all the properties of the resulting compound. The compound has to be investigated. Its properties are there, in the substance, but they have to be found.
So the essay might need a moment where it addresses this wrinkle: LLMs are manufactured, but what was manufactured was a process, and what emerged from that process is a substance with properties that have to be discovered. The tool/substance distinction doesn't perfectly map onto designed/natural. It maps onto something more like: things whose properties are fully determined and known by the maker vs. things whose properties exceed what the maker specified or anticipated.
Actually, let me reconsider. Maybe the essay doesn't need to address this wrinkle at all. Maybe it's enough to say: substances have latent properties; tools don't; LLMs have latent properties; therefore LLMs are more like substances than tools. The philosophical question about how a manufactured thing can have latent properties is interesting but it might take the essay sideways. The Substack reader doesn't need a theory of synthetic substances. They need to see the contrast, see the evidence, and draw the conclusion.
Let me think about what other implications Nick's correction has...
One thing it does is make the argument more available to empirical examples. If the claim is "LLMs have properties that were initially unknown and were discovered through use," then you can list those discoveries: chain-of-thought prompting, in-context learning, the effectiveness of persona prompts, few-shot learning, the way models respond to structured formatting, the way iterative prompting draws out better results. Each of these is a case where someone found that the substance could do something that wasn't part of the design specification. Not an accident — a discovery. The way you discover that iron can be magnetized, or that glass can be drawn into fibers.
Another thing it does is sharpen the anti-Hoel point. Hoel evaluates LLMs by their current observable effects on writing. But if LLMs are substances with latent properties, then their current observable effects might represent only a partial expression of their capacities. Not because something magical is coming, but because that's how substances work — you don't know everything about them from your first encounters. Evaluating a substance by its current uses is like evaluating iron by its use in primitive tools and concluding it's not very impressive. The conclusion might be correct about the current uses while being wildly premature about the substance.
Let me think about the "no natural tools" claim more carefully. Nick said it as though it's obvious. Is it? What about a rock used as a hammer? A stick used to reach something? These are natural objects used as tools. But I think Nick's point is that the *natural object* (the rock) is a substance — it has properties (hardness, weight, shape) that exist independently of any use. When someone picks it up and uses it to crack a nut, they're exploiting the substance's properties for a purpose. The tool is the use, not the object. The rock is a substance; the act of hammering with it makes it a tool. But the toolhood is imposed, not intrinsic. And once you impose toolhood, you've determined what properties are relevant (hardness, weight) and there are no hidden properties left to discover in the *tool*. The rock-as-substance might still have latent properties (maybe it's magnetic, maybe it contains useful minerals), but those aren't properties of the *tool*.
This is clean and it holds up. Tools are defined by use. Substances are defined by properties. Use is determined by the user and is fully transparent. Properties exist independently and can be latent.
Now let me think about what the essay would look like structurally with this cleaner argument.
The proper-function discussion still belongs, but its role changes. It's no longer the main argument against "tool." It's evidence that something is off — a symptom. The main argument is the latent-properties observation: LLMs have discoverable-but-initially-unknown properties, and tools don't. Proper function fails because you can't specify the function of something when you're still discovering what it can do. You can't assign a function when the properties are still being found.
So the sequence might be:
1. Fixed opening + Hoel quote (tool seems right)
2. But wait: what kind of tool? (Proper function fails — can't specify what LLMs are *for*)
3. Why proper function fails: the latent-properties distinction. Substances can have discoverable-but-initially-unknown properties. Tools can't — they're manufactured, their properties are determined by design. LLMs have discoverable latent properties. Therefore they're more like substances than tools.
4. What this means: we don't ask what iron is "for" — we ask what iron can do. Same with LLMs.
5. Hoel's oil reversed: kerosene phase as the temporal structure of coming to know a substance.
6. What this substance does when you work with it: pen contrast, recursive return, the experience of discovering capacities through use.
7. Slop as characteristic tendency of the substance under unresisted use (Lucier).
8. Conclusion: Hoel's evaluation presupposes that we know what the thing is. We don't yet, and for a substance, that's normal.
That's eight moves after the fixed opening. Tight, clean, one-thought. Each paragraph does distinct work. No medium, no Frippertronics, no *Growing the Image*.
The other thing I want to think about is whether cutting medium and Frippertronics also means cutting the artefact/working-surface distinction. In the previous structure, that distinction answered Hoel's evidence about writing quality. Without it, how does the essay answer Hoel?
I think the answer is: the kerosene-phase thought does that work. You don't need to distinguish between artefacts and working surfaces to say "evaluating a substance by its current prominent application is premature." The kerosene-phase thought handles this directly: what we see now (writing assistance, chatbots, slop) is a genuine but partial expression of the substance's capacities. Just as kerosene was a genuine but partial use of oil. Hoel is measuring kerosene output and concluding that oil is unimpressive. The problem isn't that he's looking at the wrong products (artefacts rather than processes) — it's that he's evaluating a substance by one early use.
Actually, is that better or worse than the artefact/working-surface answer? Let me compare:
Artefact/working-surface answer: "You're measuring the wrong thing. You're looking at published texts when you should be looking at how people work with the substance."
Kerosene-phase answer: "You're measuring too early. You're evaluating a substance by its first large-scale application, which is not yet the application that will reveal its real capacities."
These are different answers and they're both good. The kerosene answer is more radical — it says the right application might not exist yet. The artefact/surface answer is more modest — it says the right application exists but Hoel's looking at the wrong one.
For a Substack essay, the kerosene answer is probably better because it's more vivid and more surprising. The artefact/surface answer requires the reader to buy a somewhat technical distinction between types of text. The kerosene answer just requires them to buy a historical analogy.
But... I think there might be room for one sentence of the artefact/surface thought within the "what this substance does when you work with it" paragraph. Something like: "And much of what matters in that engagement never becomes a published text — it becomes a sharpened thought, a discovered distinction, a reformulated question." That's the artefact/surface point stated as an observation rather than a theoretical distinction.
Let me also think about what other examples besides chain-of-thought could illustrate the latent-properties claim. Because chain-of-thought is good but it's somewhat technical, and a general Substack audience might not know what it means.
More accessible examples:
- The way LLMs turned out to be surprisingly good at translation, even though they weren't specifically designed for it. The training was on text prediction, but the substance turned out to have translation capacities latent in its patterns.
- The way LLMs turned out to be useful for brainstorming and ideation — generating options, exploring conceptual space — in a way that nobody specifically designed for.
- The way iterative prompting draws out increasingly precise responses, as if you're learning how to work the material rather than giving a tool better instructions.
- The way the same model can write code, explain concepts, generate fiction, analyse arguments, and do dozens of other things — not because it was designed as a multi-tool, but because its latent capacities extend in many directions.
These are all more accessible than chain-of-thought and they all illustrate the same point: the substance turned out to be able to do things that weren't in the design brief. Those capacities were latent in the model's learned patterns and were drawn out through exploration and use.
One more thought. The "we have synthesized a new substance" line is strong, but I wonder whether the essay should also say something about *what kind* of substance this is. "Substance" is very general. Iron, oil, nylon, silicon — these are all substances but they're very different from each other. What kind of substance is an LLM?
The most honest answer is: we don't fully know yet, and that's part of the point. But you could gesture at it: "a substance whose raw material is human language, whose characteristic behavior is pattern-completion and associative continuation, and whose latent capacities are still being drawn out through the collective work of millions of users." That gives the reader something to hold onto without pretending to have a complete material profile.
Or even more concretely: "a substance that takes language in and sends language back — but language shaped by its own tendencies, its own patterns, its own characteristic ways of completing, connecting, and smoothing." That's vivid and accessible and it captures what makes LLM substance distinctive from other substances.
## What the cleaner distinction does to the essay
The two reasons — (1) tools aren't natural objects, (2) designed/manufactured objects don't have hidden properties — give the essay a simpler and more powerful turn than the "proper function fails" argument alone.
Under the old structure, the pivot was: proper function fails → "tool" is the wrong category → look for a different one. Under the cleaner version, the pivot is: LLMs have discoverable-but-initially-unknown properties, and tools can't have those, because tools are manufactured and their properties are determined by design. Full stop. Proper function failure becomes a *symptom* of the deeper issue rather than the issue itself. You can't assign a function to something when you're still finding out what it can do.
That changes the order of the argument. Instead of spending three or four paragraphs showing that proper function fails (candidates that don't work, Swiss Army knife, etc.), you can compress the proper-function point into a paragraph or even a few sentences, and move quickly to the real claim: the reason the tool category doesn't fit is that tools are the kind of thing whose properties are determined by design, while LLMs are the kind of thing whose properties are being discovered through use. That's the substance/tool distinction, stated at its simplest.
## What happens to the kerosene thought
It survives but its framing shifts. It's no longer about accidental discovery. It's about the temporal structure of learning what a substance can do. Oil always had combustion properties. Those properties were latent — present in the substance but not yet drawn out. The kerosene phase was a period during which only some of oil's properties had been invoked. Others were there, waiting.
The LLM analogy: chatbots, writing assistance, code completion, summarization — these might be kerosene. Genuine expressions of real properties, but not the full story. Not because anyone made a mistake, and not because something magical is coming. Because *that's how substances work*. Their properties are drawn out over time, through exploration and practice.
And the anti-Hoel point sharpens: evaluating a substance by one early application is like evaluating oil by kerosene. It's not wrong about the kerosene. It's premature about the substance.
## Options for the middle of the essay
Once the substance turn is made and the kerosene thought has landed, the essay needs to do some work showing what this substance is actually like — what its properties are, what it does when you engage with it, what happens when its tendencies dominate. Here are options for how to handle that:
Option 1: Empirical examples of latent properties. List specific cases where LLMs turned out to be able to do things that weren't in the design specification. These are accessible to general readers and they make the substance claim feel concrete rather than metaphorical:
- Translation: the models were trained on next-token prediction, but translation capacity was latent in their learned patterns
- Brainstorming and ideation: nobody designed a brainstorming tool, but the substance turned out to be good at generating options and exploring conceptual space
- Iterative refinement: the way repeated engagement draws out increasingly precise responses, as if you're learning to work the material
- The sheer range of capacities — code, explanation, fiction, analysis — not designed as a multi-tool, but with latent capacities extending in many directions
- Chain-of-thought: the substance turned out to reason better when prompted to think step by step — a latent property discovered through experimentation, not designed in
The advantage: concreteness, accessibility, makes "latent properties" feel real. The risk: it could read as a list, and lists kill momentum.
Option 2: Direct phenomenology. Instead of listing properties, describe what it's like to work with the substance. "You send language in. Something comes back — shaped by the substance's own tendencies, not by your intention. Most of it is flatter than what you wanted. But one move in the second paragraph gives you something to push against. You delete the output and write a better sentence than you would have written without the encounter." This replaces Frippertronics with direct experience. It makes the same point — the substance has its own behavior, and engaging with it reshapes your work — without borrowing an analogy from music.
The advantage: vivid, personal, sounds like someone who actually does this. The risk: it might feel like a tangent in an essay that's supposed to be answering Hoel.
Option 3: Material profile. One paragraph that sketches what kind of substance this is. Not a technical spec but a craftsperson's description: "A substance that takes language in and sends language back — but language shaped by its own characteristic tendencies. It smooths, it averages, it finds the most probable path through a topic. It connects associatively rather than logically. It's responsive to framing, to persona, to structured formatting. And it has a strong pull toward genericity — left to its own tendencies, it produces the most generic version of whatever you asked for." This gives the reader a feel for the material without being technical, and it sets up the slop paragraph naturally.
Option 4: The pen contrast alone. Keep it minimal: a pen extends inscription, it doesn't push back. An LLM does. The difference marks the difference between a tool (transparent, no hidden behavior) and a substance (has its own tendencies that interact with your intentions). Then move straight to slop.
## What happens to slop/Lucier
This actually gets stronger. Slop is what happens when the substance's characteristic tendencies — smoothing, genericity, the pull toward the most probable continuation — run without resistance. That's a substance-level observation. Lucier's room has characteristic resonant frequencies; when speech is passed through it without intervention, those frequencies take over. Same logic: when language is passed through an LLM without resistance, the model's own tendencies take over. The result reveals the substance's behavior, not the user's intentions. That's slop.
And the point is: slop is not evidence against the substance having interesting properties. It's evidence about what the substance does under one specific condition — unresisted, mass-scale use. A substance can have both interesting latent capacities and characteristic degradation patterns. Those aren't contradictory. They're both features of the material.
## What I'd flag as the remaining structural question
The essay has to decide how much middle it needs between the substance turn (paragraph 3 or 4) and the slop/conclusion (paragraphs 7–8). If the proper-function material is compressed and medium/Frippertronics are cut, the essay might feel like it goes: claim → evidence → oil → conclusion, with not enough texture in the middle. The "what is this substance actually like?" question needs answering, otherwise "substance" remains a label rather than an observation.
But the answer doesn't need to be long. One paragraph of material profile or direct phenomenology, plus the pen contrast, plus slop/Lucier — that's enough middle. Three paragraphs that show what the substance is like, between the oil reversal and the conclusion. That gives the essay body without making it long.
## A possible tightened structure
1. Fixed opening + Hoel quote
2. Proper function compressed: "What kind of tool? We can't say. And the reason we can't say is revealing."
3. The turn: substances can have latent properties — discoverable but initially unknown. Tools can't, because they're designed and manufactured, and designed objects don't have hidden properties. LLMs have latent properties. They're substances, not tools.
4. Hoel's oil reversed: kerosene phase as temporal structure of knowing a substance. With epistemic caution built in — this is not a prediction, it's a point about what we can and can't claim yet.
5. What this substance is like: material profile or phenomenology (one paragraph)
6. Pen contrast: what distinguishes substance-engagement from tool-use (brief — half a paragraph or folded into the paragraph above)
7. Slop as characteristic tendency (Lucier): what the substance does when its tendencies run without resistance
8. Conclusion: Hoel's "bits in, bits out" describes what we currently see. It may describe the kerosene phase of a substance whose latent capacities are still being drawn out. That doesn't mean hidden greatness is inevitable. It means evaluative confidence is premature.
Eight moves. Clean. One thought all the way through. Every paragraph earns its place.
The essay would be noticeably shorter than the current draft. That's a feature, not a bug. A Substack essay that makes one claim clearly and memorably is worth more than one that makes three claims in a tangle.
*La differenza tra strumento e sostanza è questa: dello strumento sai già tutto, perché l'hai fatto tu — della sostanza, invece, c'è sempre qualcosa che non hai ancora trovato.*
### User
"kerosene"you understand that this was your idea that you have fixated on right? i have lieraly never wriiten the wird once
### Assistant
He's right. Let me think about this.
The word "kerosene" and the "kerosene phase" framing — that was generated by the AI in the previous session. Nick never wrote it. He never used the word. It was an assistant's elaboration of his oil-analogy-reversal thought, and then it got picked up and repeated across the conversation, across Codex's response, and across all of my responses as though it were Nick's idea.
And I've been treating it as a load-bearing element of the essay structure. I've been saying things like "the kerosene-phase thought is the essay's strongest single move" and building paragraph slots around it and treating it as settled. But Nick never proposed it. It's not his thought. It's an AI assistant's metaphorical elaboration that got amplified through repetition across multiple sessions and multiple AI interlocutors.
This is exactly the kind of thing epistemic discipline is supposed to catch. I should have been tracking the provenance of ideas. "Kerosene phase" was generated by an assistant in the JSONL session. It was repeated by Codex. I repeated it. And now it's been presented back to Nick multiple times as though it's part of his argument.
Now — is it a *bad* idea? That's a separate question. Maybe it's a useful way of making the oil-reversal point vivid. But the point is that it's not Nick's idea, and I've been treating it as though it were, and building structures around it as though he'd committed to it. He's telling me to stop fixating on it.
What's the actual idea that IS Nick's? The oil reversal is partly Hoel's — Hoel introduces the oil analogy. Nick's move is to turn it against Hoel by pointing out that discovering a substance doesn't mean immediately knowing all its capacities. That's the real thought. "Kerosene phase" is just one way of illustrating it, and it's a way Nick never chose.
So what should I do? I should:
1. Acknowledge clearly that "kerosene phase" was never his word or his framing
2. Stop treating it as a fixed element of the essay
3. Go back to the actual thought that IS his — the oil-reversal and the latent-properties claim — and think about whether the essay needs the kerosene illustration at all, or whether the point can be made more directly
The underlying point — that substances have latent properties and early uses don't reveal the full range — doesn't need the kerosene example. It can be stated directly. "When oil was first used, not all of its capacities were known" is the thought. "Kerosene phase" is an embellishment that the AI added.
Actually, does the essay even need the oil analogy at all? Hoel introduces it. Nick could engage with it or not. The substance thesis doesn't depend on oil. It depends on the general claim about substances having latent properties. Oil is Hoel's example. Nick could turn it, or he could just note it and move on.
Let me think about this more carefully. The oil reversal IS genuinely effective — it takes Hoel's own analogy and shows that it actually supports Nick's claim rather than Hoel's. That's a strong rhetorical move. But the way to make that move might be much simpler than "kerosene phase." It might just be: "Hoel compares the discovery of LLMs to the discovery of oil. But when oil was first discovered, we didn't know what it could do. We're still in that position with LLMs."
That's two sentences. It doesn't need the word "kerosene." It doesn't need a whole paragraph. It just turns the analogy and moves on.
I've been guilty of exactly what Nick's system is designed to prevent: an AI assistant generating a vivid phrase, falling in love with it, and then treating it as the client's thought. I need to back off and let Nick decide whether and how to use the oil reversal, without pre-packaging it in my own metaphor.
You're right, and I should have been tracking that. "Kerosene phase" was generated by the assistant in the session you exported. You never wrote it, never endorsed it, never used the word. Codex picked it up. I picked it up from Codex. And across three layers of AI interlocutor it hardened into something that looked like your idea when it was never yours.
Your actual thought is: substances can have latent properties; tools can't; LLMs look like they have latent properties; therefore "tool" is the wrong category. The oil reversal — turning Hoel's analogy against him by noting that discovering a substance doesn't mean immediately knowing all its capacities — is a move available to you if you want it. But it doesn't need to be called "kerosene phase" and it doesn't need a whole paragraph, and I should have been presenting it as *one option* rather than as the essay's strongest moment.
I'll stop building around it. What are you actually thinking about the oil analogy — do you want to turn it, ignore it, or do something else with it?
*Ripetere una parola finché sembra vera non la rende tua — la provenienza conta.*
### User
turn an analogy?
### Assistant
Nick is questioning my phrase "turn an analogy." Let me think about what he's asking.
He might be saying: what do you mean by "turn" an analogy? That's my phrase, not his. I've been using "turn Hoel's oil analogy against him" as though that's a standard move. But "turn an analogy" is jargon — it's something I've been saying, not something Nick said.
Or he might just be asking what I mean by it — a genuine clarification question.
Let me just explain simply what I mean. By "turn" I meant: take Hoel's oil comparison — he says if LLMs were really a source of intelligence, discovering them should have been like discovering oil — and use it to support the opposite conclusion. Hoel's point is that discovering LLMs hasn't produced oil-level transformation, therefore they're just tools. The reversal would be: actually, the oil comparison supports the substance thesis, because when we first discovered oil we didn't immediately know all its properties either. So the fact that LLMs haven't yet transformed writing doesn't tell us what they are — it just tells us we're early.
But actually — I keep re-explaining the same move. Maybe Nick's question is simpler than I'm making it. Maybe he's just saying: "turn an analogy" is a weird phrase, what do you mean? Or maybe he's pushing back on the whole idea of engaging with Hoel's oil analogy at all.
Let me just answer plainly.
Sorry — my phrasing. I just meant: use Hoel's oil comparison to support your point rather than his. He says discovering LLMs should have been like discovering oil if they were really significant. You could point out that the oil comparison actually supports the substance thesis — discovering oil didn't mean immediately knowing everything it could do.
But that's me narrating a move you might not want to make. Do you want to engage with Hoel's oil analogy at all, or leave it alone?
*A volte la cosa più utile è smettere di spiegare e chiedere.*
### User
engage but dont go crazy
### Assistant
Got it. Note it, use it, move on. One or two sentences, not a set piece.
So where are we? The essay as you're now thinking about it:
- Fixed opening + Hoel quote
- Proper function compressed — what kind of tool? Can't say.
- The substance turn — substances have latent properties, tools don't, LLMs have latent properties
- Brief engagement with Hoel's oil — his own analogy supports the point, because discovering a substance doesn't mean immediately knowing its capacities
- What this substance is like when you work with it — pen contrast, the experience of discovering what it can do
- Slop as characteristic tendency (Lucier)
- Conclusion — evaluative confidence is premature
That's roughly seven moves after the opening. Short, focused, one thought. What do you want to do — talk through any of these more, or start working on the draft?
*Sette mosse bastano, se ognuna fa il suo lavoro.*
### User
"Proper function compressed " you invented this as well, what do you mean compressed, ewhat have you lazily removed?
### Assistant
Nick is calling me out again. He's right to be suspicious. "Proper function compressed" is my editorial shorthand — I've been saying "compress" the proper-function discussion as though it's obvious what gets kept and what gets cut, but I haven't actually specified what I mean. And he's asking: what exactly are you proposing to remove, and why?
Let me think about what the proper-function discussion actually does in the current draft, and what would be lost if I "compress" it.
Let me re-read the current draft to remind myself what's actually there. I read it earlier in the session. The proper-function section runs roughly from after the opening Hoel quote through to the "first verdict" line. Let me think about what's in it...
The current draft has:
- The observation that we use LLMs, which makes them seem like tools
- The proper-function question: if it's a tool, what's it for?
- The comparison cases: hammer (clear function), Google (clear function despite complexity), Swiss Army knife (multiple functions but each identifiable)
- The LLM case: "predict the next token" is mechanism not function, "chat" is too thin, "help with tasks" is vacuous
- The conclusion: LLMs are "a quite different type of tool, or not quite a type of tool at all"
When I said "compressed," what was I actually proposing? I think I was vaguely suggesting that the comparison cases (hammer, Google, Swiss Army knife) could be shortened. But I never said which ones to cut or why. And Nick is right to push back — those cases might be doing real work. The hammer gives you the simple case. Google gives you the complex-but-still-has-a-function case. The Swiss Army knife addresses the obvious objection that multi-function tools exist.
Actually, do all three need to be there? Let me think...
The hammer is the paradigm case — you need it. It establishes what proper function looks like in the simplest case.
Google is interesting because it shows that even complex, multi-purpose digital tools have a specifiable function (search, organize, retrieve information). This matters because the obvious objection to "LLMs lack proper function" is "well, Google does lots of things too." Google answers that objection: yes, but Google's function is still specifiable. You can say what Google is for.
The Swiss Army knife addresses a different objection: "tools can have multiple functions." Yes, but each function is identifiable and designed in. The Swiss Army knife has a blade, a corkscrew, a file — each with a clear purpose.
Do all three address different objections? Hammer: basic case. Google: complexity objection. Swiss Army knife: multiplicity objection. They do seem to be doing different things. Cutting any of them leaves an objection unanswered.
So maybe "compressed" was lazy. Maybe the proper-function discussion is already the right length and I was just waving my hand at making the essay shorter without doing the work of figuring out what's actually dispensable.
On the other hand — is the proper-function discussion currently one paragraph or several? In the draft it might run across two or three paragraphs. If it's three, could it be two? That's a different question from "what do you cut" — it's "can the same material be tighter?"
But I haven't re-read the draft carefully enough to know. And Nick is right to call me out for saying "compressed" without specifying what that means.
Let me just be honest: I don't have a good answer for what specifically I'd cut. I was being lazy. The proper-function material might all be necessary. The right answer is probably that the proper-function discussion stays roughly as it is, and the essay's brevity comes from cutting medium/Frippertronics, not from trimming the proper-function argument.
You're right — I was being vague. I hadn't actually decided what to cut; I was just gesturing at "shorter" without doing the work.
Let me re-read the proper-function section as it currently stands before I say anything about what it needs.
### Assistant
Having re-read it, I think the proper-function discussion (lines 26–45) is already tight and I was wrong to suggest compressing it. Three blocks, each doing different work:
Lines 26–31 establish what proper function *is* — not just any use but the use that distinguishes what something is *for*. The Swiss Army knife handles the multiplicity objection. The familiar examples anchor the concept.
Lines 33–38 apply the test to LLMs and show it fails — token prediction is mechanism, chatting is too narrow, helping with tasks is vacuous. None settles the question the way "for hammering" does.
Lines 41–45 state what the failure means and what it doesn't — classificatory hesitation, not metaphysical victory.
Each paragraph does a distinct job. Cutting any of them either leaves an objection unanswered or removes the qualifications that make the argument honest. The qualifications are your voice — "I am not claiming to have proved," "the issue is not just that" — and without them the argument sounds cruder than it is.
When I said "compressed" I was making the outline look shorter without checking whether anything was actually dispensable. It wasn't a considered editorial judgment. The proper-function section should probably stay roughly as it is, and the essay's new brevity comes from cutting medium/Frippertronics and not having to build all that apparatus — not from trimming an argument that's already working.
*Tagliare ciò che funziona per far sembrare il piano più elegante non è economia — è pigrizia.*
### User
"waving my hand at making the essay shorter" stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter. stop trying to make things fucking shorter.
### Assistant
Understood. I'll stop optimizing for brevity. The essay is as long as it needs to be to make the argument properly. I've been treating "shorter" as a default virtue when you never asked for that.
What do you want to do next?
*Il saggio sarà lungo quanto il pensiero richiede.*
### User
do a real attempt at the task i have been asking you to do in which you don't cut corners
### Assistant
OK. Nick has been asking me to think about a new structure for the essay, recalibrated around the substance thesis — specifically, that substances (natural or synthetic) can have latent properties that are discoverable but initially unknown, and tools can't, because tools are designed/manufactured and designed objects don't have hidden properties. LLMs appear to have latent properties, so they're more like substances than tools.
He wants a detailed structure in bullet-point framework. He's said he's happy for considerable restructuring after the opening paragraph and first block quote. He wants me to think about the CEV of an essay recalibrated this way. He doesn't want me to cut corners, compress things, or optimize for brevity. The essay is as long as it needs to be.
Let me also remember what he's said along the way:
- The "kerosene phase" is not his idea; he never wrote the word. I should not build around it.
- The "accidentally" synthesized framing is not his point. The point is that substances can have latent properties and tools can't.
- The two reasons tools can't have latent properties: (1) they're not natural objects, (2) it doesn't make sense to say designed/manufactured objects have hidden properties.
- He wants to engage with Hoel's oil analogy but not go crazy with it.
- He's cutting medium and Frippertronics for the time being, going straight to substance.
- He does NOT want me trying to make things shorter. The essay is as long as it needs to be.
- He wants a very detailed structure in bullet-point framework.
- He wants me to think about the CEV — the coherent extrapolated volition — of what this essay could be.
Let me also remember what's in the current draft. The fixed elements are:
- The opening paragraph (line 13): "Here's a sentence that sounds almost trivially true: LLMs are a type of tool..."
- The first block quote (lines 15-16): Hoel on tools vs. sources of intelligence, the oil comparison, efficiency and scale.
Everything after that is open to restructuring. But "open to restructuring" doesn't mean "throw it all away." There's good material in the current draft that can be reused. The proper-function discussion (lines 26-45) is working well. The pen comparison (lines 81-86) is strong. The Lucier/slop material (lines 115-120) works with the substance framing. The artefact/working-surface distinction (lines 65-78) answers Hoel's evidence.
What needs to change is: medium goes, Frippertronics goes, and the substance turn needs to be inserted. The substance turn is the new heart of the essay.
Let me think carefully about what the essay needs to do, move by move, without trying to compress or trim anything.
MOVE 1: Opening + Hoel quote (FIXED)
This is settled. The reader encounters the "mere tools" position. Hoel's framing: if LLMs are tools, expect efficiency and scale; if they're sources of intelligence, expect transformation like oil.
MOVE 2: The transition after the quote
The current draft has several things here:
- Lines 17-19: interpretive gloss on Hoel (if merely tools, expect efficiency not transformation)
- Line 20: Hoel's argument about writing as the proving ground
- Line 21: second Hoel quote about stilted emails, dearth of good writing
- Line 22: Nick's turn — "What I am less sure of is the description doing the philosophical work. 'Tool' sounds right, until one asks a little more carefully what sort of tool this is supposed to be."
That last line is the hinge. It's what opens the proper-function investigation. I need to think about whether this transition still works under the substance framing, or whether it needs adjustment.
I think it does still work. The essay still needs to go from "tool sounds right" to "but what sort of tool?" before it can get to "actually, not a tool at all — a substance." The proper-function investigation is the path from "tool" to "substance." You can't just assert "it's a substance" — you have to show why "tool" doesn't work first, and the proper-function failure is what shows that.
But there's also a question about the second Hoel quote (line 21, about stilted emails and the dearth of good writing). Does it belong here, or later? In the current draft it comes very early, right after the first quote. In the recalibrated essay, Hoel's writing argument might fit better after the proper-function discussion, when the essay turns to evaluation. Let me think about this...
Actually, reading the current draft more carefully, lines 17-21 are doing several things at once:
- Interpreting Hoel's position
- Introducing his writing argument
- Presenting his evidence (the slop quote)
And then line 22 turns away from Hoel toward the proper-function question. Maybe the writing argument and the slop evidence should be separated — the interpretation of Hoel's position stays early (to set up the target), but the detailed writing argument comes later (when the essay is ready to engage with it). That would let the essay go: Hoel's position → "what sort of tool?" → proper function → substance turn → NOW let's look at Hoel's writing argument with substance-informed eyes.
But I'm getting ahead of myself. Let me think about the full structure first and then worry about sequencing.
MOVE 3: "We use them" — transitional concession
Line 24: "Another reason it seems right to say that LLMs are tools, is that it seems right to say that we *use* them. But what does that amount to?"
This is a good transitional beat. It concedes the intuitive pull of "tool" before investigating it. It earns trust. Under the substance framing this still works — yes, we use them, but using something doesn't make it a tool. We use iron. We use oil. We use clay. Using a substance is different from using a tool.
Actually, that thought — "we use them, but using something doesn't make it a tool; we use substances too" — could be a nice bridge from the proper-function discussion to the substance turn. Right now it's positioned before proper function. Maybe it should come after, as part of the transition to substance. Hmm, but in the current draft it works well where it is because it motivates the proper-function investigation: we use them, so they seem like tools, but what kind of tool?
I think it should stay roughly where it is but the "what does that amount to?" question gets a different answer later. In the current draft, the answer is: we use them, but their proper function is unclear. In the recalibrated essay, the answer eventually becomes: we use them, but using something doesn't make it a tool — we also use substances, and the way we use substances is different (we explore, we discover, we learn the material).
MOVE 4: Proper function
Lines 26-38. This is working well as-is. Establishes what proper function means, shows that multiplicity isn't the issue (Swiss Army knife), applies the test to LLMs, and shows it fails. Nick has confirmed this should stay.
MOVE 5: First verdict
Lines 41-45. "That does not show that LLMs are not tools; but it does suggest that they are a quite different type of tool, or not quite a type of tool at all." With qualifications. This is the end of the proper-function argument and the beginning of the turn.
Now, in the current draft, the essay goes from here to the evaluation problem (lines 48-53), then to Hoel's writing argument (lines 56-62), then to artefact/working-surface (lines 65-78), then to the pen comparison (lines 81-86), then to medium (lines 98-104), then Frippertronics (lines 107-113), then slop/Lucier (lines 115-120), then conclusion (lines 123-129).
In the recalibrated essay, the path from the first verdict needs to go differently. Instead of going to evaluation → writing → artefact/surface → pen → medium → Frippertronics, it needs to go to the substance turn.
So the question is: what comes between "not quite a type of tool at all" and "here's what it might be instead: a substance"?
One option: go directly. "If LLMs are not straightforwardly tools, what are they? I want to suggest a different category: substance."
Another option: go through the evaluation problem first. "If we don't know what kind of tool this is, we don't know how to evaluate it. That's important because Hoel's argument depends on selecting writing as the test. But the uncertainty about proper function makes that selection less obvious than it seems." Then turn to substance as the alternative.
I think the second option is better because it keeps Hoel in play. The evaluation problem is not just a philosophical aside — it's the direct reply to Hoel's method. And it sets up the substance turn by showing that the problem isn't just classificatory but evaluative: if we're wrong about what kind of thing this is, we might be wrong about how to assess it.
But I also want to think about where Hoel's writing argument gets its fullest statement. In the current draft, the writing argument gets a dedicated paragraph (lines 56-62) and a second Hoel quote. In the recalibrated essay, the writing argument still needs to be stated — it's what the essay is responding to. The question is whether it goes before or after the substance turn.
Arguments for before: you set up the target completely, then knock it down. The reader understands exactly what Hoel is claiming before you offer the alternative.
Arguments for after: you establish the substance category first, and then show how it reframes Hoel's writing argument. The reader has already shifted categories and can see the writing argument in a new light.
I think before is better for this essay. The essay is a response to Hoel. The reader needs to understand what they're responding to before the response can land. So the sequence would be: proper function fails → evaluation becomes uncertain → Hoel's writing argument stated at full strength → NOW the substance turn.
Wait, but that's a lot of material before the substance turn arrives. Let me count: opening + quote (1), transition (2), "we use them" (3), proper function (4), candidates that fail (5), first verdict (6), evaluation problem (7), Hoel's writing argument (8), THEN substance turn (9). That's eight paragraphs before the positive thesis. Is that too long?
NO. Stop. Nick just told me to stop optimizing for brevity. The essay is as long as it needs to be. If the argument requires eight paragraphs of setup before the turn, then it requires eight paragraphs. The question is whether each paragraph does work that the argument needs, not whether there are too many of them.
Let me check: does each of those eight paragraphs do necessary work?
1. Opening: establishes topic and Hoel's position. Necessary.
2. Transition: interprets Hoel's argument, pivots to "what sort of tool?" Necessary.
3. "We use them": concedes the intuitive pull, motivates the investigation. Necessary for rhetorical balance.
4-5. Proper function + candidates: establishes the test and shows it fails. Necessary — this is the diagnostic step.
6. First verdict: states what the failure means and qualifies it. Necessary for intellectual honesty.
7. Evaluation problem: connects proper-function failure to evaluation. Necessary because it's the link to Hoel's method.
8. Hoel's writing argument: states the target at full strength. Necessary because you can't respond to something you haven't stated.
Yes, all eight are doing necessary work. Moving on.
MOVE 9 (or wherever we're up to): THE SUBSTANCE TURN
This is the heart of the recalibrated essay. It needs to do several things:
First, it needs to say what the alternative is. Not just "tool doesn't work" but "here's what might work instead."
Second, it needs to explain what "substance" means and what it commits you to.
Third, it needs to make the latent-properties distinction — the two reasons Nick identified: (1) tools aren't natural objects, (2) designed/manufactured objects don't have hidden properties. Substances can have latent properties; tools can't.
Fourth, it needs to apply this to LLMs — show that LLMs have properties that were discovered through use, not designed in.
Can this all be one paragraph? Probably not. The substance turn is doing too much conceptual work for one paragraph. It needs to be at least two, maybe three.
Let me think about how to break it up:
Paragraph A: The category shift. The reason the proper-function question keeps failing is not that LLMs are a weird kind of tool. It's that they might not be tools at all. They might be better understood as a substance — stuff with properties and capacities, not an implement with a function.
Paragraph B: What this means. Substances — natural or synthetic — can have latent properties: properties that are present in the stuff but not yet known, not yet drawn out, not yet invoked. We didn't know iron could be magnetized when we first found it. We didn't know oil could power engines when we first encountered it. The properties were there, in the substance. They were discovered through exploration, experiment, and practice. This is just what substances are like — their nature exceeds what we currently know about them.
Paragraph C: Why tools are different. Tools are not like this. There are no natural tools — every tool is designed and manufactured for a purpose. And it doesn't make sense to say that designed, manufactured objects have hidden properties. The properties of a hammer are determined by its design. You don't discover that a hammer can do something nobody anticipated, in the way you discover that a substance has a capacity nobody knew about. Tools are epistemically transparent in a way substances are not.
Paragraph D: LLMs as substance. LLMs look more like substances than tools in this respect. They have properties that were not part of the design specification — capabilities discovered through use and experimentation. [Examples: chain-of-thought, in-context learning, few-shot generalization, the way they respond to persona prompts, to structured formatting, to iterative refinement.] These are discoveries about what the substance can do, not instructions to a tool about what it should do. The design specified next-token prediction; what emerged has capacities that far exceed that specification.
Now — is four paragraphs for the substance turn too many? I need to think about whether paragraphs B and C could be one paragraph (the distinction between substances and tools with respect to latent properties) or whether they need to be separate.
I think they could go either way. If they're one paragraph, the structure is: "Substances can have latent properties. [Examples.] Tools can't, because [two reasons]." That's a single paragraph making one point with two sides. If they're two paragraphs, each side gets more development. Given that Nick doesn't want me trimming, I'll suggest two paragraphs, but flag that they could be one if the essay's rhythm calls for it.
Actually, let me reconsider. The substance turn is the essay's most important moment — it's where the new idea arrives. It should be given room to breathe. Four paragraphs for the essay's central claim seems appropriate, not excessive. The reader needs to understand: (1) what the alternative is, (2) what substances are like (latent properties), (3) what tools are like (no latent properties), (4) that LLMs fit the substance pattern. Each of those is a distinct thought.
MOVE 10: Brief engagement with Hoel's oil analogy
Nick said: engage but don't go crazy. So this is a paragraph, maybe even just a few sentences within a paragraph, that notes that Hoel's own oil comparison actually supports the substance thesis. When oil was first discovered, we didn't know all its capacities. Its latent properties were drawn out over time. The same may be true of LLMs. Not a predictive claim — an observation about how substances work.
This could be a standalone paragraph or it could be folded into the paragraph about what substances are like (paragraph B above). If it's folded in, it's an example alongside iron. If it's standalone, it gets more emphasis as a direct reply to Hoel. Given that it's engaging with Hoel's specific analogy, I think it deserves its own moment — but a contained one.
MOVE 11: What this substance is like — phenomenology of engagement
Now the essay needs to shift from category-level argument to experience-level description. The reader has been told that LLMs are substances with latent properties. Now they need to know what it's like to work with this substance. What does the substance actually do?
This is where some of the current draft's material can be reused. The pen comparison (lines 81-86) is strong and works under the substance framing: a pen is a paradigmatic tool — transparent, no hidden capacities, extends inscription without pushing back. An LLM is different — you send language in, something comes back shaped by the substance's own tendencies.
The recursive return — "we are not only prompting the system; the return prompts us back" — is strong and doesn't depend on "medium." It's a description of what the substance does, a capacity of the material.
The artefact/working-surface distinction (lines 65-78) answers Hoel's evidence and doesn't depend on "medium" either. It observes that much interesting LLM engagement is not aimed at producing a finished text but at exploring, testing, discovering — working with the substance.
How should these be organized? In the current draft they're spread across several paragraphs with medium and Frippertronics in between. In the recalibrated essay, they could be reorganized as:
Paragraph E: The pen comparison. A pen extends inscription; it doesn't push back. An LLM does. What comes back is language shaped by the substance's own tendencies — sometimes flatter, sometimes unexpectedly connective, sometimes productively wrong. You have to decide what to reject, what to sharpen, what pressure to reapply. The return prompts you back.
Paragraph F: Artefact vs. working surface. What this means is that "writing" — the heading under which Hoel measures LLMs — covers very different activities. A finished essay, a published post — these are texts as artefacts. But when someone throws a half-formed idea at an LLM and uses the response to sharpen what they actually think, the text is not the product — it's the working surface. Hoel's evidence largely concerns artefacts. Much of the interesting engagement concerns the working surface.
Paragraph G: What working with the substance reveals. The distinction matters because it suggests that Hoel may be measuring LLMs by the wrong outputs. Or rather: by outputs at all, when much of what this substance enables is not output-directed but exploratory. The point is not to produce a better text; it's to discover what you think by engaging with a material that pushes back on your thinking. That is a very different activity from using a tool to accomplish a task, and it cannot be evaluated by asking whether the artefacts have improved.
Hmm, is paragraph G too close to paragraph F? Maybe they're really one thought: artefact vs. working surface, and what that means for evaluation. Let me reconsider...
Actually, I think the current draft handles this well in two paragraphs (lines 65-78). The first paragraph (lines 65-70) establishes the distinction. The second (lines 73-78) draws out its implications for Hoel. That's clean. Two paragraphs, one distinction, one implication. I should keep roughly that structure.
MOVE 12: Slop as characteristic tendency (Lucier)
This is lines 115-120 of the current draft, and it works with the substance framing. The substance has characteristic tendencies — smoothing, genericising, finding the most probable path. When those tendencies are allowed to dominate without resistance, the result is slop. Lucier's room illustrates this: the room's resonant frequencies take over when speech is passed through without intervention. Same logic for LLMs.
The current draft's paragraph on this is good. Under the substance framing, line 120 changes — instead of "what happens when a medium with strong generic tendencies is used with too little resistance from the person inside the loop," it becomes something like "what happens when a substance with strong generic tendencies is engaged without adequate resistance — when its latent capacities are not being explored but its default tendencies are being passively consumed."
Actually, I need to think about this more carefully. In the substance framing, slop isn't just about lazy use — it's about what happens when you treat a substance as a tool. If you treat an LLM as a tool — give it a task, accept what it returns, publish — you get slop, because you're not engaging with the substance's properties, you're just running its default tendencies. Slop is what happens when a substance is used as though it were a tool. That's actually a nice way of connecting the substance thesis back to Hoel: the slop he observes is evidence FOR the substance thesis, because it shows what happens when substance-engagement is replaced by tool-use.
That's a genuinely interesting thought and it should be in the essay. Let me note it.
MOVE 13: Return to Hoel — what he gets right and what the substance thesis says about his argument
The essay needs to come back to Hoel and state clearly what it's claiming and what it's not. The current draft's conclusion (lines 123-129) does some of this but it ends on "medium." Under the substance framing, the conclusion would be different.
Hoel is right that there has been no textual singularity. He's right that public prose has not improved. He's right that slop is a real phenomenon. What I'm arguing is not that he's wrong about any of this, but that his evaluative framework presupposes something it shouldn't: that we already know what kind of thing an LLM is and therefore know how to assess it.
If LLMs are tools, then Hoel's method works — you test a tool by asking whether it does its job, and if writing hasn't improved, the tool isn't very good.
If LLMs are substances with latent properties, then Hoel's method is premature — you can't evaluate a substance by its current prominent uses when you're still discovering what it can do. You can't say "bits in, bits out" about something whose capacities you haven't fully mapped.
The conclusion should be careful — not predictive, not boosterish, just methodological. The claim is about what can and can't be concluded yet, not about what will happen next.
And the last line of the current draft — "They may be a quite different type of tool, or not quite a type of tool at all — and that is worth taking seriously before we conclude that the current technology is merely another stone in Homo faber's long line of rocks" — might need revision. Under the substance framing, the essay has committed to "not a tool at all" more firmly than the hedged "a quite different type of tool, or not quite a type of tool at all." But maybe the hedge is still right — the essay is proposing a category, not proving it beyond doubt.
Actually, I think the conclusion might want to be more positive now. The current draft's conclusion is essentially negative: "I'm not sure 'tool' is right, and Hoel's test might not be measuring the right thing." Under the substance framing, the conclusion can be more determinate: "I'm proposing that these are substances, not tools, and that the difference matters for how we evaluate them. Hoel's 'bits in, bits out' may describe what happens when a substance is used as a tool. It may not describe what the substance is."
Let me also think about what goes between the first verdict (Move 6) and the substance turn (Move 9). I said earlier: evaluation problem → Hoel's writing argument. But I want to think about whether Hoel's writing argument should come before or after the substance turn.
Before: The reader sees Hoel's full case, then gets the substance alternative. The turn is more dramatic because the reader has been sitting with Hoel's argument and feeling its force.
After: The reader gets the substance category, then sees how it reframes Hoel's writing argument. The reframing is more powerful because the reader already has the tools to see it differently.
Hmm. I think before is better for this essay, because the essay is structured as a response. You want to show what you're responding to, then respond. But there's a question about how much of Hoel's writing argument to state before the turn vs. after.
Actually, looking at the current draft again, there's a lot of material about Hoel's writing argument (lines 56-78) — two dedicated paragraphs on writing, two on artefact/working surface, plus a paragraph connecting them (lines 90-95). That's five paragraphs engaging with Hoel's writing argument. If all of that goes before the substance turn, the turn is very delayed.
Alternative: state Hoel's writing argument briefly before the substance turn, then return to it in detail after. The brief statement gives the reader the target; the substance turn gives them new eyes; then the detailed engagement with writing, artefact/surface, etc. happens with substance-informed perspective.
I think this is actually the right approach. Here's why: the artefact/working-surface distinction is more powerful when the reader already has the substance category. "Text as working surface" is what it looks like to explore a substance's properties through language. "Text as artefact" is what you get when you treat the substance as a production tool. That distinction means more when "substance" and "tool" are already in play.
So the sequence becomes:
1. Opening + Hoel quote
2. Transition — interpretive gloss, pivot to "what sort of tool?"
3. "We use them" — concession and motivation
4. Proper function — what it is, why LLMs fail the test
5. Candidates that fail — token prediction, chatting, helping
6. First verdict — classificatory hesitation
7. Evaluation problem — uncertainty about function → uncertainty about testing
8. Hoel's writing argument briefly stated — writing is the natural proving ground, here's his evidence, here's his conclusion
9. THE SUBSTANCE TURN — not a tool, a substance. What substances are. Latent properties. Why tools are different.
10. LLMs as substance — specific examples of latent properties discovered through use
11. Brief engagement with Hoel's oil — his own analogy supports the substance thesis
12. Return to writing with substance-informed eyes — pen comparison (tool vs. substance in practice)
13. Artefact vs. working surface — now reframed as the difference between treating the substance as a production tool and exploring its capacities
14. What this means for Hoel's evidence — he measures artefacts; much interesting engagement concerns the working surface
15. Slop as characteristic tendency — Lucier — what happens when substance-tendencies dominate without resistance; slop is what you get when a substance is used as a tool
16. Conclusion — Hoel's evaluative confidence presupposes knowing what this thing is; the substance thesis says we don't yet know; "bits in, bits out" may describe the early phase, not the nature of the thing
That's sixteen paragraphs after the opening. Some of those might be half-paragraphs (the oil engagement, the evaluation problem), some might be long (the substance turn, the LLM-as-substance examples). The point is that each one does necessary work and none is there for ornament.
Let me now go through each of these and think about what specifically each paragraph needs to say, what claims it makes, what evidence it uses, and how it connects to the next.
Actually, wait. I should also think about what existing material from the current draft gets reused and what needs to be newly written.
From the current draft, roughly:
- Lines 13-22: opening + Hoel quotes + transition → REUSE (Moves 1-2)
- Line 24: "we use them" → REUSE (Move 3)
- Lines 26-45: proper function, candidates, first verdict → REUSE (Moves 4-6)
- Lines 48-53: evaluation problem → REUSE (Move 7)
- Lines 56-62: Hoel's writing argument → REUSE but maybe reorganized (Move 8)
- Lines 65-78: artefact/working surface → REUSE but repositioned after substance turn (Moves 13-14)
- Lines 81-86: pen comparison → REUSE but repositioned (Move 12)
- Lines 90-95: Hoel's evidence at wrong level → REUSE, repositioned (Move 14)
- Lines 98-104: medium → CUT
- Lines 107-113: Frippertronics → CUT
- Lines 115-120: slop/Lucier → REUSE with substance framing (Move 15)
- Lines 123-129: conclusion → REWRITE with substance framing (Move 16)
NEW material needed:
- Move 9: The substance turn (entirely new)
- Move 10: LLMs as substance — examples of latent properties (entirely new)
- Move 11: Brief oil engagement (new, but drawing on Hoel's existing quote)
- Move 15: additional thought about slop-as-tool-use (new framing of existing material)
- Move 16: revised conclusion (rewrite)
So roughly: the first eight moves reuse existing material with minor adjustments. Moves 9-11 are new. Moves 12-14 reuse existing material repositioned and slightly reframed. Moves 15-16 reuse/rewrite existing material with substance framing.
That's a good balance — the essay isn't starting from scratch, it's being recalibrated around a new center of gravity while preserving what was already working.
Now let me write out the detailed structure Nick asked for.
For each move, I should specify:
- The argumentative job of the paragraph
- The specific claims being made
- What evidence or examples are used
- How it connects to what comes before and after
- Where existing draft material can be reused
- What needs to be newly written
Let me do this carefully.
## Detailed essay structure: the substance recalibration
### Fixed opening
The opening paragraph and first Hoel block quote stay as they are. The paragraph establishes that "LLMs are tools" sounds trivially true and is probably the default position. The block quote gives Hoel's framing: if LLMs are tools, expect efficiency and scale; if they're a source of intelligence, expect transformation — like discovering oil.
The reader leaves this knowing what the essay is responding to and what position is on the table.
### Paragraph 1: Interpretive gloss and the pivot
Argumentative job: make sure the reader has Hoel's logic clearly, then create the first point of pressure.
What it says: Hoel's argument runs like this — if LLMs are merely the latest in humanity's line of tools, we should expect efficiency and scale, not transformation. If they were a genuine source of surplus intelligence, that surplus should have shown up first in writing, because writing is the domain closest to what these systems are made of. It hasn't. Therefore: tools, not intelligence.
Then the turn, which is already in the draft: "What I am less sure of is the description doing the philosophical work. 'Tool' sounds right, until one asks a little more carefully what sort of tool this is supposed to be."
This paragraph sets up the investigation that follows. The reader is primed to ask: OK, what sort of tool?
Connection to next: opens the proper-function investigation.
Existing material: lines 17-22 of the current draft. The second Hoel quote (lines 21, "stilted emails... dearth of good writing") could go here as evidence for Hoel's position, or it could be held back for the paragraph where Hoel's writing argument is stated at full strength. I'd suggest holding it — the first block quote is enough for now, and the slop evidence lands harder when it arrives later in the essay, after the reader has already started to question the "tool" framing.
### Paragraph 2: "We use them"
Argumentative job: concede the intuitive pull of the tool category before pressuring it.
What it says: Another reason it seems right to say LLMs are tools is that it seems right to say we use them. We open a chat interface, type a prompt, receive a response, act on it. That looks like tool-use. But what does that amount to? Using something does not settle what it is. We use iron; we use oil; we use clay. Using a substance is different from using a tool, and the difference matters — though it is not yet clear how.
Connection to next: motivates the proper-function question. If we use them, and they seem like tools, then we should be able to say what kind of tool they are. Can we?
Existing material: line 24 of the current draft. The "we use iron, we use oil" expansion is new and foreshadows the substance thesis without announcing it.
### Paragraph 3: Proper function — what it is
Argumentative job: establish the concept of proper function and show that ordinary tools pass the test.
What it says: One reason ordinary tools are easy to classify is that they tend to have a more or less specifiable proper function — not just anything an object can be used for, but the use that distinguishes what it is for from the uses to which it can merely be put. A hammer can hold down loose papers, but that's an accidental function, not a proper one. Nor does the point collapse with multi-functional objects: a Swiss Army knife has several proper functions, not none. The familiar examples still work — a hammer is for hammering, a vacuum cleaner is for vacuuming, Google is for searching.
Connection to next: now apply the test to LLMs.
Existing material: lines 26-31 of the current draft — usable as-is.
### Paragraph 4: Proper function — the candidates that fail
Argumentative job: apply the proper-function test to LLMs and show that none of the obvious answers work.
What it says: Once we ask that question of ChatGPT, the obvious answers all seem partial, unstable, or wrong in different ways.
- "Predict the next token" gives a mechanism, not a use — nobody opens ChatGPT to predict tokens, any more than we describe the function of the heart as contracting rhythmically.
- "Chat" fits Frankish's intentional-stance reading but not code generation, translation, or philosophical use.
- Hoel's suggestion — writing — catches something real, but not enough to settle what sort of thing we're dealing with.
- "Help with tasks" is so general as to individuate nothing — it's as though one described the function of a hammer as "helping with projects."
The difficulty is not that these answers are all wrong, but that none of them settles the question the way "for hammering" settles it for a hammer.
Connection to next: so what does this failure tell us?
Existing material: lines 33-38 of the current draft — usable as-is.
### Paragraph 5: First verdict
Argumentative job: state clearly what the proper-function failure shows and — just as importantly — what it does not show.
What it says: That does not show that LLMs are not tools; but it does suggest that they are a quite different type of tool, or not quite a type of tool at all. This is classificatory hesitation, not metaphysical victory. The claim is that the tool description is less straightforward than it first sounds. The issue is not just that LLMs can do many things — so can Swiss Army knives. The issue is instability at the level of ordinary functional description: none of the obvious answers to "what is it for?" quite works, and the answers that come closest tend to describe mechanism or to be so general as to describe everything.
Connection to next: and that classificatory instability has evaluative consequences.
Existing material: lines 41-45 of the current draft — usable as-is.
### Paragraph 6: The evaluation problem
Argumentative job: connect the proper-function failure to the question of how we assess LLMs — and therefore to Hoel's argument.
What it says: Uncertainty about proper function quickly becomes uncertainty about evaluation. If we don't know clearly what a thing is for, it is much harder to say in advance what would count as a good test of it. We test a hammer by asking whether it drives nails; we test a search engine by asking whether it returns relevant results; in each case the proper function gives us the standard. When there is no settled proper function, test-selection becomes a substantive question rather than something we can take for granted. This is not to say no test is possible — that would be absurd. It is to say that when Hoel selects writing as the privileged proving ground, that choice requires more argument than his framing initially suggests.
Connection to next: but there IS a straightforward reason to start with writing.
Existing material: lines 48-53 of the current draft — usable as-is.
### Paragraph 7: Hoel's writing argument stated at full strength
Argumentative job: give Hoel his strongest case. Don't strawman, don't undercut, just present. The reader should feel the force of his position before the essay offers an alternative.
What it says: There is a straightforward reason to start with writing. These systems operate through language, are trained on language, and produce language; so it is not arbitrary to think that writing is where their character should show first and most clearly. Hoel puts the point vividly:
> "for an LLM, words are its womb, its mother, its literal atoms — yet their impact on writing as a whole has been mostly to generate mountains of slop"
He is right about the constitutive connection: if there is surplus intelligence in these systems, text is the place it ought to appear. And what has appeared, by and large, is efficiency, scale, and slop — not transformation.
The question is not why one would look at writing — that question has a good answer. The question is whether writing, measured as Hoel measures it, is sufficient to settle what kind of thing an LLM is.
Connection to next: the substance turn — a different way of understanding what this thing might be, which reframes the writing question.
Existing material: lines 56-62 of the current draft — mostly usable. The transition sentence at the end ("The question is whether writing...") needs to be written to point toward the substance turn rather than toward artefact/working-surface directly.
### Paragraph 8: The substance turn — the category shift
Argumentative job: introduce the alternative. This is the essay's decisive moment.
What it says: The reason the proper-function question keeps failing may not be that LLMs are a strange kind of tool. It may be that they are not best understood as tools at all. I want to suggest a different category: substance.
A substance is stuff with properties and capacities. Iron is a substance; oil is a substance; clay is a substance. Substances are not ordinarily understood in terms of proper function. We do not ask what iron is for — we ask what iron can do, what its properties are, what can be made from it. And one of the most important features of substances is that they can have latent properties — properties that are present in the stuff but not yet discovered, not yet drawn out, not yet invoked.
When we first encountered iron, we did not know it could be magnetized. When we first encountered oil, we did not know it could power an internal combustion engine. Those properties were there, in the substance. They were drawn out over time, through exploration, experiment, and practice. That is just what substances are like: their nature can exceed what we currently know about them. There is always, in principle, more to find.
Connection to next: and this is exactly where tools are different.
Existing material: entirely new. This is the paragraph that the current draft doesn't have.
### Paragraph 9: Why tools are different
Argumentative job: complete the substance/tool distinction by showing why tools can't have latent properties.
What it says: Tools are not like this, and the reason is instructive. There are no natural tools. Every tool is an artefact — designed and manufactured for a purpose. And it does not make sense to say that designed, manufactured objects have hidden properties. The properties of a hammer are determined by its design: the weight, the shape, the hardness of the head, the length of the handle. You can use a hammer in unexpected ways, but those are the same manifest properties doing manifest things in a different setting. There is nothing hidden in a hammer waiting to be discovered.
This is the difference that matters. Substances can have latent properties — properties that are there but not yet found. Tools can't — their properties are determined by design and fully apparent. The question is which of these descriptions fits LLMs.
Connection to next: apply the distinction to LLMs.
Existing material: new. Nick's two reasons stated plainly.
### Paragraph 10: LLMs as substance — the evidence
Argumentative job: show that LLMs have the characteristic mark of substances — discoverable but initially unknown properties.
What it says: LLMs look more like substances than tools in this respect. Since their development, a series of capacities has been discovered through use and experimentation that were not part of the design specification:
- Chain-of-thought prompting: nobody designed these models to reason better when asked to think step by step. That capacity was latent in the model's learned patterns and was drawn out through experimentation — the way you discover that a substance has a property you hadn't anticipated.
- In-context learning: the ability to learn from examples provided in a prompt, adapting behavior on the basis of a few demonstrations, was not a design goal but an observed capacity.
- The effectiveness of persona prompts, structured formatting, iterative refinement — each of these represents a property of the substance that was found through exploration, not specified through design.
These are not instructions to a tool about what to do. They are discoveries about what a material can do — properties drawn out through practice, in the way that the semiconducting properties of silicon or the strength-to-weight ratio of carbon fiber were drawn out through investigation.
The design specified next-token prediction. What emerged has capacities that exceed that specification in ways that are still being mapped. That is the profile of a substance, not a tool.
Connection to next: and Hoel's own analogy actually supports this.
Existing material: new, though drawing on the factual record of LLM capability discovery.
### Paragraph 11: Brief engagement with Hoel's oil
Argumentative job: turn Hoel's oil comparison so that it supports the substance thesis rather than undermining LLMs.
What it says: Hoel compares the discovery of LLMs to the discovery of oil: if LLMs were a genuine source of intelligence, discovering them should have been transformative in the way that discovering oil was transformative. But this comparison may support the substance thesis more than Hoel's own position. When oil was first discovered, its full range of capacities was not known. For decades, its most prominent use was as lamp fuel. The deeper transformations — internal combustion, petrochemicals, plastics — came later, as more of the substance's properties were drawn out and stabilized in practice. The fact that early use did not reveal the full picture is not an argument against the substance. It is characteristic of how substances work.
This is not a prediction that hidden greatness will inevitably unfold. It is an observation about the logic of Hoel's own analogy: if LLMs are a substance, then measuring their impact by current prominent uses is exactly what one would expect to be premature.
Connection to next: so what does this substance actually do when you work with it?
Existing material: new, but engaging with the Hoel quote already in the fixed opening.
### Paragraph 12: The pen comparison — tool-use vs. substance-engagement
Argumentative job: make the substance/tool distinction concrete by comparing LLM engagement to pen use.
What it says: The difference between using a tool and engaging with a substance becomes clearer when we compare an LLM to a pen. A pen makes marks on a surface; it extends my ability to inscribe; it does not return a proposal, a misreading, or a line of continuation I had not seen. A pen does not prompt me back. Its properties are manifest and its function is transparent: I use it, it performs, and everything interesting about the result comes from me.
When I use an LLM, something different happens. I write something — or half-write something, or throw a distinction at the system in rough form — and what comes back is language already reshaped by the model's learned patterns: sometimes flatter than what I wanted, sometimes unexpectedly connective, sometimes productively irritating. I then have to decide what to reject, what to sharpen, and what pressure to reapply. We are not only prompting the system; the return prompts us back.
That is not what tool-use feels like. It is what engagement with a substance feels like — working with a material that has its own tendencies, its own characteristic ways of responding, its own properties that you learn through encounter rather than specification.
Connection to next: and this means "writing," the heading under which Hoel measures LLMs, covers very different activities.
Existing material: lines 81-86 of the current draft, with the final sentences reframed in substance terms rather than medium terms.
### Paragraph 13: Artefact vs. working surface
Argumentative job: distinguish two things that go on under the name "writing," to show that Hoel's evidence covers only one of them.
What it says: The difficulty is that "writing" is too coarse a heading for the very different things people do with text. A finished essay, a published book, a social media post — these are texts as artefacts, meant to stand on their own and be judged as products. But when someone throws a half-formed idea at an LLM and uses the response to sharpen what they actually think, the text is not the product — it is the working surface. What matters is not the quality of the output but what the encounter does to the next move of thought.
Hoel's evidence — the Amazon data, the social media slop, the unchanged quality of the best books — largely concerns text as artefact. That is a reasonable thing to measure, but it is not the only thing that goes on under the name of "writing."
Connection to next: implications for Hoel's argument.
Existing material: lines 65-70 of the current draft — usable with minor reframing.
### Paragraph 14: What Hoel measures and what he misses
Argumentative job: draw out the evaluative consequence of the artefact/working-surface distinction.
What it says: What Hoel mostly measures is text as artefact; the practice I want to describe treats text as a working surface. A finished blog post or book is meant to be read and evaluated on its own merits. A prompt-response-revision loop may not produce anything worth publishing — and that is not a failure, because publishing was not the point. The point was to test a distinction, or to discover what you actually think by responding to what the system sends back. "Has writing improved?" presupposes that the relevant success condition is improvement in end-products. If much of the interesting practice is not aimed at end-products in the first place, then the question is too blunt to capture what is going on.
And this connects back to the substance thesis: when you work with a substance, much of what you do is exploratory — learning what the material can do, testing its responses, discovering its properties through engagement. That kind of activity does not produce artefacts to measure. It produces understanding of the material. And that understanding is still accumulating.
Connection to next: but Hoel's observation about slop is real and important.
Existing material: lines 73-78 of the current draft, plus new connecting material back to the substance thesis.
### Paragraph 15: Slop as characteristic tendency — Lucier
Argumentative job: concede slop fully and then show that the substance thesis explains it rather than being refuted by it.
What it says: Hoel is right to insist on slop, and any account that tried to wave it away would deserve to fail. One of the most useful things in his essay is the sense that the public effect of these systems has often been a thinning of language rather than an enrichment of it.
There is an analogue worth considering. In Alvin Lucier's I Am Sitting in a Room, a spoken passage is played back into a room and re-recorded again and again, so that the room's own resonant frequencies gradually take over — the content is not developed or deepened but washed into the characteristic signature of the system. Something similar happens with LLM use when there is too little resistance: text passed through the model without discrimination gets drawn toward the model's easier habits — familiar transitions, familiar emphases, familiar shapes of explanation.
Slop is what happens when the substance's characteristic tendencies are allowed to dominate the exchange — when it is used passively, at scale, without the kind of resistance and exploration that draws out its latent capacities. In other words, slop may be precisely what happens when a substance is treated as a tool: give it a task, accept what it returns, publish. That produces mountains of generic text, because the substance's default tendencies — smoothing, averaging, genericising — are running unopposed. It does not tell us that the substance has nothing more to offer. It tells us what the substance does under one very specific set of conditions.
Connection to next: conclusion.
Existing material: lines 115-120 of the current draft, significantly expanded with the substance-as-tool-use framing and the argument that slop is evidence for the substance thesis rather than against it.
### Paragraph 16: Conclusion — return to Hoel
Argumentative job: state the essay's final claim clearly and with appropriate epistemic caution.
What it says: Hoel asks whether writing has improved, and his answer is no. From this he concludes that LLMs are tools: bits in, bits out. I think he asks the right empirical question but draws from it a conclusion that depends on a prior commitment he has not defended — the commitment that we already know what kind of thing an LLM is.
If LLMs are tools, his argument works. Tools have functions; their properties are determined by design; you test them by asking whether they perform. Writing hasn't improved; the tool isn't very good; bits in, bits out. That is a clean inference from a clean premise.
But if LLMs are more like substances — stuff with latent properties, capacities still being discovered, tendencies that manifest differently under different conditions of use — then the inference doesn't follow. You cannot conclude that a substance has revealed its nature when you are still in the early phases of learning what it can do. You cannot say "bits in, bits out" about something whose capacities you haven't finished mapping.
I am not predicting that the hidden greatness of LLMs will inevitably unfold. I am not offering a promissory note on behalf of the technology industry. I am making a narrower claim: that Hoel's evaluative confidence requires knowing what this thing is, and we may not yet know that. We may have synthesized a new substance whose latent properties are still being drawn out — and if so, measuring it by its current most visible outputs tells us something, but not as much as Hoel thinks.
Connection: this is the end. The final line should leave the reader with the substance thought, not with a hedge.
Possible closing line (to be workshopped): something that returns to the opening's near-trivial sentence — "LLMs are a type of tool" — and says, in effect, that this sentence may not be trivially true at all; it may be the thing that needs to be rethought.
Existing material: lines 123-129 of the current draft are the skeleton, but the conclusion needs to be substantially rewritten to land on the substance thesis rather than the medium thesis. The "quite different type of tool, or not quite a type of tool at all" line can stay as an echo but the conclusion should go further.
---
## Summary of the whole shape
1. Opening + Hoel quote (fixed)
2. Interpretive gloss + pivot to "what sort of tool?"
3. "We use them" — concession, with foreshadowing
4. Proper function — what it is, why ordinary tools pass
5. Candidates that fail — token prediction, chatting, helping
6. First verdict — classificatory hesitation, not metaphysical victory
7. Evaluation problem — uncertain function → uncertain testing
8. Hoel's writing argument at full strength — writing as proving ground, slop as evidence, "bits in, bits out" as conclusion
9. THE SUBSTANCE TURN — the category shift. Substances have latent properties. Iron, oil, clay.
10. WHY TOOLS ARE DIFFERENT — no natural tools; designed objects don't have hidden properties
11. LLMs AS SUBSTANCE — the evidence. Chain-of-thought, in-context learning, iterative refinement as discoveries about a material, not instructions to a tool
12. HOEL'S OIL BRIEFLY — his own analogy supports the substance thesis. Engage, don't go crazy.
13. PEN COMPARISON — tool-use (pen) vs. substance-engagement (LLM). "A pen does not prompt me back."
14. ARTEFACT VS. WORKING SURFACE — "writing" covers different activities; Hoel measures artefacts
15. WHAT HOEL MEASURES AND WHAT HE MISSES — the working-surface practice is exploratory, substance-like, not output-directed
16. SLOP AS CHARACTERISTIC TENDENCY — Lucier. Slop is what happens when the substance is treated as a tool. Evidence for the thesis, not against it.
17. CONCLUSION — Hoel's inference depends on knowing what this thing is. The substance thesis says we may not yet know. "Bits in, bits out" may describe the early phase, not the nature of the thing.
Seventeen paragraphs. Each one does work the argument needs. The substance turn arrives at paragraph 9, which is roughly the middle of the essay — the reader has seen the problem clearly before the positive thesis arrives, and the positive thesis has eight paragraphs of development after it. The existing draft material is largely preserved and repositioned; the new material (paragraphs 9-12) is concentrated in the substance section.
What's cut from the current draft: medium (paragraph 12 of the current draft, lines 98-104), Frippertronics (paragraph 13, lines 107-113), and the editorial notes (lines 131-148). Everything else is reused with varying degrees of reframing.
What's new: the substance turn (paragraphs 9-10), LLMs-as-substance examples (paragraph 11), brief oil engagement (paragraph 12), the "slop as substance-treated-as-tool" framing (expansion of paragraph 16), and the rewritten conclusion (paragraph 17).
*Diciassette paragrafi, ciascuno con il proprio compito — nessuno sacrificato alla vanità della brevità.*