I suspect that for people who don't think that LLMs are a kind of mind, think they're a kind of tool. I don't think they're a kind of mind but I'm not sure they're a kind of tool either. Here I'm gonna suggest that they're more like a kind of stuff, a substance.
In a recent [essay](https://www.theintrinsicperspective.com/p/bits-in-bits-out) Erik Hoel argues for what we might call a *mere tool* position:
> If LLMs were a true source of intelligence to rival humans, then discovering them should be like discovering oil... If LLMs are tools, then we should expect their impacts to be a mirror of us, and concern efficiency and scale, rather than quality, and depend strongly on how people use them.
While Hoel is far from an ant-AI, his aim in this piece is to dampen the more **firey predictions/claims** about the capabilities of LLMs. If humanity were suddenly to discover *Oil 2* in abundance, **then a lot of line graphs would shoot straight up — productivity would rev, things would change, an epoch would shift.** What might an epoch shift in intelligence look like? Hoel points to Move 37 from AlphaGo's 2016 match against Lee Sedol, as an example of what we might expect %% can I really get away with saying that this is what Hoel is doing in this sentence? %%. At the moment it was made, Move 37 looked like an error; afterwards it was seen as a truly innovative move. And since then, the world of Go hasn't been quite the same: move 37 has had an impact on how the game is played [ one more descriptive sentence as to how it's unlocked the game. ]
If we take the arrival of surplus intellect to look something like Move 37 and its effects, then, Hoel says, it is clear that LLMs have not provided us with anything similar. If there is surplus intellect in these systems, Hoel argues, writing is where it should have appeared first: words are the LLM's native element — "its womb, its mother, its literal atoms," so if these systems can improve anything, writing is where the improvement should show. %% the authors opinion or my opinion?%% Six years after GPT-3, has writing improved? Looking at large-scale book data%%not how i write%%, Hoel finds that the best books remain what they were, the average book has gotten worse, and there has been no Move 37 moment — no creative breakthrough that could only have come from AI. For Hoel, this is exactly what tools would produce — and exactly what surplus intelligence would not. %%not how i write, and a lot of this paragraph is shit%%
Instead, he thinks we now have good reason to believe we have created a new type of *tool*. Tools, as he understands them, amplify human capacity leading to "efficiency and scale, rather than quality" and "depend strongly on how people use them." The arrival of the joseki dictionary provided Go players with exactly this kind of amplification — by compiling established corner sequences, it let players look up known patterns and study their variations systematically. It did not, however, reconfigure the game in the way that AlphaGo did. LLMs, on Hoel's view are closer to dictionary than AlphaGo
## The Substance
While I don't think that Hoel is wrong per se, his piece helped make some thoughts of I have had about LLMs and their tooliness. One reason it seems right to say that LLMs are tools is that we *use* them, and what else could something we use be, if not a tool?
But most tools have what a philosopher of technology would call a *proper function*. An object's proper function is simply the thing it is supposed to do –the reason why it was made and what you bought it to do: a hammer is for hammering, Google is for searching, etc. A lot of the time, you can use an object for a lot more things than its proper function –**a hammer can be used as a** doorstop, [this guy](https://spoon-tamago.com/tatsuo-horiuchi-excel-spreadsheet-artist/) uses Excel has been used to make art[^2]– but that doesn't take away from the fact that if you're a tool there is a thing, or things, that you're made to do.
So what's ChatGPT's proper function? I don't think there's an obvious answer. They're marketed
- I don't think this is very obvious. People use it for myriad things. LLMs meet needs for myriad things. So what is ChatGPT's proper function? Simply saying it helps or assists with tasks isn't to say very much at all –you can say it about any tool and it is going to be true. Next token prediction? No one is going to claude.ai to predict tokens. Chatting? You're going to go upstairs to get some chatting done? [^1]
- Another response here would be to say that there's nothing special about multifunctional objects. Claude does a zillion things but then so does an iPhone, the humble spork manages two. But there's a difference between multifunctional and *afunctional*. Sporks are made to do exactly two things; an iPhone, like a laptop, like a swiss army knife, has as many functions as those **apps that it houses,** add up to.
- LLMs are interestingly, amorphously afunctional: these systems are sold as *assistants*, or *intelligence*, but these are no more functions than chatting or token predicting. Nor were they designed with any specific use case in mind, any problem to solve – **the pursuit of artificial intelligence hasn't been and isn't a pursuit of something that will solve problem.**
-
-
- or things that help you, which is not the same as being sold as a tool for doing something in particular.
-
- .
- Unlike, say, the iPhone, these things were not designed with a particular function in mind. OpenAI's original mission wasn't to build something with a specific purpose, but rather to ["advance digital intelligence"](https://blog.gregbrockman.com/the-openai-mission). Neither are they marketed for particular purposes: these systems are sold as *assistants*, or *intelligence*, or things that help you, which is not the same as being sold as a tool for doing something in particular.[^3]
-
- another way you might want to push back here is by telling me that there's nothing special about multiple. Multifunctional objects. A different approach would be to say that none of this makes ChatGPT special. Couldn't the same sort of thing be said about phones, or laptops? Nah. Phones and computers are *multi-functional*, and they are multi-functional because they house and run applications which have their own functions[^1].
One thing you might say to all this is *So what?* (Or perhaps just *shut up*). But for one thing, it might make us doubt, at least a little bit more, whether Hoel is right to think that writing is a fair fight for LLMs. They weren't created with the aim of creating good writing.
- That pattern — building a thing and then discovering what it does — is familiar, but it is not how we relate to tools. It is how we relate to substances, to *stuff*.
- Oil was used as lamp fuel for a long time before anyone discovered you could run an engine on it or make plastics out of it, and the oil did not change in the meantime — the combustibility and chemical complexity were always there, waiting for practices to catch up.
- The properties came first and the uses were discovered afterwards, which is the opposite of how tools work.
- For a hammer, the purpose explains the properties: it has mass and hardness because those serve hammering.
- For a substance, the properties are just there, and what you can do with them is something you find out over time.
- ordinary tools are usually easy to classify is that they tend to have a more or less specifiable proper function.
- By this I do not mean just anything an object can be used for, but the use that distinguishes what it is *for* from the uses to which it can merely be put.
- A hammer can hold down loose papers, and a heavy book can prop open a door, but these are accidental functions rather than proper ones.
- Nor does the point collapse the moment we turn to multi-functional objects: a Swiss Army knife has several proper functions, not none.
- That is why the familiar examples still work — a hammer is for hammering, a vacuum cleaner is for vacuuming, Google is for searching.
- Once we ask that question of ChatGPT, the obvious answers all seem partial, unstable, or wrong in different ways.
- Perhaps its function is to predict the next token; but that is a description of mechanism, not of use — nobody opens ChatGPT in order to predict tokens, any more than we describe the function of the heart as contracting rhythmically.
- Perhaps it is for chatting; Keith Frankish has suggested that LLMs can be understood, from the intentional stance, as wanting to play the chat game — but that fits badly with code generation, translation, or philosophical use.
- "Helping with tasks" is so general as to individuate nothing — it is as though one described the function of a hammer as "helping with projects."
- The difficulty is not that these answers are all wrong, but that none of them settles the question the way "for hammering" settles it for a hammer.
- That does not show that LLMs are not tools; but it does suggest that they are a quite different type of tool, or not quite a type of tool at all.
- I am not claiming to have proved that LLMs cannot possibly be tools — the force of the previous step is classificatory hesitation, not metaphysical victory.
- I am claiming that the tool description is less straightforward than it first sounds.
- The issue is not just that LLMs can be used for many things — so can Swiss Army knives, and nobody finds that puzzling.
- The issue is instability at the level of ordinary functional description: none of the obvious answers to "what is it for?" quite works, and the answers that come closest tend to describe mechanism or to be so general as to describe everything.
- And that uncertainty about proper function quickly becomes uncertainty about evaluation.
- If we do not know clearly what a thing is for, it is much harder to say in advance what would count as a good test of it.
- We can test a hammer by asking whether it drives nails; we can test a search engine by asking whether it returns relevant results; in each case the proper function gives us the standard.
- When there is no settled proper function, test-selection becomes a substantive question rather than something we can take for granted.
- I am not saying that no test of LLMs is possible — that would be absurd.
- I am saying that when Hoel selects writing as the privileged proving ground, that choice requires more argument than his framing initially suggests.
- But there is a straightforward reason to start with writing.
- These systems operate through language, are trained on language, and produce language; so it is not at all arbitrary to think that writing is where their character should show first and most clearly.
- Hoel puts the point vividly:
- > "for an LLM, words are its womb, its mother, its literal atoms — yet their impact on writing as a whole has been mostly to generate mountains of slop"
- He is right about the constitutive connection: if there is surplus intelligence in these systems, text is the place it ought to appear.
- The question is not why one would look at writing — that question has a good answer.
- The question is what exactly we are measuring when we look there.
- The proper function problem may point to something deeper than a merely awkward classification.
- It is not just that we cannot easily say what ChatGPT is for; it is that the question itself may be the wrong kind of question.
- Tools are designed objects, and designed objects have their properties selected and shaped by what they are meant to do.
- A hammer has hardness, mass, and a striking surface because those properties serve hammering; the properties were reverse-engineered from the purpose.
- That is why tools do not surprise us with hidden capacities: what they can do is what they were made to do.
- Substances work differently.
- Iron has hardness, malleability, and conductivity — not because someone designed iron to have them, but because they are what iron is.
- Uses emerge from properties; the properties were not reverse-engineered from uses.
- Oil was not discovered with a list of applications attached; its capacities were latent, and practices emerged over time to elicit them.
- Silicon sat in sand for millennia before anyone found a way to make it carry computation.
- The relation between substance and use runs in the opposite direction from the relation between tool and function.
- Tools do not have discoverable-but-at-first-unknown properties, and the reason is revealing.
- They are not natural objects; they are artefacts.
- And it does not make sense to say that a manufactured object has hidden properties waiting to be found — the properties it has are the ones selected by the design brief.
- You can use a hammer in ways its designer did not anticipate, but you are not thereby discovering a latent capacity of the hammer; you are repurposing its manifest properties.
- The same hardness that drives nails can crack walnuts, but that is not a discovery about the hammer — it is a different application of what was always visible.
- What if, in building LLMs, we have synthesised something more like a substance than a tool?
- That would explain why the proper function question keeps failing: we cannot say what ChatGPT is *for* because 'for' is not the right relation to be looking for.
- The system has capacities — pattern-completion, contextual pliability, associative reach, responsiveness to iterative prompting — and uses are still emerging from those capacities.
- We do not yet know what can be elicited from this stuff, because we are still learning how to work with it.
- Hoel's oil analogy is meant to cut against the hype: if LLMs were a genuine discovery, they should have transformed things by now.
- But the analogy may cut in a different direction.
- When oil was first extracted, it was lamp fuel — useful, but hardly civilisation-altering.
- The internal combustion engine, plastics, petrochemicals: these came later, once practices had developed that could draw out what the substance could do.
- The point is not that transformation is guaranteed; it is that Hoel's argument assumes we already know what this stuff is for, and that assumption is exactly what the substance framing puts into question.
- I am not making a promissory argument, and I am not predicting that hidden greatness will inevitably unfold.
- There is a version of this thought that sounds like AI boosterism: just wait, the real applications are coming.
- That is not my claim.
- My claim is narrower: if LLMs are more like a substance than a tool, then we may be too early in the history of their use for Hoel's test to bear the weight he wants to put on it.
- That is an epistemic point, not a prediction.
- Hoel is right to insist on slop, and any account that tried to wave it away would deserve to fail.
- One of the most useful things in his essay is the sense that the public effect of these systems has often been a thinning of language rather than an enrichment of it.
- But if LLMs are a substance, slop is what we should expect when that substance is used without resistance or skill.
- Substances have characteristic tendencies: iron rusts, oil burns, this stuff smooths and genericises.
- Text passed through the model without discrimination gets drawn toward the model's easier habits — familiar transitions, familiar emphases, familiar shapes of explanation.
- That is not evidence against the substance framing; it is evidence about how this particular substance behaves when taken up carelessly or at scale.
- Remember that Hoel himself reaches for oil as his comparison — if LLMs were a genuine source of intelligence, he says, discovering them should be like discovering oil.
- He means this to cut against the hype: we discovered oil, line graphs shot up, an epoch shifted; nothing like that has happened with LLMs, so they must not be that kind of discovery.
- But oil is a substance, and the history of oil is not the history Hoel is telling.
- When oil was first extracted at scale in the 1850s, it was lamp fuel — kerosene for lamps, replacing whale oil, and that was more or less it.
- The internal combustion engine came later, plastics later still, petrochemicals later still; by the 1880s, kerosene lamps were actually in decline, being replaced by Edison's electric lights, and if you had evaluated oil at that moment by the quality of lamplight you would have concluded it was a spent resource.
- The uses that actually transformed civilisation had nothing to do with lamps.
- So if writing is the LLM's lamp fuel — the first and most obvious thing to try with something made of words — then 'has writing improved?' is a question about lamplight.
- I don't mean this as just a timing point, a way of saying 'give it a few more years and the good stuff will come'.
- The point is that if LLMs are more like a substance than a tool, then the form of the test is wrong, not just its moment.
- Tools have proper functions, and you can test a tool against its proper function — that is what proper functions are for.
- Substances don't work like that, and the question 'has writing improved?' only makes sense as a test of LLMs if you have already assumed that writing is what they are for.
- The kind of question you ask about a substance is different: not 'has it done the thing it was applied to?' but 'what practices have started to develop around it, and what do those practices look like?'
- Slop, on this picture, is not just bad news but information about the substance.
- Oil burns; that's a tendency that tells you something about what oil is, and understanding it is the first step toward doing anything useful with the stuff.
- This substance smooths, genericises, pattern-completes toward the middle of its training distribution — those are its tendencies when nothing pushes back.
- Working out how to push back, how to draw out the associative and contextual capacities while resisting the pull toward the generic, is a question about craft, and craft takes time to develop.
- The early history of any substance-craft relationship — the decades between 'we have this stuff' and 'we know how to work with it' — tends to look, from the outside, as though nothing much is happening.
- Most of the conversation about LLMs asks: are they intelligent, or are they just tools?
- Hoel is right that they are not intelligent in the Move 37 sense — there has been no moment where surplus intellect just showed up in the writing.
- But 'tool' is not the only other option, and if you start from the observation that LLMs don't have a proper function, that they were not built to do any particular thing, that their capacities were discovered after the fact and are still being discovered — then what you are looking at is not a tool with an unusually wide range of applications but a synthetic substance whose uses are still emerging.
- That is neither the hype position (these are minds, everything changes tomorrow) nor the deflationary one (these are hammers, everyone calm down).
- It is a third thing, and it suggests that the questions we should be asking are not about whether writing has improved or whether these systems can pass exams, but about what happens as people develop the craft of working with this material — and whether the practices that emerge turn out to be worth having.
[^1]: There are other ways to be multi-functional other than being a toolbox. Think of sporks.
[^3]: OpenAI's [own research into how people use ChatGPT](https://cdn.openai.com/pdf/a253471f-8260-40c6-a2cc-aa93fe9f142e/economic-research-chatgpt-usage-paper.pdf) sorts usage into categories so broad — "practical guidance", "seeking information", "writing" — that they barely tell you what the thing is actually for.