# Janus' Simulators

## Metadata
- Author: [[Scott Alexander]]
- Full Title: Janus' Simulators
- Category: #articles
- Summary: insert summary
- My notes:
- Summary: GPT is not a goal-driven agent but a simulator that predicts and mimics text patterns. It learns to act like helpful characters through human feedback but does not truly have goals or understanding. This view helps us see that future AI risks depend on whether AI simulates agents with goals, not on pure simulators like GPT.
- URL: https://www.astralcodexten.com/p/janus-simulators
## LLM Chats
## NotebookLM
## LLM Audio
## Highlights
> Elide all the differences, and the story is something like: babies are born as [pure predictive processors](https://slatestarcodex.com/2017/09/05/book-review-surfing-uncertainty/), trying to make sense of the buzzing blooming confusion of the world. But as their parents reward and punish them, they get twisted into some specific shape to better capture the reward and avoid the punishment. The mask usually looks like “having coherent beliefs, taking coherent actions, pleasing others, maintaining a high opinion of one’s self”. After maintaining this mask long enough, people identify with the mask and forget that they’re anything else. ([View Highlight](https://read.readwise.io/read/01kbwvspvtphyfdpa8mcc7gx7w))
> I propose a friendly amendment: they’re noticing that most of what they are - the vast majority of their brain - is a giant predictive model of the universe. This model is big enough that they have lived inside it their entire life, with only slight edits from lossy sensory information that help fit it to the real universe. I’ve written about this before in the context of lucid dreaming - a dreamer safe in bed can apparently wander their neighborhood, seeing each tree and car and dog in detail approximately equivalent to waking experience. No astral projection is involved - they’re wandering around their internal world-model, which contains 99% of the relevant information, with real sensory information filling in the missing 1%. Once you stop obsessing over the character you’re playing, you notice the GIANT SUPER-ACCURATE WORLD MODEL TAKING UP 99.99% OF YOUR BRAIN and you think “Huh, I guess I’m the Universe. Weird.” ([View Highlight](https://read.readwise.io/read/01kbwvvek02jemsshmxws61xjv))
> Janus dubs it a **simulator**. Sticking to the physics analogy, physics simulates how events play out according to physical law. GPT simulates how texts play out according to the rules and genres of language. ([View Highlight](https://read.readwise.io/read/01kbwvfj9d3ecy1c360emcg0nv))
> GPT doesn’t really like me. And it’s not *lying*, saying it likes me when it really doesn’t. It’s simulating a character, deciding on the fly how the character would answer this question, and then answering it. If this were Character.AI and it was simulating Darth Vader, it would answer “No, I will destroy you with the power of the Dark Side!” Darth Vader and the-character-who-likes-me-here are two different masks of GPT-3. ([View Highlight](https://read.readwise.io/read/01kbwvgevbxv6r04phq65qtdxw))
> If you reward ChatGPT for saying it’s a machine learning model, it will say it’s a machine learning model. If you reward it for saying it’s Darth Vader, it will say it’s Darth Vader. The only difference is that in the second case, you’ll understand it’s making things up. But in the first case, you might accidentally believe that it *knows* it’s a machine learning model, in the “justified true belief” sense of knowledge. ([View Highlight](https://read.readwise.io/read/01kbwvpqp3zgm1k52dnf7wpbx3))