# LLM Generates the ENTIRE Output at Once ![rw-book-cover](https://i.ytimg.com/vi/X1rD3NhlIcE/maxresdefault.jpg) ## Metadata - Author: [[Matthew Berman]] - Full Title: LLM Generates the ENTIRE Output at Once - Category: #articles - Summary: insert summary - My notes: - Summary: A new large language model called a diffusion LLM generates entire responses at once, making it ten times faster and cheaper than traditional models. This innovative approach allows for quicker and more efficient reasoning, enhancing performance in tasks like coding. The model's ability to refine outputs iteratively means it can produce high-quality results in less time. - URL: https://youtu.be/X1rD3NhlIcE?si=ddhDiwNLy0gG41vb ## LLM Chats ## NotebookLM ## LLM Audio ## Highlights > listen to this: most of the image and video generation AI tools actually work this way. This way being diffusion, and use diffusion, not autoregression. It's only text and sometimes audio that have resisted. So, it's been a bit of a mystery to me and many others why, for some reason, text prefers autoregression, but images and videos prefer diffusion. This turns out to be a fairly deep rabbit hole that has to do with the distribution of information and noise, and our own perception of them in these domains. If you look close enough, a lot of interesting connections emerge. ([View Highlight](https://read.readwise.io/read/01jnpey2j126pkkcwkr4yjcq19))