Neuromorphic
Do LLMs Need to Dream? Karpathy on Model Collapse and the Search for Entropy
Andrej Karpathy explains why language models cannot simply train on their own reflections: model samples are silently collapsed. Dreams, jokes and children show why learning systems need a steady supply of entropy.
Akmal Alif · 8 October 2026 MYT

In a short clip from his October 2025 conversation with Dwarkesh Patel, Andrej Karpathy takes on a question that sounds whimsical but sits close to a hard research problem (Dwarkesh Clips, 2025). Patel asks what the machine-learning equivalent of daydreaming or sleep might be: not inventing new problems, just reflecting on what you already know (watch from 0:00). Karpathy, who taught an influential neural-network course at Stanford, led the vision team for Tesla’s Autopilot and worked at OpenAI, answers by describing what is missing from how language models learn, and why the obvious fix quietly fails.
The clip runs under five minutes, but it connects three ideas worth separating: reading as active synthesis, the “silent collapse” of model-generated data, and a hypothesis that dreams exist to keep brains from overfitting. The full episode covers much more ground (Patel, 2025). This piece stays with the clip, which is embedded below; the timestamps throughout jump to the matching moment.
Reading is not next-token prediction
When a language model reads a book during pre-training, Karpathy notes, the text is laid out as a long sequence and the model learns to predict the next token. Some knowledge is absorbed that way. People do not read like that (watch from 0:30). For Karpathy, a book is less a stream to be memorised than a set of prompts: it provokes questions, comparisons and arguments, the kind of thing that happens in your head or at a book club with friends. The knowledge comes from manipulating the material, not only from being exposed to it.
He would like to see something similar in training: a stage in which a model thinks through what it has read, reconciles it with what it already knows and spends time on that process (watch from 1:00). Nothing equivalent exists in today’s models. In his words, it is all research.
Why not train on the model’s own reflections?
Patel asks the natural follow-up: why not generate those reflections synthetically and train on them? Karpathy’s answer is the heart of the clip (watch from 1:30). Any single synthetic reflection may look excellent. Keep training on them, though, and the model gets worse.
The reason is that model samples are silently collapsed. Each one looks reasonable on its own, but together they occupy a tiny region of the space of possible thoughts about the material. Training on that narrow region pulls the model further into it.
This matches a broader finding. Shumailov and colleagues showed that when models are trained recursively on data produced by earlier models, the tails of the original distribution disappear and quality degrades over generations, a process they called model collapse (Shumailov et al., 2024). Rare but valid content is the first casualty.
The three-joke problem
Karpathy offers a test anyone can run: ask a chatbot to tell you a joke (watch from 2:00). It seems to know about three. That is more than an anecdote. When Jentzsch and Kersting asked ChatGPT for jokes over a thousand times, more than 90% of the 1,008 responses were the same 25 jokes (Jentzsch & Kersting, 2023).
The problem is not that the jokes are bad. It is that the distribution is narrow. Humans are noisier and often less polished, but in a statistical sense they are not silently collapsed: they keep a great deal of entropy. How to make synthetic data generation work despite collapse, while keeping that entropy, is in Karpathy’s view an open research problem.
Collapse lives in the distribution, not the sample
Patel restates the problem (watch from 2:30). Ask a model to think about the same chapter ten times and you get ten versions of the same reflection. That is why scaling “reflection” over a fixed amount of prompt information does not keep paying off.
It is a useful correction for anyone evaluating synthetic data. Reviewing outputs one at a time will not reveal collapse, because each output passes inspection. Diversity is a property of the whole set, such as how often the same moves, phrasings and conclusions repeat, so it has to be measured across the distribution.
Humans collapse too
Karpathy then turns the analogy around (watch from 3:00). He suspects there may be no fundamental solution, and that people also collapse over the course of their lives. Children have not overfitted yet; they say things that surprise adults precisely because those things are not what people usually say. Adults revisit the same thoughts, say more of the same things, and learn more slowly as time goes on.
The comparison between brains and models cuts both ways. This failure mode is not unique to machines; it may be a general property of systems that learn mostly from their own recent output.
Do dreams fight overfitting?
Patel raises a paper suggesting that dreaming evolved to counter exactly this kind of overfitting (watch from 3:30). The idea matches Erik Hoel’s overfitted brain hypothesis: the strangeness of dreams helps the brain generalise by exposing it to situations unlike its daily routine, much as machine-learning practitioners inject noise or randomly drop parts of a network’s input to stop it fitting its training data too closely (Hoel, 2021).
Karpathy finds the idea interesting. Generating thoughts in your head and attending to them, he observes, is a way of training on your own samples, your own synthetic data. Do it for too long and you go off the rails.
Seek entropy
His practical conclusion is simple: you have to seek entropy in your life (watch from 4:00). Talking to other people is a great source of it. Perhaps the brain also has internal mechanisms for adding entropy to the process, and perhaps dreams are one of them.
What it means for agents that learn from experience
The clip pairs naturally with Richard Sutton’s argument that the next advances will come from agents learning from their own experience rather than from more human text, which we discussed in Two Languages of Intelligence (Silver & Sutton, 2025). Karpathy’s warning adds a condition. Experience is only as useful as it is varied. An agent that keeps generating, replaying and learning from the same narrow slice of situations is training on its own samples, and the same collapse dynamics apply.
That gives EXEPERT’s research tools a concrete principle: treat the diversity of experience as something to measure, not assume. When we review agent runs in EXEPERT World, the useful question is not only whether an agent improved, but whether it met genuinely new situations, such as different starting states, environments and perspectives, or simply rehearsed familiar ones. In the terms of The Bitter Lesson for Embodied AI, scalable learning needs scalable variety.
None of this answers Patel’s original question. There is no established machine-learning analogue of dreaming. But the clip makes the problem precise: reflection is easy to imitate one sample at a time, and hard to make useful without a reliable source of entropy.
References
Dwarkesh Clips. (2025, October 29). Do LLMs need to dream? – Andrej Karpathy [Video]. YouTube. https://www.youtube.com/watch?v=TAvxmNf2G40
Hoel, E. (2021). The overfitted brain: Dreams evolved to assist generalization. Patterns, 2(5), Article 100244. https://doi.org/10.1016/j.patter.2021.100244
Jentzsch, S., & Kersting, K. (2023). ChatGPT is fun, but it is not funny! Humor is still challenging large language models. In Proceedings of the 13th Workshop on Computational Approaches to Subjectivity, Sentiment, & Social Media Analysis (pp. 325–340). Association for Computational Linguistics. https://aclanthology.org/2023.wassa-1.29
Patel, D. (Host). (2025, October 17). Andrej Karpathy — AGI is still a decade away [Audio podcast episode]. In Dwarkesh Podcast. https://www.dwarkesh.com/p/andrej-karpathy
Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631(8022), 755–759. https://doi.org/10.1038/s41586-024-07566-y
Silver, D., & Sutton, R. S. (2025). Welcome to the era of experience [Preprint of a chapter in Designing an Intelligence, MIT Press]. Google DeepMind. https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf