EXEPERT Updates
Dwarkesh Patel Revisits Sutton: Imitation as a Prior, and the Gap Nobody Disputes
A follow-up to our reading of the Sutton–Dwarkesh interview. Patel's own reflection steelmans the Bitter Lesson case, argues imitation is short-horizon RL and a useful prior, calls the world-model debate semantic, and agrees continual learning is the real gap.
Akmal Alif · 9 October 2026 MYT

Earlier we read the Richard Sutton and Dwarkesh Patel interview (Patel, 2025a) as two vocabularies passing each other by. In our reading, prediction, goals, imitation and world models each meant one thing in Sutton's reinforcement-learning language and another in the language-model era. A week after the episode, Patel published a short follow-up that does what our post wished someone would: it translates (Patel, 2025b).
This update maps Patel's follow-up onto the sections of our original piece. Our post is linked here as a card:
Patel's video is embedded below. The timestamps jump to the matching moment. A written version is on the Dwarkesh Podcast site.
First, the steelman
Patel opens by saying they understand Sutton's view much better now than during the interview, and by stating it more strongly than they managed at the time (watch from 0:00).
Our post covered the irony that the Bitter Lesson is usually cited to defend scaling language models (Sutton, 2019). Patel's version sharpens it. The lesson is not "spend as much compute as possible". It is to use techniques that turn compute into capability most effectively and scalably (watch from 0:30). By that standard, current models look wasteful in three ways:
Deployment teaches them nothing. Most of the compute spent on a language model goes to running it, and it learns nothing while it runs.
Training is inefficient. It consumes the equivalent of tens of thousands of years of human experience (watch from 1:00).
Everything comes from human data. Pretraining text and the reinforcement-learning environments built for models are, in Patel's phrase, "human furnished playgrounds". Human data is an inelastic resource that does not scale (watch from 1:30).
Add that models learn what a person would say next, not how the world responds to action. Add that they cannot learn on the job. Sutton's conclusion follows: a new architecture with continual learning would make today's approach obsolete (watch from 2:00).
Patel's disagreement, in one line
Patel's summary is that Sutton's distinctions are real, but they are not either-or (watch from 2:42):
Imitation and reinforcement are continuous. Imitation learning is continuous with reinforcement learning and complementary to it.
Human data is a prior. A model of humans can be a prior that helps a system learn a true world model.
Continual learning may arrive incrementally. Some future form of fine-tuning at test time might deliver continual learning, the way in-context learning already gives a version of it.
Imitation: the section our post called "bootstrapping culture"
In the interview Patel argued that people learn much of their culture by copying, and Sutton replied that imitation serves goals and is refined by trial and error. The follow-up gives Patel's side three better analogies (watch from 3:22):
Fossil fuels. Ilya Sutskever has compared pretraining data to fossil fuels. Patel takes the analogy further: a finite resource can still be the necessary bridge, the way civilisation could not jump from water wheels to solar panels without coal in between (watch from 3:30).
AlphaGo and AlphaZero. AlphaGo learned first from human games and was superhuman (Silver et al., 2016). AlphaZero learned from scratch and was better (Silver et al., 2018). Patel accepts that a self-bootstrapping learner will probably win eventually. That does not mean human data plays no part in getting there. At scale it becomes unhelpful, not harmful (watch from 4:00).
Planes and birds. Humans neither predict tokens nor chase a single scalar reward. "As planes are to birds", supervised learning may be to human cultural learning (watch from 5:30).
The technical claim underneath is neat. Imitation learning is reinforcement learning with a very short horizon. Each episode is one token long, and the reward is how well the model predicted the true token (watch from 6:00).
Ground truth: our "where the conversation ended" section
Our post noted that Patel's strongest card in the interview was language models reaching gold-medal level at the International Mathematical Olympiad. The follow-up turns that into an argument. Patel concedes that predicting human text is not ground truth. Winning olympiad gold and building working applications from scratch are ground-truth tests, though, and no one could have trained a model to pass them with reinforcement learning from nothing. The human-data prior is what kick-starts the process (watch from 6:30; 7:00).
The image is pasteurising milk. Telling someone to stop boiling it because it will eventually be served cold misses that boiling is an intermediate step (watch from 7:30).
World models: our "transition model or text prior" section
This is where our translation table and Patel's follow-up meet most directly. Our post's table separated Sutton's world model, a transition model of what happens after an action, from the language-model sense of broad knowledge absorbed from text. Patel now calls the naming question semantic. What matters is whether the model of humans helps the system learn from ground truth. Refusing to call a language model's rich internal representation a world model defines the term by the process that builds it, not by the capabilities it implies (watch from 8:00).
The gap both sides agree on
On continual learning, Patel moves towards Sutton rather than away. Patel calls it a personal "hobby horse" (watch from 8:26):
The bandwidth problem. A model trained with outcome-based rewards learns only a few bits per episode, even when the episode runs to tens of thousands of tokens. Animals extract far more from what they observe.
Sutton's answer. In Sutton's newer OaK architecture, the component that learns from that observation stream is the transition model (watch from 9:00).
A possible shortcut. Patel sketches one: let the model call supervised fine-tuning as a tool, so it learns to teach itself what does not fit in its context window. Patel stresses not being an AI researcher and does not know whether it would work (watch from 9:30; 10:00).
The conclusion is generous (watch from 10:31; 11:00):
Opposite routes. Evolution ran meta-reinforcement learning to produce agents that can choose to imitate. Language models went the opposite way: imitation first, then reinforcement learning to make a coherent agent. "Maybe this won't work!"
Real gaps. Even if Sutton's route is not the first to general intelligence, Sutton names gaps the current paradigm hides because they are everywhere: no continual learning, poor sample efficiency and dependence on exhaustible human data.
The successor. If language models get there first, Patel expects the systems they build next to follow Sutton's vision (watch from 11:30).
What this changes in our reading
Our original post said neither speaker moved much. The follow-up shows one of them did. Patel does not concede the core claim. Patel still thinks human data is a sensible starting point, but has adopted Sutton's vocabulary for the problem that remains.
That makes the useful question for builders narrower and more practical. Not "are language models a dead end?" but "where does my system learn after deployment, and from what signal?" We will keep using that test in EXEPERT's own agent work, and in how we report on everyone else's.
References
Patel, D. (Host). (2025a, September 26). Richard Sutton – Father of RL thinks LLMs are a dead end [Audio podcast episode]. In Dwarkesh Podcast. https://www.dwarkesh.com/p/richard-sutton
Patel, D. (2025b, October 4). Some thoughts on the Sutton interview [Video]. YouTube. https://www.youtube.com/watch?v=u3HBJVjpXuw
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., & Hassabis, D. (2016). Mastering the game of Go with deep neural networks and tree search. Nature, 529(7587), 484–489. https://doi.org/10.1038/nature16961
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., & Hassabis, D. (2018). A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play. Science, 362(6419), 1140–1144. https://doi.org/10.1126/science.aar6404
Sutton, R. S. (2019, March 13). The bitter lesson. Incomplete Ideas. http://www.incompleteideas.net/IncIdeas/BitterLesson.html