EXEPERTAI LAB
EXEPERT / DIRECTORY
← Blog

Two Languages of Intelligence: Reading the Sutton–Dwarkesh Interview

Richard Sutton and Dwarkesh Patel used the same words—prediction, goal, imitation, world model—to mean different things. Turing Post’s breakdown shows why their viral interview was less a debate than a failed translation.

Akmal Alif · 7 October 2026 MYT

Abstract split illustration: an amber loop of action and consequence on the left, green streams of text tokens on the right, and a bridge between them that does not meet.

When Richard Sutton sat down with Dwarkesh Patel in September 2025, the conversation spread quickly and produced an unusual amount of argument (Patel, 2025). Sutton is one of the founders of reinforcement learning and, with Andrew Barto, received the 2024 ACM A.M. Turing Award for developing its conceptual and algorithmic foundations (Association for Computing Machinery, 2025). Patel hosts a long-form interview podcast that is widely followed in AI research circles. The episode carried a provocative title, Father of RL thinks LLMs are a dead end, and many listeners heard it as a debate with a winner and a loser.

A short episode of Turing Post’s Attention Span offers a more useful reading (Turing Post, 2025). Its argument is that the two speakers were not mainly disagreeing; they were using the same words with different meanings. Prediction, goal, imitation and world model each carried one sense in Sutton’s experiential vocabulary and another in the vocabulary of the large-language-model era. The conversation looked like a clash because nobody stopped to translate.

That framing is worth spelling out, because the same mistranslation appears in everyday discussions about agents, benchmarks and “reasoning” models.

Abstract split illustration: an amber loop of action and consequence on the left, green streams of text tokens on the right, and a bridge between them that does not meet.
Two vocabularies for intelligence: learning from the consequences of action, and predicting the next token. The Sutton–Dwarkesh interview showed how rarely they are translated.

Prediction: consequences or continuations?

For Sutton, prediction means anticipating what will happen in the world after an action. An agent predicts, acts and then observes the outcome. There is a ground truth, and the gap between expectation and outcome—surprise—is the signal that changes future behaviour. Touch a hot stove and the consequence is not a matter of opinion.

In the language-model framing, prediction usually means predicting the next token in a sequence of text. That objective has produced remarkable systems whose outputs can resemble reasoning and even planning. But, as the video summarises Sutton’s position, a model trained this way predicts what a person might write next, not what will happen after it replies. It does not, by default, form an expectation about the consequences of its answer and then learn from being wrong.

Both speakers said “prediction”. Only one of them meant a claim that the world could falsify.

Goals: changing the world or minimising a loss?

Sutton returned to John McCarthy’s definition of intelligence as “the computational part of the ability to achieve goals in the world” (McCarthy, 2007). On this view, having a goal is close to the essence of intelligence, and the goal refers to an outcome outside the system.

When Patel suggested that next-token prediction is itself a goal, Sutton disagreed. A training objective shapes a network’s parameters, but predicting text does not change the external world, and you cannot watch such a system act and infer the outcome it is pursuing. One speaker meant an internal optimisation target; the other meant a state of affairs the agent is trying to bring about.

Imitation: bootstrapping culture or a by-product of trial and error?

The most heated exchange concerned learning by imitation. Patel argued that people acquire much of their competence by copying others, and that cultural skills such as hunting or cultivating crops are passed on largely through imitation. That intuition supports the idea that pre-training on human text is a reasonable foundation for capable systems.

Sutton’s reply was that infant learning is dominated by trial and error and by observing consequences (Patel, 2025). Even when a child copies a word or a gesture, there is usually a purpose behind it: attention, food, a reaction, a test of a boundary. In Sutton’s sense, imitation serves goals; it is not a separate mechanism that can stand in for experience. The video’s presenter adds a parent’s observation: children repeat things over and over, but rarely for no reason.

World models: a transition model or a text prior?

The term world model caused similar confusion. In Sutton’s usage it is a transition model: if I do this, what happens next? It belongs to an agent that acts, and its predictions can be checked against what actually follows.

The video highlights Patel’s suggestion that language models may be the best world models built so far. In Sutton’s vocabulary that does not hold, because a model of what people say about the world is not the same as a model of how the world responds to the agent’s actions. According to the video, Sutton even preferred to call language models networks rather than models, to avoid borrowing the meaning he reserves for the word.

Sutton’s four-part agent

The clearest way to translate between the two vocabularies is Sutton’s own description of an agent. In The Quest for a Common Model of the Intelligent Decision Maker, he argues that psychology, artificial intelligence, economics, control theory and neuroscience share a picture of the decision maker with four internal components: perception, a policy, a value function and a transition model of the world (Sutton, 2022). He sketched the same picture in the interview.

Term

Sutton’s experiential sense

Common language-model sense

Prediction

The expected consequence of acting, checked against what happens

The probability of the next token in text

Goal

An outcome in the world that the agent tries to bring about

The training objective, such as next-token loss

Imitation

Behaviour that serves a goal and is refined by trial and error

Learning the distribution of human text or demonstrations

World model

A transition model: if I act, what follows?

Broad knowledge absorbed from text, often called a prior

Policy

What to do in the current situation

The model itself, tuned as a policy during post-training

Value function

How well things are going, learned with temporal-difference methods

Rarely explicit; reward models appear mainly in post-training

Seen this way, Sutton’s critique is not that language models are useless. It is that they lack the loop in which an agent’s own experience—actions, consequences and rewards—keeps updating all four components. He calls this the experiential paradigm, the direction he and David Silver describe as the era of experience (Silver & Sutton, 2025).

The irony of the Bitter Lesson

Sutton’s 2019 essay The Bitter Lesson argued that general methods that scale with computation, chiefly search and learning, eventually beat approaches built around human knowledge (Sutton, 2019). It is often cited as an argument for scaling language models. EXEPERT discussed what it means for agents in The Bitter Lesson for Embodied AI.

The interview complicates that use. Sutton suggested that language models still lean heavily on human knowledge through their training data, and that systems learning directly from experience could prove more scalable—which would make them another instance of the Bitter Lesson rather than its confirmation. He also admitted surprise at how effective neural networks have become at language tasks. And, as the video points out with some amusement, he treated the essay itself as an empirical observation about one period of history, not a law guaranteed to hold for the next seventy years.

Where the conversation ended

The interview closed with Sutton’s argument that a succession from humanity to digital intelligence or augmented humans is inevitable. He offered four steps: there is no unified human governance that could halt the work; researchers will eventually understand how intelligence works; that understanding will not stop at human level but reach superintelligence; and over time the most intelligent entities tend to gain resources and power (Patel, 2025). The video’s presenter reads this hopefully, as a picture of AI augmenting people rather than replacing them.

On the central disagreement, neither speaker moved much. Patel pointed to language models reaching gold-medal performance at the International Mathematical Olympiad as evidence that they are a good foundation to build on. Sutton held that learning from experience is the path that will scale. The video’s conclusion is that this was less a debate than two vocabularies passing each other by.

Why the translation matters

For anyone building or evaluating AI systems, the practical lesson is to say which language you are speaking. When a product claims that a model “plans” or “has goals”, ask whether those words refer to a training objective and a fluent output, or to outcomes in an environment that could prove the system wrong.

That distinction shapes EXEPERT’s research tools. EXEPERT World records what an agent observed, the action it chose and the state that followed, so that a prediction can be checked against a consequence instead of being judged by how plausible a description sounds. That does not settle the argument between Sutton and Patel, and it is not a claim that experience alone is sufficient. It is a way of keeping the experiential vocabulary measurable.

The full interview is worth hearing, and so is the ten-minute breakdown: Richard Sutton and Dwarkesh Patel – speaking two different languages on Turing Post’s YouTube channel. Taken together, they stop sounding like a shouting match and start reading like a dictionary that has not been written yet.

References

Association for Computing Machinery. (2025, March 5). ACM A.M. Turing Award honors two researchers who led the development of cornerstone AI technology. https://awards.acm.org/about/2024-turing

McCarthy, J. (2007). What is artificial intelligence? Stanford University. https://www-formal.stanford.edu/jmc/whatisai.pdf

Patel, D. (Host). (2025, September 26). Richard Sutton – Father of RL thinks LLMs are a dead end [Audio podcast episode]. In Dwarkesh Podcast. https://www.dwarkesh.com/p/richard-sutton

Silver, D., & Sutton, R. S. (2025). Welcome to the era of experience [Preprint of a chapter in Designing an Intelligence, MIT Press]. Google DeepMind. https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf

Sutton, R. S. (2019, March 13). The bitter lesson. Incomplete Ideas. http://www.incompleteideas.net/IncIdeas/BitterLesson.html

Sutton, R. S. (2022). The quest for a common model of the intelligent decision maker (arXiv:2202.13252). arXiv. https://arxiv.org/abs/2202.13252

Turing Post. (2025, September 29). Richard Sutton and Dwarkesh Patel – speaking two different languages [Video]. YouTube. https://www.youtube.com/watch?v=EwhHUQuqiaQ