Neuromorphic
Reading Ilya Sutskever’s 2021 LLM Forecast in 2026
A year before ChatGPT, Ilya Sutskever explained why prediction leads to understanding, why scaling needs data, and why code would change first. Read in 2026, his forecast is strongest where it was most concrete.
Akmal Alif · 8 October 2026 MYT

In December 2021, almost a year before ChatGPT turned large language models into a household topic, Ilya Sutskever sat down with Scale AI’s Alexandr Wang to talk about where the technology was heading (Scale AI, 2021). Sutskever was then OpenAI’s co-founder and chief scientist and one of the architects of the GPT models. He left OpenAI in May 2024 and co-founded Safe Superintelligence a month later (Fortune, 2024).
Watching the conversation again in 2026 is a useful exercise. Some of its claims now sound obvious, which is a measure of how much they shaped the field. Others remain open questions, and a few look more prescient than they did at the time. The interview is embedded below; the timestamps throughout jump to the matching moment.
Prediction as a path to understanding
The idea at the centre of the interview is simple: if a system can make very good guesses about what comes next, it must have understood something (watch from 11:11). Sutskever illustrates it with a mystery novel. Near the end, a careful reader can narrow down who the culprit is before the sentence finishes, and that guess is only possible because the reader followed the plot.
OpenAI tested the principle before transformers existed. A small network trained only to predict the next character of Amazon product reviews developed a single unit that tracked sentiment, without ever being told what sentiment was (watch from 12:11; Radford et al., 2017). The GPT models took the same bet further with a better architecture, and then with scale.
The framing has aged into the default story of modern AI, and also into its most contested one. Richard Sutton’s objection, discussed in Two Languages of Intelligence, is precisely that predicting text is not the same as predicting the consequences of acting in the world.
Scale needs data, not just compute
Sutskever’s most practical point is about data. Academic machine learning grew up around fixed benchmarks, which made methods comparable but quietly fixed the amount of data everyone could use. The GPT results showed that scaling works when compute and data grow together (watch from 15:14), the relationship OpenAI had formalised in its scaling-laws work (Kaplan et al., 2020).
He also named the limit. Language is abundant, but specialised domains such as law have far less data, so a model can converse brilliantly and still fall short as a lawyer (watch from 16:14). His advice was not to bet against deep learning, because every year people had declared its limits and every year they were wrong, while admitting that limits might eventually arrive.
Three years later, Sutskever said so himself. In his Test of Time talk at NeurIPS 2024, he argued that pre-training as practised would end because compute keeps growing while the supply of human-written data does not: there is only one internet (Sutskever, 2024). The 2021 interview already pointed at the next question, how to use the same compute more intelligently when data runs short (watch from 18:15).
A Moore’s law for data
Wang, whose company sells training data, pushes on how to make each hour of human expertise go further. Sutskever offers two routes: improve methods so they need less data, or make the human teachers more efficient, for example by asking them for help only on the hardest cases (watch from 20:15).
Both routes became industries. Human feedback on model outputs became the standard way to shape assistants, and synthetic data became a research field of its own, with its own failure modes. The limits of training on a model’s own output are the subject of Do LLMs Need to Dream?.
Codex and the computer as an actuator
The section on Codex is the one that has aged best. Codex was a GPT model trained on code instead of prose, and Sutskever’s first observation was that it was remarkable it worked at all (watch from 21:15). The accompanying paper reported that it solved 28.8% of the HumanEval programming problems, where GPT-3 solved none (Chen et al., 2021).
His framing was that code models can control the computer: the machine becomes their actuator (watch from 23:20). He also described their knowledge as encyclopedic rather than deep, which makes them most useful with libraries and interfaces a programmer does not know, and he was clear that their output must be checked when the code matters.
He predicted that programming would change the way it always has, by moving to a higher level of abstraction, from assembly to Fortran to C to Python and now to natural language (watch from 25:21). Coding assistants and agents have since become one of the most widely used applications of language models, and “you still need to check it” remains the right advice.
The inversion: creative and white-collar work first
One of the most quoted ideas from this period appears here in compact form. Many people expected automation to reach simple physical tasks first; instead, generative models reached images, writing and code (watch from 26:22). Sutskever urged economists to watch the trend closely so that good ideas would be ready as the technology improved.
Wang attributes the pattern to data: there is a vast digital record of text, images and code, and very little of people setting tables. Sutskever agrees but adds a second reason. Generative models produce new, plausible data, which makes them naturally suited to creative work (watch from 28:22).
Generalisation: the student who memorises
The most useful idea for understanding today’s models may be his analogy about generalisation (watch from 29:23). One student prepares for an exam by memorising every exercise in the textbook. Another reads the fundamentals and stops. If both score well, the second student did something harder.
Neural networks, he said, are like the first student. They generalise impressively for computers, but not yet at a human level, so they compensate with enormous amounts of data. Better generalisation would make data scarcity matter much less. That remains one of the central open problems of the field.
Efficiency, neurons and alignment
Sutskever was confident that methods would keep getting more efficient. A decade earlier, the only way to use huge amounts of compute was embarrassingly parallel work like MapReduce; deep learning found another, but not the best one (watch from 38:32). Because halving the compute needed for a model is equivalent to doubling a computer, the incentive to find better methods is enormous (watch from 40:32).
Asked whether simple artificial neurons are the wrong abstraction, he thought it extremely unlikely. The worst case, he suggested, is that one biological neuron behaves like a small supercomputer and needs something like a million artificial ones to match it (watch from 42:32). The answer is a reminder that the brain is a source of rough estimates and inspiration, not a blueprint to copy.
On alignment, his example was GPT-3 itself, which understood language but often did not do what it was asked. The instruction-following models were built to fulfil a user’s intent faithfully, and he noted that the more aligned model was also the more useful one (watch from 44:32). The InstructGPT paper later showed people preferring a 1.3-billion-parameter instruction-tuned model to the 175-billion-parameter GPT-3 (Ouyang et al., 2022). He also argued that larger models are easier to steer, not harder (watch from 49:39).
What held up
Read in 2026, the forecast is strongest where it was most concrete. Scaling with data held, and its data limit became real enough that Sutskever named it publicly. Code became one of the technology’s defining uses. Alignment as usefulness became the template for every assistant. His closing bet on biology and medicine (watch from 51:42) was rewarded sooner than most expected: protein structure prediction shared the 2024 Nobel Prize in Chemistry (Nobel Prize Outreach, 2024).
The open items are the ones he flagged as open. Generalisation still trails human learning, and the question of what to do when data runs out has moved from footnote to centre stage. His advice to the audience still fits: build applications that solve real problems, and work on reducing the real harms (watch from 53:42).
References
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. de O., … Zaremba, W. (2021). Evaluating large language models trained on code (arXiv:2107.03374). arXiv. https://arxiv.org/abs/2107.03374
Fortune. (2024, June 20). Ilya Sutskever left OpenAI after mutinying against Sam Altman—now he’s launching his own startup for safe AI. https://fortune.com/2024/06/20/openai-ilya-sutskever-sam-altman-safe-superintelligence/
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., & Amodei, D. (2020). Scaling laws for neural language models (arXiv:2001.08361). arXiv. https://arxiv.org/abs/2001.08361
Nobel Prize Outreach. (2024, October 9). The Nobel Prize in Chemistry 2024 [Press release]. https://www.nobelprize.org/prizes/chemistry/2024/press-release/
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback (arXiv:2203.02155). arXiv. https://arxiv.org/abs/2203.02155
Radford, A., Jozefowicz, R., & Sutskever, I. (2017). Learning to generate reviews and discovering sentiment (arXiv:1704.01444). arXiv. https://arxiv.org/abs/1704.01444
Scale AI. (2021, December 9). OpenAI co-founder Ilya Sutskever: What’s next for large language models (LLMs) [Video]. YouTube. https://www.youtube.com/watch?v=UHSkjro-VbE
Sutskever, I. (2024, December 13). Sequence to sequence learning with neural networks: What a decade [Test of Time Award talk]. NeurIPS 2024, Vancouver, Canada. https://blog.neurips.cc/2024/11/27/announcing-the-neurips-2024-test-of-time-paper-awards/