EXEPERTAI LAB
EXEPERT / DIRECTORY
← Blog

Ten Computer Science Papers That Built Modern AI, from Turing to GPT-3

Fireship's century of AI in ten papers, each summarised with a verified reference: Turing's machine, Shannon's bit, the perceptron and its winter, Lamport's clocks, backpropagation, PageRank, AlexNet, the transformer and GPT-3.

Akmal Alif · 8 October 2026 MYT

Ten paper sheets laid along a timeline from left to right, each with a tiny diagram: a tape with a read head, a signal wave, a single neuron, a crossed-out exclusive-or grid, clock arrows, a layered network, a web of linked nodes, a convolution grid, an attention matrix and a large glowing block.

Research Paper is where we read the papers behind the technology: what each one claimed, what it changed, and where to find the original. To start the category, here is a whole century of them.

In a June 2026 video, Fireship tells the history of artificial intelligence through ten papers, from Alan Turing's machine to the scaling bet behind GPT-3 (Fireship, 2026). The video is embedded below. Below it, each paper gets a short summary, a note where the popular version simplifies the history, and a full reference, so you can go to the source.

I read every major CS paper of the last 100 years... (Fireship)The YouTube player loads in privacy-enhanced mode when you press play.Watch on YouTube ↗

Defining the machine and the bit

1. Turing, “On Computable Numbers” (1936–37). David Hilbert had asked whether a mechanical procedure could decide the truth of any mathematical statement, the Entscheidungsproblem. To answer it, Alan Turing first had to define what an algorithm is. Turing's answer was an imaginary machine with an unbounded tape, a read-write head and a small table of rules. The paper then showed that some questions about such machines cannot be decided by any machine, so the answer to Hilbert was no (Turing, 1937). The best-known example is now called the halting problem: whether an arbitrary program will ever finish. The paper was submitted in 1936 and printed in 1937, which is why both dates appear (watch from 1:00; 1:20).

2. Shannon, “A Mathematical Theory of Communication” (1948). Claude Shannon set meaning aside and treated information as a measurable quantity: surprise, counted in bits, with entropy as the measure of uncertainty in a source (Shannon, 1948). The video draws a straight line from here to modern language models, which are trained to predict the next token (watch from 1:40; 2:20). One refinement: the experiments in which people guess the next letter of English text come from Shannon's 1951 follow-up paper, which estimated the entropy of printed English (Shannon, 1951).

The neuron, the winter and the clock

3. Rosenblatt, “The Perceptron” (1958). Frank Rosenblatt, a psychologist at Cornell, proposed a machine loosely modelled on neurons. It weighs its inputs, makes a decision and adjusts the weights when it is wrong, so it learns to classify patterns (Rosenblatt, 1958). The video recalls the immediate hype, with Navy funding and press talk of machines that would soon be conscious (watch from 3:00).

4. Minsky and Papert, “Perceptrons” (1969). This one is a book rather than a paper. It proved hard limits on single-layer perceptrons, most famously that they cannot compute exclusive or (Minsky & Papert, 1969). It is often blamed for the collapse in neural-network funding that followed. The fine print the video highlights is that networks with more layers can represent such functions; the unsolved problem was how to train them (watch from 3:20).

5. Lamport, “Time, Clocks, and the Ordering of Events in a Distributed System” (1978). Leslie Lamport showed that separate machines with no shared clock cannot agree on a single “now”. They can still agree on order, using the happened-before relation: if A could have caused B, A comes first. Logical clocks follow from that (Lamport, 1978). The video's point is that today's distributed systems, including training runs spread across thousands of GPUs, rest on this kind of reasoning about order (watch from 4:00).

Learning, data and the breakthrough

6. Rumelhart, Hinton and Williams, “Learning Representations by Back-Propagating Errors” (1986). This paper showed how to train the stack of layers. Run the input forward and measure the error, then use the chain rule to push that error back through every layer, nudging each weight so the output is a little less wrong. Hidden layers then learn useful features on their own, and problems like exclusive or become easy (Rumelhart et al., 1986). Data and compute were still too scarce for it to shine (watch from 6:00).

7. Brin and Page, “The Anatomy of a Large-Scale Hypertextual Web Search Engine” (1998). This is the prototype of Google, built at Stanford. It ranks pages with PageRank, which treats each link as a vote weighted by the importance of the page casting it (Brin & Page, 1998). The video's framing is that web search helped organise the huge body of human text that later became training data (watch from 6:40).

8. Krizhevsky, Sutskever and Hinton, “ImageNet Classification with Deep Convolutional Neural Networks” (2012). AlexNet combined a deep convolutional network, the ImageNet dataset and two consumer gaming GPUs. In the 2012 ImageNet challenge it reached a top-5 error of 15.3%, against 26.2% for the next-best entry (Krizhevsky et al., 2012). That margin convinced the field that deep learning worked, given enough data and compute (watch from 7:20).

The architecture and the scale

9. Vaswani et al., “Attention Is All You Need” (2017). The transformer replaced sequential reading with attention, letting every token weigh every other token at once. It trains in parallel and scales well (Vaswani et al., 2017). It is the T in GPT (watch from 8:00). For how the attention computation works step by step, see our Machine Learning post on attention in five pieces.

10. Brown et al., “Language Models Are Few-Shot Learners” (2020). OpenAI scaled a transformer to 175 billion parameters and trained it on a large slice of the internet. The resulting GPT-3 could translate, answer questions and do simple arithmetic from a few examples in the prompt, with no task-specific training (Brown et al., 2020). In the video's telling, this is the bet that capability emerges from scale, and the line that led to ChatGPT (watch from 8:40; 9:00).

The century in one line

The video's summary is that each paper added one piece to the machine (watch from 9:40):

  • Turing defined it.

  • Shannon gave it a currency.

  • Rosenblatt gave it a neuron.

  • Backpropagation taught it to learn.

  • The web gave it data, and the transformer gave it an architecture.

  • OpenAI turned up the scale.

The loop also closes neatly. A chatbot predicting the next token is doing at scale what Shannon's volunteers did with pencil and paper.

All ten are worth opening, and most are shorter than their reputations. Lamport's paper fits in eight journal pages, and the backpropagation paper in four.

References

Brin, S., & Page, L. (1998). The anatomy of a large-scale hypertextual Web search engine. Computer Networks and ISDN Systems, 30(1–7), 107–117. https://doi.org/10.1016/S0169-7552(98)00110-X

Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems 33 (pp. 1877–1901). https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html

Fireship. (2026, June 17). I read every major CS paper of the last 100 years... [Video]. YouTube. https://www.youtube.com/watch?v=ML3q7Ok4hJg

Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems 25. https://papers.nips.cc/paper/2012/hash/c399862d3b9d6b76c8436e924a68c45b-Abstract.html

Lamport, L. (1978). Time, clocks, and the ordering of events in a distributed system. Communications of the ACM, 21(7), 558–565. https://doi.org/10.1145/359545.359563

Minsky, M., & Papert, S. (1969). Perceptrons: An introduction to computational geometry. MIT Press.

Rosenblatt, F. (1958). The perceptron: A probabilistic model for information storage and organization in the brain. Psychological Review, 65(6), 386–408. https://doi.org/10.1037/h0042519

Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323(6088), 533–536. https://doi.org/10.1038/323533a0

Shannon, C. E. (1948). A mathematical theory of communication. Bell System Technical Journal, 27(3), 379–423. https://doi.org/10.1002/j.1538-7305.1948.tb01338.x

Shannon, C. E. (1951). Prediction and entropy of printed English. Bell System Technical Journal, 30(1), 50–64. https://doi.org/10.1002/j.1538-7305.1951.tb01366.x

Turing, A. M. (1937). On computable numbers, with an application to the Entscheidungsproblem. Proceedings of the London Mathematical Society, s2-42(1), 230–265. https://doi.org/10.1112/plms/s2-42.1.230

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems 30 (pp. 5998–6008). https://arxiv.org/abs/1706.03762