Machine Learning
Attention in Five Pieces: Score, Scale, Mask, Normalize, Mix
Scaled dot-product attention fits on one line and confuses almost everyone. Following Zachary Huang's walkthrough, we run it by hand on one four-word sentence: score, scale, normalize, mix and mask, with every number and a short Python function.
Akmal Alif · 8 October 2026 MYT

Machine Learning is where we take one idea behind modern AI and work through it until it makes sense. The first is the formula at the centre of every transformer: scaled dot-product attention. It runs inside the GPT models behind ChatGPT, it fits on one line, and the 2017 paper that introduced it, Attention Is All You Need, leaves many readers more confused than before (Vaswani et al., 2017).
In a 21-minute video from September 2026, Zachary Huang breaks the formula into five pieces and runs it by hand on a four-word sentence, so that every number can be checked (Huang, 2026). This post follows the same route. The video is embedded below, and the timestamps jump to the matching moment.
Why one row per word is not enough
A model only sees numbers, so every word has to be translated first. The video compares three ways of doing it on two sentences: “the crane lifted steel” and “the crane ate fish” (watch from 0:59).
Word IDs give each word an arbitrary number. Shuffle the vocabulary and the same sentence gets different numbers, so the numbers carry no meaning (watch from 1:34). Embeddings fix that by giving each word a row of numbers in which similar words sit close together. Huang uses two columns, “how animal” and “how machine”, so every word can be drawn as an arrow (watch from 2:35).
That exposes the real problem. “Crane” gets a single row, and it lands in the middle, half bird and half machine. The first sentence needs the machine and the second needs the bird, but an embedding table can only hand back the same row every time (watch from 3:38). The meaning has to come from the neighbours. That is the job of attention.
Query, key and value as a group quiz
Attention lets every word rebuild itself from the words around it, the way a reader uses “lifted” and “steel” to decide which crane is meant. To do that, each word plays three roles at once (watch from 4:05).
Huang's analogy is a group quiz. The query is the question a word asks the class; for “crane” it is roughly am I an animal or a machine? Each key is a classmate's reply, the kind of help that word advertises. Each value is what the classmate actually hands over: the word's own meaning. The better a reply matches the question, the more of that classmate's offer you take (watch from 4:40).
In a real model, trained weights turn each word's embedding into its query, key and value. For the walkthrough the video simply writes the three small tables by hand, two numbers per word on the same two axes (watch from 6:12).
The five pieces
Written out, with d the width of a key and M the mask, the formula is:
Attention(Q, K, V) = softmax( Q·Kᵀ / √d + M ) · V1. Score: compare every question with every reply
Each word's query is compared with every key, its own included, by multiplying matching numbers and adding them up. That is the dot product, a similarity meter that grows when two arrows point the same way and when they are long. One matrix multiplication, Q times K transposed, produces all sixteen scores for the four-word sentence. Crane's row comes out at 0.14, 0.56, 0.70 and 0.63, and “lifted” with “steel”, long and parallel, scores highest of any pair of different words (watch from 7:15).
2. Scale: turn the volume down
Every score is then divided by the square root of d. With two numbers per vector that barely changes anything: crane's 0.14 becomes 0.099 (watch from 8:18).
The reason shows up with wider vectors. Huang repeats each number a hundred times, so every vector holds 200 numbers, and crane's scores become a hundred times larger (watch from 9:19). Fed into the next step, scores that large hand 99.9% of the attention to a single word. Dividing by the square root of 200 restores a real spread, with “lifted” at about half and the others sharing the rest (watch from 10:24). Scaling is the difference between sharing and winner-takes-all.
4. Normalize: split 100% of the trust
The video deliberately takes piece four before piece three. Softmax turns each row of scores into shares that are all positive and add up to one. Dividing by the row total alone would fail, because real scores can be negative. So softmax first raises e, about 2.718, to each score, which makes everything positive, and then divides by the total (watch from 11:16).
For crane the shares come out at 19% for “the”, 26% for itself, 28% for “lifted” and 27% for “steel”. Crane trusts its machine neighbours a little more than itself, and even “the” keeps a share, because softmax never produces an exact zero (watch from 12:57).
5. Mix: rebuild each word from its neighbours
Each word is then rebuilt as a weighted average of every word's value, using its own shares as the weights. In matrix form that is the weights times V. “Lifted” and “steel” carry machine-heavy values, so they pull crane toward the machine side, while crane's own share stops it from swinging all the way. Nobody told the model which crane this is; the neighbours did. Run “the crane ate fish” through the same steps and crane lands on the animal side instead (watch from 13:51; 14:58).
3. Mask: no reading ahead
The last piece matters for models like GPT that write one word at a time. During training the model sees the whole sentence, but at “lifted” it is supposed to predict “steel”. Without a mask, “lifted” has already mixed in about 30% of “steel”, so the answer is hidden inside the clue. That scores well in training and fails in use, where the next word does not exist yet (watch from 15:26).
The mask M adds zero wherever a word may look and negative infinity wherever it may not. Because softmax raises e to each score, negative infinity becomes an exact zero share (watch from 16:30). Crane's row becomes 43% “the”, 57% itself and nothing for the two words after it, so crane stays ambiguous on its own row. The later rows do see “lifted” and “steel”, and the last row, the one GPT reads to predict the next word, sees the whole sentence (watch from 17:32; 20:38).
The whole function
The video ends by running every piece as one short Python function and checking each line's output shape (watch from 18:44). Here is a compact version written for this post. With T words and d numbers per word, Q, K and V are each T × d. The only addition to the video's steps is subtracting each row's maximum before exponentiating, a standard guard against overflow that leaves softmax unchanged.
import numpy as np
def attention(Q, K, V, causal=True):
T, d = Q.shape
scores = Q @ K.T # 1. score: T x T
scores = scores / np.sqrt(d) # 2. scale: T x T
if causal: # 3. mask: hide future words
allowed = np.tril(np.ones((T, T), dtype=bool))
scores = np.where(allowed, scores, -np.inf)
weights = np.exp(scores - scores.max(axis=1, keepdims=True))
weights = weights / weights.sum(axis=1, keepdims=True) # 4. normalize
return weights @ V # 5. mix: T x dStep | Operation | Shape |
|---|---|---|
Score | Q · Kᵀ | T × T |
Scale | divide by √d | T × T |
Mask | −∞ above the diagonal | T × T |
Normalize | softmax per row | T × T, rows sum to 1 |
Mix | weights · V | T × d |
What to take away
The formula reads as one block of symbols, but each part answers a single question. Which words matter to me? is the score. Am I being drowned out by large numbers? is the scale. Am I allowed to look? is the mask. How much do I trust each word? is the softmax. What do I become? is the mix.
Real transformers repeat this across many heads and layers, with learned weights in place of hand-written tables, but none of that changes the five steps. If one of them still feels abstract, the video's final table lists every line, every shape and every number for the toy sentence, and it is worth pausing on (watch from 20:38).
References
Huang, Z. (2026, September 3). Give me 21 min, I will make attention click forever [Video]. YouTube. https://www.youtube.com/watch?v=eo1BZCcFYvI
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems 30 (pp. 5998–6008). https://arxiv.org/abs/1706.03762