Board contents
Text extracted from this public whiteboard for search and accessibility.
S11 · Open
Session 11 · Scaling and multi-head attention
Live board · deepen in matheion MA-LLM · Session 11
Today’s story Two upgrades on Sessions 8–9: shrink huge dot products so softmax stays soft, and run several attention ‘views’ in parallel.
You should leave able to… • Explain why we divide by √(key width) • Motivate several heads with the dog/food sentence • List the four steps of one attention block
With matheion Use this board to teach. Open matheion → MA-LLM → Session 11 for full prose, diagrams, and the auto-quiz.
S11 · Why scale
Why divide the scores?
Key width = how many numbers are in each key/query vector (paper often uses 64)
Attention(Q, K, V) = softmax( (Q Kᵀ) / √(key_width) ) × V Problem: as key_width grows, raw dots get huge → softmax becomes nearly one-hot → gradients die. Fix: divide by √(key_width) so the lottery stays soft enough to learn.
S11 · Multi-head
Several heads, then combine
Why several? Sentence: “A dog ate the food because it was hungry.” One head might bind ‘it’ → ‘dog’. Another might track ‘ate’ → ‘food’. A single head averages those jobs into mush.
Recipe 1) Split into several smaller attention views (heads). 2) Each head has its own ask/match/deliver maps. 3) Concatenate the head outputs. 4) One final linear mix. Paper base model: 8 heads, each key width 64, model width 512 (= 8×64).
S11 · Four steps
One attention block · mental model
1 · Project Build queries, keys, values (per head).
2 · Score Dots, then divide by √(key width).
3 · Mask Hide future scores if next-token LM.
4 · Mix Softmax → weighted values; concat heads; final mix.
S11 · Glossary
Glossary · key concepts this session
Keep this frame visible while teaching · say the term, then the plain line
Scaled dot-product Softmax( (Q Kᵀ) / √key_width ) × V — scaling keeps softmax soft.
Key width Number of dimensions in each key/query vector (paper often uses 64).
Multi-head attention Several attention views in parallel, then concatenate and mix.
Head One attention view with its own Q/K/V maps.
Model width Main embedding size of the stack (paper base: 512 = 8 heads × 64).
Saturation Softmax nearly one-hot when scores are huge — gradients die.
S11 · Check
Exit ticket
Say these aloud · then quiz in matheion MA-LLM · Session 11
Prompt 1 What goes wrong without dividing by √(key width)?
Prompt 2 Why several heads on the dog/food sentence?
Prompt 3 8 heads, model width 512 → each key width is?