Board contents
Text extracted from this public whiteboard for search and accessibility.
S13 · Open
Session 13 · The Transformer stack
Live board · deepen in matheion MA-LLM · Session 13
Today’s story Encoder and decoder are repeated blocks: attention, a small per-token network, and residual+normalise. Three attention uses appear in one translation model.
You should leave able to… • Sketch encoder vs decoder blocks in words • Name the three attention uses • Say what the per-token network adds
With matheion Use this board to teach. Open matheion → MA-LLM → Session 13 for full prose, diagrams, and the auto-quiz.
S13 · Blocks
Encoder block vs decoder block
Encoder (repeat N times) 1) Self-attention (full — every source token reads every source token) 2) Add residual + normalise 3) Small per-token network (two linear layers + nonlinearity) 4) Add residual + normalise Builds a rich representation of the source.
Decoder (repeat N times) 1) Self-attention with causal mask (Session 12) 2) Residual + normalise 3) Cross-attention: queries from decoder, keys/values from encoder 4) Residual + normalise 5) Per-token network + residual + normalise
S13 · Three attentions
Three uses of multi-head attention
Encoder self Ask/match/deliver all from the source side. Every source token reads every source token.
Decoder self (masked) Target side only, with future hidden. Keeps generation honest.
Cross (encoder–decoder) Queries from the target side; keys/values from the source memory. ‘What to say’ looks up ‘what was read’.
S13 · Per-token net
The small per-token network
Often called a feed-forward network (FFN): same weights at every position, different from attention
For each token vector x (separately): hidden = relu( x × W1 + b1 ) # expand width (paper: 512 → 2048) out = hidden × W2 + b2 # back to model width (2048 → 512) Attention mixes information ACROSS tokens. This network mixes features INSIDE one token. Residual: output_of_step = x + sublayer(x) then normalise — keeps deep stacks trainable. Paper base: model width 512, 6 encoder + 6 decoder layers, 8 heads.
S13 · Ends
In and out of the stack
In Token embeddings + position vectors (Sessions 4 & 12) feed the first layer.
Out Final linear layer → raw scores over the dictionary → softmax lottery (Session 2). Paper often reuses the embedding table weights here.
S13 · Glossary
Glossary · key concepts this session
Keep this frame visible while teaching · say the term, then the plain line
Encoder Stack that reads the source with full self-attention.
Decoder Stack that generates the target with masked self-attention + cross-attention.
Residual connection output = x + sublayer(x) — keeps an identity path for training deep nets.
Layer normalisation Stabilises activations after sublayers (often written with the residual).
Feed-forward / per-token net (FFN) Same small MLP at every position; mixes features inside one token.
Encoder–decoder attention Decoder queries look into encoder keys/values (the source memory).
Weight tying Reuse embedding table weights for the final vocabulary projection.
S13 · Check
Exit ticket
Say these aloud · then quiz in matheion MA-LLM · Session 13
Prompt 1 Sketch encoder vs decoder steps in words.
Prompt 2 Name where queries vs keys/values come from in all three attentions.
Prompt 3 What does the per-token network do that attention doesn’t?