Board contents
Text extracted from this public whiteboard for search and accessibility.
S12 · Open
Session 12 · Positional encodings
Live board · deepen in matheion MA-LLM · Session 12
Today’s story Bare attention doesn’t know order. Add a position vector to each token embedding — the classic paper uses sines and cosines.
You should leave able to… • Show order blindness with a swap • Write the sin/cos recipe in words + formula • Contrast fixed sinusoids vs learned positions
With matheion Use this board to teach. Open matheion → MA-LLM → Session 12 for full prose, diagrams, and the auto-quiz.
S12 · Problem
Attention alone doesn’t see order
Shuffle the tokens Without position info, ‘Dog bites man’ and ‘Man bites dog’ can look the same to the mechanism — it only sees a bag of vectors, not who came first.
S12 · Sinusoids
Add a position vector
Each position gets a pattern of sines/cosines; add it to the token’s embedding (same width)
For position pos and dimension index i: even dims: sin( pos / 10000^{2i / model_width} ) odd dims: cos( pos / 10000^{2i / model_width} ) Then: input_vector = token_embedding + position_vector Why sines? For a fixed jump k, the pattern at pos+k is a simple transform of the pattern at pos — handy for relative distance.
S12 · Alternatives
Learned positions?
Alternative Train one vector per position index (clip at a max length). Paper: nearly the same quality. Sinusoids kept partly to help lengths longer than those seen in training.
Modern note Newer models use other position recipes. Same problem (order), newer parameterisations. Sin/cos is enough for this session.
S12 · Glossary
Glossary · key concepts this session
Keep this frame visible while teaching · say the term, then the plain line
Positional encoding A vector added so the model knows token order.
Permutation blindness Bare attention treats a shuffled bag of tokens the same.
Sinusoidal PE Fixed sin/cos pattern per position and dimension (Attention Is All You Need).
Learned positions Train one vector per position index instead of using formulas.
Relative position Information about distance/offset between tokens, not only absolute index.
S12 · Check
Exit ticket
Say these aloud · then quiz in matheion MA-LLM · Session 12
Prompt 1 Why must we inject position at all?
Prompt 2 Where do you add the position vector?
Prompt 3 One pro of sinusoids vs a learned table.