chalkline

Public teaching whiteboard

Attention as soft lookup — LLM building block

Attention as a soft dictionary lookup over values — concrete numbers on a free teaching whiteboard. · by zlu

Topics: LLM & Transformers

Attention as soft lookup — LLM building block
S8 · Open
S8 · Metaphor
S8 · Three steps
S8 · Formula
S8 · Hand example
S8 · Glossary
S8 · Check

Session 8 · Attention as soft lookup

Live board · deepen in matheion MA-LLM · Session 8

Today’s story Attention is a soft dictionary lookup: ask with a query, match keys, return a blend of values — smooth enough to train with gradients.

You should leave able to… • Tell identical → nearest → soft • Write weights from query·key dots • Hand-compute a 3-key example

With matheion Use this board to teach. Open matheion → MA-LLM → Session 8 for full prose, diagrams, and the auto-quiz.

Lookup you already know

Dictionary Keys = words Values = definitions Query = the word you want

Phone book Keys = names Values = numbers Query = who you’re calling

Exact match fails in nets Vectors almost never equal exactly. ‘She’ points toward ‘Alice’ — it isn’t identical.

Identical → nearest → soft

1 · Identical Return the value whose key equals the query. Otherwise fail. Never fires with continuous vectors.

2 · Nearest Return the value of the closest key. Always defined — but jumps when the winner flips → no useful gradient.

3 · Soft (attention) Give every value a weight between 0 and 1 (weights sum to 1). Output = weighted mix of all values. Tiny move in the query → tiny move in the output.

Where the weights come from

query q · keys k_j · values v_j — same softmax shape as Session 2, candidates are keys

text
similarity of query to key j  =  q · k_j     (dot product)
weight_j = exp(q · k_j) / sum_t exp(q · k_t)
output   = sum_j  weight_j * v_j

Weights look like probabilities — but we do NOT draw a key.
We average the values.

Roles Query and key must have the same width so dots make sense. Value can have any width — what’s returned needn’t match what’s searched.

Worked example · one query, three keys

text
q = (1.0, 0.0)
key1=(1.0,0.2)  value1=(5, 1)    →  q·key1 = 1.00
key2=(0.9,0.1)  value2=(4.5,1.2) →  q·key2 = 0.90
key3=(-0.5,0.5) value3=(-1, 3)   →  q·key3 = -0.50

exp ≈ 2.72, 2.46, 0.61    sum ≈ 5.79
weights ≈ 0.47, 0.425, 0.105
output ≈ (4.15, 1.30)   # mostly value1 + value2

Read it Most mass on key1, strong secondary on key2, little on key3. Soft lookup mixes — it doesn’t snap.

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Query What you’re looking for (the ‘ask’ vector).

Key What can be matched against the query.

Value What gets returned / mixed into the output.

Dot product Similarity score between query and a key (alignment of two vectors).

Soft lookup / attention Weighted average of values; weights from softmax over query·key scores.

Identical lookup Exact key match only — useless with continuous vectors.

Nearest lookup Hard closest key — jumps, not differentiable.

Differentiable Tiny input change → tiny output change; needed for gradient training.

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 8

Prompt 1 Why nearest lookup blocks learning.

Prompt 2 Write the weight formula and say: average, don’t sample.

Prompt 3 Recompute the 3-key toy from memory.

Drag to pan · scroll to zoom · read-only

Board contents

Text extracted from this public whiteboard for search and accessibility.

S8 · Open

Session 8 · Attention as soft lookup

Live board · deepen in matheion MA-LLM · Session 8

Today’s story Attention is a soft dictionary lookup: ask with a query, match keys, return a blend of values — smooth enough to train with gradients.

You should leave able to… • Tell identical → nearest → soft • Write weights from query·key dots • Hand-compute a 3-key example

With matheion Use this board to teach. Open matheion → MA-LLM → Session 8 for full prose, diagrams, and the auto-quiz.

S8 · Metaphor

Lookup you already know

Dictionary Keys = words Values = definitions Query = the word you want

Phone book Keys = names Values = numbers Query = who you’re calling

Exact match fails in nets Vectors almost never equal exactly. ‘She’ points toward ‘Alice’ — it isn’t identical.

S8 · Three steps

Identical → nearest → soft

1 · Identical Return the value whose key equals the query. Otherwise fail. Never fires with continuous vectors.

2 · Nearest Return the value of the closest key. Always defined — but jumps when the winner flips → no useful gradient.

3 · Soft (attention) Give every value a weight between 0 and 1 (weights sum to 1). Output = weighted mix of all values. Tiny move in the query → tiny move in the output.

S8 · Formula

Where the weights come from

query q · keys k_j · values v_j — same softmax shape as Session 2, candidates are keys

similarity of query to key j = q · k_j (dot product) weight_j = exp(q · k_j) / sum_t exp(q · k_t) output = sum_j weight_j * v_j Weights look like probabilities — but we do NOT draw a key. We average the values.

Roles Query and key must have the same width so dots make sense. Value can have any width — what’s returned needn’t match what’s searched.

S8 · Hand example

Worked example · one query, three keys

q = (1.0, 0.0) key1=(1.0,0.2) value1=(5, 1) → q·key1 = 1.00 key2=(0.9,0.1) value2=(4.5,1.2) → q·key2 = 0.90 key3=(-0.5,0.5) value3=(-1, 3) → q·key3 = -0.50 exp ≈ 2.72, 2.46, 0.61 sum ≈ 5.79 weights ≈ 0.47, 0.425, 0.105 output ≈ (4.15, 1.30) # mostly value1 + value2

Read it Most mass on key1, strong secondary on key2, little on key3. Soft lookup mixes — it doesn’t snap.

S8 · Glossary

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Query What you’re looking for (the ‘ask’ vector).

Key What can be matched against the query.

Value What gets returned / mixed into the output.

Dot product Similarity score between query and a key (alignment of two vectors).

Soft lookup / attention Weighted average of values; weights from softmax over query·key scores.

Identical lookup Exact key match only — useless with continuous vectors.

Nearest lookup Hard closest key — jumps, not differentiable.

Differentiable Tiny input change → tiny output change; needed for gradient training.

S8 · Check

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 8

Prompt 1 Why nearest lookup blocks learning.

Prompt 2 Write the weight formula and say: average, don’t sample.

Prompt 3 Recompute the 3-key toy from memory.