Board contents
Text extracted from this public whiteboard for search and accessibility.
S5 · Open
Session 5 · Convolutional networks for sequences
Live board · deepen in matheion MA-LLM · Session 5
Today’s story Mix neighbouring token vectors with a shared sliding window — a 1D convolution. Local, parallel, and parameter-efficient — but the window size is fixed.
You should leave able to… • Hand-compute a width-3 convolution on five numbers • Say how stacking grows the receptive field • Give one local win and one long-range miss
With matheion Use this board to teach. Open matheion → MA-LLM → Session 5 for full prose, diagrams, and the auto-quiz.
S5 · Idea
Sliding window on a sequence
Vision: mix nearby pixels · Text: mix nearby token vectors
Kernel width k = 3 At position i, look at tokens i−1, i, i+1 Mix them with shared weights to produce a new vector y_i. Same weights at every i — that reuse is the convolution.
Why try this for language? • Local syntax is often local (not good, adjective + noun) • Shared kernels → few parameters • All positions update together (unlike RNNs next session) Session 7 will ask: is local enough?
S5 · Hand calc
Hand calc · five numbers, width-3 kernel
x = [2, 0, 1, 3, 1] · weights (0.5, 1.0, 0.5) · missing edge neighbours = 0
y_i = 0.5·x_{i-1} + 1.0·x_i + 0.5·x_{i+1} pos 1: window (0, 2, 0) → 0.5·0 + 2 + 0.5·0 = 2.0 pos 2: window (2, 0, 1) → 0.5·2 + 0 + 0.5·1 = 1.5 pos 3: window (0, 1, 3) → 0 + 1 + 1.5 = 2.5 pos 4: window (1, 3, 1) → 0.5 + 3 + 0.5 = 4.0 pos 5: window (3, 1, 0) → 1.5 + 1 + 0 = 2.5 Same three weights every row. Real models: each x_i is a vector (width H); each weight is a small matrix — same slide.
S5 · Receptive field
Stacking grows what you can see
Rule of thumb After L layers of width k=3: receptive field ≈ 1 + L·(k−1) = 1 + 2L L=4 → field ≈ 9 positions. Dependency 40 tokens back → need ~40 layers of k=3. Depth buys reach — linearly.
Dilation (optional) Dilation 2, k=3: look at i−2, i, i+2 Field grows faster. Still a FIXED offset schedule — not “pick Alice because She is female”.
S5 · Win and miss
Local win · long-range miss
Win · “not good” Width-3 on “good” sees “not”. Sentiment flip is local — convolution helps.
Miss · Alice … She “She” may sit 40 tokens after “Alice”. No small window contains the name. Stacking can reach it only by paying for depth — and still cannot choose the target by meaning.
Shapes Input: (# tokens) × width. One padded conv layer: same shape. Params ~ k × width × width per layer (shared across positions).
S5 · Glossary
Glossary · key concepts this session
Keep this frame visible while teaching · say the term, then the plain line
1D convolution Slide a shared kernel along the token sequence; each output mixes a local window.
Kernel width k How many neighbouring positions one filter looks at (e.g. 3).
Weight sharing The same kernel weights are reused at every position.
Receptive field How far back/forward a position can depend on after stacking layers.
Dilation Skip gaps in the window (i−2,i,i+2) to grow reach faster — still fixed offsets.
Local dependency A link that sits inside a small neighbourhood (e.g. not → good).
S5 · Check
Exit ticket
Say these aloud · then quiz in matheion MA-LLM · Session 5
Prompt 1 Compute y_2 for x=[2,0,1,3,1] with weights (0.5,1,0.5).
Prompt 2 About how many k=3 layers until position 10 can see position 1?
Prompt 3 One local win and one long-range miss for convolutions on text.