chalkline

Public teaching whiteboard

CNNs for sequences — NLP before Transformers

Using convolutional networks on text sequences: strengths and limits, taught on a free public whiteboard. · by zlu

Topics: LLM & Transformers

CNNs for sequences — NLP before Transformers
S5 · Open
S5 · Idea
S5 · Hand calc
S5 · Receptive field
S5 · Win and miss
S5 · Glossary
S5 · Check

Session 5 · Convolutional networks for sequences

Live board · deepen in matheion MA-LLM · Session 5

Today’s story Mix neighbouring token vectors with a shared sliding window — a 1D convolution. Local, parallel, and parameter-efficient — but the window size is fixed.

You should leave able to… • Hand-compute a width-3 convolution on five numbers • Say how stacking grows the receptive field • Give one local win and one long-range miss

With matheion Use this board to teach. Open matheion → MA-LLM → Session 5 for full prose, diagrams, and the auto-quiz.

Sliding window on a sequence

Vision: mix nearby pixels · Text: mix nearby token vectors

Kernel width k = 3 At position i, look at tokens i−1, i, i+1 Mix them with shared weights to produce a new vector y_i. Same weights at every i — that reuse is the convolution.

Why try this for language? • Local syntax is often local (not good, adjective + noun) • Shared kernels → few parameters • All positions update together (unlike RNNs next session) Session 7 will ask: is local enough?

Hand calc · five numbers, width-3 kernel

x = [2, 0, 1, 3, 1] · weights (0.5, 1.0, 0.5) · missing edge neighbours = 0

text
y_i = 0.5·x_{i-1} + 1.0·x_i + 0.5·x_{i+1}

pos 1: window (0, 2, 0) → 0.5·0 + 2 + 0.5·0 = 2.0
pos 2: window (2, 0, 1) → 0.5·2 + 0 + 0.5·1 = 1.5
pos 3: window (0, 1, 3) → 0 + 1 + 1.5     = 2.5
pos 4: window (1, 3, 1) → 0.5 + 3 + 0.5   = 4.0
pos 5: window (3, 1, 0) → 1.5 + 1 + 0     = 2.5

Same three weights every row.
Real models: each x_i is a vector (width H); each weight is a small matrix — same slide.

Stacking grows what you can see

Rule of thumb After L layers of width k=3: receptive field ≈ 1 + L·(k−1) = 1 + 2L L=4 → field ≈ 9 positions. Dependency 40 tokens back → need ~40 layers of k=3. Depth buys reach — linearly.

Dilation (optional) Dilation 2, k=3: look at i−2, i, i+2 Field grows faster. Still a FIXED offset schedule — not “pick Alice because She is female”.

Local win · long-range miss

Win · “not good” Width-3 on “good” sees “not”. Sentiment flip is local — convolution helps.

Miss · Alice … She “She” may sit 40 tokens after “Alice”. No small window contains the name. Stacking can reach it only by paying for depth — and still cannot choose the target by meaning.

Shapes Input: (# tokens) × width. One padded conv layer: same shape. Params ~ k × width × width per layer (shared across positions).

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

1D convolution Slide a shared kernel along the token sequence; each output mixes a local window.

Kernel width k How many neighbouring positions one filter looks at (e.g. 3).

Weight sharing The same kernel weights are reused at every position.

Receptive field How far back/forward a position can depend on after stacking layers.

Dilation Skip gaps in the window (i−2,i,i+2) to grow reach faster — still fixed offsets.

Local dependency A link that sits inside a small neighbourhood (e.g. not → good).

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 5

Prompt 1 Compute y_2 for x=[2,0,1,3,1] with weights (0.5,1,0.5).

Prompt 2 About how many k=3 layers until position 10 can see position 1?

Prompt 3 One local win and one long-range miss for convolutions on text.

Drag to pan · scroll to zoom · read-only

Board contents

Text extracted from this public whiteboard for search and accessibility.

S5 · Open

Session 5 · Convolutional networks for sequences

Live board · deepen in matheion MA-LLM · Session 5

Today’s story Mix neighbouring token vectors with a shared sliding window — a 1D convolution. Local, parallel, and parameter-efficient — but the window size is fixed.

You should leave able to… • Hand-compute a width-3 convolution on five numbers • Say how stacking grows the receptive field • Give one local win and one long-range miss

With matheion Use this board to teach. Open matheion → MA-LLM → Session 5 for full prose, diagrams, and the auto-quiz.

S5 · Idea

Sliding window on a sequence

Vision: mix nearby pixels · Text: mix nearby token vectors

Kernel width k = 3 At position i, look at tokens i−1, i, i+1 Mix them with shared weights to produce a new vector y_i. Same weights at every i — that reuse is the convolution.

Why try this for language? • Local syntax is often local (not good, adjective + noun) • Shared kernels → few parameters • All positions update together (unlike RNNs next session) Session 7 will ask: is local enough?

S5 · Hand calc

Hand calc · five numbers, width-3 kernel

x = [2, 0, 1, 3, 1] · weights (0.5, 1.0, 0.5) · missing edge neighbours = 0

y_i = 0.5·x_{i-1} + 1.0·x_i + 0.5·x_{i+1} pos 1: window (0, 2, 0) → 0.5·0 + 2 + 0.5·0 = 2.0 pos 2: window (2, 0, 1) → 0.5·2 + 0 + 0.5·1 = 1.5 pos 3: window (0, 1, 3) → 0 + 1 + 1.5 = 2.5 pos 4: window (1, 3, 1) → 0.5 + 3 + 0.5 = 4.0 pos 5: window (3, 1, 0) → 1.5 + 1 + 0 = 2.5 Same three weights every row. Real models: each x_i is a vector (width H); each weight is a small matrix — same slide.

S5 · Receptive field

Stacking grows what you can see

Rule of thumb After L layers of width k=3: receptive field ≈ 1 + L·(k−1) = 1 + 2L L=4 → field ≈ 9 positions. Dependency 40 tokens back → need ~40 layers of k=3. Depth buys reach — linearly.

Dilation (optional) Dilation 2, k=3: look at i−2, i, i+2 Field grows faster. Still a FIXED offset schedule — not “pick Alice because She is female”.

S5 · Win and miss

Local win · long-range miss

Win · “not good” Width-3 on “good” sees “not”. Sentiment flip is local — convolution helps.

Miss · Alice … She “She” may sit 40 tokens after “Alice”. No small window contains the name. Stacking can reach it only by paying for depth — and still cannot choose the target by meaning.

Shapes Input: (# tokens) × width. One padded conv layer: same shape. Params ~ k × width × width per layer (shared across positions).

S5 · Glossary

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

1D convolution Slide a shared kernel along the token sequence; each output mixes a local window.

Kernel width k How many neighbouring positions one filter looks at (e.g. 3).

Weight sharing The same kernel weights are reused at every position.

Receptive field How far back/forward a position can depend on after stacking layers.

Dilation Skip gaps in the window (i−2,i,i+2) to grow reach faster — still fixed offsets.

Local dependency A link that sits inside a small neighbourhood (e.g. not → good).

S5 · Check

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 5

Prompt 1 Compute y_2 for x=[2,0,1,3,1] with weights (0.5,1,0.5).

Prompt 2 About how many k=3 layers until position 10 can see position 1?

Prompt 3 One local win and one long-range miss for convolutions on text.