Board contents
Text extracted from this public whiteboard for search and accessibility.
S6 · Open
Session 6 · Recurrent networks for sequences
Live board · deepen in matheion MA-LLM · Session 6
Today’s story Keep a memory vector and update it as each token arrives: h_t = f(h_{t−1}, x_t). Unbounded reach in principle — but sequential, bottlenecked, and long-pathed.
You should leave able to… • Unroll three toy hidden-state steps by hand • Count hops from Alice to She • Name the three structural costs
With matheion Use this board to teach. Open matheion → MA-LLM → Session 6 for full prose, diagrams, and the auto-quiz.
S6 · Idea
A memory that walks left to right
Update rule h_t = f(h_{t−1}, x_t) h_0 = zeros (usually) x_t = embedding at step t h_t = new memory after reading x_t Same f at every step — sharing across time.
Unrolled picture x1 → f → h1 ↓ x2 → f → h2 ↓ x3 → f → h3 Each arrow to the next f is a sequential dependency. You cannot fill h1,h2,h3 of one layer all at once.
S6 · Hand calc
Hand calc · scalar memory
Toy: h_t = tanh(0.5·h_{t−1} + x_t) · h_0 = 0 · x = [1, 0, 1]
t=1: input = 0.5·0 + 1 = 1 → h1 = tanh(1) ≈ 0.76 t=2: input = 0.5·0.76 + 0 ≈ 0.38 → h2 = tanh(0.38) ≈ 0.36 t=3: input = 0.5·0.36 + 1 ≈ 1.18 → h3 = tanh(1.18) ≈ 0.83 Zero x1 and rerun → h3 changes. So step 1 still influences step 3 through the chain — “memory”. Real models: h_t is a vector of width hundreds/thousands; f is a small net. LSTM/GRU = same skeleton with gates that learn to keep or forget.
S6 · Three costs
Three structural costs
These are why Session 7 will say RNNs struggle at language scale
1 · Sequential h_t needs h_{t−1}. S tokens ⇒ S serial steps. No parallel fill across the sequence.
2 · Bottleneck All of the past must fit in one vector h_t. Many entities compete for the same slots — easy to overwrite Alice.
3 · Long path Alice at pos 1 → She at pos 7: h1 → h2 → … → h7 (six hops). Gradients shrink along the chain. Far links are fragile to learn.
S6 · LSTM note
LSTM / GRU in one card
Still recurrent Gates learn to copy or erase pieces of memory. Better long memory in practice. Does NOT remove: sequential steps, one-vector funnel, hop count ≈ distance.
What they get right • Unbounded reach in principle • One f for any length • Natural for streaming text Next session: score CNN vs RNN on Alice/She.
S6 · Glossary
Glossary · key concepts this session
Keep this frame visible while teaching · say the term, then the plain line
RNN Recurrent net: update a hidden state one token at a time with a shared f.
Hidden state h_t The memory vector after reading token t.
Unrolling Drawing the same f once per time step left to right.
Sequential bottleneck Cannot compute all positions of a layer in parallel.
Path length Number of hidden-state hops between two positions (≈ distance).
LSTM / GRU Gated RNNs that learn keep/forget — still sequential recurrence.
S6 · Check
Exit ticket
Say these aloud · then quiz in matheion MA-LLM · Session 6
Prompt 1 With h_t=tanh(0.5 h_{t−1}+x_t), h_0=0, x=[0,1] — what is h_2≈?
Prompt 2 Alice at 2, She at 9 — about how many hops?
Prompt 3 Name the three structural costs of RNNs.