chalkline

Public teaching whiteboard

LLM tokenization explained — from text to tokens

How language models turn text into tokens. Free public LLM teaching whiteboard with concrete string examples. · by zlu

Topics: LLM & Transformers

LLM tokenization explained — from text to tokens
S1 · Open
S1 · Same sentence ×3
S1 · Subword examples
S1 · Merge recipe
S1 · Merge by hand
S1 · IDs & round-trip
S1 · Why this design wins
S1 · Glossary
S1 · Check

Session 1 · From text to tokens

Live board · deepen in matheion MA-LLM · Session 1

Today’s story Nets need numbers. We’ll cut one sentence three ways, then build a tiny ‘merge frequent neighbours’ recipe by hand — that’s how modern tokenisers invent pieces like est.

You should leave able to… • Cut one sentence three ways and compare token counts • Merge frequent letter-pairs by hand until ‘est’ appears • Predict how ‘unhappiness’ splits, and why that beats ‘unknown’

With matheion Use this board to teach. Open matheion → MA-LLM → Session 1 for full prose, diagrams, and the auto-quiz.

Same sentence, three cuttings

Running example: "lowest prices"

text
TEXT = "lowest prices"

1) CHARACTERS — cut every letter (and the space)
   l o w e s t _ p r i c e s
   → 13 pieces. Model must learn that l-o-w-e-s-t means lowest.

2) WHOLE WORDS — cut on spaces only
   lowest | prices
   → 2 pieces. But the dictionary must list every form:
     low, lower, lowest, prices, price, priced, typos…

3) SUBWORDS — reuse shared pieces
   low | est | prices
   → 3 pieces. The bit ‘est’ can be reused in newest, highest…

What to say aloud Characters → many pieces, hard learning. Whole words → few pieces, huge dictionary. Subwords → share pieces across words so rare forms still get signal.

What subword cuts look like in practice

These are the kinds of pieces a trained tokeniser actually emits

text
word            → pieces                      why it helps
────────────────────────────────────────────────────────────
the             →  the                        very common → keep whole
playing         →  play | ing                 ‘ing’ reused everywhere
newest          →  new | est                  ‘est’ shared with highest…
unhappiness     →  un | happi | ness          prefix + stem + suffix
tokenization    →  token | ization            long rare stem, shared ending
ChatGPT         →  Chat | G | PT              rare name → fall back to pieces
12345           →  12 | 345                   digits often chunked

Notice: common stuff stays short; rare stuff reuses known bits.

How the vocabulary gets built

Idea first: repeatedly glue the most common neighbouring pair

1 · Start from letters Every character (or byte) is already a legal piece. Any string can be written.

2 · Count neighbours Scan lots of text. Which two pieces sit next to each other most often?

3 · Glue the winner Make that pair a new single piece. Rewrite the text with the glue applied.

4 · Repeat until big enough Stop at a chosen dictionary size (tens of thousands). This merge-by-frequency recipe is what people call BPE (byte-pair encoding). Closely related systems (e.g. WordPiece) use the same ‘build pieces from parts’ idea with a different scoring rule.

Worked example · invent ‘est’

Toy corpus: lowest · newest · highest · testing

text
Start (letters only):
  l o w e s t    n e w e s t    h i g h e s t    t e s t i n g

Neighbour ‘e s’ and then ‘s t’ show up a lot → glue them.

After gluing e+s → es, then es+t → est:
  l o w est     n e w est    h i g h est     t est i n g

Later glues may build ‘low’, ‘new’, ‘high’, ‘test’…
At cut time, ‘newest’ becomes  new | est  (2 pieces), not 6 letters.

Student check If we never glued further for ‘best’, how does it cut? b | est Same ‘est’ piece as in newest — that’s why subwords beat a frozen word list.

Each piece gets an integer ID

text
text = "newest prices"

Suppose the dictionary rows are:
  1042 → ‘new’
  881  → ‘est’
  55   → ‘ prices’   (leading space kept as part of the piece)

encode → [1042, 881, 55]
decode → ‘newest prices’   (must match exactly)

Vocabulary size = number of rows ≈ 30,000–100,000
(GPT-2’s dictionary had 50,257 entries).

Token count for this string = 3
English word count was 2 — never confuse the two.

Pitfall · teach this hard ‘How many tokens?’ is NOT ‘how many words?’ ‘I love ChatGPT’ might become 5–7 tokens depending on the tokeniser. Always run encode before claiming a length.

Design win · one example

Rare word: ‘unhappiness’ If the dictionary only stores whole words and has never seen this one, the model gets an ‘unknown’ placeholder — almost no meaning. With subwords: un | happi | ness Those bits already appeared in unhappy, happiness, kindness… so the pieces already carry signal.

Common word: ‘the’ Stays ONE piece. No point splitting the most frequent English word into t|h|e. Rule of thumb: frequent → keep whole; rare → reuse known bits.

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Token A piece of text the model treats as one unit — often a subword, not always a whole word.

Tokeniser The program that cuts a string into tokens and maps them to integer IDs (and back).

Vocabulary / dictionary The fixed list of allowed tokens. Size ≈ tens of thousands of rows.

Subword A reusable fragment (ing, est, …). Rare words are built from known fragments.

BPE Byte-pair encoding: repeatedly glue the most frequent neighbouring pair to grow the vocabulary.

Encode / decode String → list of IDs, and IDs → string. Must round-trip exactly.

Token count vs word count How many tokens ≠ how many English words. Always ask the tokeniser.

Unknown placeholder What a whole-word system does for never-seen words — almost no meaning.

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 1

Prompt 1 Cut ‘lowest prices’ three ways. Who creates the most pieces? Who needs the biggest dictionary?

Prompt 2 From the toy corpus, show how ‘est’ gets invented by gluing neighbours.

Prompt 3 Predict pieces for ‘unhappiness’ and why that’s better than an unknown placeholder.

Drag to pan · scroll to zoom · read-only

Board contents

Text extracted from this public whiteboard for search and accessibility.

S1 · Open

Session 1 · From text to tokens

Live board · deepen in matheion MA-LLM · Session 1

Today’s story Nets need numbers. We’ll cut one sentence three ways, then build a tiny ‘merge frequent neighbours’ recipe by hand — that’s how modern tokenisers invent pieces like est.

You should leave able to… • Cut one sentence three ways and compare token counts • Merge frequent letter-pairs by hand until ‘est’ appears • Predict how ‘unhappiness’ splits, and why that beats ‘unknown’

With matheion Use this board to teach. Open matheion → MA-LLM → Session 1 for full prose, diagrams, and the auto-quiz.

S1 · Same sentence ×3

Same sentence, three cuttings

Running example: "lowest prices"

TEXT = "lowest prices" 1) CHARACTERS — cut every letter (and the space) l o w e s t _ p r i c e s → 13 pieces. Model must learn that l-o-w-e-s-t means lowest. 2) WHOLE WORDS — cut on spaces only lowest | prices → 2 pieces. But the dictionary must list every form: low, lower, lowest, prices, price, priced, typos… 3) SUBWORDS — reuse shared pieces low | est | prices → 3 pieces. The bit ‘est’ can be reused in newest, highest…

What to say aloud Characters → many pieces, hard learning. Whole words → few pieces, huge dictionary. Subwords → share pieces across words so rare forms still get signal.

S1 · Subword examples

What subword cuts look like in practice

These are the kinds of pieces a trained tokeniser actually emits

word → pieces why it helps ──────────────────────────────────────────────────────────── the → the very common → keep whole playing → play | ing ‘ing’ reused everywhere newest → new | est ‘est’ shared with highest… unhappiness → un | happi | ness prefix + stem + suffix tokenization → token | ization long rare stem, shared ending ChatGPT → Chat | G | PT rare name → fall back to pieces 12345 → 12 | 345 digits often chunked Notice: common stuff stays short; rare stuff reuses known bits.

S1 · Merge recipe

How the vocabulary gets built

Idea first: repeatedly glue the most common neighbouring pair

1 · Start from letters Every character (or byte) is already a legal piece. Any string can be written.

2 · Count neighbours Scan lots of text. Which two pieces sit next to each other most often?

3 · Glue the winner Make that pair a new single piece. Rewrite the text with the glue applied.

4 · Repeat until big enough Stop at a chosen dictionary size (tens of thousands). This merge-by-frequency recipe is what people call BPE (byte-pair encoding). Closely related systems (e.g. WordPiece) use the same ‘build pieces from parts’ idea with a different scoring rule.

S1 · Merge by hand

Worked example · invent ‘est’

Toy corpus: lowest · newest · highest · testing

Start (letters only): l o w e s t n e w e s t h i g h e s t t e s t i n g Neighbour ‘e s’ and then ‘s t’ show up a lot → glue them. After gluing e+s → es, then es+t → est: l o w est n e w est h i g h est t est i n g Later glues may build ‘low’, ‘new’, ‘high’, ‘test’… At cut time, ‘newest’ becomes new | est (2 pieces), not 6 letters.

Student check If we never glued further for ‘best’, how does it cut? b | est Same ‘est’ piece as in newest — that’s why subwords beat a frozen word list.

S1 · IDs & round-trip

Each piece gets an integer ID

text = "newest prices" Suppose the dictionary rows are: 1042 → ‘new’ 881 → ‘est’ 55 → ‘ prices’ (leading space kept as part of the piece) encode → [1042, 881, 55] decode → ‘newest prices’ (must match exactly) Vocabulary size = number of rows ≈ 30,000–100,000 (GPT-2’s dictionary had 50,257 entries). Token count for this string = 3 English word count was 2 — never confuse the two.

Pitfall · teach this hard ‘How many tokens?’ is NOT ‘how many words?’ ‘I love ChatGPT’ might become 5–7 tokens depending on the tokeniser. Always run encode before claiming a length.

S1 · Why this design wins

Design win · one example

Rare word: ‘unhappiness’ If the dictionary only stores whole words and has never seen this one, the model gets an ‘unknown’ placeholder — almost no meaning. With subwords: un | happi | ness Those bits already appeared in unhappy, happiness, kindness… so the pieces already carry signal.

Common word: ‘the’ Stays ONE piece. No point splitting the most frequent English word into t|h|e. Rule of thumb: frequent → keep whole; rare → reuse known bits.

S1 · Glossary

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Token A piece of text the model treats as one unit — often a subword, not always a whole word.

Tokeniser The program that cuts a string into tokens and maps them to integer IDs (and back).

Vocabulary / dictionary The fixed list of allowed tokens. Size ≈ tens of thousands of rows.

Subword A reusable fragment (ing, est, …). Rare words are built from known fragments.

BPE Byte-pair encoding: repeatedly glue the most frequent neighbouring pair to grow the vocabulary.

Encode / decode String → list of IDs, and IDs → string. Must round-trip exactly.

Token count vs word count How many tokens ≠ how many English words. Always ask the tokeniser.

Unknown placeholder What a whole-word system does for never-seen words — almost no meaning.

S1 · Check

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 1

Prompt 1 Cut ‘lowest prices’ three ways. Who creates the most pieces? Who needs the biggest dictionary?

Prompt 2 From the toy corpus, show how ‘est’ gets invented by gluing neighbours.

Prompt 3 Predict pieces for ‘unhappiness’ and why that’s better than an unknown placeholder.