chalkline

Public teaching whiteboard

LLM sampling: temperature, top-k, and top-p

How sampling turns a next-token distribution into text — temperature and nucleus sampling on a public whiteboard. · by zlu

Topics: LLM & Transformers

LLM sampling: temperature, top-k, and top-p
S3 · Open
S3 · Greedy vs sample
S3 · Temperature
S3 · Trim the tail
S3 · Glossary
S3 · Check

Session 3 · Sampling from the distribution

Live board · deepen in matheion MA-LLM · Session 3

Today’s story Session 2 left us a lottery over next tokens. Now we choose: always take the winner, or draw — and how to reshape that lottery.

You should leave able to… • Contrast always-pick-top vs draw • Show temperature on a worked 3-way example • Cut a ranked list with top-k and top-p

With matheion Use this board to teach. Open matheion → MA-LLM → Session 3 for full prose, diagrams, and the auto-quiz.

Same lottery, different next words

Prompt ends with “Once upon a” · toy probabilities: time 40%, night 25%, midnight 15%, …

Always pick the top Always choose ‘time’. Run 1–3: Once upon a time … Deterministic. Fine for short facts. Dull for stories.

Draw from the lottery Roll the dice using those %. Run 1: … time Run 2: … night Run 3: … midnight Different stories appear. That’s the product.

Temperature · reshape the lottery

Temperature is a knob T. Divide each raw score by T, then softmax. Do not divide the probabilities.

text
Raw scores:  time=2.0  night=1.0  day=0.0

T = 1 (no change):
  probabilities ≈ (0.67, 0.24, 0.09)

T = 0.5 (sharper — scores effectively stretched):
  probabilities ≈ (0.84, 0.14, 0.02)  → almost always ‘time’

T = 2 (flatter — scores compressed):
  probabilities ≈ (0.51, 0.31, 0.18)  → ‘day’ shows up more

Knob feel Low T → safer, repetitive. High T → wilder, more nonsense risk. Remember: act on raw scores, then softmax.

Trim bad options before drawing

Two common recipes — same ranked list

text
Sorted lottery:
  time .40 | night .25 | midnight .15 | morning .10 | forever .05 | zebra .05

Keep only the top 3 (often called top-k with k=3):
  keep {time, night, midnight}, drop the rest,
  renormalise those three so they sum to 1, then draw.

Keep enough to cover 80% mass (often called top-p / nucleus with p=0.8):
  time .40
  + night → .65
  + midnight → .80  ← stop
  morning / forever / zebra never eligible this step

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Decoding Turning the next-token lottery into an actual chosen token (and repeating).

Greedy decoding Always pick the top probability. Safe, repetitive.

Temperature (T) Divide each logit by T, then softmax. Low T → sharp; high T → flat.

Top-k Keep only the k highest-probability tokens, renormalise, then sample.

Top-p / nucleus Keep the smallest set whose probabilities sum to at least p, then sample.

Renormalise After dropping options, rescale remaining probabilities so they sum to 1 again.

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 3

Prompt 1 Temperature divides raw scores — why not the probabilities?

Prompt 2 When is always-pick-top the right default?

Prompt 3 On the ranked list above, what does top-3 keep vs 80%-mass keep?

Drag to pan · scroll to zoom · read-only

Board contents

Text extracted from this public whiteboard for search and accessibility.

S3 · Open

Session 3 · Sampling from the distribution

Live board · deepen in matheion MA-LLM · Session 3

Today’s story Session 2 left us a lottery over next tokens. Now we choose: always take the winner, or draw — and how to reshape that lottery.

You should leave able to… • Contrast always-pick-top vs draw • Show temperature on a worked 3-way example • Cut a ranked list with top-k and top-p

With matheion Use this board to teach. Open matheion → MA-LLM → Session 3 for full prose, diagrams, and the auto-quiz.

S3 · Greedy vs sample

Same lottery, different next words

Prompt ends with “Once upon a” · toy probabilities: time 40%, night 25%, midnight 15%, …

Always pick the top Always choose ‘time’. Run 1–3: Once upon a time … Deterministic. Fine for short facts. Dull for stories.

Draw from the lottery Roll the dice using those %. Run 1: … time Run 2: … night Run 3: … midnight Different stories appear. That’s the product.

S3 · Temperature

Temperature · reshape the lottery

Temperature is a knob T. Divide each raw score by T, then softmax. Do not divide the probabilities.

Raw scores: time=2.0 night=1.0 day=0.0 T = 1 (no change): probabilities ≈ (0.67, 0.24, 0.09) T = 0.5 (sharper — scores effectively stretched): probabilities ≈ (0.84, 0.14, 0.02) → almost always ‘time’ T = 2 (flatter — scores compressed): probabilities ≈ (0.51, 0.31, 0.18) → ‘day’ shows up more

Knob feel Low T → safer, repetitive. High T → wilder, more nonsense risk. Remember: act on raw scores, then softmax.

S3 · Trim the tail

Trim bad options before drawing

Two common recipes — same ranked list

Sorted lottery: time .40 | night .25 | midnight .15 | morning .10 | forever .05 | zebra .05 Keep only the top 3 (often called top-k with k=3): keep {time, night, midnight}, drop the rest, renormalise those three so they sum to 1, then draw. Keep enough to cover 80% mass (often called top-p / nucleus with p=0.8): time .40 + night → .65 + midnight → .80 ← stop morning / forever / zebra never eligible this step

S3 · Glossary

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Decoding Turning the next-token lottery into an actual chosen token (and repeating).

Greedy decoding Always pick the top probability. Safe, repetitive.

Temperature (T) Divide each logit by T, then softmax. Low T → sharp; high T → flat.

Top-k Keep only the k highest-probability tokens, renormalise, then sample.

Top-p / nucleus Keep the smallest set whose probabilities sum to at least p, then sample.

Renormalise After dropping options, rescale remaining probabilities so they sum to 1 again.

S3 · Check

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 3

Prompt 1 Temperature divides raw scores — why not the probabilities?

Prompt 2 When is always-pick-top the right default?

Prompt 3 On the ranked list above, what does top-3 keep vs 80%-mass keep?