chalkline

Public teaching whiteboard

From language model to AI assistant

How base LMs become assistants: instruction tuning and chat formatting — free public whiteboard. · by zlu

Topics: LLM & Transformers

From language model to AI assistant
S15 · Open
S15 · Stages
S15 · In context
S15 · Limits
S15 · Glossary
S15 · Check

Session 15 · From next-token prediction to assistants

Live board · deepen in matheion MA-LLM · Session 15

Today’s story Same Transformer stack at every stage — change the data (and a bit of training) so a next-token machine behaves like a helpful assistant.

You should leave able to… • Name pre-train / fine-tune / align • Explain in-context examples vs weight updates • State two limits that follow from the mechanism

With matheion Use this board to teach. Open matheion → MA-LLM → Session 15 for full prose, diagrams, and the auto-quiz.

Same stack, new data

1 · Pre-train Next-token prediction on huge text corpora. Learns language, facts, patterns — not yet “be a helpful assistant”.

2 · Fine-tune Show (prompt → good answer) pairs. Teaches the format of assistance on top of the same architecture.

3 · Align Steer with human (or AI) preferences: prefer helpful / harmless / honest replies over worse ones. Still next-token underneath.

Examples inside the prompt

What it is Put worked examples in the prompt. The model steers via attention — weights in the network stay frozen.

What it isn’t Not permanent learning. Clear the prompt, and those “lessons” vanish. Long prompts also cost more (Session 14’s quadratic bill).

Limits the mechanism imposes

Fluent ≠ true A sharp next-token lottery can sound confident while wrong. That’s hallucination — not a calibrated truth oracle.

Context window Only the tokens inside the window are visible to attention. Outside = invisible.

Seminar arc Tokens → lottery → soft lookup → stack → product. You can now place every piece on that arc.

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Pre-training Next-token learning on huge raw text — language and patterns, not ‘assistant’ yet.

Fine-tuning Train on (prompt → good answer) pairs to teach helpful formats.

Alignment Steer toward preferred behaviour (helpful / harmless / honest) still via next-token training.

In-context learning Examples in the prompt steer attention; network weights stay frozen.

Hallucination Fluent, confident next-token output that isn’t true.

Context window limit Tokens outside the window are invisible to attention.

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 15

Prompt 1 Pre-train vs fine-tune vs align in one line each.

Prompt 2 In-prompt examples: do weights change?

Prompt 3 Name two failure modes that follow from next-token lotteries.

Drag to pan · scroll to zoom · read-only

Board contents

Text extracted from this public whiteboard for search and accessibility.

S15 · Open

Session 15 · From next-token prediction to assistants

Live board · deepen in matheion MA-LLM · Session 15

Today’s story Same Transformer stack at every stage — change the data (and a bit of training) so a next-token machine behaves like a helpful assistant.

You should leave able to… • Name pre-train / fine-tune / align • Explain in-context examples vs weight updates • State two limits that follow from the mechanism

With matheion Use this board to teach. Open matheion → MA-LLM → Session 15 for full prose, diagrams, and the auto-quiz.

S15 · Stages

Same stack, new data

1 · Pre-train Next-token prediction on huge text corpora. Learns language, facts, patterns — not yet “be a helpful assistant”.

2 · Fine-tune Show (prompt → good answer) pairs. Teaches the format of assistance on top of the same architecture.

3 · Align Steer with human (or AI) preferences: prefer helpful / harmless / honest replies over worse ones. Still next-token underneath.

S15 · In context

Examples inside the prompt

What it is Put worked examples in the prompt. The model steers via attention — weights in the network stay frozen.

What it isn’t Not permanent learning. Clear the prompt, and those “lessons” vanish. Long prompts also cost more (Session 14’s quadratic bill).

S15 · Limits

Limits the mechanism imposes

Fluent ≠ true A sharp next-token lottery can sound confident while wrong. That’s hallucination — not a calibrated truth oracle.

Context window Only the tokens inside the window are visible to attention. Outside = invisible.

Seminar arc Tokens → lottery → soft lookup → stack → product. You can now place every piece on that arc.

S15 · Glossary

Glossary · key concepts this session

Keep this frame visible while teaching · say the term, then the plain line

Pre-training Next-token learning on huge raw text — language and patterns, not ‘assistant’ yet.

Fine-tuning Train on (prompt → good answer) pairs to teach helpful formats.

Alignment Steer toward preferred behaviour (helpful / harmless / honest) still via next-token training.

In-context learning Examples in the prompt steer attention; network weights stay frozen.

Hallucination Fluent, confident next-token output that isn’t true.

Context window limit Tokens outside the window are invisible to attention.

S15 · Check

Exit ticket

Say these aloud · then quiz in matheion MA-LLM · Session 15

Prompt 1 Pre-train vs fine-tune vs align in one line each.

Prompt 2 In-prompt examples: do weights change?

Prompt 3 Name two failure modes that follow from next-token lotteries.