This commit is contained in:
2026-06-10 19:45:12 +08:00
parent 1183d10849
commit a3ebb28ce0
12 changed files with 35469 additions and 491 deletions
+276
View File
@@ -0,0 +1,276 @@
import json
import uuid
from pathlib import Path
cells = []
def _src(text):
lines = text.split("\n")
out = [line + "\n" if i < len(lines) - 1 else line for i, line in enumerate(lines)]
if out and out[-1] == "":
out.pop()
return out
def md(text):
cells.append({"cell_type": "markdown", "id": uuid.uuid4().hex[:8],
"metadata": {}, "source": _src(text)})
def code(text):
cells.append({"cell_type": "code", "execution_count": None, "id": uuid.uuid4().hex[:8],
"metadata": {}, "outputs": [], "source": _src(text)})
# ----------------------------------------------------------------------------
md("""# The GPT tokenizer — Byte Pair Encoding from scratch
Turn raw text into tokens the model can read — **concept -> code -> Your turn** each step.
This is notebook **07**, the final piece of the pipeline. Every notebook so far fed the model
either single characters (Shakespeare) or single letters (names). Real GPTs use **subword**
tokens built by **Byte Pair Encoding (BPE)**. Here we build a BPE tokenizer from scratch, then
see why tokenization quietly causes many famous LLM quirks. Follows Karpathy's
*"Let's build the GPT Tokenizer"*.""")
md("""## Prologue — in plain English
A neural net only eats numbers, so text must be chopped into pieces (**tokens**) and each piece
mapped to an id. Two extremes:
- **One token per character** (what we did): tiny vocabulary, but sequences are very long and
the model must relearn common spellings everywhere.
- **One token per word**: short sequences, but the vocabulary explodes and unseen words break it.
**BPE** is the practical middle: start from raw bytes, then repeatedly **merge the most common
adjacent pair** into a new token. Frequent chunks like `the`, `ing`, or ` and` become single
tokens; rare text still survives as smaller pieces.
Real-life picture: inventing shorthand. You start writing letter by letter, notice you keep
writing "t-h-e", so you invent one squiggle for "the." Then you notice " a-n-d" and give it a
squiggle too. BPE does this automatically, keeping the most useful shorthands.""")
# ----------------------------------------------------------------------------
md("""### 1.1 Text -> UTF-8 bytes (the 256 starting tokens)
Computers already store text as **bytes** (numbers 0-255) via UTF-8. So our alphabet starts
with exactly 256 tokens — every possible byte. Plain English letters are one byte each;
accented or non-Latin characters take several bytes (which is why they cost more tokens later).""")
code("""sample = "Hello world! tokenization is fun."
the_bytes = sample.encode("utf-8")
print("text :", sample)
print("bytes :", list(the_bytes))
print("length :", len(the_bytes), "bytes ->", len(the_bytes), "starting tokens")
multi = "café na\u00efve" # accented letters take >1 byte each
print("\\nnon-ASCII example:", multi)
print("bytes:", list(multi.encode("utf-8")), "(note: more bytes than characters)")""")
# ----------------------------------------------------------------------------
md("""### 2.1 Two tiny helpers: count pairs, and merge a pair
- `get_stats` counts how often each adjacent pair appears.
- `merge` replaces every occurrence of a chosen pair with a single new token id.
That is the entire mechanism of BPE.""")
code("""def get_stats(ids):
counts = {}
for pair in zip(ids, ids[1:]):
counts[pair] = counts.get(pair, 0) + 1
return counts
def merge(ids, pair, idx):
new_ids = []
i = 0
while i < len(ids):
if i < len(ids) - 1 and ids[i] == pair[0] and ids[i + 1] == pair[1]:
new_ids.append(idx)
i += 2
else:
new_ids.append(ids[i])
i += 1
return new_ids
# demo on a toy list: merge the most common pair
demo = [1, 2, 3, 1, 2, 3, 1, 2]
st = get_stats(demo)
top = max(st, key=st.get)
print("counts:", st)
print("most common pair:", top, "->", st[top], "times")
print("after merging", top, "into 99:", merge(demo, top, 99))""")
md("""### 2.2 Train BPE — learn the merges
We train on a few KB of real text (Shakespeare). Starting from raw bytes, we repeatedly find
the most common pair and merge it, recording each merge. After `num_merges` rounds we have a
vocabulary of `256 + num_merges` tokens.""")
code("""train_text = open("input.txt", "r", encoding="utf-8").read()[:20000]
ids = list(train_text.encode("utf-8"))
print("training bytes:", len(ids))
num_merges = 60
merges = {} # (a, b) -> new_id
work = list(ids)
for i in range(num_merges):
stats = get_stats(work)
pair = max(stats, key=stats.get)
idx = 256 + i
work = merge(work, pair, idx)
merges[pair] = idx
print("merges learned:", len(merges))
print("tokens after training:", len(work), "(was", len(ids), "bytes)")
print(f"compression: {len(ids) / len(work):.2f}x")""")
# ----------------------------------------------------------------------------
md("""### 3.1 encode and decode
- **decode**: turn token ids back into text. We build each token's byte string by stitching the
merges back together.
- **encode**: turn new text into token ids by applying the learned merges, always doing the
*earliest-learned* applicable merge first.""")
code("""# build the id -> bytes table
vocab = {idx: bytes([idx]) for idx in range(256)}
for (p0, p1), idx in merges.items():
vocab[idx] = vocab[p0] + vocab[p1]
def decode(ids):
data = b"".join(vocab[idx] for idx in ids)
return data.decode("utf-8", errors="replace")
def encode(text):
tokens = list(text.encode("utf-8"))
while len(tokens) >= 2:
stats = get_stats(tokens)
# pick the pair whose merge was learned earliest
pair = min(stats, key=lambda p: merges.get(p, float("inf")))
if pair not in merges:
break
tokens = merge(tokens, pair, merges[pair])
return tokens
trial = "the king and the queen"
enc = encode(trial)
print("text :", trial)
print("tokens:", enc)
print("count :", len(enc), "tokens for", len(trial), "characters")
print("decode round-trip ok:", decode(enc) == trial)""")
md("""### 3.2 See what the learned tokens are
The merges discovered the common chunks of English on their own. Let's print a few learned
tokens (ids >= 256) as the text they stand for.""")
code("""print("a few learned tokens:")
for idx in range(256, 256 + min(20, num_merges)):
piece = vocab[idx].decode("utf-8", errors="replace")
print(f" id {idx}: {piece!r}")""")
# ----------------------------------------------------------------------------
md("""### 4.1 Regex pre-splitting (what real GPTs add)
A subtlety: plain BPE might merge across spaces and punctuation (e.g. glue `dog.` or `the the`
into weird tokens). GPT-2 first **splits** the text into sensible chunks with a regex (words,
runs of spaces, punctuation), then runs BPE **inside** each chunk only.
Real GPT uses the `regex` library with Unicode classes; here is a simplified version with the
standard `re` module to show the idea.""")
code("""import re
# simplified GPT-2-style splitter (real one uses the `regex` lib with \\p{L}, \\p{N})
pat = re.compile(r"'s|'t|'re|'ve|'m|'ll|'d| ?\\w+| ?[^\\s\\w]+|\\s+")
demo_text = "Hello, world! It's 2026 already."
chunks = re.findall(pat, demo_text)
print("split into chunks (BPE then runs inside each one):")
print(chunks)""")
md("""### 4.2 Special tokens
Real tokenizers also reserve **special tokens** that are not learned from text but inserted to
mark structure, for example `<|endoftext|>` between documents, or chat markers like
`<|im_start|>` / `<|im_end|>`. They get their own ids above the learned vocabulary.
Real-life picture: punctuation the model never "spells" — single reserved symbols that mean
"new document starts here" or "the user is speaking now."
```
<|endoftext|> -> id 50256 in GPT-2 (separates documents)
<|im_start|> -> chat role marker in chat models
```""")
# ----------------------------------------------------------------------------
md("""### 5.1 Why tokenization causes LLM quirks
Many puzzling LLM behaviours trace back to how text is tokenized. We measure a few.
- **Spelling / reversing words is hard**: a word is often a *single* token, so the model does
not "see" its letters.
- **Arithmetic is fragile**: numbers split into arbitrary chunks, not clean digits.
- **Non-English costs more**: it falls back to many byte-tokens, so the same meaning uses more
tokens (and more money / context).""")
code("""def n_tokens(s):
return len(encode(s))
print("spelling: 'extraordinary' ->", n_tokens("extraordinary"), "token(s)")
print(" the model sees a chunk, not the 13 letters\\n")
print("numbers : '1234567' ->", n_tokens("1234567"), "tokens (split oddly, not per-digit)")
print(" '127 + 677' ->", n_tokens("127 + 677"), "tokens\\n")
eng = "hello how are you"
non = "\u4f60\u597d\u4f60\u597d\u4f60\u597d" # non-English of similar visible length
print("non-English costs more tokens for similar content:")
print(" english :", repr(eng), "->", n_tokens(eng), "tokens")
print(" non-latin :", repr(non), "->", n_tokens(non), "tokens")""")
md("""**Your turn 7** — Raise `num_merges` (e.g. to 300) in section 2.2, retrain, and re-check the
compression ratio and the token counts in section 5.1. More merges = better compression, but a
bigger vocabulary. This trade-off is exactly what real tokenizer designers tune.""")
# ----------------------------------------------------------------------------
md("""## Series wrap-up
You built the whole stack, from the ground up:
- `00_micrograd.ipynb` — backprop from scratch (the engine)
- `01_build_gpt.ipynb` — the bigram baseline
- `02_makemore_mlp.ipynb` — an MLP with embeddings
- `03_batchnorm_activations.ipynb` — keeping deep nets healthy
- `04_backprop_ninja.ipynb` — gradients by hand
- `05_wavenet.ipynb` — a deeper, hierarchical model
- `06_build_gpt_attention.ipynb` — the transformer (GPT)
- `07_gpt_tokenizer.ipynb` — turning text into tokens (this notebook)
Together these cover Andrej Karpathy's *Neural Networks: Zero to Hero* course, on real data,
with runnable code at every step.
**Checklist**
- [ ] Text -> UTF-8 bytes (256 base tokens)
- [ ] `get_stats` + `merge` = the BPE mechanism
- [ ] Trained merges; measured compression
- [ ] encode / decode round-trips
- [ ] Regex pre-splitting and special tokens
- [ ] Explained tokenization-caused LLM quirks""")
# ----------------------------------------------------------------------------
nb = {
"cells": cells,
"metadata": {
"kernelspec": {"display_name": ".venv (3.12.10)", "language": "python", "name": "python3"},
"language_info": {"name": "python", "version": "3.12.10"},
},
"nbformat": 4,
"nbformat_minor": 5,
}
out_path = Path(r"d:\BeastProjects\gpt_from_scratch\07_gpt_tokenizer.ipynb")
out_path.write_text(json.dumps(nb, ensure_ascii=False, indent=2), encoding="utf-8")
print("wrote", out_path, "with", len(cells), "cells")