The unglamorous front door of every large language model.
"How does an LLM turn text into tokens? Walk me through how a tokenizer works — and how would you build one?"
A tokenizer is a compression scheme wearing a job-interview costume. Neural nets only eat numbers, so text has to become a list of integers first. Byte-pair encoding (BPE) learns which character chunks appear together most often in a big corpus, merges them into "tokens," and freezes that rulebook. From then on, encoding is just applying the frozen merges, greedily, in order — no learning at inference time, ever.
The single most interview-useful sentence you can say: "Tokenization is lossy-looking but must be lossless — decode(encode(x)) == x for every input, including emoji and binary garbage — and the vocab is fixed at train time, so inference just looks things up."
"A tokenizer splits text into words." Nope. It splits text into statistically frequent chunks. unbelievable becomes un + believ + able — pieces that aren't words in any dictionary but were frequent neighbors in the training corpus. Spaces, punctuation, and even parts of emoji get the same treatment. If your mental model is "fancy word splitter," rebuild it as "a zip file's dictionary that learned English."
encode(text) → int[]: text in, token IDs out. Fast — it runs on every request.decode(ids) → text: IDs back to text, byte-for-byte identical to the input.Analogy first: every token ID is a row number in a giant lookup table (the embedding matrix). Bigger vocab = shorter sequences but a fatter table. The math tells you where the pain lives.
| Quantity | Assumption | Result |
|---|---|---|
| Embedding matrix | 100k vocab × 12,288 dims × fp16 (2 bytes) | 2.46 GB — per matrix. Untied input/output = ~5 GB just to name the tokens. |
| English density | ~4 characters per token | 1M tokens ≈ 4M chars ≈ 700k words ≈ ~2,800 pages |
| Cost of a prompt | $2.50 / 1M input tokens, 500-token request | 500 / 1e6 × $2.50 = $0.00125 — a tenth of a cent |
| Cost of a corpus | 1 GB of text ≈ 250M tokens at $2.50/1M | ~$625 to run one GB through a frontier model |
| Merge table size | 100k merges × ~2 short strings | A few MB — the rulebook is tiny; the embedding table is the monster |
Takeaway after the math: vocab size is a three-way tug — fewer tokens per request (cheaper inference, longer context fit) vs. embedding memory vs. rare-token quality. That's the whole trade-off; everything else is details.
Think of it like a post office sorting line: normalize the envelope, slice it into rough chunks, then stamp each chunk with its ID number from a fixed codebook.
flowchart LR
A["Raw text"] --> B["Normalizer
unicode NFC, strip control chars"]
B --> C["Pre-tokenizer
regex split on spaces and punctuation"]
C --> D["BPE merge engine
apply learned merges in order"]
E[("merges.txt
frozen rulebook")] --> D
F[("vocab.json
token to id")] --> D
D --> G["Token IDs"]
G --> H["Embedding lookup
id to vector"]
The regex that does the rough split is doing real semantic work. GPT-4's famous pre-tokenizer regex deliberately splits numbers into ≤3-digit chunks (123456 → 123 456) and keeps contractions together ('ll, 're). Change this regex and you change what the model finds easy: three-digit number chunks are part of why models do okay-ish at arithmetic. It's the cheapest hyperparameter in the whole pipeline — one line of regex, enormous downstream effects.
Training is the only "smart" part, and it's just counting. Split the corpus into words, count every adjacent pair of symbols, merge the most frequent pair into a new symbol, repeat. Each merge is recorded in order — that ordered list is the tokenizer.
sequenceDiagram
participant U as You
participant T as Trainer
participant C as Corpus
participant V as Vocab store
U->>T: train(corpus, num_merges=120)
T->>C: split into words, then chars
loop 120 times
T->>T: count every adjacent pair
T->>T: best pair (a, b) becomes token ab
T->>V: append merge rule in order
end
T->>U: merges.txt + vocab.json (frozen)
Worked mini-example on the word low appearing often next to lower: early merges learn l+o → lo, then lo+w → low, then low+e → lowe. Frequent words collapse into single tokens; rare words stay shattered into pieces. That's the whole trick — frequency-proportional compression.
Encoding learns nothing. Take the input, split into words/chars, then walk the merge list in the exact order learned and fuse every occurrence of each pair. Order matters: merge #3 can only fire on the output of merges #1–2.
sequenceDiagram
participant C as Client
participant T as Tokenizer
participant M as Merge table
C->>T: encode("hello world")
T->>T: pre-tokenize into words
T->>M: load merges in learned order
loop for each merge (a, b)
T->>T: fuse every adjacent a b
end
T->>C: [1045, 995] token IDs
Some tokenizers (WordPiece) do use greedy longest-match at encode time. BPE's merge-order application is subtly different and guarantees consistency with training: the merges you apply are exactly the merges training discovered, in the same order. Greedy longest-match can produce segmentations training never saw, which slightly shifts the distribution the model was trained on. In practice both work; BPE's way is more faithful to its own training.
train(corpus_paths, vocab_size) -> None
# writes merges.txt, vocab.json, config.json
encode(text: str) -> list[int]
decode(ids: list[int]) -> str
# invariant: decode(encode(x)) == x
encode_batch(texts) -> list[list[int]]
# for throughput: parallelize here
merges.txtordered pairsvocab.jsontoken → idconfig.jsonregex + versionThe merge file is append-only and ordered — line 1 fired before line 2 at training, and must fire first at encoding. Version all three together; a v2 tokenizer reading v1 merges is a silent-corruption bug.
| Decision | Option A | Option B | Verdict |
|---|---|---|---|
| Vocab size | Small (30k): tiny embeddings, but text shatters into many tokens — longer sequences, pricier inference | Large (200k): compact sequences, but GBs of embedding memory | Match the domain: code/multilingual wants bigger |
| Base symbols | Characters: clean display, but any unseen char is OOV | Bytes (256 base): zero OOV ever — any input encodable | Bytes for production (this is what GPT does) |
| Algorithm | BPE: simple, fast, merge-order encode | Unigram: probabilistic, better compression, slower training | BPE unless you can cite a measured win |
| Pre-tokenizer | Aggressive split: cleaner tokens, more of them | Lax split: fewer tokens, weirder boundaries | Steal GPT-4's regex; it's battle-tested |
solidGoldMagikarp) that the model barely saw in training and that make it behave erratically. Mitigation: audit low-frequency tokens; some teams prune them.Opinionated and concrete — the stack I'd reach for tomorrow:
tokenizers) — regex pre-tokenization + merge application parallelized across cores; Python bindings for training scripts.tokenizer/v7/{merges,vocab,config}.json + SHA-256 pinned in the model checkpoint.decode(encode(x)) == x, plus a golden set of known tokenizations that may never change.run/running share pieces), no OOV.This isn't a mockup. When the page loaded, JavaScript actually trained a byte-pair tokenizer on a small built-in corpus — counting pairs, merging the most frequent, 120 times. Type anything below and watch the frozen merge rules chop it into tokens with real IDs.
This playground uses character-level BPE (the original 2016 Sennrich formulation) so tokens display cleanly. Production tokenizers (GPT-2 and later) use byte-level BPE: the base vocabulary is all 256 byte values, which makes out-of-vocabulary input mathematically impossible — any text, any emoji, any binary blob can always be represented as bytes. The trade-off: merges can split a multi-byte UTF-8 character in half, so decoding must reassemble bytes before interpreting them as text. Same algorithm, different alphabet.
Built as a single self-contained file · diagrams render locally, no CDN · back to the takeaway ↑