System design · build it to learn it

Build a Tokenizer

The unglamorous front door of every large language model.

"How does an LLM turn text into tokens? Walk me through how a tokenizer works — and how would you build one?"

On this page: takeaway · requirements · capacity math · architecture · deep dives · API + data model · trade-offs · failure modes · what I'd actually build · interview tips · live BPE playground

🎯 The takeaway, first

A tokenizer is a compression scheme wearing a job-interview costume. Neural nets only eat numbers, so text has to become a list of integers first. Byte-pair encoding (BPE) learns which character chunks appear together most often in a big corpus, merges them into "tokens," and freezes that rulebook. From then on, encoding is just applying the frozen merges, greedily, in order — no learning at inference time, ever.

The single most interview-useful sentence you can say: "Tokenization is lossy-looking but must be lossless — decode(encode(x)) == x for every input, including emoji and binary garbage — and the vocab is fixed at train time, so inference just looks things up."

🚫 Misconception, busted

"A tokenizer splits text into words." Nope. It splits text into statistically frequent chunks. unbelievable becomes un + believ + able — pieces that aren't words in any dictionary but were frequent neighbors in the training corpus. Spaces, punctuation, and even parts of emoji get the same treatment. If your mental model is "fancy word splitter," rebuild it as "a zip file's dictionary that learned English."

01Requirements

Functional

  • encode(text) → int[]: text in, token IDs out. Fast — it runs on every request.
  • decode(ids) → text: IDs back to text, byte-for-byte identical to the input.
  • Handle any Unicode: emoji, CJK, zalgo text, binary junk. Never crash, never silently drop a character.
  • Deterministic: same text → same IDs, on every machine, forever.

Non-functional

  • Vocab ~50k–200k entries. Each entry costs embedding memory (see math below).
  • Encode throughput: megabytes per second, single thread; p99 under ~50ms for a page of text.
  • The tokenizer is immutable after training — version it like a database migration.
  • Train-time and inference-time tokenizers must be bit-identical, or the model reads gibberish.

02Back-of-envelope math

Analogy first: every token ID is a row number in a giant lookup table (the embedding matrix). Bigger vocab = shorter sequences but a fatter table. The math tells you where the pain lives.

QuantityAssumptionResult
Embedding matrix100k vocab × 12,288 dims × fp16 (2 bytes)2.46 GB — per matrix. Untied input/output = ~5 GB just to name the tokens.
English density~4 characters per token1M tokens ≈ 4M chars ≈ 700k words ≈ ~2,800 pages
Cost of a prompt$2.50 / 1M input tokens, 500-token request500 / 1e6 × $2.50 = $0.00125 — a tenth of a cent
Cost of a corpus1 GB of text ≈ 250M tokens at $2.50/1M~$625 to run one GB through a frontier model
Merge table size100k merges × ~2 short stringsA few MB — the rulebook is tiny; the embedding table is the monster

Takeaway after the math: vocab size is a three-way tug — fewer tokens per request (cheaper inference, longer context fit) vs. embedding memory vs. rare-token quality. That's the whole trade-off; everything else is details.

03Architecture

Think of it like a post office sorting line: normalize the envelope, slice it into rough chunks, then stamp each chunk with its ID number from a fixed codebook.

flowchart LR
    A["Raw text"] --> B["Normalizer
unicode NFC, strip control chars"] B --> C["Pre-tokenizer
regex split on spaces and punctuation"] C --> D["BPE merge engine
apply learned merges in order"] E[("merges.txt
frozen rulebook")] --> D F[("vocab.json
token to id")] --> D D --> G["Token IDs"] G --> H["Embedding lookup
id to vector"]
Go deeper: why the pre-tokenizer regex matters

The regex that does the rough split is doing real semantic work. GPT-4's famous pre-tokenizer regex deliberately splits numbers into ≤3-digit chunks (123456 → 123 456) and keeps contractions together ('ll, 're). Change this regex and you change what the model finds easy: three-digit number chunks are part of why models do okay-ish at arithmetic. It's the cheapest hyperparameter in the whole pipeline — one line of regex, enormous downstream effects.

04Component deep-dives

Training: learning the merges

Training is the only "smart" part, and it's just counting. Split the corpus into words, count every adjacent pair of symbols, merge the most frequent pair into a new symbol, repeat. Each merge is recorded in order — that ordered list is the tokenizer.

sequenceDiagram
    participant U as You
    participant T as Trainer
    participant C as Corpus
    participant V as Vocab store
    U->>T: train(corpus, num_merges=120)
    T->>C: split into words, then chars
    loop 120 times
        T->>T: count every adjacent pair
        T->>T: best pair (a, b) becomes token ab
        T->>V: append merge rule in order
    end
    T->>U: merges.txt + vocab.json (frozen)

Worked mini-example on the word low appearing often next to lower: early merges learn l+o → lo, then lo+w → low, then low+e → lowe. Frequent words collapse into single tokens; rare words stay shattered into pieces. That's the whole trick — frequency-proportional compression.

Encoding: applying the frozen rulebook

Encoding learns nothing. Take the input, split into words/chars, then walk the merge list in the exact order learned and fuse every occurrence of each pair. Order matters: merge #3 can only fire on the output of merges #1–2.

sequenceDiagram
    participant C as Client
    participant T as Tokenizer
    participant M as Merge table
    C->>T: encode("hello world")
    T->>T: pre-tokenize into words
    T->>M: load merges in learned order
    loop for each merge (a, b)
        T->>T: fuse every adjacent a b
    end
    T->>C: [1045, 995] token IDs
Go deeper: why not just take the longest match greedily?

Some tokenizers (WordPiece) do use greedy longest-match at encode time. BPE's merge-order application is subtly different and guarantees consistency with training: the merges you apply are exactly the merges training discovered, in the same order. Greedy longest-match can produce segmentations training never saw, which slightly shifts the distribution the model was trained on. In practice both work; BPE's way is more faithful to its own training.

05API + data model

API surface (it's a library, not a service)

train(corpus_paths, vocab_size) -> None
  # writes merges.txt, vocab.json, config.json

encode(text: str) -> list[int]
decode(ids: list[int]) -> str
  # invariant: decode(encode(x)) == x

encode_batch(texts) -> list[list[int]]
  # for throughput: parallelize here

Data model (three small files)

merges.txtordered pairs
vocab.jsontoken → id
config.jsonregex + version

The merge file is append-only and ordered — line 1 fired before line 2 at training, and must fire first at encoding. Version all three together; a v2 tokenizer reading v1 merges is a silent-corruption bug.

06Trade-offs

DecisionOption AOption BVerdict
Vocab sizeSmall (30k): tiny embeddings, but text shatters into many tokens — longer sequences, pricier inferenceLarge (200k): compact sequences, but GBs of embedding memoryMatch the domain: code/multilingual wants bigger
Base symbolsCharacters: clean display, but any unseen char is OOVBytes (256 base): zero OOV ever — any input encodableBytes for production (this is what GPT does)
AlgorithmBPE: simple, fast, merge-order encodeUnigram: probabilistic, better compression, slower trainingBPE unless you can cite a measured win
Pre-tokenizerAggressive split: cleaner tokens, more of themLax split: fewer tokens, weirder boundariesSteal GPT-4's regex; it's battle-tested

07Failure modes

08What I'd actually build

Opinionated and concrete — the stack I'd reach for tomorrow:

09Interview tips

10🔬 Live widget: a real BPE tokenizer, trained in your browser

This isn't a mockup. When the page loaded, JavaScript actually trained a byte-pair tokenizer on a small built-in corpus — counting pairs, merging the most frequent, 120 times. Type anything below and watch the frozen merge rules chop it into tokens with real IDs.

training…
–
vocab size (chars + merges)
–
merges learned

Tokens (green border = learned merge · red border = unseen char)

Press "Tokenize →"
–
characters
–
tokens
–
chars per token
–
cost @ $2.50/1M tokens

Learned merges (first 40 of 120, in learned order)

Go deeper: char-level vs byte-level BPE

This playground uses character-level BPE (the original 2016 Sennrich formulation) so tokens display cleanly. Production tokenizers (GPT-2 and later) use byte-level BPE: the base vocabulary is all 256 byte values, which makes out-of-vocabulary input mathematically impossible — any text, any emoji, any binary blob can always be represented as bytes. The trade-off: merges can split a multi-byte UTF-8 character in half, so decoding must reassemble bytes before interpreting them as text. Same algorithm, different alphabet.

v2026.10.03-02