Tokenization: Text into Tokens

Before a model can do arithmetic on language, text is chopped into tokens from a fixed vocabulary. Subword tokenization is the quiet compromise that makes the vocabulary both small and complete.

Large Language Models: From Transformers to Frontier Models

To make this module concrete, we are going to follow one token on its journey through the whole model — picking it up as raw text here, and setting it down four chapters from now as a predicted next word. Think of it as watching a single passenger pass through an airport: first they are checked in and given a boarding pass, then a uniform of numbers, then they board the long train of transformer blocks, and finally they are read out at the gate. This chapter is check-in.

A neural network multiplies numbers; it has no idea what a letter is. So the very first thing that happens to your prompt, before any attention or matrix multiply, is that the text is cut into tokens and each token is looked up as a number. This chapter is about that cut — why it is done the way it is, and why the choice matters more than it first appears.

Three ways to cut text

Imagine you must turn text into a sequence of items drawn from a fixed list — the vocabulary. There are three natural granularities.

  • Characters. The vocabulary is tiny (a few hundred symbols) and can spell anything, but sequences become very long and each token carries almost no meaning on its own. The model spends its effort re-learning spelling.
  • Words. Each token is meaningful and sequences are short, but the vocabulary is enormous and can never be complete. Every new name, typo, or joined word ("catdog") is an out-of-vocabulary hole the model cannot represent.
  • Subwords. A middle path: common words stay whole, and rare words break into reusable pieces. "tokenization" might become token + ization; "unhappiness" into un + happ + iness. The vocabulary stays a manageable size (typically 30,000–150,000) yet can spell any string, because in the worst case it falls back to single characters.

Every modern language model uses subword tokenization. It is the compromise that makes the vocabulary both bounded and complete.

How the pieces are chosen

The vocabulary is not hand-written — it is learned from a large text corpus before training, by a family of algorithms (Byte-Pair Encoding and its relatives). The intuition is simple and worth holding onto:

Start from individual characters. Repeatedly find the pair of adjacent symbols that occurs most often together, and merge it into a new single token. Stop when you have as many tokens as you wanted.

Frequent sequences like th, ing, the, and whole common words get merged early and become single tokens. Rare sequences never get merged and stay as small pieces. The result is that common text is cheap (few tokens) and rare text is expensive (many tokens) — which is exactly what you want.

Spaces live inside tokens

In most tokenizers the leading space is part of the token: Paris (with a space) and Paris (without) are different tokens. This is why you often see tokens written with a visible leading space. It lets the model handle word boundaries without a separate "space" token everywhere.

From tokens to IDs

Once the text is split, each token is replaced by its index in the vocabulary — just a row number. The sentence

The cat sat

might become the token list [" The", " cat", " sat"] and then the integer list [464, 3797, 3332]. Those integers are all the model ever sees of your text. Everything downstream operates on them.

A sentence split into subword tokens, each token mapped to an integer ID from the vocabulary
Text is cut into subword tokens, then each token becomes its row number in the vocabulary. The model only ever sees the integers.

Why the cut matters

Tokenization is easy to dismiss as plumbing, but it quietly shapes everything:

  • Context length is counted in tokens, not words. A model with a 128k-token window holds roughly 128k tokens ≈ 90,000–100,000 English words, but far fewer for a language whose script the tokenizer handled poorly.
  • Cost and speed scale with tokens. A prompt that tokenizes into more pieces costs more and runs slower, and the same text can tokenize very differently across languages.
  • Some tasks are hard because of it. Character-level questions ("how many r's are in strawberry?") are genuinely awkward for a model that sees straw + berry, not letters. The difficulty is in the tokenizer, not the reasoning.

We now have a sequence of integers. The next chapter turns each integer into a vector the network can actually compute with.

Try it yourself
Transformer Lab: tokenization →

Train a tiny byte-pair encoder merge by merge and see how common and rare words are split.

EasyTokenization

Why do modern models use subword tokens instead of whole words?

MediumTokenization

A tokenizer merges the most frequent adjacent pairs first. What does this imply about how many tokens common vs rare words take?

MediumTokenization

Why is 'how many r's are in strawberry?' unusually hard for a language model?