Our token has been checked in and handed a boarding pass — its ID. But an ID number carries no meaning, so the next step is to give the token its real identity: a vector that encodes what the token is like. This is where the token puts on its uniform.
A token ID like 3797 is just a row number — it has no size, no neighbours, no meaning. The number 3797 is not "bigger" or "closer" to 3798 in any useful sense. To compute with a token, the model replaces its ID with a vector of real numbers it can add, scale and multiply. That vector is the token's embedding.
The embedding table
The model holds one big lookup table: a matrix with one row per vocabulary token and one column per dimension of the model. If the vocabulary has tokens and the model works in dimensions, has shape . A common size is and somewhere from 768 to many thousands.
"Embedding a token" is nothing more than selecting its row:
To make the sizes concrete: GPT-2's smallest model used and a vocabulary of about 50,000, so its embedding table held roughly 38 million numbers in that one matrix; its largest version used . Bigger models go wider still.
Crucially, the numbers in are learned parameters, trained by the same gradient descent as the rest of the network. The model discovers, over training, what each token's vector should be so that next-token prediction works well.
One intuition helps here. Imagine each of the dimensions as a question asked of every token: are you a noun? a verb? something to do with people? with places? with emotion? do you usually end a sentence? The token's vector is its list of answers — and because there are hundreds of such questions, the vector captures a rich, many-sided picture of the token's role. No one writes these questions down; training invents whatever questions turn out to be useful, and we rarely know what any single dimension "means." But the picture — every token described by its answers to a long, learned questionnaire — is a faithful one.
Why vectors, and why they carry meaning
Once tokens are points in a -dimensional space, two things become possible that integers never allowed:
- Similarity has a meaning. Tokens used in similar ways drift to nearby points. The vectors for
kingandqueen, orParisandLondon, end up close, because treating them similarly helped prediction. Distance in this space is learned semantic similarity. - Directions can carry features. Because the vectors are added and transformed linearly throughout the model, particular directions in the space come to stand for properties (tense, plurality, sentiment, topic). The model manipulates meaning by moving vectors along these directions.
None of this is hand-designed. It emerges because a good geometry lowers the loss.
The vector produced here is the token's first entry in what is often called the residual stream — the running -dimensional vector at each position that every later block reads from and adds to. Attention and feed-forward layers do not replace this vector; they add corrections to it. Keep that picture: the embedding is the starting value of a vector that the whole stack gradually refines.
A whole sequence at once
A prompt of tokens becomes embedding vectors — an matrix, one row per position. This matrix is what enters the first transformer block. Everything from here on operates on these vectors in parallel, mixing information between them (attention) and refining each one (feed-forward).
One detail to flag now
Notice that the embedding of cat is the same vector wherever cat appears — first word or last, the lookup returns the identical row. The embedding table knows what the token is but nothing about where it sits. Attention, as we will see, is also position-blind. So a sequence of embeddings alone cannot tell "the cat sat" from "sat cat the". Fixing that is the job of positional information, which we add to these vectors in Module 5.
These embeddings — one learned vector per token — are the raw material. The next chapter opens the block that transforms them.
Tap words to compare their similarity and try vector analogies such as king − man + woman.