Embeddings: Tokens into Vectors

Each token ID is swapped for a learned vector. Those vectors form the 'residual stream' the whole model reads from and writes to, and their geometry carries meaning.

Large Language Models: From Transformers to Frontier Models

Our token has been checked in and handed a boarding pass — its ID. But an ID number carries no meaning, so the next step is to give the token its real identity: a vector that encodes what the token is like. This is where the token puts on its uniform.

A token ID like 3797 is just a row number — it has no size, no neighbours, no meaning. The number 3797 is not "bigger" or "closer" to 3798 in any useful sense. To compute with a token, the model replaces its ID with a vector of real numbers it can add, scale and multiply. That vector is the token's embedding.

The embedding table

The model holds one big lookup table: a matrix EE with one row per vocabulary token and one column per dimension of the model. If the vocabulary has VV tokens and the model works in dd dimensions, EE has shape V×dV \times d. A common size is V≈100,000V \approx 100{,}000 and dd somewhere from 768 to many thousands.

"Embedding a token" is nothing more than selecting its row:

embedding(x)=E[x]⟶a vector in Rd.\text{embedding}(x) = E[x] \quad\longrightarrow\quad \text{a vector in } \mathbb{R}^{d}.

To make the sizes concrete: GPT-2's smallest model used d=768d = 768 and a vocabulary of about 50,000, so its embedding table held roughly 38 million numbers in that one matrix; its largest version used d=1600d = 1600. Bigger models go wider still.

Crucially, the numbers in EE are learned parameters, trained by the same gradient descent as the rest of the network. The model discovers, over training, what each token's vector should be so that next-token prediction works well.

One intuition helps here. Imagine each of the dd dimensions as a question asked of every token: are you a noun? a verb? something to do with people? with places? with emotion? do you usually end a sentence? The token's vector is its list of answers — and because there are hundreds of such questions, the vector captures a rich, many-sided picture of the token's role. No one writes these questions down; training invents whatever questions turn out to be useful, and we rarely know what any single dimension "means." But the picture — every token described by its answers to a long, learned questionnaire — is a faithful one.

Why vectors, and why they carry meaning

Once tokens are points in a dd-dimensional space, two things become possible that integers never allowed:

  • Similarity has a meaning. Tokens used in similar ways drift to nearby points. The vectors for king and queen, or Paris and London, end up close, because treating them similarly helped prediction. Distance in this space is learned semantic similarity.
  • Directions can carry features. Because the vectors are added and transformed linearly throughout the model, particular directions in the space come to stand for properties (tense, plurality, sentiment, topic). The model manipulates meaning by moving vectors along these directions.

None of this is hand-designed. It emerges because a good geometry lowers the loss.

Embeddings are where the residual stream begins

The vector produced here is the token's first entry in what is often called the residual stream — the running dd-dimensional vector at each position that every later block reads from and adds to. Attention and feed-forward layers do not replace this vector; they add corrections to it. Keep that picture: the embedding is the starting value of a vector that the whole stack gradually refines.

A whole sequence at once

A prompt of nn tokens becomes nn embedding vectors — an n×dn \times d matrix, one row per position. This matrix is what enters the first transformer block. Everything from here on operates on these nn vectors in parallel, mixing information between them (attention) and refining each one (feed-forward).

Token IDs indexing rows of an embedding table to produce a matrix of vectors, one row per token position
Each token ID selects a learned row from the embedding table. A sequence of n tokens becomes an n × d matrix — the input to the first block.

One detail to flag now

Notice that the embedding of cat is the same vector wherever cat appears — first word or last, the lookup returns the identical row. The embedding table knows what the token is but nothing about where it sits. Attention, as we will see, is also position-blind. So a sequence of embeddings alone cannot tell "the cat sat" from "sat cat the". Fixing that is the job of positional information, which we add to these vectors in Module 5.

These embeddings — one learned vector per token — are the raw material. The next chapter opens the block that transforms them.

Try it yourself
Transformer Lab: embeddings →

Tap words to compare their similarity and try vector analogies such as king − man + woman.

EasyEmbeddings

What exactly does 'embedding a token' compute?

MediumEmbeddings

Why can distance between embedding vectors represent meaning, when distance between token IDs cannot?

MediumEmbeddings

The embedding of ' cat' is identical no matter where ' cat' appears in the sentence. What problem does that create?