Inside a Transformer Block

The tall stack is one block repeated. Each block does two things — mix information across positions (attention) and refine each position (a feed-forward network) — wrapped in residual connections and normalization that keep a deep stack trainable.

Large Language Models: From Transformers to Frontier Models

Our token now has its uniform — a single vector that says what it means and where it sits. It is ready to board the long train. Each carriage is a transformer block, and the token must ride through every one of them, changing a little in each, before it reaches the far end. This chapter opens up a single carriage.

We have a sequence of vectors, one per token, entering the stack. The stack is not a zoo of different layers — it is one block repeated dozens of times, each with its own weights. Understand the block and you understand the model. It has exactly two working parts and two supporting ones.

The two working parts

1. Attention — mixing across positions. On its own, a token's vector only knows about itself. Attention is the one place in the whole architecture where positions talk to each other: it lets each token look at the others and pull in what is relevant. When it needs to know what it refers to, attention is how it reaches back to cat. Module 3 builds this mechanism from scratch; for now, treat it as "each position gathers a weighted mix of information from the other positions."

2. Feed-forward network (FFN) — refining each position. After gathering context, each token's vector is passed, independently, through a small two-layer network:

FFN(h)=W2 σ(W1h+b1)+b2.\text{FFN}(h) = W_2\,\sigma(W_1 h + b_1) + b_2.

It expands the vector to a larger inner width, applies a non-linear activation σ\sigma, and projects back. The same FFN weights are applied at every position, but each position is processed alone — no mixing happens here. If attention decides what to combine, the FFN decides what to make of it. This is an ordinary neural network layer; the activation and its role are the same ones covered in the Deep Learning course's activation functions lesson.

A useful slogan: attention moves information between tokens; the feed-forward network thinks about each token.

The two supporting parts

Stack fifty of those working parts naively and the model will not train — the signal explodes or vanishes as it passes through. Two devices make depth survivable.

Residual connections. Instead of replacing its input, each sub-layer adds a correction to it:

h←h+Attention(h),h←h+FFN(h).h \leftarrow h + \text{Attention}(h), \qquad h \leftarrow h + \text{FFN}(h).

This is the residual stream from the last chapter made literal. The + matters enormously: it gives gradients a clean, uninterrupted path from the top of the stack all the way back to the embeddings, so even a very deep model can be trained. (The Deep Learning course's backpropagation lesson is why this path matters.) It also means each block only has to learn a small adjustment to the running vector, not rebuild it from nothing.

Normalization. Before each sub-layer, the vector is normalized — rescaled so its numbers have a controlled size. This keeps the inputs to attention and the FFN in a stable range no matter how deep the block sits. Modern models normalize before each sub-layer ("pre-norm"), which makes training especially stable.

The block, assembled

Putting the four parts together, one block does:

h←h+Attention(Norm(h))h \leftarrow h + \text{Attention}(\text{Norm}(h)) h←h+FFN(Norm(h))h \leftarrow h + \text{FFN}(\text{Norm}(h))

and then the next block does the same with its own weights, and so on up the stack. (During training a third, minor device, dropout, randomly zeroes a fraction of values inside each sub-layer to discourage over-reliance on any one path; it is switched off when the model is actually used. Many recent models drop it entirely, so we treat it as an optional training aid rather than part of the architecture.)

How tall is the stack? GPT-2 came in four sizes — 12, 24, 36, and 48 blocks — and large modern models run to the high dozens or more. Our token must pass through every block in turn, so a 48-block model puts it through this same carriage 48 times, each with different weights.

One transformer block: the residual stream passing through a norm then attention with an add, then a norm then a feed-forward network with an add
One block. Norm → attention → add, then norm → feed-forward → add. The stack is this, repeated, each copy with its own weights.

Why this shape works

  • Attention gives the model an all-to-all communication step: any token can, in principle, draw on any other.
  • The FFN gives it per-token computation and most of its raw parameter count — it is where a great deal of the model's stored knowledge lives.
  • Residuals turn a deep stack into a sequence of small, learnable refinements to a shared vector.
  • Normalization keeps every one of those refinements numerically well-behaved.

Depth lets the model build meaning in stages: early blocks resolve local structure (whose it is whose, basic syntax), later blocks assemble higher-level meaning. No single block has to do everything.

The one picture to remember

A transformer block reads the residual stream, computes a correction (via attention, then the FFN), and adds it back. The residual stream is a wide conversation the whole stack contributes to, one small edit at a time.

The vector at each position is now refined by the whole stack. The last chapter of this module turns the top-of-stack vector back into an actual next-token prediction.

EasyArchitecture

Attention and the feed-forward network divide one labour between them. What does each do?

MediumArchitecture

Each sub-layer adds its output to its input rather than replacing it. Why is that '+' so important for deep models?

MediumArchitecture

Roughly where in a block does most of the model's parameter count and stored knowledge sit?