The transformer architecture from 2017 still sits at the centre of every large language model. What has changed since is a set of upgrades, each one an answer to a practical pressure that shows up only when you make the model big and actually serve it to users. This chapter names those pressures, so that every later module has an obvious reason to exist.
The base everyone starts from
The original transformer block has two moving parts, stacked many times:
- Attention — every position looks at every other position and pulls in what is relevant.
- A feed-forward network (FFN) — each position is transformed on its own by a small two-layer network.
Around each part sits a residual connection and a normalization step, which keep the signal and gradients stable as the stack grows deep. That is the whole base model, and Modules 2–4 build it.
If you stopped there, you would have a working language model. You would also have one that is slow to generate from, expensive to scale, blind to word order unless you help it, and wasteful in its arithmetic. The rest of the field's progress addresses those four problems in turn.
Pressure 1 — Order: attention is blind to position
Attention treats its input as a set, not a sequence. Shuffle the words and, without help, the raw attention computation gives the same answer. Since language is all about order, we must inject position into the inputs. Module 5 traces the journey from simple counters to sinusoidal encodings and finally rotary positional encodings (RoPE), the scheme most modern models use.
Pressure 2 — Memory at inference: the KV cache grows
When a model generates text token by token, it can avoid recomputing the past by storing the keys and values of every previous token — the KV cache. This makes generation fast, but the cache grows with every token and every layer, and at long context lengths it, not the model's weights, becomes the thing that fills your GPU memory.
A whole family of attention variants exists to shrink that cache:
- Multi-Query Attention (MQA) — all heads share one set of keys and values.
- Grouped-Query Attention (GQA) — a middle ground: groups of heads share.
- Multi-Head Latent Attention (MLA) — store a small compressed vector per token and reconstruct keys and values on the fly.
Module 6 is devoted to this family. It is the first place the course crosses from "the classic transformer" into "modern advancements".
Pressure 3 — Scale: making the FFN bigger without paying for it
The cheapest way to make a model more capable is to give it more parameters — but more parameters normally means proportionally more compute for every token. Mixture-of-Experts (MoE) breaks that link. It replaces the single FFN with many expert FFNs and a router that sends each token to only a couple of them. The model can hold ten times the parameters while each token still only pays for a small slice. Module 7 covers routing, the load-balancing tricks that stop a few experts from hogging all the traffic, and the fine-grained and shared-expert designs used today.
Pressure 4 — Efficiency of training signal and arithmetic
Two more upgrades round out a modern system:
- Multi-Token Prediction (MTP) — instead of training the model to predict only the very next token, also have it predict a couple of tokens ahead. This gives a richer training signal and enables faster generation. Module 8 covers it.
- Quantization and low precision — storing and multiplying numbers in 16, 8, or even fewer bits instead of 32. Done carefully, this roughly halves or quarters memory and speeds up every matrix multiply with little loss in quality. Module 9 covers mixed precision, fine-grained quantization, and the precision choices that keep training stable.
The map
Keep this map in mind. Whenever a later module introduces something that looks like a clever trick, it is really an answer to one of these four pressures: order, inference memory, scale, or efficiency.
Modules 2–5 give you a complete, working transformer language model. Modules 6–9 are the upgrades that make it fast, big and cheap enough to be a frontier system. If you only want to understand how an LLM works, the first half suffices; if you want to understand how a competitive one is built, the second half is where the real engineering lives.