How a Modern LLM Is Built

A map of the whole course: the base transformer everyone shares, and the four pressures — memory, compute, scale and precision — that produced every modern upgrade to it.

Large Language Models: From Transformers to Frontier Models

The transformer architecture from 2017 still sits at the centre of every large language model. What has changed since is a set of upgrades, each one an answer to a practical pressure that shows up only when you make the model big and actually serve it to users. This chapter names those pressures, so that every later module has an obvious reason to exist.

The base everyone starts from

The original transformer block has two moving parts, stacked many times:

  • Attention — every position looks at every other position and pulls in what is relevant.
  • A feed-forward network (FFN) — each position is transformed on its own by a small two-layer network.

Around each part sits a residual connection and a normalization step, which keep the signal and gradients stable as the stack grows deep. That is the whole base model, and Modules 2–4 build it.

If you stopped there, you would have a working language model. You would also have one that is slow to generate from, expensive to scale, blind to word order unless you help it, and wasteful in its arithmetic. The rest of the field's progress addresses those four problems in turn.

Pressure 1 — Order: attention is blind to position

Attention treats its input as a set, not a sequence. Shuffle the words and, without help, the raw attention computation gives the same answer. Since language is all about order, we must inject position into the inputs. Module 5 traces the journey from simple counters to sinusoidal encodings and finally rotary positional encodings (RoPE), the scheme most modern models use.

Pressure 2 — Memory at inference: the KV cache grows

When a model generates text token by token, it can avoid recomputing the past by storing the keys and values of every previous token — the KV cache. This makes generation fast, but the cache grows with every token and every layer, and at long context lengths it, not the model's weights, becomes the thing that fills your GPU memory.

A whole family of attention variants exists to shrink that cache:

  • Multi-Query Attention (MQA) — all heads share one set of keys and values.
  • Grouped-Query Attention (GQA) — a middle ground: groups of heads share.
  • Multi-Head Latent Attention (MLA) — store a small compressed vector per token and reconstruct keys and values on the fly.

Module 6 is devoted to this family. It is the first place the course crosses from "the classic transformer" into "modern advancements".

Pressure 3 — Scale: making the FFN bigger without paying for it

The cheapest way to make a model more capable is to give it more parameters — but more parameters normally means proportionally more compute for every token. Mixture-of-Experts (MoE) breaks that link. It replaces the single FFN with many expert FFNs and a router that sends each token to only a couple of them. The model can hold ten times the parameters while each token still only pays for a small slice. Module 7 covers routing, the load-balancing tricks that stop a few experts from hogging all the traffic, and the fine-grained and shared-expert designs used today.

Pressure 4 — Efficiency of training signal and arithmetic

Two more upgrades round out a modern system:

  • Multi-Token Prediction (MTP) — instead of training the model to predict only the very next token, also have it predict a couple of tokens ahead. This gives a richer training signal and enables faster generation. Module 8 covers it.
  • Quantization and low precision — storing and multiplying numbers in 16, 8, or even fewer bits instead of 32. Done carefully, this roughly halves or quarters memory and speeds up every matrix multiply with little loss in quality. Module 9 covers mixed precision, fine-grained quantization, and the precision choices that keep training stable.

The map

A base transformer block in the centre, with four labelled arrows pointing to its upgrades: positional encodings for order, KV-cache attention variants for inference memory, mixture-of-experts for scale, and quantization and multi-token prediction for efficiency
The base transformer and the four pressures that produced its modern upgrades. Each pressure is a module of this course.

Keep this map in mind. Whenever a later module introduces something that looks like a clever trick, it is really an answer to one of these four pressures: order, inference memory, scale, or efficiency.

You do not need all of it to start

Modules 2–5 give you a complete, working transformer language model. Modules 6–9 are the upgrades that make it fast, big and cheap enough to be a frontier system. If you only want to understand how an LLM works, the first half suffices; if you want to understand how a competitive one is built, the second half is where the real engineering lives.

EasyArchitecture

Attention treats its input as a set. What problem does that create for language, and which module fixes it?

MediumArchitecture

Mixture-of-Experts lets a model have far more parameters without proportionally more compute per token. How?

MediumInference

Why does the KV cache, rather than the model's weights, often become the memory bottleneck for long conversations?