Anatomy of a Frontier Model

A walk through a modern large language model with every part named — from a token entering the tokenizer to the trained, quantized, sparse, long-context system that predicts the next word. Everything in the course, assembled into one picture.

Large Language Models: From Transformers to Frontier Models

We have built a modern language model one idea at a time. This chapter assembles them. We will follow a single token through a frontier-scale model and name every component as we pass it — and you should recognise all of them now. If any name is unfamiliar, the module it came from is noted beside it.

The journey, end to end

Start with raw text and watch a token travel the whole system.

  1. Tokenization (Module 2). The text is cut into subword tokens and each becomes an integer ID. Nothing downstream sees letters — only tokens.
  2. Embedding (Module 2). Each ID selects a learned vector from the embedding table. A sequence becomes a matrix of vectors — the start of the residual stream.
  3. The block stack (Modules 2–9). The vectors rise through dozens of identical transformer blocks. Inside each block:
    • Normalization, then attention, added back to the residual stream; then normalization, then a feed-forward layer, added back (Module 2).
    • The attention is multi-head (Module 4), made causal by masking the future (Module 3), with rotary positional encoding rotating the queries and keys so order is felt (Module 5).
    • At inference the attention reads and writes a KV cache, kept small by latent attention with its content/position split (Module 6).
    • The feed-forward layer is not one network but a mixture-of-experts: a router sends the token to a few fine-grained experts plus always-on shared experts, kept evenly used by bias-based load balancing (Module 7).
    • The heavy matrix multiplies run in FP8, with sensitive parts and master weights kept high precision, scaled in fine-grained blocks and accumulated carefully (Module 9).
  4. Final norm and unembedding (Module 2). The top-of-stack vector is projected to a score for every vocabulary token — the logits.
  5. Softmax and sampling (Module 2). Logits become probabilities; a token is chosen (greedy, or with temperature and top-p). It is appended, and the loop runs again.
A labelled stack from tokenizer through embedding, a transformer block annotated with RoPE attention, latent KV cache, and a mixture-of-experts feed-forward, up to the unembedding and softmax
A modern model with every part named. The spine is the original transformer; the annotations are the course's efficiency and scaling ideas layered onto it.

How the pieces relate

Seen together, the upgrades are not a grab-bag — each answers a distinct pressure, and they compose cleanly because each lives in a different part of the block:

  • Order is handled inside attention, by RoPE (Module 5).
  • Inference memory is handled inside attention too, by compressing the KV cache (Module 6).
  • Scale is handled in the feed-forward layer, by sparse experts (Module 7).
  • Training signal is handled at the output, by predicting several tokens (Module 8).
  • Arithmetic cost cuts across all of them, by low precision (Module 9).

Because attention efficiency lives in the attention sub-layer, sparsity in the feed-forward sub-layer, and quantization in the number format, they stack without interfering: a single model is long-context and sparse and low-precision at once. That composability is why a frontier model can be, simultaneously, enormous in capacity and cheap per token.

What stayed the same, and what changed

Step back and the most striking thing is how much of the original 2017 transformer survives. The residual stream, the alternation of attention and feed-forward, the normalization, the softmax output — all unchanged. A researcher from the transformer's debut would recognise the skeleton immediately.

What changed is everything around the edges, and all of it in service of one goal: making the model affordable at scale. Latent attention, grouped and multi-query before it, rotary encodings, mixture-of-experts, multi-token training, FP8 — none of them alters what a language model fundamentally does. They alter what it costs to train and run one large enough to be good. That is the real story of modern LLM architecture: the core idea was right, and a decade of engineering has been spent making it cheap enough to scale to the point where its latent capabilities emerge.

The whole course in one sentence

A large language model is a next-token predictor built from a stack of attention-and-feed-forward blocks — and the art of a modern one is a set of efficiency moves (compressed attention memory, sparse experts, low precision, richer training) layered onto that stack so it can be scaled far enough to become genuinely capable.

You now understand a frontier model's architecture end to end. The final chapter points to the one large piece this course deliberately left out — turning this raw predictor into a helpful, aligned assistant — and where to learn it.

MediumSynthesis

RoPE, latent attention, mixture-of-experts, multi-token prediction and quantization all coexist in one model. Why don't they interfere?

MediumSynthesis

What part of the original transformer survives in a modern model, and what is the purpose of everything that changed?

EasySynthesis

Trace, in order, the five main stages a token passes through from text to a predicted next token.