Course Overview

Build a modern large language model from the ground up: tokens and embeddings, the attention mechanism, positional encodings, and the efficiency and scaling ideas behind today's frontier systems — KV caching, grouped-query and latent attention, mixture-of-experts, multi-token prediction and low-precision training.

Large Language Models: From Transformers to Frontier Models

A large language model is, underneath, a next-word guesser trained at enormous scale. This course takes that plain idea and builds it out into the full machinery of a modern system: how text becomes numbers, how the attention mechanism lets a model weigh every word against every other, how a model knows the order of its input, and — once the basic transformer is in place — the efficiency and scaling ideas that separate a textbook model from a frontier one.

We start from first principles and code each idea up conceptually before layering on the next. By the end you will be able to read a modern model's architecture description and recognise every block in it: the residual stream, multi-head attention and its memory-saving variants (multi-query, grouped-query and latent attention), rotary positional encodings, the mixture-of-experts feed-forward layer, multi-token prediction, and the low-precision arithmetic that makes training affordable.

The course is written for a wide audience. An undergraduate who has met gradient descent and matrix multiplication can follow every chapter; a postgraduate or working engineer will find the later modules on efficient attention, sparsity and quantization go all the way to what current systems actually do.

What you should already know

You will be comfortable if you have seen:

  • Vectors and matrices — multiplication, dot products, transposes.
  • The basics of neural networks — a neuron, a weight, a non-linear activation, and training by gradient descent and backpropagation.

If any of that is rusty, our Deep Learning course covers it from scratch, and we link to the exact lessons where they are needed — for example the backpropagation algorithm, output and loss functions and activation functions.

How the course is organised

The modules move in one straight line, each building on the last:

  1. Orientation — what a language model is and how a modern one is put together.
  2. The transformer: journey of a token — the end-to-end path from text to a predicted next word.
  3. Attention from first principles — self-attention, scaling, and causal masking.
  4. Multi-head attention — many attention heads working in parallel.
  5. Positional information — why order must be injected, up to rotary encodings (RoPE).
  6. Efficient attention at inference — the KV cache and the family of tricks that shrink it.
  7. Mixture-of-experts — scaling the feed-forward layer without scaling the cost.
  8. Beyond next-token training — multi-token prediction.
  9. Numerical precision and quantization — training and serving in low precision.
  10. Putting it together — the anatomy of a frontier model, and where to go next.

Every chapter ends with a few questions to check understanding. Many chapters also link to an experiment in the Transformer Lab, where you can run the idea yourself: tokenization, embeddings, attention, positional encodings, sampling, the KV cache, mixture-of-experts and quantization.

About the source

The structure of this course was inspired by the excellent open lecture series "Build DeepSeek from Scratch" by Vizuara, which walks through a modern model end to end. We have re-taught the material in our own words and generalised it: the ideas here — attention variants, mixture-of-experts, multi-token prediction, quantization — belong to the whole field of transformer-based language models, not to any single system. Where a particular model popularised an idea, we say so, but the focus stays on the concept.