Where to Go Next

This course built the architecture — the model that predicts the next token. Turning that raw predictor into a helpful, honest, reasoning assistant is a second stage, post-training, sketched here with pointers so you know what to learn next.

Large Language Models: From Transformers to Frontier Models

You have reached the end of the architecture. You can now read a modern model's technical description and recognise every block — attention and its efficient variants, positional encodings, mixture-of-experts, multi-token training, quantization. That is a genuinely advanced place to be. But it is worth being clear about what this course did not cover, because a frontier assistant is more than its architecture, and knowing the gap is the start of closing it.

The gap: a predictor is not yet an assistant

Everything we built produces one thing: a good next-token predictor, trained on an ocean of text. Left there, the model will happily continue any text — including rudely, falsely, or unhelpfully — because "continue plausibly" is all it was ever asked to do. The polite, instruction-following, careful assistant you interact with is the result of a second stage of training applied on top of the pretrained model. The architecture is the engine; this stage is the steering.

This second stage is called post-training (or alignment / fine-tuning), and it is a large field in its own right. We name its main ideas here so the map is complete.

Instruction tuning

The first post-training step is supervised fine-tuning on examples of the behaviour you want: prompts paired with good responses. Show the model many "here is a request, here is a helpful answer" pairs and it learns to respond in that shape rather than merely continuing text. The architecture and the training objective (next-token prediction) are unchanged — only the data changes, from raw text to curated demonstrations. This alone turns a raw predictor into something that follows instructions.

Learning from preferences

Demonstrations only go so far; often it is easier to say which of two answers is better than to write the perfect one. Preference-based methods use exactly that signal. The classic version, reinforcement learning from human feedback (RLHF), trains a reward model to score responses the way humans would, then adjusts the language model to produce responses the reward model rates highly. Newer methods — such as direct preference optimization (DPO) — reach a similar end more simply, optimising directly from preferred-versus-rejected answer pairs without a separate reward model. Either way, the model is shaped toward responses people actually prefer: helpful, honest, harmless.

Why this is a whole separate subject

Preference optimization brings in machinery this course did not touch — reward modelling, reinforcement-learning objectives, and the stability tricks that keep such training from degrading the model. Modern "reasoning" systems push it furthest, using large-scale reinforcement learning with automatically-checkable rewards (did the code run? is the maths correct?) to teach the model to work through problems step by step. It deserves its own course; here we only mark that it exists and sits after everything we built.

Reasoning models

The most recent frontier is teaching models to reason — to produce a long internal chain of thought before answering, and to get better at it through reinforcement learning against rewards that can be checked automatically. This is why some models visibly "think" before responding. It is built entirely on the architecture of this course — the same transformer, attention, experts and all — with the reasoning ability layered on in post-training. The engine you now understand is exactly the engine those systems run on.

How to go deeper

A few honest suggestions for continuing:

  • Read a real model's technical report. You now have the vocabulary. Pick a recent open model's paper and read its architecture section — you will recognise the pieces, and the notation will finally make sense.
  • Implement one block from scratch. Understanding and building are different. Coding a small attention layer, a mixture-of-experts router, or RoPE in a few dozen lines turns recognition into real knowledge.
  • Study post-training next. Instruction tuning, RLHF/DPO, and reasoning are the natural sequel to this course and the other half of how a frontier assistant is made.
  • Follow the prerequisites back when needed. If any step felt shaky, the Deep Learning course covers the foundations — gradients, backpropagation, loss functions, optimizers — that everything here rests on.

A closing word

The throughline of this whole course has been a single, almost humble idea: predict the next token, at scale. Everything else — the attention that lets tokens share information, the positional encodings that restore order, the cache tricks and sparse experts and low-precision arithmetic — exists to make that one idea scale far enough to become remarkable. The architecture is not magic; it is a stack of understandable parts, each solving a nameable problem, composed with care. You now understand that stack. The rest is practice.

EasyPost-training

Why is a pretrained model not yet a usable assistant, and what stage fixes that?

MediumPost-training

How do preference-based methods like RLHF differ from plain instruction tuning?

MediumPost-trainingreasoning

Reasoning models 'think' before answering. How does this relate to the architecture built in this course?