What a Language Model Is

A large language model is a next-token predictor: given some text, it puts a probability on every possible next token. Everything else is built on that one idea.

Large Language Models: From Transformers to Frontier Models

Strip away the marketing and a large language model does exactly one thing: it reads some text and predicts what comes next. Not a paragraph, not an answer — just the next token, one at a time. The fluency you see in a chatbot is this single prediction, run over and over, each new token fed back in as part of the input.

Next-token prediction

Suppose the model has read:

The capital of France is

A language model does not "look up" the answer. It produces a probability distribution over its entire vocabulary — every token it knows — for what comes next:

Candidate next tokenProbability
Paris0.91
the0.03
located0.02
a0.01
… (tens of thousands more)…

The model picks a token from this distribution (here, almost certainly Paris), appends it to the text, and runs again on the longer input to get the token after that. This loop — predict, append, predict again — is called autoregressive generation. "Autoregressive" just means each output depends on the outputs so far.

If we write the text as a sequence of tokens x1,x2,…,xtx_1, x_2, \dots, x_t, the model computes

P(xt+1∣x1,x2,…,xt).P(x_{t+1} \mid x_1, x_2, \dots, x_t).

Everything in this course — attention, positional encodings, mixture-of-experts — exists to make this one conditional probability accurate and cheap to compute.

Why "token" and not "word"

Models do not work with words directly. Text is first chopped into tokens: common words become a single token, rare words split into pieces, and spaces and punctuation are tokens too. The word "tokenization" might become token + ization; "Paris" is likely one token Paris (with its leading space). We devote the next module's first chapter to this, but for now hold the picture: a token is a chunk of text, and the model's vocabulary is a fixed list of maybe 30,000 to 150,000 of them.

Where the "learning" happens

The model is a function with billions of adjustable numbers — its parameters or weights. Training shows it an ocean of text and, for every position, nudges the weights so that the probability it assigns to the token that actually came next goes up. That is ordinary supervised learning with a cross-entropy loss, the same loss you met for classification — except the "class" is the next token, and there are tens of thousands of classes.

If you want the mechanics of that loss and how gradients flow back through the weights, they are exactly the output and loss functions and backpropagation lessons from the Deep Learning course. Nothing about the loss changes for a language model; only the network in the middle does.

Two phases: pre-training and post-training

Reading the raw internet to become a good next-token predictor is pre-training. It produces a model that can continue text but does not necessarily answer politely or follow instructions. A second, much smaller phase — post-training (instruction tuning and preference optimization such as RLHF) — shapes that raw predictor into a helpful assistant. This course is about the architecture that both phases train; we point to post-training at the end.

Why "large" earns its name

The large in large language model is not marketing. Somewhere on the way from small models to enormous ones, something strange happens: abilities appear that were simply absent before. A small model trained on next-token prediction learns grammar and short-range patterns. Scale it up — more parameters, more data, more compute — and it begins, without being asked, to do arithmetic, unscramble words, translate, summarise and follow instructions. These are called emergent abilities: they switch on past a size threshold rather than fading in gradually.

The numbers tell the story of the race. GPT-2's largest version had about 1.5 billion parameters; GPT-3 jumped to 175 billion; frontier models today are estimated in the trillions. Model size has grown roughly exponentially for years, because each time the field crossed a threshold, new capabilities appeared "for free." That is why so much of this course is about scaling cheaply — a model is only as useful as the scale you can afford to train and run, which is exactly the pressure the later modules relieve.

Capability is a by-product of prediction

Nothing in the training objective says "learn to translate" or "learn arithmetic." The model is only ever rewarded for predicting the next token. But to predict the next token across the whole internet well, it turns out you must implicitly learn grammar, facts, reasoning patterns and a great deal more. Those skills are a by-product of becoming an excellent next-token predictor at scale — which is why this one humble objective produces such general systems.

The shape of the whole model

At the largest scale, every model in this course has the same three-part shape:

  1. In: turn tokens into vectors (an embedding lookup).
  2. Middle: a tall stack of identical transformer blocks that mix information between positions (attention) and transform it (a feed-forward network).
  3. Out: turn the final vector at each position back into a probability over the vocabulary.
High-level diagram: tokens flow into an embedding layer, up through a stack of transformer blocks, and out through a projection to next-token probabilities
Every model in this course has this shape. The rest of the course is about what lives inside the middle stack.

The genius and the difficulty are both in the middle stack. The next module walks a single token all the way up it.

EasyFundamentals

A model outputs a probability for every token in its vocabulary at each step. If the vocabulary has 100,000 tokens, how many numbers does it produce for one prediction?

EasyFundamentals

Why is generating a sentence described as running the same prediction many times, rather than once?

MediumTraining

Training a language model uses cross-entropy loss, the same loss as classification. What plays the role of the 'correct class' at each position?