The last three modules made models cheaper by changing what they compute — compressing the cache, activating few experts, enriching the objective. This module changes something more basic: the precision of the numbers themselves. A model is billions of numbers multiplied and added; store and multiply them in fewer bits and everything gets smaller and faster. That is quantization, and it is the last major efficiency lever in the course.
A number is a budget of bits
Every parameter, and every activation, is a number held in some format, and the format decides how many bits it occupies. The default in deep learning has long been 32-bit floating point (FP32): one sign bit, several exponent bits (which set the range — how large or small the number can be), and the rest as mantissa bits (which set the precision — how finely values are distinguished).
The memory cost is just arithmetic. A 70-billion-parameter model stored in FP32 (4 bytes each) needs about 280 GB just for weights; in a 16-bit format, about 140 GB; in an 8-bit format, about 70 GB. Halving the bits halves the memory — and smaller numbers also move through the hardware faster, so the matrix multiplies speed up too. For models this large, the format is not a detail; it is the difference between fitting on the hardware and not.
The common formats, and what they trade
A few formats recur throughout this module, and the whole game is the balance between range and precision for a given bit budget:
- FP32 — 32 bits. Wide range, high precision. The safe default; expensive.
- FP16 — 16 bits. Half the memory, but both range and precision shrink — it can overflow or underflow on values FP32 handled easily.
- BF16 (brain float 16) — also 16 bits, but spends them differently: it keeps FP32's range and gives up precision instead. Because training blows up when values fall out of range more than when they lose a little precision, BF16 is usually the better 16-bit choice.
- FP8 — 8 bits. A quarter of FP32's memory and very fast, but little range and little precision. Usable only with care.
- INT8 — 8-bit integers, no fractional part; values map onto a small fixed range like −127…127.
Notice the pattern: at a fixed number of bits you choose how to split them between range and precision, and at fewer bits you have less of both. The lower you go, the more the format needs help to stay usable.
Why giving up precision is acceptable
The obvious worry: if we represent each number less precisely, won't the model get worse? A little — but far less than you would fear, because a neural network is an average of a vast number of small operations, and rounding errors in different places tend to wash out rather than compound. The model does not need every weight to 7 decimal places; it needs the overall computation to come out about right.
A picture makes it concrete. Take a photograph and reduce it to a handful of colours. Viewed from across the room it looks the same; only on close inspection do you see the banding. Quantization does this to a model's numbers: coarser values, nearly the same behaviour, a fraction of the storage. Quality dips slightly; memory and speed improve a lot. For large models that is a trade well worth making.
How a number is actually quantized: scaling
To squeeze a set of real-valued numbers into a low-bit format, you scale them. Take the group's largest magnitude, divide every value by it (so everything lands in −1…1), then multiply by the largest value the target format can hold. The values now span the format's full range and can be rounded to it; to read them back you divide by the same factor. This scaling step is the heart of quantization — and, as the next chapter shows, how you choose the scale is exactly where the clever engineering lives, because one careless scale can wreck the precision of a whole tensor.
Quantization cuts across everything earlier in the course. The weights of the attention and feed-forward layers, the activations flowing between them, even the KV cache and the experts — all are numbers that can be stored and multiplied in low precision. So this lever compounds with the others: a mixture-of-experts model with a compressed cache also quantized is smaller and faster on all three counts at once.
The promise is clear — big savings for a small quality cost. But FP8, the aggressive format that delivers the biggest savings, is genuinely fragile: too little range and precision to use naively across a whole model. The next chapter is the set of techniques that make low precision actually work — using it only where it is safe, and scaling it carefully where it is not.