Making Low Precision Work

FP8 is too fragile to use naively. Four engineering ideas make it safe: run only the heavy matrix multiplies in low precision while keeping sensitive parts high, scale in small blocks so one outlier can't spoil a tensor, accumulate sums in high precision, and compute scales on the fly.

Large Language Models: From Transformers to Frontier Models

The last chapter promised big savings from low precision and ended with a warning: the most aggressive format, FP8, is genuinely fragile — barely any range, barely any precision. Using it blindly across a model wrecks training. This chapter is the set of techniques that make it work anyway. They are engineering, not new theory, but they are what turns "8-bit would be nice" into a model that actually trains. Four ideas, each fixing a specific way naive FP8 fails.

Idea 1 — Mixed precision: low where it's safe, high where it's not

The first idea is to simply not quantize everything. A transformer's computation is dominated, by far, by big matrix multiplications — weights times activations, in the attention and feed-forward layers, both forward and backward. Those multiplies are where the time and memory go, so running them in FP8 captures almost all the savings.

Everything else stays in higher precision. In particular, a handful of components are sensitive — small errors in them hurt the whole model — and are kept in BF16 or FP32: the embeddings, the final output projection, the normalization steps, the attention's softmax, and the mixture-of-experts router. And crucially, the master copy of the weights and the optimizer's state are kept in FP32 throughout training; the FP8 versions used in the matmuls are made on the fly from those master weights and then discarded. The accumulated result of an FP8 matmul is also computed in high precision and only then stored in BF16.

The name for this — using low precision for the bulk and high precision for the sensitive parts — is mixed precision. It is the single most important idea here: you get most of FP8's speed and memory win while protecting the few places that cannot tolerate the error.

A matrix multiply with FP8 inputs and weights feeding a high-precision accumulator, while master weights and optimizer state sit in FP32 and sensitive modules stay in BF16
Mixed precision: the heavy matmuls run in FP8, but master weights, optimizer state, and sensitive modules (embeddings, output head, norms, attention, router) stay in high precision.

Idea 2 — Fine-grained scaling: one outlier shouldn't spoil the tensor

Recall how quantization scales a group of numbers by its single largest magnitude. That is fine until the group contains an outlier. Suppose most values are around 2–4 but one is 500. Divide everything by 500 to fit the range, and the small values become tiny fractions that FP8 cannot distinguish — they all round to nearly the same low-precision value and lose their information. One big number has ruined the precision of all the small ones.

The fix is to scale in small blocks rather than across the whole tensor. Split an activation vector into chunks (say 128 values each) and give each chunk its own scale factor; split a weight matrix into tiles (say 128×128) and scale each tile independently. Now an outlier only affects the scale of its own block — the other blocks keep scales matched to their own modest values, and their precision survives. This per-block scaling is called fine-grained quantization, and it is what makes FP8 robust to the uneven, outlier-prone distributions real models produce.

Why blocks help, in one line

A single scale per tensor is hostage to its largest value. Many small scales, one per block, mean each group is quantized to its own range — so a huge value in one corner can't drown out the fine detail everywhere else.

Idea 3 — Accumulate in high precision

A matrix multiply is a sum of many products. In FP8, each product is fine, but adding up hundreds of them in FP8 is not: the running total grows while each new term stays small, and in low precision the small terms fall off the bottom of the sum and are lost. Do the whole accumulation in FP8 and the result drifts badly.

So the multiplies are done in FP8 — fast — but the partial sums are accumulated in higher precision (promoted to FP32 as they add up). The hardware multiplies cheaply and keeps a precise running total on the side. This "high-precision accumulation" costs little and removes a large source of error, letting the low-precision multiplies be trusted.

Idea 4 — Choose the format, and scale on the fly

Two finishing touches round it out:

  • Spend the 8 bits wisely. Within FP8 there is still a choice of how many bits go to exponent (range) versus mantissa (precision). For the quantities in a model, leaning toward more mantissa (precision) where range allows gives better results — a small format choice that matters at the margins.
  • Compute scales dynamically. The scale factor for each block is best recomputed on the fly from the block's current values, rather than reusing a stale, pre-set scale. As training shifts the distribution of values, an up-to-date ("online") scale always matches the data, avoiding slow drift between the numbers and the scale meant to fit them.

Putting it together

None of these ideas is dramatic alone; together they are what lets a frontier-scale model train largely in 8-bit. Run the heavy matmuls in FP8 for speed and memory (mixed precision), but keep master weights, optimizer state and sensitive modules high (mixed precision again); scale in small blocks so outliers stay local (fine-grained); accumulate sums in high precision so nothing falls off the bottom; pick a precision-leaning format and keep the scales current. The payoff is roughly halved memory and markedly faster training versus BF16, with quality close enough that the trade is worth it at scale.

That completes the efficiency story. Over four modules the model has been made cheaper to run (efficient attention), cheaper to scale (mixture-of-experts), richer to train (multi-token prediction) and cheaper to compute (quantization). The final module steps back and assembles all of it into the picture of a modern frontier model — and points to what lies beyond this course.

Try it yourself
Transformer Lab: quantization →

Quantize weights to INT8, FP8 and INT4 and see what one outlier does, with and without block-wise scales.

MediumQuantizationmixed precision

What is mixed precision, and which parts of the model are deliberately kept in high precision?

HardQuantizationfine-grained

Why can a single outlier ruin a whole tensor's precision under one global scale, and how does fine-grained scaling fix it?

MediumQuantizationaccumulation

Why must the accumulation in an FP8 matrix multiply be done in higher precision?