Why Mixture-of-Experts

The feed-forward layer holds most of a model's parameters and runs in full for every token. Mixture-of-experts breaks that: many expert networks, but a router sends each token to only a couple — so capacity grows while the work per token stays small.

Large Language Models: From Transformers to Frontier Models

The last module made attention cheap to run. This one turns to the other half of the transformer block — the feed-forward network (FFN) — and a different pressure: not inference memory, but scale. How do you make a model much more capable without making every token much more expensive? The answer, mixture-of-experts (MoE), is the second great efficiency idea in modern models, and it rests on one word: sparsity.

The feed-forward layer is where the parameters live

Recall the FFN from Module 2: it expands each token's vector to a wider inner dimension (classically 4× the model width), applies a non-linearity, and contracts it back. That expand-and-contract is where a transformer does much of its "thinking" — and where most of its parameters sit. Two big weight matrices per block, each roughly 4d24d^2 in size, add up fast: across all layers, the feed-forward networks typically account for the majority of a model's parameter count.

That matters because, in a plain ("dense") model, every token passes through all of those parameters. More capability means a wider FFN means more parameters means more arithmetic per token, in both training and inference. Capability and cost are chained together. Mixture-of-experts breaks the chain.

The idea: many experts, a few used at a time

MoE replaces the single FFN in a block with many FFNs, called experts — say 8, 64, or more. But it does not run them all. A small router (a gating network) looks at each token and picks just a few experts — often two — to process it. The token flows only through its chosen experts; the rest sit idle for that token.

The consequence is a split between two very different numbers:

  • Total parameters — all the experts together — can be huge, giving the model enormous capacity to store knowledge.
  • Active parameters — the few experts any single token actually uses — stay small, so the compute per token stays low.

A model can hold, say, ten times the parameters of a dense model while each token still only pays for a small slice. This is sparse activation: most of the model is switched off for any given token, like a building full of specialists where each visitor sees only the two they need.

A transformer block's single feed-forward network replaced by several expert networks with a router selecting two of them for a token
Mixture-of-experts replaces the block's one FFN with many experts and a router. Each token is sent to only a couple — total capacity is large, but work per token stays small.

Why sparsity buys capability for free

Why should switching most of the model off not hurt? Because the experts specialise. When researchers inspect a trained MoE, different experts reliably take on different kinds of token — one gravitates to punctuation, another to numbers, another to verbs, another to a topic. The router learns to send each token to the experts suited to it.

Given that, running all experts on every token is wasteful: a number token has little to gain from the punctuation expert. Sparse routing simply skips the irrelevant specialists. The model gets the breadth of many experts (more total knowledge) without the cost of consulting all of them (low active compute). That is the whole bargain.

An old idea, newly essential

Mixture-of-experts is not new — the core idea of training several specialist networks with a gate to pick among them dates to the early 1990s. What changed is scale: once models grew large enough that the FFN dominated cost, sparse experts became the natural way to keep growing. Today it is mainstream — many of the largest open models are MoE models.

Experts live in every block, and routing is per-block

One picture to keep straight: the experts are not a single bank the token visits once. Every transformer block has its own set of experts and its own router. So a token is routed afresh at each layer, and the experts it uses in block 1 need not be the experts it uses in block 2 — specialisation is per-layer. A token for the word "one" might take the number-ish experts in an early block and quite different ones later, as the model builds up meaning in stages.

This raises an obvious worry. If a router is free to send tokens wherever it likes, what stops it from falling in love with a few experts and starving the rest? A handful of overworked experts and a crowd of idle ones would waste the model's capacity entirely. Keeping the experts balanced is the central engineering problem of MoE — and the subject of this module's middle chapters, after the next one makes the router itself precise.

MediumMoE

Why does scaling a dense model's capability also scale its cost per token, and how does MoE break that link?

EasyMoE

What is the difference between a model's total parameters and its active parameters under MoE?

MediumMoE

Why doesn't skipping most experts for each token cripple the model?