Keeping the Experts Balanced

Left alone, routing collapses onto a few experts. The classic cures add a balancing loss and cap each expert's intake; a newer cure nudges the router with a per-expert bias and no extra loss at all — balancing the load without disturbing the real objective.

Large Language Models: From Transformers to Frontier Models

The last chapter ended with a threat: routers drift into a rut where a few experts do all the work. If balance is not enforced, most of an MoE's capacity — and most of the memory it costs — is wasted. This chapter is the set of techniques that keep every expert busy. They matter because the whole bargain of MoE assumes the experts are actually used.

Measuring imbalance: importance and load

To fix imbalance we first need to name it, and there are two distinct quantities, easy to confuse.

Expert importance asks: across all tokens, how much total routing weight did each expert receive? Sum each expert's softmax weights over all tokens. An expert that is picked often, with high weights, has high importance; one rarely picked has low importance. Balanced importance means no expert is the router's darling.

Expert load asks something different: how many tokens were actually dispatched to each expert? This is a count, not a sum of weights.

The subtlety — and it is the crux — is that equal importance does not guarantee equal load. One expert might get a single token with weight 1.0; another might get four tokens with weight 0.25 each. Both have importance 1.0, yet one handled one token and the other handled four. Balancing the weights is not the same as balancing the work. A good scheme has to watch both.

The classic cure: an auxiliary loss (and a capacity cap)

The traditional fix adds a second term to the training objective — an auxiliary load-balancing loss — that is small when experts are evenly used and large when they are not.

One version penalises uneven importance directly: compute how much the per-expert importances vary (their spread relative to their average) and add that as a loss, so training is pushed toward uniform importance. But because importance and load differ, the more effective version penalises the product of load and importance per expert, summed over experts. That product is smallest exactly when tokens are spread evenly and the router's confidence is spread evenly — so minimising it aligns the two and flattens both. This balancing loss is added to the next-token loss, scaled by a tunable weight.

A second guardrail is a hard cap: expert capacity. Each expert is allowed at most a fixed number of tokens per batch — roughly (tokens ÷ experts) × a capacity factor. If more tokens try to pile into one expert, the overflow is dropped or pushed elsewhere. A capacity factor slightly above 1 gives a little slack; too high and one expert can still hog traffic, too low and tokens get dropped. It is a blunt backstop against the worst pile-ups.

The auxiliary loss has a real cost

The balancing loss works, but it fights the model. The next-token objective and the balancing objective want different things, and their gradients mix. Weight the balancing term too little and experts collapse; too much and you blunt the model's actual language ability to buy balance. You are forced to trade quality for balance along a knob that has no comfortable setting. That dissatisfaction motivates the newer approach.

A cleaner cure: balance with a bias, no extra loss

A more recent idea removes the balancing loss entirely and balances the load a different way — by steering the router directly. The trick is a small per-expert bias added to the router's scores before the top-kk selection.

It works as a feedback controller, updated each step:

  1. Track how loaded each expert has been relative to the average.
  2. For an overloaded expert, nudge its bias down a touch — lowering its scores so it is chosen a little less next time.
  3. For an underloaded expert, nudge its bias up — raising its scores so it is chosen a little more.

Over many steps this self-corrects toward even load: whenever an expert gets greedy it is gently demoted, whenever one is neglected it is promoted. Crucially, the bias only affects which experts are selected — it is not part of the training loss, so it adds no competing gradient to the next-token objective. The model optimises purely for language, while a side mechanism keeps the load even.

This is another "best of both worlds" moment, in the same spirit as latent attention: you get good balance and an undisturbed training objective at once, rather than trading one for the other. In practice this loss-free balancing has been found to deliver both better balance and better model quality than the auxiliary-loss approach, and it is the method used in the most recent large MoE models.

Per-expert bias added to router scores: overloaded experts biased down, underloaded experts biased up, steering selection toward even load
Loss-free balancing: a per-expert bias is nudged down for overloaded experts and up for underloaded ones, steering the top-k selection toward even load — with no extra term in the training loss.

Where this leaves us

With balancing solved, MoE delivers what it promised: many experts, all of them used, each token paying for only a few. But there is still headroom in how the experts are organised. The plain design — a modest number of equal, independent experts — turns out to leave specialisation on the table. The final chapter of this module covers two refinements that sharpen it: making the experts smaller and more numerous, and giving a few experts a permanent, shared role.

Try it yourself
Transformer Lab: balancing the experts →

Watch an overloaded expert drop tokens, then run bias-based balancing steps.

MediumMoEbalancing

Why can two experts have equal importance but very different load?

MediumMoEbalancing

What is the drawback of enforcing balance with an auxiliary loss?

HardMoEbalancing

How does bias-based (loss-free) balancing avoid that drawback?