The Router: Choosing Experts

A small gating network scores every expert for each token, keeps the top few, turns their scores into weights, and blends only those experts' outputs. That tiny mechanism is what makes sparse activation work.

Large Language Models: From Transformers to Frontier Models

Mixture-of-experts stands or falls on one small component: the router — also called the gating network. It is what decides, for each token, which experts to wake up. It is tiny compared to the experts themselves, but it is where the sparsity actually happens. This chapter makes it precise.

Scoring the experts

The router is itself a little learned layer — in the simplest form, a single weight matrix. For a token's vector it produces one score per expert: a measure of how well-suited each expert is to this token. If there are 64 experts, the router emits 64 numbers.

Concretely, the token vector is multiplied by the router's matrix to give a vector of raw scores, one entry per expert. These scores are learned: during training the router's weights are nudged so that tokens end up at experts that handle them well. The router is trained jointly with everything else, by the same gradient descent — no separate objective picks the experts for it.

Keep only the top few

Raw scores over all experts are dense — every expert has some score. Sparsity comes from a hard choice: keep only the top-kk highest-scoring experts and discard the rest. Typically kk is small — often 2. For this token, those kk experts are "on"; the other 62 are "off" and do no work.

This top-kk selection is the single step that turns a bank of experts into a sparse model. Choose k=1k = 1 and each token uses exactly one expert; choose kk equal to the number of experts and you are back to a dense model that runs everything. Small kk is the whole point: it is what keeps active parameters low.

From scores to weights

The chosen experts should not all count equally — one may be a much better fit than another. So the router converts the surviving top-kk scores into weights with a softmax, giving positive numbers that sum to 1. A token routed to experts 5 and 23 might weight them 0.8 and 0.2, meaning "mostly expert 5, a little expert 23."

Each chosen expert processes the token independently, producing its own output vector. The block's output for that token is the weighted sum of those outputs, using the router's weights:

output=∑e ∈ top-kwe⋅Experte(token).\text{output} = \sum_{e \,\in\, \text{top-}k} w_e \cdot \text{Expert}_e(\text{token}).

So the router does two jobs at once: it selects which experts act (via top-kk) and it weights how much each contributes (via the softmax). The output has the same shape as a plain FFN's would — a drop-in replacement for the block's feed-forward layer — but computed from only a couple of experts instead of one big network.

A token scored against all experts by the router, the top two kept and softmaxed into weights, their expert outputs blended into one vector
The router scores every expert, keeps the top-k, softmaxes those into weights, and blends only the chosen experts' outputs. Selection and weighting in one small layer.

Why routing is also what makes MoE hard

The router is elegant, but notice the tension built into it. The choice of experts is a hard, discrete pick — an expert is either in the top-kk or not. Discreteness is awkward for gradient descent, which prefers smooth changes, and it opens the door to a failure mode: nothing in what we have described forces the router to spread tokens around.

Left to itself, a router often finds a rich-get-richer rut. Early in training a few experts happen to be chosen slightly more, so they improve faster, so the router prefers them more, so they are chosen even more. The result is routing collapse: a few experts do almost all the work and the rest sit idle, never learning anything. A model with 64 experts that effectively uses 4 has thrown away most of its capacity — and most of the memory it is paying for.

So the router cannot simply be left to optimise next-token prediction; it must also be nudged to use the experts evenly. How to enforce that balance — without wrecking the model's real objective — is the problem the next chapter takes on.

Try it yourself
Transformer Lab: mixture of experts →

Route a sentence to top-1 or top-2 experts and see the gate weights.

EasyMoErouting

What two jobs does the router perform for each token?

MediumMoErouting

Which single step converts a bank of experts into a sparse model, and how does the choice of k matter?

MediumMoErouting

What is routing collapse, and why does it arise on its own?