Mixture-of-experts stands or falls on one small component: the router — also called the gating network. It is what decides, for each token, which experts to wake up. It is tiny compared to the experts themselves, but it is where the sparsity actually happens. This chapter makes it precise.
Scoring the experts
The router is itself a little learned layer — in the simplest form, a single weight matrix. For a token's vector it produces one score per expert: a measure of how well-suited each expert is to this token. If there are 64 experts, the router emits 64 numbers.
Concretely, the token vector is multiplied by the router's matrix to give a vector of raw scores, one entry per expert. These scores are learned: during training the router's weights are nudged so that tokens end up at experts that handle them well. The router is trained jointly with everything else, by the same gradient descent — no separate objective picks the experts for it.
Keep only the top few
Raw scores over all experts are dense — every expert has some score. Sparsity comes from a hard choice: keep only the top- highest-scoring experts and discard the rest. Typically is small — often 2. For this token, those experts are "on"; the other 62 are "off" and do no work.
This top- selection is the single step that turns a bank of experts into a sparse model. Choose and each token uses exactly one expert; choose equal to the number of experts and you are back to a dense model that runs everything. Small is the whole point: it is what keeps active parameters low.
From scores to weights
The chosen experts should not all count equally — one may be a much better fit than another. So the router converts the surviving top- scores into weights with a softmax, giving positive numbers that sum to 1. A token routed to experts 5 and 23 might weight them 0.8 and 0.2, meaning "mostly expert 5, a little expert 23."
Each chosen expert processes the token independently, producing its own output vector. The block's output for that token is the weighted sum of those outputs, using the router's weights:
So the router does two jobs at once: it selects which experts act (via top-) and it weights how much each contributes (via the softmax). The output has the same shape as a plain FFN's would — a drop-in replacement for the block's feed-forward layer — but computed from only a couple of experts instead of one big network.
Why routing is also what makes MoE hard
The router is elegant, but notice the tension built into it. The choice of experts is a hard, discrete pick — an expert is either in the top- or not. Discreteness is awkward for gradient descent, which prefers smooth changes, and it opens the door to a failure mode: nothing in what we have described forces the router to spread tokens around.
Left to itself, a router often finds a rich-get-richer rut. Early in training a few experts happen to be chosen slightly more, so they improve faster, so the router prefers them more, so they are chosen even more. The result is routing collapse: a few experts do almost all the work and the rest sit idle, never learning anything. A model with 64 experts that effectively uses 4 has thrown away most of its capacity — and most of the memory it is paying for.
So the router cannot simply be left to optimise next-token prediction; it must also be nudged to use the experts evenly. How to enforce that balance — without wrecking the model's real objective — is the problem the next chapter takes on.
Route a sentence to top-1 or top-2 experts and see the gate weights.