Sharper Experts: Fine-Grained and Shared

Two refinements make experts specialise better: slice them smaller and more numerous so each can own a narrow skill, and set a few experts to always-on so they absorb the common knowledge every token needs — freeing the rest to specialise.

Large Language Models: From Transformers to Frontier Models

Balanced MoE works, but the plain design — a modest number of equal, independent experts — leaves specialisation half-finished. This chapter covers two refinements that push experts toward being genuinely specialised, each aimed at a specific way the plain design falls short. Both keep the parameter count and compute per token the same; they only change how the experts are carved up and assigned.

Problem one: experts forced to be generalists

With only a handful of experts — say 8 or 16 — each one is, by necessity, a jack-of-all-trades. There are far more distinct kinds of knowledge in language than there are experts, so every expert must cram many unrelated skills into its parameters: a little grammar, a little arithmetic, a little geography. Call this knowledge hybridity — each expert is a blurry mixture rather than a sharp specialist, and blurry experts are harder for the router to use well.

The fix: fine-grained experts. Slice each expert into several smaller ones. Instead of 16 wide experts, have 64 narrow ones, each with a proportionally smaller inner dimension — and route to proportionally more of them (say the top 8 instead of top 2). The arithmetic is unchanged: more experts, each smaller, with the same total parameters and the same active compute per token. What changes is granularity. With many narrow experts, each one can own a tighter slice of knowledge, and the router can assemble a more precise combination of specialists for each token. Finer experts means sharper specialisation at no extra cost.

Problem two: everyone re-learning the basics

A different waste hides in the plain design. Some knowledge is needed by almost every token — basic syntax, common words, general facts. With independent experts, each one has to learn this common material on its own, because any expert might be the one handling a given token. So the same general knowledge gets duplicated across many experts. Call this knowledge redundancy — capacity spent storing the same basics many times over, capacity that could have gone to specialisation.

The fix: shared experts. Set aside a few experts that are always on — every token passes through them, no routing involved — alongside the usual routed experts. The shared experts absorb the common, everybody-needs-it knowledge. That relieves the routed experts of having to carry the basics, so they are free to specialise hard on their niches. The block's output simply adds the shared experts' contribution to the routed experts' contribution.

A token passing through always-on shared experts plus a few router-selected routed experts, their outputs combined
Shared experts (always on) absorb common knowledge for every token; routed experts (top-k selected) specialise. Their outputs are added — so specialists need not re-learn the basics.

Two problems, two complementary fixes

It is worth holding the pairing straight, because the two refinements attack different wastes:

  • Fine-grained segmentation → cures knowledge hybridity. More, smaller experts means each can be a tighter specialist instead of a generalist.
  • Shared experts → cure knowledge redundancy. Always-on experts hold the common knowledge once, so routed experts stop duplicating it.

They stack cleanly: a modern MoE layer typically has many fine-grained routed experts (of which a token uses a handful) plus a small number of shared experts (which every token uses). Both operate within the same parameter and compute budget as a plain MoE — they are reorganisations, not enlargements.

Why these help, concretely

When models with these refinements are compared against a plain MoE of the same active size, they match or beat much larger dense models on quality while activating only a small fraction of their parameters. Ablations show both pieces pulling their weight: remove the shared expert and quality drops; cut the number of fine-grained experts and it drops again. Sharper, better-organised experts are simply a better use of the same budget.

Module recap

That completes mixture-of-experts. The through-line: the feed-forward layer dominates a model's parameters and, in a dense model, runs in full for every token. MoE replaces it with many experts and a router that activates only a few — decoupling capacity from per-token cost (chapters 1–2). Making that work means keeping the experts evenly used, which a per-expert bias can do without disturbing the training objective (chapter 3). And organising the experts well — many fine-grained ones plus a few shared ones — sharpens their specialisation within the same budget (this chapter).

Attention efficiency (Module 6) and sparse scaling (Module 7) are the two largest levers behind modern frontier models. The next module turns to a smaller but clever one that changes the training objective itself: instead of predicting only the next token, predict several — multi-token prediction.

MediumMoEexperts

What is knowledge hybridity, and how do fine-grained experts address it?

MediumMoEexperts

What is a shared expert, and which problem does it solve?

EasyMoEexperts

Both refinements keep the parameter and compute budget fixed. Why is that important?