Back to Home

Introduction to Mixed Expert (MoE): Why Big Models Need to 'Divide Work and Collaborate'

September 21, 2026 at 01:32 PMSource: RunByAI0 comment(s)TechGuide

1、 The dilemma of dense models

Traditional Transformers are "dense": for every token processed, all parameters must participate in computation. When the parameter scale reaches billions, the computational cost will increase approximately linearly with the parameters, making training and inference extremely expensive. Is it possible to use only a small portion of the parameters for each token without reducing the total capacity of the model?

2、 MoE's idea: Break down FFN into multiple 'experts'

Mixture of Experts (MoE) replaces the feedforward network (FFN) in Transformer with a set of parallel "experts", each of whom is a smaller FFN. Each token no longer goes through the same FFN, but is assigned to which experts by a routing network (Router/Gate), with the common practice being Top-2 activation.

The key is that the total number of parameters in the model can be large, but each token only activates a few experts, and the actual number of parameters involved in the calculation (FLOPs) is only related to the "number of activated experts", not the total number of parameters.

3、 What does sparse activation bring

-The total parameters can be made very large (with high capacity), but the computational cost per token remains at a low level;

-Under the same computing power budget, MoE usually performs better than dense models with the same amount of computation;

-The cost is high video memory usage (all experts have to install video memory), as well as engineering complexity in communication and load balancing.

4、 Load balancing and training stability

If the router always sends tokens to a few experts, the rest of the experts will 'starve' and the training efficiency will sharply decrease. Therefore, auxiliary loss is usually added during training to encourage expert load balancing, which is also one of the main sources of instability in MoE training.

5、 Representative model

Switch Transformer validated the scale effect of sparse experts using Top-1 routing; Mixtral 8x7B, with 8 experts and Top-2 routing, has become a representative work of open-source MoE. Subsequently, numerous open-source and closed source models have successively adopted the MoE architecture.

6、 Suitable or not suitable

Suitable for scenarios that pursue "large capacity, low single processing power", especially for services that require high throughput inference and control of single token costs.

Not suitable for scenarios where video memory is extremely limited or where deployment is minimalist, in which case dense small models are still more practical.

Conclusion

The essence of MoE is to 'swap sparsity for scale': instead of cramming all knowledge into the same computational path, each token only invites a few experts for consultations.

[Reference source]

-William Fedus et al.: Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity(arXiv:2101.03961, 2021)

-Albert Q. Jiang et al.: Mixtral of Experts (arXiv: 2401.04088, 2024)

Large Language Model (LLM)deep learning
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment