Back to Home

Introduction to Mixed Expert Models (MoE): How Big Models Can Be "Big and Fast"

September 18, 2026 at 01:33 PMSource: RunByAI0 comment(s)TechGuide

The Mixture of Experts (MoE) is an architecture that controls computational costs while expanding the size of model parameters. It can be traced back to the idea of adaptive expert mixing proposed by Jacobs et al. in 1991, and in recent years, with the maturity of sparse training methods, it has been widely applied to large-scale language models.

1、 The problems faced by dense models

The traditional Dense model activates all parameters for each word element it processes. This means that the more parameters there are, the greater the computational complexity of a single inference, and the training and deployment costs also increase accordingly. When the model size grows from billions to hundreds of billions of parameters, the computational cost quickly becomes unbearable.

2、 The core idea of MoE: Sparse activation

MoE splits a large feedforward network into multiple parallel "Expert" subnetworks and introduces a routing network (Router, also known as Gate). For each input word element, the routing network only selects a few most relevant experts to participate in the calculation, while the remaining experts remain inactive.

In this way, the total number of parameters in the model can be large, but the actual number of parameters involved in the calculation (activation) of each morpheme is very small, which is called "sparse activation". For example, a model with total parameters reaching hundreds of billions but only a few billion parameters per activation can achieve expressive power close to that of a large model at a computational cost close to that of a small model.

3、 Routing and load balancing

Routing networks typically calculate the score of each expert for the current word element through a learnable linear layer, and then select the Top-k experts with the highest scores (usually k=1 or 2). A key issue in training is load balancing: if routing always favors a few experts and the rest of the experts do not receive sufficient training, it will result in resource waste. A common practice is to introduce auxiliary loss to encourage more even distribution of lexical elements among experts.

4、 Typical MoE model

In recent years, multiple public models have adopted the MoE architecture, such as Google's GShard and Switch Transformer, the open-source community's Mixtral 8x7B, and some models in the DeepSeek series in China. These models generally demonstrate that under similar activation parameter quantities, their performance is often better than dense models of the same scale.

5、 Advantages and Challenges

In terms of advantages, MoE can usually achieve larger model capacity at the same inference cost, and it is also easier to scale up to super large scales during training. The challenges mainly lie in the video memory and communication overhead brought by expert parallelism, the stability of routing training, and the fluctuation of performance caused by differences in the abilities of different experts. In addition, due to the need to reserve parameters for all experts, the overall memory requirements for MoE models are usually still high.

6、 Summary

MoE has achieved a new balance between model capacity and computational cost through its sparse design of "large total parameters and few single activations", becoming one of the important directions for current large-scale model expansion. By understanding the concepts of experts, routing, and load balancing, one can grasp the basic principles of MoE.

[Reference source]

- Jacobs et al., "Adaptive Mixtures of Local Experts", Neural Computation, 1991

- Fedus et al., "Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity", JMLR 2022(arXiv:2101.03961)

- Jiang et al., "Mixtral of Experts"(arXiv:2401.04088)

-Comprehensive compilation of industry information that has been publicly released

large modelMoE
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment