Back to Home

Mixed Expert Models (MoE): Why Large Models are Moving towards Sparsification

September 14, 2026 at 08:03 AMSource: RunByAI0 comment(s)Tech

If we observe the large models released in the past two years, we will find a clear trend: the parameter size is getting larger and larger, but the parameters that are truly "activated" for each inference are only a small part of them. The key architecture behind this is the Mixture of Experts (MoE) model.

1、 From 'full activation' to 'sparse activation'

Traditional Dense Transformers involve every parameter in each layer during a forward computation. This means that the larger the model, the higher the computational cost of a single inference, and the cost is almost proportional to the number of parameters.

The approach of MoE is completely different: it splits the original single feedforward network (FFN) into several parallel "experts", each of whom is also a small feedforward network. When an input reaches this layer, a "router/gathering network" decides which experts to distribute the input to for processing, usually only selecting top-k (such as top-2) experts to participate in the computation.

In this way, the total number of parameters in the model can be made very large, but the actual activated parameters for each token only account for a small part, thus achieving a better balance between "model capacity" and "computational cost".

2、 Router: The Brain of MoE

The router is the core of MoE. It receives the representation vector of each token, outputs the probability distribution on all experts, and then selects the top k experts based on probability, and summarizes their outputs weighted by weight.

A major challenge in training is "load balancing": if the router always favors a few experts, the rest of the experts will not receive sufficient training, resulting in a "winner takes all" situation. To this end, researchers have introduced mechanisms such as auxiliary loss to encourage more even distribution of tokens among experts.

3、 Why do big companies prefer MoE

Firstly, training and reasoning are more economical: under the same computational budget, sparse models can have a larger total number of parameters, thereby obtaining stronger expressive power.

Secondly, the inference cost is lower: only a portion of experts are activated each time, and the floating-point operation of a single token is significantly lower than that of a dense model with the same total number of parameters.

Thirdly, it is easy to expand: increasing the number of experts can expand the model capacity without increasing the computational load proportionally.

As a result, many open source and closed source models have adopted the MoE architecture, becoming a representative of "high parameter, low activation".

4、 Challenges and trade-offs

MoE is not a silver bullet. In addition to load balancing, it also faces issues such as high memory usage (all expert parameters need to be loaded), high communication overhead (experts need to communicate across devices when distributed across multiple cards), and unstable training. In engineering, it is often alleviated through expert parallelism, capacity factor control, and more refined routing strategies.

5、 Summary

MoE represents an important idea in the development of large models: not blindly pursuing the use of all parameters every time, but allowing the model to call the most suitable part when needed. As inference cost gradually becomes a key constraint for AI implementation, "sparsity" is likely to occupy a more important position in future model design.

【 Reference Source 】 Comprehensive compilation of industry information released publicly

MoElarge model
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment