MoE (Mixture of Experts) is an architecture design that expands model capacity by sparsely activating multiple sub networks, and has become a key technology for improving efficiency in current large language models.
##The core principle of MoE
The core idea of MoE architecture is to split a large model into multiple "expert" subnetworks, activating only a portion of them during each inference. A gate controlled network (Router) dynamically decides which experts to activate based on input content, and Top-2 routing strategy is usually the most commonly used. In this way, the total number of parameters in the model can be significantly increased, but the computational cost of each inference is only related to the number of activated experts.
##Practice of Mixral 8x7B
Mistral AI's open-source Mixral 8x7B in 2023 is the benchmark implementation of the MoE architecture. The model consists of 8 experts, each with approximately 7 billion parameters, for a total of approximately 47 billion parameters. However, each inference only activates two experts (approximately 13 billion parameters), significantly reducing inference costs while maintaining high performance. Mixtral has performed excellently in multiple benchmark tests, even surpassing dense models with the same computational complexity.
##The Challenge of MoE
Although MoE is efficient, it also faces challenges such as imbalanced expert load, difficult model convergence, and high communication overhead for distributed training. The 2017 publication by Shazeer et al. titled "Externally Large Neural Networks: The Sparely Gated Mixture of Experts Layer" (ICLR 2017) laid the theoretical foundation for MoE in the field of NLP.
【 Reference source 】 Shazeer et al., "Externally Large Neural Networks", ICLR 2017; Mistral AI Official Blog; Related open-source documents.