AI chips are the infrastructure of the artificial intelligence industry, and their performance directly determines the upper limit of AI applications. Since 2024, the global competition landscape for AI chips has undergone significant changes, gradually shifting from a "training computing power arms race" to a new stage of "optimization of inference efficiency".
In the field of training chips, NVIDIA still maintains an absolute leading position. The H100 and B200 GPUs, with their CUDA ecosystem and NVLink interconnect technology, have become the preferred hardware for large model training. But the competitive landscape is diversifying: AMD's MI300X demonstrates competitiveness in terms of cost-effectiveness, while Google's TPU v5p continues to iterate in the field of self-developed AI chips. More noteworthy is that a group of AI chip startups are emerging - Groq's LPU architecture, Cerebras' wafer level chips, SambaNova's RDU architecture, etc., all demonstrating unique advantages in specific scenarios.
The competition for inference chips is even more intense. As large models move from the training phase to large-scale deployment, inference cost becomes a key bottleneck. Various specialized inference chips have emerged: mobile chip manufacturers such as Qualcomm and MediaTek have accelerated the integration of AI acceleration units (NPUs); Apple's M-series chips have accumulated profound experience in end-to-end deployment; Domestic manufacturers such as Cambrian and Horizon Robotics continue to make breakthroughs in the NPU field.
In terms of architectural innovation, emerging technologies such as Compute in Memory and optical computing are moving from the laboratory to engineering validation. By eliminating the von Neumann bottleneck, the integration of storage and computing has achieved a higher energy efficiency ratio - in specific AI inference tasks, the energy efficiency can reach 10-100 times that of traditional chips. Although these technologies are still far from large-scale commercialization, they represent the future direction of AI chips.
The software ecosystem is also an indispensable part of competition. The moat effect of CUDA is still significant, but open-source projects such as OpenAI's Triton and PyTorch 2.0 compilation optimization are reducing migration costs. The competition for future AI chips will not only be a hardware confrontation, but also a comprehensive competition of "hardware+software+ecosystem". This article is a comprehensive compilation of technical white papers and industry analysis reports publicly released in the semiconductor industry.