In recent years, the deployment focus of large models has been shifting from the cloud to the terminal. On device AI refers to the direct use of AI on mobile phones PC、 Running AI models on terminal devices such as cars does not require uploading data to the cloud. Behind this trend is the leap in chip computing power and the maturity of model miniaturization technology.
From the chip side, Qualcomm Snapdragon series and Apple M-series chips have integrated neural network processing units (NPUs) to provide dedicated computing power for end-to-end inference. In 2024, Apple released Apple Intelligence at the WWDC Developer Conference, which deeply integrates generative AI capabilities into iPhones, iPads, and Macs, and emphasizes that most processing is done on the device side to protect user privacy. Google's Gemini Nano model is also built into some Pixel phones in an end-to-end form. The PC side is also bustling, with Microsoft defining Copilot+PC as the new generation of AI PCs, requiring devices to have NPU computing power of 40 TOPS or more.
The advantages of end-to-end AI are very clear: firstly, privacy and security, sensitive data does not need to leave the device; Secondly, it has low latency, with inference completed locally and faster response times; Thirdly, it is available offline and can still work without a network environment. But challenges also exist - terminal memory and power consumption are limited, and the number of model parameters is strictly constrained; At present, there is still a gap between end-to-end models and cloud based large-scale models in complex inference tasks.
The consensus in the industry is "end-to-end cloud collaboration": simple, high-frequency, and privacy sensitive tasks are handed over to small models on the client-side, while complex reasoning and knowledge intensive tasks are handed over to large models in the cloud. Meta's open-source Llama series small parameter versions, as well as domestic open-source models such as Qwen, provide a wide range of model choices for end-to-end deployment. It can be foreseen that with the continuous improvement of NPU computing power and the advancement of model compression technologies such as quantification, distillation, and pruning, end-to-end AI will become a key part of AI inclusiveness - allowing billions of users to enjoy the productivity improvement brought by AI without relying on expensive cloud services.
This article is a comprehensive compilation of product information and industry reports publicly released by companies such as Apple, Google, and Microsoft.