In the past two years, the ability of large models mainly existed in the cloud - users asked questions, data was uploaded to the data center, and the model completed inference on the server and then transmitted the results back. However, cloud inference faces three major challenges: latency, cost, and privacy. Network fluctuations affect user experience, massive requests drive up computing power bills, and sensitive data leaving the domain also raises concerns for businesses.
On Device AI is designed to address these issues. It deploys the model to terminal devices such as mobile phones, computers, cars, and home appliances after compression, quantization, and pruning, allowing inference to be completed locally. Nowadays, mainstream flagship smartphones are capable of smoothly running small language models with billions of parameters, enabling offline translation, intelligent summarization, local voice assistants, and other functions; The new generation of AI PCs comes standard with NPU (Neural Network Processor) to provide localized document processing and content generation capabilities for office scenarios.
The advantages of end-to-end AI are obvious: faster response (no need for network round trips), more secure privacy (data does not leave the device), and lower usage costs (no reliance on cloud computing power and pay by volume). Of course, limited by the memory and computing power of the terminal, there is still a gap between the end-to-end model and the cloud based large model in complex inference. Therefore, "end-to-end cloud collaboration" has become the mainstream architecture - simple tasks are processed locally, while complex tasks are handed over to the cloud.
It can be foreseen that with the continuous advancement of model miniaturization technology and the continuous improvement of terminal computing power, end-to-end AI will become an important pole in the popularization of AI. The big model is no longer a "black box" in the data center, but a "portable intelligence" that truly enters daily devices.
[Reference source] This article comprehensively compiles product and technology information publicly released by industry media, chip, and terminal manufacturers.