The past intelligent voice assistants were basically a pipeline: speech recognition (ASR) converted sound into text, language models understood the text and generated reply text, and speech synthesis (TTS) read the text out. This link is mature and interpretable, but there are two unavoidable losses: delay accumulates segment by segment, and paralingual information (tone, emotion, pause) is lost in the step of "converting to text".
1、 The overlooked 'text bottleneck'
The output of ASR is standardized text, and the speaker's emotions, hesitation, stress, and pace are no longer retained. For downstream models, a choked voice is no different from a calm recitation. This is the root cause of many voice assistants' hearing clearly but not understanding '.
2、 What did the end-to-end speech model do
The end-to-end (speech to speech) route no longer forces intermediate results to fall into text: the model directly models a mixture of speech tokens and text tokens, with input audio and output audio. There are three direct benefits:
Lower latency: eliminates the serial waiting of "waiting for the entire paragraph to be finished → recognition → generation → synthesis", allowing for close to natural alternation of turns, including interruption, agreement, and buzzer.
Retain paralingual information: Tone and emotion can be used as conditions to participate in generation, and the tone and rhythm of the response are more natural.
Unified Multimodality: Speech, text, and images can be aligned in the same model, leaving space for "watching while speaking" interaction.
3、 Costs and difficulties
End to end is not free.
One is the scarcity of data. High quality "speech to speech" parallel data is much less than text data, and training usually relies on large-scale self supervised pre training, plus a small amount of aligned data.
The second is a decrease in controllability. There is no text in the middle, making content review, sensitive information filtering, and log tracking more difficult than text links.
The third challenge is difficulty in evaluation. The quality of voice interaction is not only related to transcription accuracy, but also to naturalness, timing of turns, and appropriateness of tone, and currently lacks a universally recognized standard.
The fourth is stability. When generating audio directly, the model may produce tone that does not match the semantics, and even result in saying the wrong thing.
4、 Practical suggestions for engineering selection
For scenarios that require high accuracy of content (customer service work orders, medical inquiries), retaining an auditable link of "voice text inference text" is often more secure. An end-to-end model can be used in the interaction layer to hand over key conclusions to the text link for review.
For scenarios that prioritize real-time performance and user experience (such as real-time translation, companion dialogue, and in car voice), the end-to-end solution has a significant advantage in low latency.
Regardless of which path is chosen, interrupt handling, mute detection, and error fallback need to be designed as first-class citizens - users have much lower tolerance for voice interaction than typing.
5、 Summary
In the next stage of voice interaction, the competition is shifting from "inaccurate recognition" to "smooth dialogue". Behind this is a change in architecture: from a pipeline that concatenates multiple specialized models to a unified modeling of speech models. But the controllability and evaluation challenges it brings also need to be planned in advance before landing.
[Reference source] Comprehensive compilation of industry information publicly released.