End to end voice model: Why AI voice interaction is moving from "assembly line" to "one model"
The past intelligent voice assistants were basically a pipeline: speech recognition (ASR) converted sound into text, language models understood the text and generated reply text, and speech synthesis