Controlled comparison of cascaded and end-to-end voice-agent architectures
Determine which real-time voice-agent architecture—cascaded or end-to-end—offers the preferable overall trade-off between preserving paralinguistic content and providing self-hostability and controllability, using controlled evaluations on matched tasks and conditions.
References
This survey organised 38 primary sources on real-time voice agents into six application-centric categories, stating for each work the problem it targets, its mechanism, and its reported evidence. Two trajectories dominate. First, the architectural question remains genuinely open: end-to-end models preserve paralinguistic content that cascades discard, yet cascades presently win on self-hostability and controllability, and duplex behaviour has been shown to be obtainable within a cascade.