Controlled comparison of cascaded and end-to-end voice-agent architectures

Determine which real-time voice-agent architecture—cascaded or end-to-end—offers the preferable overall trade-off between preserving paralinguistic content and providing self-hostability and controllability, using controlled evaluations on matched tasks and conditions.

Background

The paper identifies an unresolved architectural trade-off in real-time voice agents. End-to-end speech-to-speech systems can preserve paralinguistic information that is lost when speech is converted into text, whereas cascaded systems currently offer stronger self-hostability and controllability. The survey also reports that duplex turn-taking can be implemented within a cascade, indicating that full-duplex behaviour does not by itself require an end-to-end architecture.

Because the surveyed cascaded and end-to-end systems are generally evaluated by different groups on different tasks, the evidence does not establish which architecture is superior overall. Resolving the problem requires controlled comparisons that measure paralinguistic fidelity, self-hosting feasibility, controllability, latency, and interactive behaviour under comparable conditions.

References

This survey organised 38 primary sources on real-time voice agents into six application-centric categories, stating for each work the problem it targets, its mechanism, and its reported evidence. Two trajectories dominate. First, the architectural question remains genuinely open: end-to-end models preserve paralinguistic content that cascades discard, yet cascades presently win on self-hostability and controllability, and duplex behaviour has been shown to be obtainable within a cascade.

— Evaluating Real-Time Voice Agents: From Component Quality to Grounded Outcomes  (2609.30798 - Negi et al., 25 Sep 2026) in Section 6, “Open Challenges and Conclusion,” opening paragraph