Streamable causal architectures for discrete audio tokenizers
Develop causal, streamable architectures for discrete audio tokenizers that can operate in real time while maintaining high perceptual quality and computational efficiency, overcoming the current reliance of many self-supervised learning–based tokenizers on non-causal encoders.
References
Thus, achieving streamability with high-quality and efficient causal architectures remains an open research challenge.
— Discrete Audio Tokens: More Than a Survey!
(2506.10274 - Mousavi et al., 12 Jun 2025) in Section 2.5 (Streamability and Domain Categorization) – Streamability paragraph
Yet the first design leaves three questions open: whether the quantizer uses its finite vocabulary effectively, whether the computation is causal enough for streaming, and whether speaker identity should be exposed to the planner at all.
— The Evolving Bottleneck in Speech Generation: Interface Co-design and Staged Alignment from CosyVoice to Qwen-Audio-3.0-TTS
(2609.16514 - Chen et al., 15 Sep 2026) in Section 4.1, “CosyVoice: from acoustic compression to supervised semantics”