Spoken Language Modeling for Spontaneous Conversational Speech
Investigate spoken language modeling with StreamAlign units on spontaneous conversational speech, extending beyond the read-style and web-scale speech settings evaluated in the paper.
References
Second, our experiments primarily target read-style and web-scale speech data. While we verify that speech reconstruction remains robust on spontaneous conversational speech (Appendix~\ref{app:robustness}), we do not investigate spoken language modeling in these settings, which we leave as an important direction for future work.
— StreamAlign: Streaming Text-Aligned Speech Tokenization
(2609.09719 - Kim et al., 9 Sep 2026) in Section ‘Limitations’