Spoken Language Modeling for Spontaneous Conversational Speech

Investigate spoken language modeling with StreamAlign units on spontaneous conversational speech, extending beyond the read-style and web-scale speech settings evaluated in the paper.

Background

The experiments primarily evaluate read-style and web-scale speech data. The paper additionally tests speech reconstruction on spontaneous conversational speech and reports robustness without retraining, but it does not evaluate the spoken language modeling component in that domain.

This unresolved issue concerns whether StreamAlign-SLM can preserve its semantic, acoustic, and continuation performance when trained or evaluated on spontaneous conversational speech, whose disfluencies, interactional structure, and distribution differ from audiobook-style data.

References

Second, our experiments primarily target read-style and web-scale speech data. While we verify that speech reconstruction remains robust on spontaneous conversational speech (Appendix~\ref{app:robustness}), we do not investigate spoken language modeling in these settings, which we leave as an important direction for future work.

StreamAlign: Streaming Text-Aligned Speech Tokenization  (2609.09719 - Kim et al., 9 Sep 2026) in Section ‘Limitations’