Minimum achievable frame rate for streaming speech codecs

Determine how far the frame rate of a streaming neural speech codec that jointly captures semantic and acoustic information can be reduced.

Background

The paper explains that existing streaming codecs combining semantic and acoustic information, particularly Mimi, operate at 12.5 Hz, whereas codecs operating at lower frame rates are either offline or rely on additional side information. This creates an unresolved question about the lower limit of frame rate compatible with streaming operation while retaining both semantic and acoustic information. ZipCodec advances the state of the art by demonstrating operation at 6.25 Hz, but the general lower bound remains undetermined.

References

To the best of our knowledge, no such codec has been demonstrated below 12.5 Hz, leaving open how far the frame rate of a streaming codec can be reduced.

ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding  (2609.11642 - Libera et al., 10 Sep 2026) in Section 1, Introduction