Reduce first-packet latency without sacrificing accuracy

Reduce the first-output startup delay of VibeVoice-ASR-Streaming below the current full-chunk-plus-lookahead latency while preserving the accuracy benefits of the larger 22-frame chunks.

Background

VibeVoice-ASR-Streaming reports a steady-state speaker-attribution latency of 2.00 seconds for the 22-frame configuration and 1.53 seconds for the 15-frame configuration. However, the first output incurs an additional startup delay because the system must receive one complete chunk together with its lookahead: 3.5 seconds at 22 frames and 2.5 seconds at 15 frames.

The paper identifies reducing this initial delay while retaining the recognition and speaker-attribution accuracy associated with larger chunks as unresolved future work. This is a concrete deployment problem for real-time voice assistants and agents, where startup latency affects responsiveness even when subsequent outputs meet the target streaming latency.

References

Reducing this startup delay without sacrificing the accuracy of larger chunks remains future work.

VibeVoice-ASR-Streaming Technical Report  (2609.02812 - Tu et al., 2 Sep 2026) in Section 6, Conclusion and Limitations, subsection “First-packet latency”