Reliable objective metrics for perceived audio quality
Develop objective evaluation metrics that reliably correlate with human judgments of perceived audio quality across architectures and training objectives, overcoming the poor alignment observed between existing metrics such as VisQOL and MOSNet and subjective MUSHRA ratings.
References
This observation underscores the open challenge of designing reliable objective proxies for perceived quality.
— Moshi: a speech-text foundation model for real-time dialogue
(2410.00037 - Défossez et al., 2024) in Section 5.2, Audio Tokenization (Discussion)
A standardized causal-baseline suite with matched end-to-end latency budgets and consistent objective measurements is left for future work.
— CSAVocoder: A Causal Spatial Audio Vocoder Towards Real-Time Spatial Audio Generation
(2608.25404 - Zhu et al., 26 Aug 2026) in Section “Limitations”