Jointly optimize localization and quality estimation
Determine whether jointly optimizing the distortion-localization model and perceptual speech-quality estimation can enable the two objectives to reinforce each other beyond the separately trained DAMOS components.
References
Despite these promising results, several directions remain open. First, the current framework relies on synthetically generated distortion masks for both training the localization model and conducting our inference-time robustness analysis; evaluating DAMOS under the localization model’s naturally occurring error distribution on real, non-synthetic distortions would provide a more direct measure of its practical robustness. Second, the localization and quality prediction models are trained separately, with the localization backbone kept frozen during MOS training; jointly optimizing distortion localization and perceptual quality estimation may allow the two objectives to reinforce each other further.
Despite these promising results, several directions remain open. First, the current framework relies on synthetically generated distortion masks for both training the localization model and conducting our inference-time robustness analysis; evaluating DAMOS under the localization model’s naturally occurring error distribution on real, non-synthetic distortions would provide a more direct measure of its practical robustness. Second, the localization and quality prediction models are trained separately, with the localization backbone kept frozen during MOS training; jointly optimizing distortion localization and perceptual quality estimation may allow the two objectives to reinforce each other further. Finally, extending the framework toward fully frame-level, interpretable quality assessment—where predicted distortion locations are directly exposed as human-interpretable evidence for a quality score—represents a natural next step toward more transparent and diagnostic speech quality assessment.