Jointly optimize localization and quality estimation

Determine whether jointly optimizing the distortion-localization model and perceptual speech-quality estimation can enable the two objectives to reinforce each other beyond the separately trained DAMOS components.

Background

The localization and quality-prediction models in DAMOS are trained separately, and the localization backbone remains frozen during MOS training. The authors leave unresolved whether joint optimization would improve the interaction between distortion localization and perceptual quality estimation.

References

Despite these promising results, several directions remain open. First, the current framework relies on synthetically generated distortion masks for both training the localization model and conducting our inference-time robustness analysis; evaluating DAMOS under the localization model’s naturally occurring error distribution on real, non-synthetic distortions would provide a more direct measure of its practical robustness. Second, the localization and quality prediction models are trained separately, with the localization backbone kept frozen during MOS training; jointly optimizing distortion localization and perceptual quality estimation may allow the two objectives to reinforce each other further.

DAMOS: Learning Distortion-Aware Speech Quality Assessment through Explicit Distortion Localization  (2608.21176 - Li et al., 21 Aug 2026) in Conclusion, p. 11

Despite these promising results, several directions remain open. First, the current framework relies on synthetically generated distortion masks for both training the localization model and conducting our inference-time robustness analysis; evaluating DAMOS under the localization model’s naturally occurring error distribution on real, non-synthetic distortions would provide a more direct measure of its practical robustness. Second, the localization and quality prediction models are trained separately, with the localization backbone kept frozen during MOS training; jointly optimizing distortion localization and perceptual quality estimation may allow the two objectives to reinforce each other further. Finally, extending the framework toward fully frame-level, interpretable quality assessment—where predicted distortion locations are directly exposed as human-interpretable evidence for a quality score—represents a natural next step toward more transparent and diagnostic speech quality assessment.

DAMOS: Learning Distortion-Aware Speech Quality Assessment through Explicit Distortion Localization  (2608.21176 - Li et al., 21 Aug 2026) in Conclusion, p. 11