Improve Discriminative Flow Matching for Speech Dereverberation

Determine whether using a better-defined training objective or a non-masking discriminative model can improve Discriminative Flow Matching performance for speech dereverberation and realize its benefits more fully.

Background

The paper evaluates Discriminative Flow Matching (DFM) for speech dereverberation using a mask-based DCCRN discriminative model to estimate dry speech. Although DFM outperforms the DCCRN baseline and achieves results comparable to or better than the conventional Conditional Flow Matching baseline, the reported gains—particularly for the SCOREQ metric—are statistically negligible. The authors attribute this limitation to the inferior quality of the discriminative model, which causes the Discriminative Flow-State Representations to diverge from the true Flow-State. They specifically identify improved training objectives and discriminative models that do not rely on masking as possible ways to improve dereverberation performance, but leave whether these changes would yield the anticipated benefits unresolved.

References

Furthermore, estimating dry speech with a mask-based method like DCCRN is inherently difficult; therefore, we hypothesize that with a better-defined training objective or a better discriminative model without masking, performance can be improved, allowing the true benefits of \ac{DFM} to be realized for this speech dereverberation task.

Discriminative Flow Matching: Beyond Time-Conditioning in Generative Restoration via Flow-State Representations  (2609.04525 - Shetu et al., 3 Sep 2026) in Appendix, subsection “Speech Dereverberation,” subsection “Results”