Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
102 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

WPD++: An Improved Neural Beamformer for Simultaneous Speech Separation and Dereverberation (2011.09162v1)

Published 18 Nov 2020 in eess.AS and cs.SD

Abstract: This paper aims at eliminating the interfering speakers' speech, additive noise, and reverberation from the noisy multi-talker speech mixture that benefits automatic speech recognition (ASR) backend. While the recently proposed Weighted Power minimization Distortionless response (WPD) beamformer can perform separation and dereverberation simultaneously, the noise cancellation component still has the potential to progress. We propose an improved neural WPD beamformer called "WPD++" by an enhanced beamforming module in the conventional WPD and a multi-objective loss function for the joint training. The beamforming module is improved by utilizing the spatio-temporal correlation. A multi-objective loss, including the complex spectra domain scale-invariant signal-to-noise ratio (C-Si-SNR) and the magnitude domain mean square error (Mag-MSE), is properly designed to make multiple constraints on the enhanced speech and the desired power of the dry clean signal. Joint training is conducted to optimize the complex-valued mask estimator and the WPD++ beamformer in an end-to-end way. The results show that the proposed WPD++ outperforms several state-of-the-art beamformers on the enhanced speech quality and word error rate (WER) of ASR.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (7)
  1. Zhaoheng Ni (32 papers)
  2. Yong Xu (432 papers)
  3. Meng Yu (65 papers)
  4. Bo Wu (144 papers)
  5. Shixiong Zhang (11 papers)
  6. Dong Yu (329 papers)
  7. Michael I Mandel (12 papers)
Citations (7)

Summary

We haven't generated a summary for this paper yet.