Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Neural Spatio-Temporal Beamformer for Target Speech Separation (2005.03889v5)

Published 8 May 2020 in eess.AS and cs.SD

Abstract: Purely neural network (NN) based speech separation and enhancement methods, although can achieve good objective scores, inevitably cause nonlinear speech distortions that are harmful for the automatic speech recognition (ASR). On the other hand, the minimum variance distortionless response (MVDR) beamformer with NN-predicted masks, although can significantly reduce speech distortions, has limited noise reduction capability. In this paper, we propose a multi-tap MVDR beamformer with complex-valued masks for speech separation and enhancement. Compared to the state-of-the-art NN-mask based MVDR beamformer, the multi-tap MVDR beamformer exploits the inter-frame correlation in addition to the inter-microphone correlation that is already utilized in prior arts. Further improvements include the replacement of the real-valued masks with the complex-valued masks and the joint training of the complex-mask NN. The evaluation on our multi-modal multi-channel target speech separation and enhancement platform demonstrates that our proposed multi-tap MVDR beamformer improves both the ASR accuracy and the perceptual speech quality against prior arts.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (7)
  1. Yong Xu (432 papers)
  2. Meng Yu (65 papers)
  3. Shi-Xiong Zhang (48 papers)
  4. Lianwu Chen (14 papers)
  5. Chao Weng (61 papers)
  6. Jianming Liu (13 papers)
  7. Dong Yu (329 papers)
Citations (39)

Summary

We haven't generated a summary for this paper yet.