Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
80 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

End-to-End Dereverberation, Beamforming, and Speech Recognition with Improved Numerical Stability and Advanced Frontend (2102.11525v1)

Published 23 Feb 2021 in eess.AS and cs.SD

Abstract: Recently, the end-to-end approach has been successfully applied to multi-speaker speech separation and recognition in both single-channel and multichannel conditions. However, severe performance degradation is still observed in the reverberant and noisy scenarios, and there is still a large performance gap between anechoic and reverberant conditions. In this work, we focus on the multichannel multi-speaker reverberant condition, and propose to extend our previous framework for end-to-end dereverberation, beamforming, and speech recognition with improved numerical stability and advanced frontend subnetworks including voice activity detection like masks. The techniques significantly stabilize the end-to-end training process. The experiments on the spatialized wsj1-2mix corpus show that the proposed system achieves about 35% WER relative reduction compared to our conventional multi-channel E2E ASR system, and also obtains decent speech dereverberation and separation performance (SDR=12.5 dB) in the reverberant multi-speaker condition while trained only with the ASR criterion.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (10)
  1. Wangyou Zhang (35 papers)
  2. Christoph Boeddeker (36 papers)
  3. Shinji Watanabe (416 papers)
  4. Tomohiro Nakatani (50 papers)
  5. Marc Delcroix (94 papers)
  6. Keisuke Kinoshita (44 papers)
  7. Tsubasa Ochiai (43 papers)
  8. Naoyuki Kamo (13 papers)
  9. Reinhold Haeb-Umbach (60 papers)
  10. Yanmin Qian (96 papers)
Citations (32)