Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
4 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Deep Xi as a Front-End for Robust Automatic Speech Recognition (1906.07319v2)

Published 18 Jun 2019 in eess.AS, cs.SD, and eess.SP

Abstract: Current front-ends for robust automatic speech recognition(ASR) include masking- and mapping-based deep learning approaches to speech enhancement. A recently proposed deep learning approach toa prioriSNR estimation, called DeepXi, was able to produce enhanced speech at a higher quality and intelligibility than current masking- and mapping-based approaches. Motivated by this, we investigate Deep Xi as a front-end for robust ASR. Deep Xi is evaluated using real-world non-stationary and coloured noise sources at multiple SNR levels. Our experimental investigation shows that DeepXi as a front-end is able to produce a lower word error rate than recent masking- and mapping-based deep learning front-ends. The results presented in this work show that Deep Xi is a viable front-end, and is able to significantly increase the robustness of an ASR system. Availability: Deep Xi is available at:https://github.com/anicolson/DeepXi

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (2)
  1. Aaron Nicolson (13 papers)
  2. Kuldip K. Paliwal (8 papers)
Citations (12)

Summary

We haven't generated a summary for this paper yet.