Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Audio Adversarial Examples for Robust Hybrid CTC/Attention Speech Recognition (2007.10723v1)

Published 21 Jul 2020 in eess.AS and cs.SD

Abstract: Recent advances in Automatic Speech Recognition (ASR) demonstrated how end-to-end systems are able to achieve state-of-the-art performance. There is a trend towards deeper neural networks, however those ASR models are also more complex and prone against specially crafted noisy data. Those Audio Adversarial Examples (AAE) were previously demonstrated on ASR systems that use Connectionist Temporal Classification (CTC), as well as attention-based encoder-decoder architectures. Following the idea of the hybrid CTC/attention ASR system, this work proposes algorithms to generate AAEs to combine both approaches into a joint CTC-attention gradient method. Evaluation is performed using a hybrid CTC/attention end-to-end ASR model on two reference sentences as case study, as well as the TEDlium v2 speech recognition task. We then demonstrate the application of this algorithm for adversarial training to obtain a more robust ASR model.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (5)
  1. Ludwig Kürzinger (10 papers)
  2. Edgar Ricardo Chavez Rosas (2 papers)
  3. Lujun Li (30 papers)
  4. Tobias Watzel (4 papers)
  5. Gerhard Rigoll (49 papers)
Citations (4)

Summary

We haven't generated a summary for this paper yet.