Papers
Topics
Authors
Recent
Search
2000 character limit reached

Unimodal Aggregation for CTC-based Speech Recognition

Published 15 Sep 2023 in cs.CL, cs.SD, and eess.AS | (2309.08150v2)

Abstract: This paper works on non-autoregressive automatic speech recognition. A unimodal aggregation (UMA) is proposed to segment and integrate the feature frames that belong to the same text token, and thus to learn better feature representations for text tokens. The frame-wise features and weights are both derived from an encoder. Then, the feature frames with unimodal weights are integrated and further processed by a decoder. Connectionist temporal classification (CTC) loss is applied for training. Compared to the regular CTC, the proposed method learns better feature representations and shortens the sequence length, resulting in lower recognition error and computational complexity. Experiments on three Mandarin datasets show that UMA demonstrates superior or comparable performance to other advanced non-autoregressive methods, such as self-conditioned CTC. Moreover, by integrating self-conditioned CTC into the proposed framework, the performance can be further noticeably improved.

Authors (2)
Definition Search Book Streamline Icon: https://streamlinehq.com
References (18)
  1. “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in ICASSP, 2016, pp. 4960–4964.
  2. “Towards end-to-end speech recognition with recurrent neural networks,” in International conference on machine learning, 2014, pp. 1764–1772.
  3. “Intermediate loss regularization for CTC-based speech recognition,” in ICASSP, 2021, pp. 6224–6228.
  4. “Relaxing the conditional independence assumption of CTC-based ASR by conditioning on intermediate predictions,” in Interspeech, 2021.
  5. “Mask CTC: Non-autoregressive end-to-end asr with CTC and mask predict,” in Interspeech, 2020.
  6. “Fast end-to-end speech recognition via non-autoregressive models and cross-modal knowledge transferring from BERT,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 1897–1911, 2021.
  7. Linhao Dong and Bo Xu, “CIF: Continuous integrate-and-fire for end-to-end speech recognition,” in ICASSP, 2020, pp. 6079–6083.
  8. “Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition,” in Interspeech, 2022.
  9. “Improving non-autoregressive end-to-end speech recognition with pre-trained acoustic and language models,” in ICASSP, 2022, pp. 8522–8526.
  10. “A comparative study on non-autoregressive modelings for speech-to-text generation,” in 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2021, pp. 47–54.
  11. “Conformer: Convolution-augmented Transformer for Speech Recognition,” in Interspeech, 2020, pp. 5036–5040.
  12. “E-Branchformer: Branchformer with enhanced merging for speech recognition,” in 2022 IEEE Spoken Language Technology Workshop (SLT), 2023, pp. 84–91.
  13. “AISHELL-1: An open-source mandarin speech corpus and a speech recognition baseline,” in 20th conference of the oriental chapter of the international coordinating committee on speech databases and speech I/O systems and assessment (O-COCOSDA), 2017, pp. 1–5.
  14. “AISHELL-2: Transforming mandarin asr research into industrial scale,” arXiv:1808.10583, 2018.
  15. “HKUST/MTS: A very large scale mandarin telephone speech corpus,” in Chinese Spoken Language Processing: 5th International Symposium, ISCSLP 2006, Singapore, December 13-16, 2006. Proceedings, 2006, pp. 724–735.
  16. “ESPnet: End-to-end speech processing toolkit,” in Interspeech, 2018, pp. 2207–2211.
  17. “Hybrid CTC/attention architecture for end-to-end speech recognition,” IEEE Journal of Selected Topics in Signal Processing, vol. 11, no. 8, pp. 1240–1253, 2017.
  18. “A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Translation, and Understanding Tasks,” in Interspeech, 2023, pp. 2208–2212.
Citations (1)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 2 tweets with 3 likes about this paper.