Unimodal Aggregation for CTC-based Speech Recognition
Abstract: This paper works on non-autoregressive automatic speech recognition. A unimodal aggregation (UMA) is proposed to segment and integrate the feature frames that belong to the same text token, and thus to learn better feature representations for text tokens. The frame-wise features and weights are both derived from an encoder. Then, the feature frames with unimodal weights are integrated and further processed by a decoder. Connectionist temporal classification (CTC) loss is applied for training. Compared to the regular CTC, the proposed method learns better feature representations and shortens the sequence length, resulting in lower recognition error and computational complexity. Experiments on three Mandarin datasets show that UMA demonstrates superior or comparable performance to other advanced non-autoregressive methods, such as self-conditioned CTC. Moreover, by integrating self-conditioned CTC into the proposed framework, the performance can be further noticeably improved.
- “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in ICASSP, 2016, pp. 4960–4964.
- “Towards end-to-end speech recognition with recurrent neural networks,” in International conference on machine learning, 2014, pp. 1764–1772.
- “Intermediate loss regularization for CTC-based speech recognition,” in ICASSP, 2021, pp. 6224–6228.
- “Relaxing the conditional independence assumption of CTC-based ASR by conditioning on intermediate predictions,” in Interspeech, 2021.
- “Mask CTC: Non-autoregressive end-to-end asr with CTC and mask predict,” in Interspeech, 2020.
- “Fast end-to-end speech recognition via non-autoregressive models and cross-modal knowledge transferring from BERT,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 1897–1911, 2021.
- Linhao Dong and Bo Xu, “CIF: Continuous integrate-and-fire for end-to-end speech recognition,” in ICASSP, 2020, pp. 6079–6083.
- “Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition,” in Interspeech, 2022.
- “Improving non-autoregressive end-to-end speech recognition with pre-trained acoustic and language models,” in ICASSP, 2022, pp. 8522–8526.
- “A comparative study on non-autoregressive modelings for speech-to-text generation,” in 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2021, pp. 47–54.
- “Conformer: Convolution-augmented Transformer for Speech Recognition,” in Interspeech, 2020, pp. 5036–5040.
- “E-Branchformer: Branchformer with enhanced merging for speech recognition,” in 2022 IEEE Spoken Language Technology Workshop (SLT), 2023, pp. 84–91.
- “AISHELL-1: An open-source mandarin speech corpus and a speech recognition baseline,” in 20th conference of the oriental chapter of the international coordinating committee on speech databases and speech I/O systems and assessment (O-COCOSDA), 2017, pp. 1–5.
- “AISHELL-2: Transforming mandarin asr research into industrial scale,” arXiv:1808.10583, 2018.
- “HKUST/MTS: A very large scale mandarin telephone speech corpus,” in Chinese Spoken Language Processing: 5th International Symposium, ISCSLP 2006, Singapore, December 13-16, 2006. Proceedings, 2006, pp. 724–735.
- “ESPnet: End-to-end speech processing toolkit,” in Interspeech, 2018, pp. 2207–2211.
- “Hybrid CTC/attention architecture for end-to-end speech recognition,” IEEE Journal of Selected Topics in Signal Processing, vol. 11, no. 8, pp. 1240–1253, 2017.
- “A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Translation, and Understanding Tasks,” in Interspeech, 2023, pp. 2208–2212.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Collections
Sign up for free to add this paper to one or more collections.