A Comparative Study on Non-Autoregressive Modelings for Speech-to-Text Generation (2110.05249v1)

Published 11 Oct 2021 in eess.AS, cs.CL, and cs.SD

Abstract: Non-autoregressive (NAR) models simultaneously generate multiple outputs in a sequence, which significantly reduces the inference speed at the cost of accuracy drop compared to autoregressive baselines. Showing great potential for real-time applications, an increasing number of NAR models have been explored in different fields to mitigate the performance gap against AR models. In this work, we conduct a comparative study of various NAR modeling methods for end-to-end automatic speech recognition (ASR). Experiments are performed in the state-of-the-art setting using ESPnet. The results on various tasks provide interesting findings for developing an understanding of NAR ASR, such as the accuracy-speed trade-off and robustness against long-form utterances. We also show that the techniques can be combined for further improvement and applied to NAR end-to-end speech translation. All the implementations are publicly available to encourage further research in NAR speech processing.

PDF Abstract

Summarize Bookmark Chat (Pro)

Authors (9)

Yosuke Higuchi (23 papers)
Nanxin Chen (30 papers)
Yuya Fujita (16 papers)
Hirofumi Inaguma (42 papers)
Tatsuya Komatsu (29 papers)
Jaesong Lee (8 papers)
Jumon Nozaki (4 papers)
Tianzi Wang (37 papers)
Shinji Watanabe (416 papers)

Citations (41)

View on Semantic Scholar

A Comparative Study on Non-Autoregressive Modelings for Speech-to-Text Generation (2110.05249v1)

Related Papers