Improved Multi-Stage Training of Online Attention-based Encoder-Decoder Models (1912.12384v1)

Published 28 Dec 2019 in eess.AS, cs.LG, cs.SD, eess.SP, and stat.ML

Abstract: In this paper, we propose a refined multi-stage multi-task training strategy to improve the performance of online attention-based encoder-decoder (AED) models. A three-stage training based on three levels of architectural granularity namely, character encoder, byte pair encoding (BPE) based encoder, and attention decoder, is proposed. Also, multi-task learning based on two-levels of linguistic granularity namely, character and BPE, is used. We explore different pre-training strategies for the encoders including transfer learning from a bidirectional encoder. Our encoder-decoder models with online attention show 35% and 10% relative improvement over their baselines for smaller and bigger models, respectively. Our models achieve a word error rate (WER) of 5.04% and 4.48% on the Librispeech test-clean data for the smaller and bigger models respectively after fusion with long short-term memory (LSTM) based external LLM (LM).

Authors (6)

Abhinav Garg (11 papers)
Dhananjaya Gowda (16 papers)
Ankur Kumar (16 papers)
Kwangyoun Kim (18 papers)
Mehul Kumar (7 papers)
Chanwoo Kim (68 papers)

Citations (15)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Improved Multi-Stage Training of Online Attention-based Encoder-Decoder Models (1912.12384v1)

Summary

Related Papers