Attentive Contextual Carryover for Multi-Turn End-to-End Spoken Language Understanding (2112.06743v1)

Published 13 Dec 2021 in cs.CL and cs.AI

Abstract: Recent years have seen significant advances in end-to-end (E2E) spoken language understanding (SLU) systems, which directly predict intents and slots from spoken audio. While dialogue history has been exploited to improve conventional text-based natural language understanding systems, current E2E SLU approaches have not yet incorporated such critical contextual signals in multi-turn and task-oriented dialogues. In this work, we propose a contextual E2E SLU model architecture that uses a multi-head attention mechanism over encoded previous utterances and dialogue acts (actions taken by the voice assistant) of a multi-turn dialogue. We detail alternative methods to integrate these contexts into the state-ofthe-art recurrent and transformer-based models. When applied to a large de-identified dataset of utterances collected by a voice assistant, our method reduces average word and semantic error rates by 10.8% and 12.6%, respectively. We also present results on a publicly available dataset and show that our method significantly improves performance over a noncontextual baseline

PDF Abstract

Summarize Bookmark Chat (Pro)

Authors (11)

Kai Wei (30 papers)
Thanh Tran (52 papers)
Feng-Ju Chang (15 papers)
Kanthashree Mysore Sathyendra (10 papers)
Thejaswi Muniyappa (4 papers)
Jing Liu (525 papers)
Anirudh Raju (20 papers)
Ross McGowan (4 papers)
Nathan Susanj (12 papers)
Ariya Rastrow (55 papers)
Grant P. Strimel (21 papers)

Citations (10)

View on Semantic Scholar

Attentive Contextual Carryover for Multi-Turn End-to-End Spoken Language Understanding (2112.06743v1)

Related Papers