Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
41 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
41 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

On joint training with interfaces for spoken language understanding (2106.15919v3)

Published 30 Jun 2021 in cs.CL, cs.SD, and eess.AS

Abstract: Spoken language understanding (SLU) systems extract both text transcripts and semantics associated with intents and slots from input speech utterances. SLU systems usually consist of (1) an automatic speech recognition (ASR) module, (2) an interface module that exposes relevant outputs from ASR, and (3) a natural language understanding (NLU) module. Interfaces in SLU systems carry information on text transcriptions or richer information like neural embeddings from ASR to NLU. In this paper, we study how interfaces affect joint-training for spoken language understanding. Most notably, we obtain the state-of-the-art results on the publicly available 50-hr SLURP dataset. We first leverage large-size pretrained ASR and NLU models that are connected by a text interface, and then jointly train both models via a sequence loss function. For scenarios where pretrained models are not utilized, the best results are obtained through a joint sequence loss training using richer neural interfaces. Finally, we show the overall diminishing impact of leveraging pretrained models with increased training data size.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (9)
  1. Anirudh Raju (20 papers)
  2. Milind Rao (13 papers)
  3. Gautam Tiwari (7 papers)
  4. Pranav Dheram (7 papers)
  5. Bryan Anderson (2 papers)
  6. Zhe Zhang (181 papers)
  7. Chul Lee (21 papers)
  8. Bach Bui (2 papers)
  9. Ariya Rastrow (55 papers)
Citations (11)