Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
41 tokens/sec
GPT-4o
60 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
8 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Towards Generalized Models for Task-oriented Dialogue Modeling on Spoken Conversations (2203.04045v1)

Published 8 Mar 2022 in cs.CL

Abstract: Building robust and general dialogue models for spoken conversations is challenging due to the gap in distributions of spoken and written data. This paper presents our approach to build generalized models for the Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations Challenge of DSTC-10. In order to mitigate the discrepancies between spoken and written text, we mainly employ extensive data augmentation strategies on written data, including artificial error injection and round-trip text-speech transformation. To train robust models for spoken conversations, we improve pre-trained LLMs, and apply ensemble algorithms for each sub-task. Typically, for the detection task, we fine-tune \roberta and ELECTRA, and run an error-fixing ensemble algorithm. For the selection task, we adopt a two-stage framework that consists of entity tracking and knowledge ranking, and propose a multi-task learning method to learn multi-level semantic information by domain classification and entity selection. For the generation task, we adopt a cross-validation data process to improve pre-trained generative LLMs, followed by a consensus decoding algorithm, which can add arbitrary features like relative \rouge metric, and tune associated feature weights toward \bleu directly. Our approach ranks third on the objective evaluation and second on the final official human evaluation.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (10)
  1. Ruijie Yan (2 papers)
  2. Shuang Peng (11 papers)
  3. Haitao Mi (56 papers)
  4. Liang Jiang (277 papers)
  5. Shihui Yang (1 paper)
  6. Yuchi Zhang (13 papers)
  7. Jiajun Li (66 papers)
  8. Liangrui Peng (2 papers)
  9. Yongliang Wang (36 papers)
  10. Zujie Wen (21 papers)
Citations (3)