Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
41 tokens/sec
GPT-4o
60 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
8 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Assessing Dialogue Systems with Distribution Distances (2105.02573v3)

Published 6 May 2021 in cs.CL

Abstract: An important aspect of developing dialogue systems is how to evaluate and compare the performance of different systems. Existing automatic evaluation metrics are based on turn-level quality evaluation and use average scores for system-level comparison. In this paper, we propose to measure the performance of a dialogue system by computing the distribution-wise distance between its generated conversations and real-world conversations. Specifically, two distribution-wise metrics, FBD and PRD, are developed and evaluated. Experiments on several dialogue corpora show that our proposed metrics correlate better with human judgments than existing metrics.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (6)
  1. Jiannan Xiang (11 papers)
  2. Yahui Liu (40 papers)
  3. Deng Cai (181 papers)
  4. Huayang Li (26 papers)
  5. Defu Lian (142 papers)
  6. Lemao Liu (62 papers)
Citations (16)