Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
102 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

xCOMET: Transparent Machine Translation Evaluation through Fine-grained Error Detection (2310.10482v1)

Published 16 Oct 2023 in cs.CL

Abstract: Widely used learned metrics for machine translation evaluation, such as COMET and BLEURT, estimate the quality of a translation hypothesis by providing a single sentence-level score. As such, they offer little insight into translation errors (e.g., what are the errors and what is their severity). On the other hand, generative LLMs are amplifying the adoption of more granular strategies to evaluation, attempting to detail and categorize translation errors. In this work, we introduce xCOMET, an open-source learned metric designed to bridge the gap between these approaches. xCOMET integrates both sentence-level evaluation and error span detection capabilities, exhibiting state-of-the-art performance across all types of evaluation (sentence-level, system-level, and error span detection). Moreover, it does so while highlighting and categorizing error spans, thus enriching the quality assessment. We also provide a robustness analysis with stress tests, and show that xCOMET is largely capable of identifying localized critical errors and hallucinations.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (6)
  1. Nuno M. Guerreiro (27 papers)
  2. Ricardo Rei (34 papers)
  3. Daan van Stigt (3 papers)
  4. Luisa Coheur (33 papers)
  5. Pierre Colombo (48 papers)
  6. André F. T. Martins (113 papers)
Citations (82)