---
title: WMT25 Translation Evaluation Shared Task
url: https://www.emergentmind.com/topics/wmt25-translation-evaluation-shared-task
type: topic
---

# WMT25 Translation Evaluation Shared Task

The WMT25 Translation Evaluation Shared Task constitutes a centralized benchmarking initiative within the Workshop on Statistical Machine Translation (WMT), designed to catalyze improvements in both machine translation (MT) systems and their automated evaluation. It serves as the definitive annual testbed for multilingual MT, quality estimation, and evaluation methods, encompassing low-resource, high-resource, and novel language directions. At WMT25, the evaluation shared task focused on both the ranking of MT system outputs and the advancement of automatic and metric-based evaluation technologies, notably incorporating human-in-the-loop protocols and LLM-based assessment.

## 1. Task Structure and Participation

The WMT25 shared task was characterized by unprecedented scale and breadth, with 43 teams submitting systems (36 remaining after withdrawals/disqualifications), spanning both constrained (e.g., ≤20B parameters, public data only) and unconstrained categories (no model size/data restrictions) [2508.14909]. Participants included academic groups, commercial entities, and large foundation model providers. The task comprised two primary foci:

- **General MT Evaluation:** 32 language pairs, including high-resource (e.g., English↔Czech, English↔Chinese) and low-resource directions (e.g., English→Bhojpuri, English→Maasai), with test domains such as news, social media, speeches, and literary texts.
- **Translation Quality Estimation and Error Span Detection:** Automatic and reference-based system-level and segment-level evaluation; dedicated subtracks for the prediction of error spans according to MQM or ESA annotation schemes [2510.24707, 2509.13980].

Systems had to translate ~37,000 words per test direction, organized as 100-word paragraphs grouped into documents. Output collection adopted a document-first approach, falling back to paragraphs if necessary.

## 2. Data Resources and Preprocessing Pipelines

Both the development of MT systems and the construction of evaluation pipelines leveraged carefully curated multilingual corpora:

- **Parallel Data:** For high-resource languages, large-scale parallel corpora were available (e.g., OPUS, ParaCrawl, CCMatrix, WikiMatrix, EU-bookshop, MultiUN; up to billions of sentence pairs) [2508.12774]. For low-resource settings (Slavic-language and regional tracks), parallel data were scarce, necessitating aggressive data filtering, similarity-based sentence retrieval, and generation of synthetic parallel data via back-translation or LLM-based translation [2509.22490].
- **Monolingual and Synthetic Data:** To augment data for scarce language pairs (e.g., Upper Sorbian, Lower Sorbian, and Ukrainian), monolingual corpora and synthetic parallel data generated by LLM translation were crucial. Synthetic pairs were produced by translating monolingual sentences using parameter-efficiently fine-tuned LLMs (e.g., LoRA-adapted Qwen2.5-3B) [2509.22490]. Quality and domain adaptation were maintained via retrieval-augmented filtering and in-domain finetuning.
- **Preprocessing:** Corpus construction incorporated extensive cleaning—language detection, deduplication, normalization (e.g., punctuation, Unicode), and tokenization. Sentence embeddings (e.g., mean-pooled Qwen2.5-3B hidden states) facilitated similarity retrieval-based data curation, especially for low-resource development sets [2509.22490].

The preprocessing and data curation pipeline enabled robust system development even for challenging language pairs with minimal native resources.

## 3. Evaluation Metrics, Aggregation, and Human Judgement

Evaluation methodology at WMT25 prioritized metric diversity and robust statistical treatment, employing both reference-based and referenceless paradigms:

- **Reference-based Metrics:**
  - **BLEU:** Standard n-gram overlap metric with brevity penalty,
    $$\text{BLEU} = BP \cdot \exp\left(\sum_{n=1}^N w_n \log p_n\right)$$
    where $p_n$ is n-gram precision and $w_n$ are typical uniform weights [2509.22490, 2508.14909].
  - **chrF/ChrF++:** Generalized F-score over character n-grams (augmented with word n-grams for ChrF++),
    $$\text{chrF} = \frac{(1+\beta^2) \cdot \text{Precision}_{\text{char}} \cdot \text{Recall}_{\text{char}}}{\beta^2 \cdot \text{Precision}_{\text{char}} + \text{Recall}_{\text{char}}}$$
    with β configurable, typically β=2 [2509.22490, 2508.14909].
  - **COMET and Derivatives:** Supervised reference-based regression models trained to predict human adequacy scores; COMETKiwi variants support referenceless quality estimation paradigms [2510.24707, 2508.12774, 2508.14909].
- **LLM-based Judgement:**
  - **GEMBA-ESA:** LLM-as-judge scoring pipeline prompting GPT-4.1 and CommandA to rate adequacy and fluency in a reference-free mode, resulting in a 0–100 scale judgement [2508.14909].
- **Metric Aggregation:**
  - Disparate scales managed via median–interpercentile scaling and cross-metric averaging:
    $$
    z^\text{(m)}_s = \frac{x^\text{(m)}_s - \mu_m}{D_m}; \quad \bar{z}_s = \frac{1}{|M|}\sum_m z^\text{(m)}_s
    $$
    where $x^\text{(m)}_s$ is system s's average under metric m, $\mu_m$ is the median, $D_m$ the IQR, and $M$ is the set of metrics [2508.14909]. Final ranks are linearly mapped such that best = 1, worst = N.
- **Human-in-the-loop Evaluation:**
  - The final official ranking is determined by human judges (error-span annotation) rather than automatic scoring, due to the known limitations and biases of metric optimization, re-ranking, and metric overfitting [2508.14909].

This diverse and statistically grounded evaluation protocol ensures resilience to metric-specific overfitting and enhances cross-system comparability.

## 4. System Architectures and Training Paradigms

WMT25 submissions encompassed a spectrum from traditional NMT to advanced instruction-tuned large language models, with a strong emphasis on:

- **Foundation Model Adaptation:** Systems such as SALAMANDRATA (2B/7B), In2x (LLaMA-3/70B style), Gemma 3 (12B/27B) were adapted via continual pretraining (CPT) on parallel data followed by supervised instruction tuning [2508.12774, 2508.14472, 2510.24707].
- **Parameter-efficient Fine-tuning:** Widespread use of LoRA (Low-Rank Adaptation) enabled efficient training on large LLM backbones even in resource-constrained environments, by parameterizing updates with $\Delta W = AB,\, A \in \mathbb{R}^{d\times r},\, B \in \mathbb{R}^{r\times k},\, r\ll\min(d,k)$ [2509.22490].
- **Joint Task Learning:** Several systems, such as JGU Mainz’s Qwen2.5-3B-Instruct adaptation, performed joint multi-task training (e.g., MT + QA), leveraging a shared LLM backbone [2509.22490].
- **Curriculum and Synthetic Data:** Instructional datasets were constructed via clustering, labeling, creative rewriting, and chain-of-thought enhancement. Synthetic data generation pipelines, such as Magpie/Self-Instruct with critic LLM validation, bootstrapped resources for low-resource languages [2508.14472].
- **Reinforcement Learning (RL):** Fine-tuning objectives included custom RL reward models, such as rule-based or generative principle satisfaction (e.g., r_rule, r_gen), optimized via GRPO (Generalized PPO) [2508.14472].
- **Quality-aware Decoding:** Minimum Bayes Risk (MBR) decoding and re-ranking with QE models (e.g., COMET-KIWI) were prominent strategies for output selection, offering systematic gains in COMET and other human-correlated metrics [2508.12774].

Through these methodologies, system performance improved markedly in both challenging low-resource and general settings.

## 5. Results, Ablations, and Analytical Insights

WMT25 produced extensive quantitative insights and methodological lessons:

- **Performance Gains:**
  - Parameter-efficient finetuning (LoRA) on small LLMs coupled with synthetic/back-translated data resulted in absolute ChrF++ gains exceeding 50 points for Slavic low-resource targets [2509.22490].
  - Instruction-tuning and quality-aware decoding (MBR, TRR) delivered up to +3.4 COMET points over CPT alone and a further +1.9 with MBR in the SALAMANDRATA pipeline [2508.12774].
  - BLEU, chrF, and COMET improvements versus proprietary LLMs were observed in Japanese-related directions for In2x (e.g., BLEU en→ja: 35.2 vs. GPT-4.1’s 33.8; COMET en→ja: 0.72 vs. Gemini 0.68) [2508.14472].
- **Metric and Evaluation Analysis:**
  - Automatic metric aggregation reliably ranked systems but demonstrably favored those employing re-ranking and minimum-risk strategies, highlighting the importance of human evaluation for final system selection [2508.14909].
  - Context-rich (long-span) training and evaluation (concatenated segments, weighted score aggregation) led to substantially higher Pearson correlations with human judgment across DA, SQM, and MQM labels, especially for MQM (e.g., +0.328 Pearson gain for COMET-22-LS vs. COMET-22) [2509.13980].
  - Character-level F1 for generative error-span detection models approached or exceeded strong encoder-based baselines (e.g., GemSpanEval vs. XCOMET-XXL) in several language pairs [2510.24707].
- **Ablations and Error Analyses:**
  - Adding in-domain QA finetuning can erode MT performance if not tuned with care (ChrF++ drop of –0.9 on DSB vs. gain of +0.1 on HSB) [2509.22490].
  - For zero-shot transfer to Bhojpuri, including Hindi data in CPT was found to be critical (BLEU: 9.32 with EN–HI data, 0.35 without) [2508.12774].
  - LLM-based metric overfitting, paragraph-level scoring, and reference quality remain active challenges for metric reliability [2508.14909].

These analyses informed both methodological refinements and best practices.

## 6. Best Practices and Recommendations for Future WMT Evaluation

Key lessons and forward-looking guidance synthesized from WMT25 include:

- Leverage parameter-efficient adaptation (e.g., LoRA) and synthetic data generation (10–100k examples) as first-order strategies for low-resource MT [2509.22490].
- Apply similarity-based retrieval over large monolingual and parallel corpora to construct high-yield, domain-tailored training sets [2509.22490].
- Utilize hybrid metric pipelines, combining reference-based, reference-less, and LLM-as-judge scores via robust scaling and cross-metric aggregation [2508.14909].
- Conduct joint fine-tuning for multi-task LLMs (MT, QA, etc.), but iteratively tune in-domain adaptation rates to preserve task-specific performance [2509.22490].
- For evaluation, always report metric configurations, including SACREBLEU-tokenized BLEU and ChrF, and document the use of re-ranking or MBR decoding [2508.14909, 2508.12774].
- Emphasize human-evaluation results (e.g., error-span annotation) in system papers, as automatic metrics may not reflect true translation quality [2508.14909].

Collectively, these principles from WMT25 substantiate an increasingly sophisticated, multi-level approach to MT system development and validation, emphasizing both methodological rigor and practical adaptability across resource environments.

Source: https://www.emergentmind.com/topics/wmt25-translation-evaluation-shared-task