Papers
Topics
Authors
Recent
Search
2000 character limit reached

AfriMTE: MT Quality Evaluation for African Languages

Updated 30 June 2026
  • AfriMTE is a comprehensive dataset providing span-level error annotations and direct assessment scores for both adequacy and fluency in African machine translation.
  • It leverages a simplified MQM taxonomy and rigorous dual-annotator protocols, including statistical normalization and disagreement handling, to ensure reliable evaluation.
  • The dataset covers 13 diverse language pairs and underpins the development of learned MT metrics like AfriCOMET, advancing quality assessment in low-resource settings.

AfriMTE Dataset

AfriMTE is a large-scale, human evaluation dataset designed to address the lack of sentence-level, span-annotated, high-quality machine translation (MT) evaluation data for under-resourced African languages. Developed as part of the efforts reported by Muzny et al. (“AfriMTE and AfriCOMET: Enhancing COMET to Embrace Under-resourced African Languages”) (Wang et al., 2023), AfriMTE provides both adequacy and fluency judgments—paired with error-span annotations based on a simplified Multidimensional Quality Metrics (MQM) taxonomy—across 13 typologically and geographically diverse African languages. The resource underpins learned MT evaluation metrics, especially for languages and domains underserved by traditional n-gram-based metrics.

1. Purpose, Design, and Objectives

AfriMTE was created to fill a critical gap in MT evaluation for African languages, where previous work was hampered by the following obstacles:

  • Reliance on surface-form metrics (e.g., BLEU) with weak correlation to human judgments for morphologically rich, low-resource languages.
  • Absence of human-judged, error-labeled, sentence-level MT evaluation data.
  • Complexity and inapplicability of existing MQM guidelines for local annotators.
  • Limited human resources with bilingual proficiency in African–English or African–French pairs.

AfriMTE is structured to:

  • Provide span-level error labels and direct assessment (DA) scoring (0–100 scale) for both adequacy (faithfulness to source meaning) and fluency (well-formedness in the target language).
  • Support robust and interpretable training of learned reference-based and reference-free evaluation metrics (e.g., AfriCOMET).
  • Ensure coverage of multiple linguistic families and translation directions.

2. Language Coverage and Translation Directions

AfriMTE includes 13 language pairs (“LPs”) selected for typological breadth and data availability. These span Afro-Asiatic (Darija, Egyptian Arabic, Hausa, Somali), Atlantic-Congo (Igbo, Kikuyu, Swahili, Twi, isiXhosa, Yoruba), Nilo-Saharan (Luo), and Indo-European (English, French) families. The LPs and directions are:

Language Pair (LP) Direction
ary–fra Darija → French
eng–arz English → Egyptian Arabic
eng–fra English → French (control)
eng–hau English → Hausa
eng–ibo English → Igbo
eng–kik English → Kikuyu
eng–luo English → Luo
eng–som English → Somali
eng–swh English → Swahili
eng–twi English → Twi
eng–xho English → isiXhosa
eng–yor English → Yoruba
yor–eng Yoruba → English

“Eng–yor” also features four domain-specific devtest sets (IT, Movies, News, TED).

3. Dataset Construction: Source Data, Annotation, and Protocols

Sentences were selected from the FLORES-200 multi-way dev (270 sentences) and devtest (250 sentences) sets, yielding up to 520 source sentences per direction (with slightly reduced counts for certain languages due to post-hoc outlier / disagreement filtering).

Annotation involves:

  • Two or more native bilingual annotators per sentence.
  • Adequacy protocol: annotators highlight error spans in the MT output (categories: Addition, Omission, Mistranslation, Untranslated) and score adequacy (0–100).
  • Fluency protocol: annotators highlight errors (Grammar, Spelling, Typography, Unintelligible) and score fluency (0–100).

Aggregated, bias-corrected sentence scores are computed as

zij=rij−μjσjz_{ij} = \frac{r_{ij} - \mu_j}{\sigma_j}

where rijr_{ij} is the raw DA score, μj\mu_j and σj\sigma_j the mean and std. deviation per annotator. Final segment score: z‾i=1K∑j=1Kzij\overline{z}_i = \frac{1}{K} \sum_{j=1}^K z_{ij} where K≥2K \geq 2.

Disagreement handling removes annotations differing by >34 points between annotators. Inter-annotator agreement after this protocol was 0.797 (adequacy), 0.748 (fluency) using a leave-one-out Pearson correlation scheme.

4. Annotation Taxonomies and Guidelines

A simplified MQM-aligned error taxonomy enables reliable human annotation without trained professional linguists. Main error classes:

Adequacy (source + MT visible):

  • Addition
  • Omission
  • Mistranslation
  • Untranslated

Fluency (MT only):

  • Grammar
  • Spelling
  • Typography
  • Unintelligible

Each error is marked as a character span with its category.

Annotation guidelines reduce subjectivity by specifying anchors:

  • DA 100: perfect meaning/fluent, natural
  • DA 0: nonsense/no meaning/incomprehensible Scores are bucketed in coarse bands to minimize calibration drift.

5. Dataset Size, Coverage, and Statistics

After all quality control stages, per-language “qualified” adequacy annotation counts are as follows:

LP Qualified Annotations
ary–fra 394
eng–arz 518
eng–fra 515
eng–hau 490
eng–ibo 240
eng–kik 410
eng–luo 499
eng–som 434
eng–swh 352
eng–twi 516
eng–xho 494
eng–yor 484
yor–eng 439

Across all directions:

  • Median adequacy DA scores: 58.67 (“eng–swh”) to 100.00 (“eng–xho”)
  • Median fluency DA scores: 68.83 to 100.00
  • Average errors: 1–2 per sentence (mistranslation most frequent adequacy error)
  • Spearman correlation (error counts vs. DA): —0.675 (“Mistranslation”); —0.791 (“Total Errors”)

6. Data Formats, Metadata, and Accessibility

AfriMTE is distributed as triple-parallel files:

  • <source, MT output, reference>
  • JSON or CSV files enumerate annotator IDs, span-level error annotations, raw and normalized DA scores, and quality-control flags

Annotation guidelines and tool screenshots accompany the dataset for full replicability. There is no training/validation/test split; dev is intended for metric development, devtest for held-out evaluation.

7. Applications, Significance, and Relationship to Metrics

AfriMTE serves as the foundation for:

  • Training and evaluation of neural learned metrics such as AfriCOMET, optimized for robust, human-aligned MT quality estimation in African language settings.
  • Fine-grained MT error analysis for African language pairs, including in low-resource or zero-shot scenarios.
  • Research on annotator agreement, error typology prevalence, and cross-linguistic translation phenomena.

The dataset directly enables development of evaluation metrics demonstrating higher rank correlation with human judgments (e.g., AfriCOMET’s Spearman ρ = 0.441 with AfriMTE judgments) compared to BLEU, advancing the state of the art for African MT quality assessment (Wang et al., 2023).

8. Impact and Connections to Broader Evaluation Ecosystem

AfriMTE introduces the first systematic MQM-style, span-annotated, human-evaluation testbed for the diverse MT landscape of Africa. By concentrating on both adequacy and fluency, and providing robust agreement and documentation, it distinguishes itself from legacy “test sets” repurposed from translation or classification benchmarks. Its influence is documented in subsequent works leveraging AfriMTE for model selection, error analysis, and evaluation resource enrichment for African and related low-resource language tasks (Wang et al., 2023).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to AfriMTE Dataset.