---
title: 'AfriMTE: MT Quality Evaluation for African Languages'
url: https://www.emergentmind.com/topics/afrimte-dataset
type: topic
---

# AfriMTE: MT Quality Evaluation for African Languages

AfriMTE Dataset

AfriMTE is a large-scale, human evaluation dataset designed to address the lack of sentence-level, span-annotated, high-quality machine translation (MT) evaluation data for under-resourced African languages. Developed as part of the efforts reported by Muzny et al. (“AfriMTE and AfriCOMET: Enhancing COMET to Embrace Under-resourced African Languages”) [2311.09828], AfriMTE provides both adequacy and fluency judgments—paired with error-span annotations based on a simplified Multidimensional Quality Metrics (MQM) taxonomy—across 13 typologically and geographically diverse African languages. The resource underpins learned MT evaluation metrics, especially for languages and domains underserved by traditional n-gram-based metrics.

## 1. Purpose, Design, and Objectives

AfriMTE was created to fill a critical gap in MT evaluation for African languages, where previous work was hampered by the following obstacles:
- Reliance on surface-form metrics (e.g., BLEU) with weak correlation to human judgments for morphologically rich, low-resource languages.
- Absence of human-judged, error-labeled, sentence-level MT evaluation data.
- Complexity and inapplicability of existing MQM guidelines for local annotators.
- Limited human resources with bilingual proficiency in African–English or African–French pairs.

AfriMTE is structured to:
- Provide span-level error labels and direct assessment (DA) scoring (0–100 scale) for both adequacy (faithfulness to source meaning) and fluency (well-formedness in the target language).
- Support robust and interpretable training of learned reference-based and reference-free evaluation metrics (e.g., AfriCOMET).
- Ensure coverage of multiple linguistic families and translation directions.

## 2. Language Coverage and Translation Directions

AfriMTE includes 13 language pairs (“LPs”) selected for typological breadth and data availability. These span Afro-Asiatic (Darija, Egyptian Arabic, Hausa, Somali), Atlantic-Congo (Igbo, Kikuyu, Swahili, Twi, isiXhosa, Yoruba), Nilo-Saharan (Luo), and Indo-European (English, French) families. The LPs and directions are:

| Language Pair (LP) | Direction              |
|---------------------|------------------------|
| ary–fra             | Darija → French        |
| eng–arz             | English → Egyptian Arabic |
| eng–fra             | English → French (control) |
| eng–hau             | English → Hausa        |
| eng–ibo             | English → Igbo         |
| eng–kik             | English → Kikuyu       |
| eng–luo             | English → Luo          |
| eng–som             | English → Somali       |
| eng–swh             | English → Swahili      |
| eng–twi             | English → Twi          |
| eng–xho             | English → isiXhosa     |
| eng–yor             | English → Yoruba       |
| yor–eng             | Yoruba → English       |

“Eng–yor” also features four domain-specific devtest sets (IT, Movies, News, TED).

## 3. Dataset Construction: Source Data, Annotation, and Protocols

Sentences were selected from the FLORES-200 multi-way dev (270 sentences) and devtest (250 sentences) sets, yielding up to 520 source sentences per direction (with slightly reduced counts for certain languages due to post-hoc outlier / disagreement filtering).

Annotation involves:
- Two or more native bilingual annotators per sentence.
- Adequacy protocol: annotators highlight error spans in the MT output (categories: Addition, Omission, Mistranslation, Untranslated) and score adequacy (0–100).
- Fluency protocol: annotators highlight errors (Grammar, Spelling, Typography, Unintelligible) and score fluency (0–100).

Aggregated, bias-corrected sentence scores are computed as

\[
z_{ij} = \frac{r_{ij} - \mu_j}{\sigma_j}
\]
where \(r_{ij}\) is the raw DA score, \(\mu_j\) and \(\sigma_j\) the mean and std. deviation per annotator. Final segment score:
\[
\overline{z}_i = \frac{1}{K} \sum_{j=1}^K z_{ij}
\]
where \(K \geq 2\).

Disagreement handling removes annotations differing by >34 points between annotators. Inter-annotator agreement after this protocol was 0.797 (adequacy), 0.748 (fluency) using a leave-one-out Pearson correlation scheme.

## 4. Annotation Taxonomies and Guidelines

A simplified MQM-aligned error taxonomy enables reliable human annotation without trained professional linguists. Main error classes:

**Adequacy (source + MT visible):**
- Addition
- Omission
- Mistranslation
- Untranslated

**Fluency (MT only):**
- Grammar
- Spelling
- Typography
- Unintelligible

Each error is marked as a character span with its category.

Annotation guidelines reduce subjectivity by specifying anchors:
- DA 100: perfect meaning/fluent, natural
- DA 0: nonsense/no meaning/incomprehensible
Scores are bucketed in coarse bands to minimize calibration drift.

## 5. Dataset Size, Coverage, and Statistics

After all quality control stages, per-language “qualified” adequacy annotation counts are as follows:

| LP        | Qualified Annotations |
|-----------|----------------------|
| ary–fra   | 394                  |
| eng–arz   | 518                  |
| eng–fra   | 515                  |
| eng–hau   | 490                  |
| eng–ibo   | 240                  |
| eng–kik   | 410                  |
| eng–luo   | 499                  |
| eng–som   | 434                  |
| eng–swh   | 352                  |
| eng–twi   | 516                  |
| eng–xho   | 494                  |
| eng–yor   | 484                  |
| yor–eng   | 439                  |

Across all directions:
- Median adequacy DA scores: 58.67 (“eng–swh”) to 100.00 (“eng–xho”)
- Median fluency DA scores: 68.83 to 100.00
- Average errors: 1–2 per sentence (mistranslation most frequent adequacy error)
- Spearman correlation (error counts vs. DA): —0.675 (“Mistranslation”); —0.791 (“Total Errors”)

## 6. Data Formats, Metadata, and Accessibility

AfriMTE is distributed as triple-parallel files:
- <source, MT output, reference>
- JSON or CSV files enumerate annotator IDs, span-level error annotations, raw and normalized DA scores, and quality-control flags

Annotation guidelines and tool screenshots accompany the dataset for full replicability. There is no training/validation/test split; dev is intended for metric development, devtest for held-out evaluation.

## 7. Applications, Significance, and Relationship to Metrics

AfriMTE serves as the foundation for:
- Training and evaluation of neural learned metrics such as AfriCOMET, optimized for robust, human-aligned MT quality estimation in African language settings.
- Fine-grained MT error analysis for African language pairs, including in low-resource or zero-shot scenarios.
- Research on annotator agreement, error typology prevalence, and cross-linguistic translation phenomena.

The dataset directly enables development of evaluation metrics demonstrating higher rank correlation with human judgments (e.g., AfriCOMET’s Spearman ρ = 0.441 with AfriMTE judgments) compared to BLEU, advancing the state of the art for African MT quality assessment [2311.09828].

## 8. Impact and Connections to Broader Evaluation Ecosystem

AfriMTE introduces the first systematic MQM-style, span-annotated, human-evaluation testbed for the diverse MT landscape of Africa. By concentrating on both adequacy and fluency, and providing robust agreement and documentation, it distinguishes itself from legacy “test sets” repurposed from translation or classification benchmarks. Its influence is documented in subsequent works leveraging AfriMTE for model selection, error analysis, and evaluation resource enrichment for African and related low-resource language tasks [2311.09828].

Source: https://www.emergentmind.com/topics/afrimte-dataset