Papers
Topics
Authors
Recent
Search
2000 character limit reached

DIVERS-Bench: Robust LID Evaluation Suite

Updated 12 July 2026
  • DIVERS-Bench is a multilingual evaluation suite that assesses sentence-level language identification systems for domain robustness, language coverage, and code-switching sensitivity.
  • It aggregates diverse data from professional translations, social media, children’s stories, and speech transcripts to expose performance trade-offs and resource imbalances.
  • The benchmark shows that top-2 extraction from monolingual models fails on code-switched inputs, urging the development of multi-label methods for real-world applications.

Searching arXiv for the benchmark and closely related papers.

DIVERS-Bench is a comprehensive evaluation suite for language identification (LID) designed to stress-test sentence-level systems under domain shift and code-switching rather than only on clean, monolingual benchmarks. Introduced together with DIVERS-CS, its code-switching component, it aggregates multilingual data from professionally translated text, children’s stories, social media, speech transcripts, and diverse web text, then evaluates off-the-shelf LID models with standardized metrics for domain robustness, language coverage, resource imbalance, and multi-label failure modes (Ojo et al., 22 Sep 2025).

1. Concept and objectives

DIVERS-Bench, denoted $\csbench$ in the paper, is organized around three explicit goals: domain robustness, coverage and resource imbalance, and code-switching robustness. The benchmark addresses a recurrent pattern in LID evaluation: systems trained on clean, formal, monolingual corpora often appear strong on curated test sets but degrade on noisy, informal, and multilingual inputs. In this setting, LID is not treated as a solved preprocessing problem, but as a robustness problem whose difficulty depends strongly on domain and label space (Ojo et al., 22 Sep 2025).

The benchmark’s scope is sentence-level LID. For monolingual data, the task is standard multiclass classification over a language set L\mathcal{L}. For code-switched inputs, the benchmark reuses sentence-level softmax classifiers and evaluates their top-2 outputs against a two-language gold set. This design is deliberately diagnostic: it exposes the extent to which models trained for monolingual classification can or cannot be repurposed for realistic multilingual sentences.

A central implication of the benchmark design is that DIVERS-Bench measures both representational breadth and environmental robustness. Its evaluation is not limited to high-resource languages or standard web text, and it includes language groups mapped to resource classes $0$–$5$ using Joshi et al.’s taxonomy. This makes the benchmark simultaneously a domain-shift benchmark and a coverage benchmark.

2. Benchmark composition and data sources

The monolingual portion of DIVERS-Bench combines five datasets drawn from sharply different textual domains. These range from highly curated professional translations to informal microblogs and speech-related text. Labels from heterogeneous source conventions are normalized with pycountry to ISO codes, scripts are preserved, and evaluation uses standard provided test splits or the full dataset as appropriate (Ojo et al., 22 Sep 2025).

Dataset Domain and source Languages / size
FLORES-200 Professionally translated sentences from English Wikipedia 201 languages; ≈215,000\approx 215{,}000 sentences
SMOL Professionally translated diverse web and text sources 111 languages; ≈49,000\approx 49{,}000 sentences
Bloom Library Stories Community-created children’s story text 101 languages; ≈1,000\approx 1{,}000 story segments
TweetLID Twitter microblogs 8 languages; ≈17,000\approx 17{,}000 tweets
Common Voice sentences Text prompts from read-speech data 102 languages; ≈495,000\approx 495{,}000 sentence prompts

Each source plays a distinct evaluative role. FLORES-200 acts as the cleanest benchmark and is close to an upper-bound setting for many models. SMOL remains clean but is more heterogeneous and more heavily weighted toward under-represented languages. Bloom Library provides clean but informal narrative text. TweetLID stress-tests short, noisy, multilingual social media inputs. Common Voice sentences represent speech-related text with community-contributed variation, including local names and mixed registers.

This composition is methodologically important because it decouples multilinguality from cleanliness. A system that performs well on multilingual translation text is not thereby shown to be robust on community-generated stories, speech prompts, or microblogs. DIVERS-Bench formalizes that distinction rather than treating all sentence-level LID data as equivalent.

3. Task formulation and evaluation protocol

For monolingual inputs, DIVERS-Bench formulates LID as a multiclass mapping

f:X→L,f : X \to \mathcal{L},

with prediction

L\mathcal{L}0

and, for probabilistic models,

L\mathcal{L}1

The standard decision rule is

L\mathcal{L}2

For DIVERS-CS, the benchmark retains the same single-label classifiers but extracts the top-2 predicted labels,

L\mathcal{L}3

to compare against the gold two-language set. This is a constrained multi-label evaluation rather than a native multi-label model.

The monolingual evaluation uses three principal metrics. Macro F1 averages per-language F1 scores across the label set. False positive rate (FPR) measures spurious assignment of languages the input does not belong to. Resource-class-specific F1 averages F1 within each resource class, enabling direct comparison between high-resource and low-resource language groups.

The code-switching evaluation adds Full Match (FM) and Partial Match (PM). FM is the proportion of instances for which both languages are correctly identified:

L\mathcal{L}4

PM is the proportion for which at least one of the two languages is identified:

L\mathcal{L}5

The benchmark also reports Most Frequent Language (MFL), the language most often predicted among the top-2 outputs, to expose systematic bias toward dominant languages (Ojo et al., 22 Sep 2025).

A common misconception is that strong monolingual sentence classifiers can be evaluated on code-switched data simply by reading off their top-2 probabilities. DIVERS-Bench operationalizes exactly that procedure and shows that it usually fails badly at the FM level. The benchmark therefore treats top-2 extraction as a diagnostic baseline, not as an adequate solution to code-switching.

4. DIVERS-CS and sentence-level code-switching evaluation

DIVERS-CS is the code-switching component of DIVERS-Bench. It is a consolidated benchmark assembled from naturally occurring corpora with intra-sentential or fine-grained code-switching, clear language-pair annotations, and a preference for manually annotated data. Each instance is labeled at sentence level with a set of two languages, while token-level labels from the source corpora are used only to ensure that the examples are genuinely code-switched (Ojo et al., 22 Sep 2025).

Language pair Source(s) Train / test
en-de TongueSwitcher 30,000 / 253
en-hi SemEval-2020 Task 9, HinGE 17,360 / 7,440
zh-en CroCoSum, ASCEND 15,938 / 5,940
en-ta Dravidian-CodeMix 11,237 / 4,270
en-ml Dravidian-CodeMix 4,839 / 1,710
arz-en ArzEn-ST 2,160 / 1,812
eu-es BaSCo 1,087 / 412
arq-en Sabty (2021) 738 / 317
en-id Barik et al. (2019) 577 / 248
en-tr Yirmibeşoğlu and Eryiğit (2018) 260 / 112

The ten pairs span high-resource and low-resource settings, including dialectal Arabic varieties and Dravidian languages. Original train–test splits are preserved where available; otherwise, a L\mathcal{L}6 split is used. The benchmark’s sentence-level formulation intentionally ignores token-level switching boundaries during evaluation. This makes DIVERS-CS a test of whether global sentence classifiers can recover both constituent languages, not a token-tagging benchmark.

This design has a precise interpretive consequence. A model can achieve moderate PM by guessing one dominant language correctly while almost never identifying both languages. DIVERS-CS therefore distinguishes “detecting some multilinguality” from “correctly characterizing a code-switched sentence.” The reported MFL values, often dominated by English or another high-resource language, are central to that interpretation.

5. Evaluated models and empirical findings

DIVERS-Bench evaluates eight pre-existing LID systems without retraining: CLD3, FastText LID, Franc, LangDetect, LangID.py, GlotLID, OpenLID, and ConLID. These span n-gram neural models, trigram-frequency methods, Naive Bayes over n-grams, and large-coverage FastText-based systems, including ConLID’s supervised contrastive learning formulation (Ojo et al., 22 Sep 2025).

The headline result is domain-sensitive degradation. On the five monolingual datasets, ConLID records L\mathcal{L}7 F1 on FLORES, L\mathcal{L}8 on Common Voice, L\mathcal{L}9 on Bloom, $0$0 on SMOL, and $0$1 on TweetLID, for an average of $0$2. GlotLID records $0$3 on FLORES, $0$4 on Bloom, $0$5 on SMOL, $0$6 on TweetLID, and $0$7 on Common Voice, for an average of $0$8. OpenLID records $0$9 on FLORES, $5$0 on Bloom, $5$1 on SMOL, $5$2 on TweetLID, and $5$3 on Common Voice, averaging $5$4. These numbers show that near-ceiling results on clean translation text do not transfer uniformly to noisy or low-resource-heavy domains.

The benchmark also exposes a coverage–performance trade-off. On the supported-language subset used in the partial evaluation, narrow-coverage systems improve sharply: LangDetect rises from $5$5 F1 in full evaluation to $5$6 F1 on the supported subset, and CLD3 rises from about $5$7 to $5$8. Broad-coverage systems such as ConLID and GlotLID improve by only about $5$9 points. This indicates that narrow models are highly accurate where they apply, but unsuitable as general multilingual LID systems.

Resource imbalance is equally pronounced. On SMOL class-≈215,000\approx 215{,}0000 languages, CLD3 and LangDetect score ≈215,000\approx 215{,}0001 F1; Franc scores ≈215,000\approx 215{,}0002; FastText scores ≈215,000\approx 215{,}0003; OpenLID scores ≈215,000\approx 215{,}0004; ConLID and GlotLID reach ≈215,000\approx 215{,}0005 and ≈215,000\approx 215{,}0006, respectively. High-resource classes ≈215,000\approx 215{,}0007–≈215,000\approx 215{,}0008 often exceed ≈215,000\approx 215{,}0009–≈49,000\approx 49{,}0000 F1 for the strongest broad-coverage models, while low-resource classes ≈49,000\approx 49{,}0001–≈49,000\approx 49{,}0002 may fall below ≈49,000\approx 49{,}0003–≈49,000\approx 49{,}0004 or even to zero. The benchmark therefore makes low-resource brittleness visible even when aggregate macro scores appear acceptable.

The code-switching results are more severe. FM is near zero on many pairs. On en-de, OpenLID reaches ≈49,000\approx 49{,}0005 FM, GlotLID ≈49,000\approx 49{,}0006, ConLID ≈49,000\approx 49{,}0007, Franc ≈49,000\approx 49{,}0008, and LangDetect ≈49,000\approx 49{,}0009. On eu-es, OpenLID reaches ≈1,000\approx 1{,}0000 FM, the highest reported value for that pair. On en-id, GlotLID reaches ≈1,000\approx 1{,}0001 FM and OpenLID ≈1,000\approx 1{,}0002. On many pairs, including arz-en, zh-en, en-hi, en-ml, and en-ta, FM is effectively ≈1,000\approx 1{,}0003–≈1,000\approx 1{,}0004 for most models. PM, by contrast, is typically around ≈1,000\approx 1{,}0005–≈1,000\approx 1{,}0006, as in zh-en where ConLID reaches ≈1,000\approx 1{,}0007 PM, OpenLID ≈1,000\approx 1{,}0008, Franc ≈1,000\approx 1{,}0009, GlotLID ≈17,000\approx 17{,}0000, and LangDetect ≈17,000\approx 17{,}0001. The combination of very low FM and roughly ≈17,000\approx 17{,}0002 PM indicates that models often guess one constituent language but rarely recover both.

These findings support two precise conclusions. First, sentence-level softmax LID models are poorly matched to multi-label code-switched inputs. Second, robustness improvements aimed at out-of-domain monolingual text, such as ConLID’s supervised contrastive learning, do not by themselves solve code-switching.

6. Significance, limitations, and relation to adjacent benchmark traditions

DIVERS-Bench’s main contribution is methodological unification. It evaluates the same set of models across five monolingual domains with standardized metrics, adds resource-class analysis, and extends sentence-level evaluation to code-switching with FM and PM. This suggests a shift in LID benchmarking from “How accurate is the classifier on a clean multilingual test set?” to “How stable is the classifier across domain shifts, resource disparities, and multilingual mixing?” (Ojo et al., 22 Sep 2025).

The benchmark also clarifies several limitations of current practice. Training on clean monolingual corpora such as Wikipedia or curated web data does not ensure robustness on social media, speech transcripts, or community-generated content. Top-2 extraction from single-label classifiers is an inadequate approximation to multi-label sentence understanding. Evaluation restricted to supported-language subsets can conceal the operational costs of limited language coverage. A plausible implication is that future LID systems will need multi-domain training data, explicit validation on low-resource language groups, and architectures that support multi-label or token-level reasoning for code-switching.

Within the broader landscape of evaluation suites, DIVERS-Bench should be distinguished from several adjacent but separate “diversity benchmark” efforts. FairDiverse standardizes fairness- and diversity-aware ranking in search and recommendation rather than LID (Xu et al., 17 Feb 2025). EnsembleBench evaluates diversity and accuracy in model ensembles (Wu et al., 2020). The DUST work studies novelty-driven discovery of diverse unionable tuples in data lakes (Khatiwada et al., 31 Aug 2025). DivBench evaluates under- and over-diversification in text-to-image generation (Friedrich et al., 2 Jul 2025). These projects share a benchmark-oriented philosophy, but DIVERS-Bench’s technical object is specifically multilingual sentence-level language identification under domain shift and code-switching.

In that narrower sense, DIVERS-Bench is best understood as an evaluation suite for stress-testing the operational assumptions of LID. Its empirical record shows that strong performance on FLORES or similarly clean datasets is not sufficient evidence of real-world robustness, and that multilingual sentence classification remains fragile precisely where deployed systems are most likely to encounter noisy, low-resource, and mixed-language text.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to DIVERS-BENCH.