DIVERS-CS: Code-Switching LID Benchmark
- DIVERS-CS is a benchmark for sentence-level identification of naturally occurring code-switched text, emphasizing diverse language pairs and domains.
- It standardizes datasets from 10 language pairs with fine-grained annotations, revealing that monolingual LID models often underperform on mixed inputs.
- Empirical results show very low full match accuracy, highlighting biases and the need for multilabel or span-based approaches in real-world LID.
Searching arXiv for the benchmark paper and closely related code-switching LID work.
DIVERS-CS is a code-switching benchmark dataset introduced within the broader DIVERS-BENCH evaluation for language identification (LID) under domain shift and multilingual mixing. It is designed for the evaluation of sentence-level identification of naturally occurring, intra-sentential code-switched text, with an emphasis on diversity in language pairs, domains, and resource conditions. The benchmark aggregates and standardizes high-quality code-switching data spanning 10 language pairs, and its empirical use exposes a marked mismatch between strong performance on curated monolingual LID benchmarks and weak performance on mixed-language inputs, especially when systems must identify both languages present in a single sentence (Ojo et al., 22 Sep 2025).
1. Definition and scope
DIVERS-CS, expanded as DIVERSe Code-Switching, is a unified benchmark resource for code-switched language identification. Its target problem is not conventional monolingual classification, but the identification of the two languages present in a code-switched sentence. The benchmark focuses primarily on intra-sentential code-switching, i.e., switching within a single sentence, which the source paper describes as the most complex setting for LID (Ojo et al., 22 Sep 2025).
The benchmark is positioned as an evaluation resource rather than a new pretraining corpus. Its principal role is to standardize assessment across diverse code-switching conditions, including informal and noisy text. The underlying analysis argues that existing LID systems often overfit to clean, monolingual data and that DIVERS-CS makes this limitation measurable in a controlled, cross-pair setting (Ojo et al., 22 Sep 2025).
A central methodological feature is that evaluation is framed as sentence-level set prediction. Because many existing LID models are trained with a softmax objective for monolingual inputs, the benchmark evaluates them on code-switched sentences by taking the top-2 predicted languages as the predicted language set. This design directly probes whether a model can detect multiple languages within the same sentence rather than only assigning a dominant language label (Ojo et al., 22 Sep 2025).
2. Dataset construction and selection criteria
The construction of DIVERS-CS proceeds by collecting and consolidating naturally occurring code-switched datasets from prior work and multiple domains. The source paper states four explicit selection criteria: the included segments must be naturally occurring, must have reliable annotation with explicit fine-grained language-pair labels, must pass quality filtering that excludes sources with excessive noise or unclear labeling, and must contribute to diversity across language pairs, including settings beyond English-centric data (Ojo et al., 22 Sep 2025).
After manual filtering, the retained material is aggregated into a single benchmark with standardized train/test splits per language pair. When a source dataset did not already provide a split, a 70/30 split was used. This standardized aggregation is one of the benchmark’s core contributions, because prior code-switching resources are heterogeneous in provenance, annotation granularity, and partitioning (Ojo et al., 22 Sep 2025).
The benchmark draws from several previously used code-switching corpora. The paper names sources including SemEval-2020 Task 9, Dravidian-CodeMix, CroCoSum, ASCEND, ArzEn-ST, and BasCo. It also notes that some source corpora provide fine-grained word- or span-level labels, but that the benchmark evaluation presented here is on the sentence-level set-prediction task rather than token- or span-level tagging (Ojo et al., 22 Sep 2025).
This suggests that DIVERS-CS is intended less as a replacement for token-level code-switching datasets than as a common evaluation layer across heterogeneous source corpora. A plausible implication is that the benchmark is especially useful for stress-testing off-the-shelf LID systems that were not designed for structured multilingual segmentation.
3. Linguistic coverage and composition
DIVERS-CS spans 10 language pairs. The benchmark table in the source paper reports each pair’s code, resource class, and train/test counts. The non-English language’s resource class follows the Joshi et al. (2020) taxonomy when available (Ojo et al., 22 Sep 2025).
| Language pair | Code | #Train / #Test |
|---|---|---|
| English–German | en-de | 30,000 / 253 |
| English–Hindi | en-hi | 17,360 / 7,440 |
| Chinese–English | zh-en | 15,938 / 5,940 |
| English–Tamil | en-ta | 11,237 / 4,270 |
| English–Malayalam | en-ml | 4,839 / 1,710 |
| Egyptian–English | arz-en | 2,160 / 1,812 |
| Basque–Spanish | eu-es | 1,087 / 412 |
| Arabic–English | arq-en | 738 / 317 |
| English–Indonesian | en-id | 577 / 248 |
| English–Turkish | en-tr | 260 / 112 |
The reported resource classes are 5 for en-de, 4 for en-hi, 5 for zh-en, 3 for en-ta, 1 for en-ml, 3 for arz-en, 4 for eu-es, 3 for en-id, and 4 for en-tr, while arq-en is marked with no resource class in the reproduced table (Ojo et al., 22 Sep 2025).
The composition is diverse in several senses. First, it includes both high- and low-resource language pairs. Second, it includes pairs involving distinct scripts and orthographic systems, such as Chinese–English and Arabic varieties paired with English. Third, it extends beyond highly studied English-dominant pairs through the inclusion of Basque–Spanish and Arabic–English, among others. The paper also characterizes the data as spanning domains such as Twitter or microblogs, chatbots, interviews, YouTube comments, spoken dialog transcripts, and short written comments (Ojo et al., 22 Sep 2025).
The benchmark sentences are often conversational and informal, with switches occurring at phrase or word boundaries. The source text further notes the presence of non-standard spelling, colloquialisms, user-generated content, and code-mixed named entities. These properties are important because they move evaluation closer to real-world multilingual usage than curated monolingual LID test sets (Ojo et al., 22 Sep 2025).
4. Evaluation protocol and metrics
DIVERS-CS uses two main evaluation metrics: Full Match (FM) and Partial Match (PM). FM is the strict metric: it measures the percentage of instances for which both languages in a code-switched sentence are correctly identified. PM is more permissive: it measures the fraction of overlap between the true and predicted language sets, and therefore credits predictions that recover at least one of the languages present (Ojo et al., 22 Sep 2025).
The paper gives the following formulas:
Here, is the number of code-switched sentences for language pair , is the set of true languages for instance , and or denotes the model prediction (Ojo et al., 22 Sep 2025).
The paper evaluates several existing LID systems, including ConLID, OpenLID, GlotLID, FastText LID, Franc, LangDetect, LangID.py, and CLD3. Because many of these systems are architecturally monolingual classifiers, the benchmark operationalizes code-switching evaluation by interpreting the top-2 model probabilities as a predicted language set. This evaluation design does not modify the underlying models; rather, it reveals how poorly monolingual decision rules transfer to mixed-language input (Ojo et al., 22 Sep 2025).
A methodological significance of DIVERS-CS is therefore that it changes what counts as success in sentence-level LID. Under monolingual evaluation, predicting the dominant or most salient language may suffice. Under DIVERS-CS, success requires recognizing multilinguality itself.
5. Empirical findings
The central empirical result is that Full Match performance is extremely low across language pairs and models. For most pairs and most evaluated systems, FM scores are close to zero. Even on pairs that would ordinarily be regarded as comparatively favorable, such as English–German, English–Turkish, and Chinese–English, the evaluated models often fail to identify both languages in a single sentence (Ojo et al., 22 Sep 2025).
The paper reproduces FM scores for several representative models:
| Pair | ConLID | OpenLID | Franc | GlotLID | LangDetect |
|---|---|---|---|---|---|
| arz-en | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| zh-en | 0.00 | 0.38 | 0.00 | 0.00 | 0.00 |
| en-id | 0.00 | 9.27 | 0.81 | 10.08 | 0.00 |
| en-tr | 0.00 | 4.24 | 0.00 | 4.24 | 2.54 |
| en-de | 0.00 | 4.35 | 1.19 | 6.32 | 5.14 |
| en-hi | 0.44 | 0.00 | 0.00 | 0.00 | 0.01 |
| eu-es | 0.26 | 16.02 | 0.73 | 11.89 | 0.00 |
| en-ml | 0.31 | 0.29 | 0.00 | 0.29 | 0.06 |
| en-ta | 0.18 | 0.02 | 0.00 | 0.05 | 0.05 |
| eg-en | 0.14 | 2.24 | 0.00 | 1.09 | 0.00 |
The highest FM value reported in the summary is 16.02% for OpenLID on Basque–Spanish. Most other values are far lower, often effectively zero (Ojo et al., 22 Sep 2025).
By contrast, Partial Match scores are much higher, typically around 48–50% across pairs. The paper interprets this discrepancy as evidence that current systems often detect one language, usually English, but miss the fact that two languages are present. It also reports a “Most Frequent Language” analysis indicating a strong bias toward predicting English irrespective of the other language in the pair (Ojo et al., 22 Sep 2025).
This FM–PM gap is one of the benchmark’s most consequential observations. It isolates the failure mode precisely: the models are not merely uncertain; they are systematically mis-specified for multi-label language identification.
6. Challenges exposed by the benchmark
DIVERS-CS is explicitly presented as a diagnostic instrument for several technical weaknesses in current LID systems. The first is the incompatibility between monolingual softmax classification and sentence-level multi-label detection. A softmax objective presumes one correct class per input, whereas a code-switched sentence requires identification of more than one language (Ojo et al., 22 Sep 2025).
The second is a bias toward English. The benchmark analysis states that models generally predict only English or only one language per sentence, rarely both. This is notable because the failure persists not only for low-resource or orthographically complex pairs, but also for mid- and high-resource partners such as German, Hindi, and Spanish. The paper therefore argues that lack of data alone does not explain the observed deficiencies (Ojo et al., 22 Sep 2025).
The third is sensitivity to domain and script variation. DIVERS-CS includes speech transcripts, web text, social media texts, children’s stories, and code-switched text within the broader DIVERS-BENCH framework, while the DIVERS-CS subset specifically includes messy, noisy, informal, and spoken domains, as well as pairs with distinct scripts and orthographies. This heterogeneity appears to amplify failures that may remain hidden on curated monolingual benchmarks (Ojo et al., 22 Sep 2025).
The fourth is the challenge of real-world code-switching structure. The dataset includes word- and phrase-level switching, colloquial forms, non-standard spelling, and user-generated content. These properties complicate coarse language-level decisions and suggest that sentence-level classification alone may be insufficient for robust multilingual analysis (Ojo et al., 22 Sep 2025).
A common misconception would be that code-switched LID is largely solved for high-resource or cross-script pairs because the languages are “easy to tell apart.” The reported Full Match results directly contradict that view.
7. Research implications
The benchmark’s main conclusion is that code-switched LID is far from solved. Existing models with broad linguistic coverage still fail to recognize two languages in code-switched sentences at useful rates under strict evaluation (Ojo et al., 22 Sep 2025).
The paper argues that future systems should move beyond softmax-based, monolingual, single-label classification toward multilabel or span-based prediction. It explicitly points toward word-level or attention- and masking-based approaches, with MaskLID mentioned as an example direction in the summary text. It also emphasizes the need for training data that includes natural code-switched text across many language pairs, scripts, and domains (Ojo et al., 22 Sep 2025).
DIVERS-CS also has an infrastructural significance. By assembling and publicly releasing a unified evaluation benchmark, it enables reproducible comparison across code-switching settings that were previously fragmented across corpora and task definitions. This suggests a shift in evaluation norms for LID: performance on curated monolingual datasets can no longer be treated as sufficient evidence of robustness in multilingual, real-world environments.
Within the broader landscape of multilingual NLP, DIVERS-CS functions as both a benchmark and a critique. It demonstrates that sentence-level LID systems can exhibit superficially strong performance under conventional testing while failing at a core multilingual capability: recognizing that multiple languages may be present in the same utterance (Ojo et al., 22 Sep 2025).