Papers
Topics
Authors
Recent
Search
2000 character limit reached

TidyVoice 2026 Challenge Evaluation Plan

Published 29 Jan 2026 in eess.AS and cs.SD | (2601.21960v1)

Abstract: The performance of speaker verification systems degrades significantly under language mismatch, a critical challenge exacerbated by the field's reliance on English-centric data. To address this, we propose the TidyVoice Challenge for cross-lingual speaker verification. The challenge leverages the TidyVoiceX dataset from the novel TidyVoice benchmark, a large-scale, multilingual corpus derived from Mozilla Common Voice, and specifically curated to isolate the effect of language switching across approximately 40 languages. Participants will be tasked with building systems robust to this mismatch, with performance primarily evaluated using the Equal Error Rate on cross-language trials. By providing standardized data, open-source baselines, and a rigorous evaluation protocol, this challenge aims to drive research towards fairer, more inclusive, and language-independent speaker recognition technologies, directly aligning with the Interspeech 2026 theme, "Speaking Together."

Summary

  • Tidyvoice 2026 challenge highlights the impact of language mismatch on cross-lingual speaker verification (ASV) systems, with performance metrics such as Equal Error Rate (EER) and minDCF to a large extend attributable to language mismatch conditions.
  • Training condition is open, allowing the use of external resources while enforcing disclosure of all external data, the facilitator uses a language-disjoint evaluation, drawing from TidyVoiceX 2026 dataset and splitting it into separate language groups and SPEAKER LANGUAGE with unseen languages evaluated for performance.
  • The baseline model uses SimAM-ResNet34, achieving an average EER of 3.07% on seen-language development data. EER rises to 9.79-11.59% Euclidean error deviation on unseen languages demonstrating performance degradation and solidifying need to focus on linguistic embeddings.

Motivation and task definition

The TidyVoice 2026 Challenge, organized as part of Interspeech 2026 under the theme "Speaking Together," addresses a well-documented weakness of modern speaker verification (ASV) systems: performance degradation under language mismatch between enrollment and test speech. The challenge is built on TidyVoiceX (Farhadipour et al., 22 Jan 2026), a curated multilingual partition of Mozilla Common Voice spanning approximately 40 languages, and is designed to isolate the effect of language switching from other sources of acoustic variability.

The task is standard speaker verification: given enrollment segments for a target speaker and a test segment, the system outputs a trial-level similarity score—formally framed as a log-likelihood ratio LLR(s)=log⁡(P(s∣H0)/P(s∣H1))\text{LLR}(s) = \log(P(s|H_0)/P(s|H_1)), though any threshold-compatible score is accepted. Trials are processed independently, with no cross-gender trials permitted.

Training condition and generalization design

The challenge uses an open training condition: participants may use any public or proprietary data, pre-trained models (VoxCeleb, VoxBlink2, wav2vec2/HuBERT/WavLM), and non-speech augmentation corpora such as MUSAN. The single strict restriction is that only the official TidyVoiceX training partition may be drawn from Mozilla Common Voice; all other MCV data is forbidden. All external resources must be fully disclosed in the system description.

The most consequential design choice is the language-disjoint evaluation. Training and development data cover 40 languages, while the final evaluation set (TidyVoiceX2) contains 38 entirely unseen languages. Two trial lists are scored:

  • tv26_eval-A: enrollment from seen languages, test from unseen languages (~4.0M trials)
  • tv26_eval-U: both enrollment and test from unseen languages (~1.28M trials)

Language labels are withheld during evaluation to prevent systems from exploiting language-specific calibration. This design directly tests whether speaker embeddings generalize across languages rather than merely interpolating within seen languages—a substantially harder requirement than prior multilingual benchmarks.

Metrics

Ranking is by Equal Error Rate (EER) as the primary metric, computed on pooled trials and on four trial-pair subsets crossing target/non-target status with same/different language. The secondary metric is minDCF with Cmiss=Cfa=1.0C_{\text{miss}} = C_{\text{fa}} = 1.0 and Ptar=0.01P_{\text{tar}} = 0.01, used for tie-breaking. The four-way subset analysis is central to the challenge's diagnostic purpose: it exposes whether a system exploits language coincidence as a discriminative cue rather than speaker identity.

Baseline results

The official baseline is a SimAM-ResNet34 implemented in WeSpeaker, pre-trained on VoxBlink2 + VoxCeleb2 and fine-tuned on TidyVoiceX Train with large-margin training. On the development set it achieves an overall EER of 3.07% (minDCF 0.82). The condition-wise breakdown is revealing:

Target pair language Non-target pair language EER (%)
Same Different 0.88
Same Same 2.97
Different Different 1.79
Different Same 5.19

The asymmetry is stark: the baseline performs best (0.88% EER) when targets share a language but non-targets do not—i.e., when language mismatch itself acts as a discriminative cue—and worst (5.19% EER) when targets are cross-lingual but non-targets share a language. This nearly six-fold spread demonstrates that the baseline relies substantially on language characteristics rather than purely speaker identity, which is precisely the failure mode the challenge is designed to penalize. On the unseen-language evaluation set, baseline EER rises to 9.06% on tv26_eval-A and 11.59% on tv26_eval-U, confirming that current embedding extractors generalize poorly to languages absent from training. These figures establish both a reproducible reference point and a clear headroom signal for participants.

Rules, protocol, and integrity safeguards

Evaluation proceeds through two phases: an offline development phase using released training/dev data (~50 GB via the Mozilla Data Collective), followed by online ranking on CodaBench during the evaluation phase. Key rules include:

  • Submission limit: at most three distinct systems per trial list; submissions must be complete (missing segments are rejected).
  • Eligibility: final ranking requires a system description paper submitted to the Interspeech 2026 special session.
  • Reproducibility: top-ranked teams must release model checkpoints and inference scripts sufficient to reproduce reported scores; unreproducible results lead to disqualification.
  • Privacy protections: re-identification of speakers or linkage of challenge data to external records ("data recombination") is strictly prohibited, with disqualification as the sanction.
  • Scope: verification only; closed- or open-set identification is out of scope.
  • Manual relabeling of challenge data is prohibited.

System descriptions must document front-end and back-end components, data partitions, external resources, dev-set performance, and per-trial CPU/GPU execution time and memory footprint.

Limitations and open questions

Several constraints bound what conclusions can be drawn from this evaluation. First, the corpus consists exclusively of read speech, which minimizes stylistic variability by design; results will not necessarily transfer to conversational or far-field conditions. Second, the open training condition means participants' absolute scores are not comparable across teams unless external data usage is identical—disclosure requirements mitigate but do not eliminate this confound. Third, because language labels are withheld at evaluation time, organizers cannot verify whether strong unseen-language performance reflects genuine language-independence or incidental robustness of particular pre-training mixtures. Finally, the baseline analysis shows language leakage in a single architecture; whether this pattern holds for self-supervised front-ends or disentanglement approaches (e.g., prefix-tuned cross-attention methods) remains an empirical question the challenge is positioned to answer but does not itself resolve.

Conclusion

The TidyVoice 2026 Challenge provides a standardized, reproducible benchmark for cross-lingual speaker verification, distinguished by its language-disjoint evaluation design and its diagnostic four-way trial-pair analysis. The baseline results—3.07% EER on seen-language development data versus 9–12% EER on unseen languages, with a 0.88% to 5.19% spread attributable to language-matching patterns—quantify the extent to which contemporary ASV embeddings conflate language and speaker information. By enforcing disclosure, prohibiting re-identification, and requiring artifact release, the evaluation plan aims to produce comparable, verifiable progress toward language-independent speaker recognition.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.