Papers
Topics
Authors
Recent
Search
2000 character limit reached

Libri2Mix: 2-Speaker Benchmark

Updated 14 July 2026
  • Libri2Mix is a simulated 2-speaker mixture corpus derived from LibriSpeech and WHAM, serving as a standardized benchmark for separation, TSE, and TS-ASR tasks.
  • It offers diverse configurations—including clean, noisy, different sampling rates, and overlap modes—to address various experimental setups and evaluation protocols.
  • The benchmark supports extensive evaluation metrics like SI-SDRi, PESQ, WER, and others, making it a versatile test bed for both signal-level and end-task performance studies.

Libri2Mix is a simulated 2-speaker mixture corpus within the LibriMix family, derived from LibriSpeech and WHAM, and used as a standard benchmark for monaural speech separation, target speaker extraction (TSE), target-speaker automatic speech recognition (TS-ASR), and related multi-speaker tasks. Across the literature, Libri2Mix appears in clean and noisy conditions, 8 kHz and 16 kHz releases, and both min and max mixture modes, with train-100, train-360, development, and test splits serving as recurring evaluation conventions (Ravenscroft et al., 2023, Huang et al., 2022, Wang et al., 2024).

1. Definition and corpus construction

Libri2Mix is repeatedly described as a simulated 2-speaker mixture corpus derived from LibriSpeech and WHAM. In one formulation, “speech samples come from the LibriSpeech corpus and noise samples come from the WHAM corpus,” and LibriMix uses LUFS rather than SSR to set loudness: speaker loudness lies between 25-25 and 33-33 LUFS, while noise loudness lies between 38-38 and 30-30 LUFS (Ravenscroft et al., 2023). In another formulation oriented toward separation, the benchmark includes a sep_clean condition with two speakers only and a sep_noisy condition with additive WHAM! noise, using the “16kHz min” version of LibriMix (Huang et al., 2022).

Published descriptions of corpus scale depend on which LibriMix subset is being used. A broader clean two-speaker description reports 64,700 training mixtures (about 270 hours), 3,000 development mixtures (about 11 hours), and 3,000 test mixtures (about 11 hours), with average mixture duration 14.8 s (Ge et al., 2022). By contrast, several train-100 protocols report 13,900 training mixtures, 3,000 validation mixtures, and 3,000 evaluation or test mixtures, reflecting the smaller Libri2Mix-100 configuration frequently used in TSE and TS-ASR studies (Zhang et al., 2023, Zhang et al., 2024).

The corpus is therefore not a single immutable release in experimental practice. It is better understood as a standardized mixture-generation framework whose instantiations differ by sampling rate, overlap convention, noise condition, and training subset.

2. Configurations, subsets, and protocol variation

A common misconception is that “Libri2Mix” denotes one fixed protocol. In practice, papers use materially different Libri2Mix configurations, and cross-paper comparisons are only meaningful when those choices are aligned.

Configuration Characteristics Representative use
Clean, min, 16 kHz Two speakers, no added noise Separation, TSE, SSL probing (Huang et al., 2022, Peng et al., 2024, Zhang et al., 2024)
Noisy, min, 16 kHz Two speakers plus WHAM! noise Noise-robust separation, TS-ASR, generative TSE (Zhang et al., 2023, Wang et al., 25 May 2025)
8 kHz protocols Often used to match earlier Conv-TasNet/DPRNN/SepFormer settings Separation and TSE baselines (Wang et al., 2024, Chen et al., 2023)
max, 100% overlap Full-length overlap, chosen for joint SD/SS/ASR and long-form separation studies Unified modeling and chunk-wise inference (Shakeel et al., 28 Aug 2025, Zorkina et al., 7 Jul 2026)
MC-Libri2Mix 4-channel reverberant extension of Libri2Mix Localized TSE and spatial modeling (Ge et al., 2022)

Subset usage is similarly heterogeneous. In one noisy-separation protocol, train-360 is about 212 hours and train-100 about 58 hours, with development and test each about 11 hours (Wang et al., 2024). In several 16 kHz train-100 protocols, the same split is identified instead by utterance count rather than hours, namely 13,900 training mixtures (Zhang et al., 2023, Zhang et al., 2024).

The sampling-strategy literature adds an important caveat: Libri2Mix train-100 and test do not have matched signal-length statistics. One study reports a difference in mean mixture length of 6.2 s and a difference in standard deviation of 1.79 s between train-100 and test (Ravenscroft et al., 2023). This has direct consequences for cropping, batching, and claims about training-efficiency trade-offs.

3. Task formulations built on Libri2Mix

Libri2Mix supports several distinct problem formulations. In noisy two-speaker separation, the observation model is written as

y=n+k=1Ksk,K=2,\mathbf{y} = \mathbf{n} + \sum_{k=1}^{K} \mathbf{s}_k,\qquad K=2,

with single-channel mixtures and clean source references available for supervised training and intrusive evaluation (Wang et al., 2024). In blind separation studies, the benchmark is therefore used to estimate both speakers from one waveform, usually with PIT-based objectives (Huang et al., 2022, Yip et al., 2023).

In TSE, the same corpus is reinterpreted as a speaker-conditioned extraction problem. A typical formulation is

y=x+i,\mathbf{y} = \mathbf{x} + \mathbf{i},

where x\mathbf{x} is the target speaker and i\mathbf{i} is the interfering speech or interference mixture, and an auxiliary enrollment utterance identifies which speaker should be recovered (Peng et al., 2024). In several TSE protocols, each 2-speaker mixture is used twice, once with speaker A as target and once with speaker B as target (Ge et al., 2022, Zhang et al., 2024).

Libri2Mix also underpins TS-ASR. In that setting, the model consumes the mixture waveform and a same-speaker enrollment utterance, then predicts only the target speaker’s transcription, sometimes without any explicit separation module (Zhang et al., 2023). Later work adapts foundation models such as Whisper to this setting and continues to use noisy Libri2Mix as the main simulated benchmark (Guo et al., 2024).

The benchmark has also been extended structurally. MC-Libri2Mix converts the original anechoic mixtures into a 4-channel reverberant corpus using a linear 4-microphone array with 5 cm spacing, pyroomacoustics-generated room impulse responses, room dimensions sampled from [5,10][5,10] m in length and width and [3,4][3,4] m in height, and 33-330 between 200 and 600 ms (Ge et al., 2022). More recent work further embeds Libri2Mix into joint diarization-separation-ASR training or into multi-condition PSE-style tasks such as mix_single, mix_clean, and mix_both (Shakeel et al., 28 Aug 2025, Huang, 4 Dec 2025).

4. Evaluation methodology and common metrics

Because Libri2Mix provides clean source references, it supports a broad range of intrusive metrics. Separation and TSE papers routinely report SI-SDRi or SI-SNRi, PESQ, and ESTOI or STOI (Wang et al., 2024, Zhang et al., 2024). Some TSE work also defines extraction accuracy as the proportion of utterances with SI-SDRi greater than 1 dB (Zhang et al., 2024).

Non-intrusive quality measures are also common. NISQA MOS is used in noisy separation studies on Libri2Mix to quantify reference-free speech quality (Wang et al., 2024), while DNSMOS P.835 appears in recent generative TSE work (Wang et al., 25 May 2025, Li et al., 11 Mar 2026). When Libri2Mix is paired with downstream applications, additional task-specific metrics enter the protocol: WER for TS-ASR (Zhang et al., 2023, Guo et al., 2024), cpWER for chunk-wise blind separation with downstream ASR (Zorkina et al., 7 Jul 2026), EER for speaker verification after separation (Zorkina et al., 7 Jul 2026), and DER for diarization in joint SD/SS/ASR systems (Shakeel et al., 28 Aug 2025).

Enrollment handling is another recurring methodological variable. Several TSE and TS-ASR works follow the SpeakerBeam-style informed protocol, using a clean utterance from the same speaker as enrollment, with official or pre-defined enrollment lists for development and test, and random same-speaker enrollment selection during training (Zhao et al., 2022, Zhang et al., 2023, Peng et al., 2024). This matters because Libri2Mix scores can depend not only on the mixture condition but also on how enrollment audio is sampled, cropped, or augmented.

5. Representative empirical results

Reported Libri2Mix numbers span multiple tasks and are not directly interchangeable across configurations. Even so, they illustrate why the corpus has become a reference point across separation, extraction, ASR, and unified multi-speaker modeling.

Setting Reported result on Libri2Mix Source
Blind speech separation 20.4 dB SI-SDRi (Yip et al., 2023)
Noisy separation with generative correction 12.98 dB SI-SNRi on Libri2Mix noisy test (Wang et al., 2024)
Clean TSE with multi-level speaker representation 15.91 dB SI-SDRi and 97.02% extraction accuracy (Zhang et al., 2024)
8 kHz TSE with MC-SpEx 14.61 dB SI-SDR, PESQ 3.195, ESTOI 84.9% (Chen et al., 2023)
Generative TSE with SoloSpeech PESQ 1.89, ESTOI 0.78, SI-SNR 11.12 dB, WER 0.15, SIM 0.96 (Wang et al., 25 May 2025)
TS-ASR with SQ-Whisper 14.6% WER on noisy Libri2Mix test (Guo et al., 2024)
Joint SD/SS/ASR with UME 1.37% DER on Libri2Mix evaluation set (Shakeel et al., 28 Aug 2025)

Several trends emerge from these results. First, Libri2Mix remains a strong discriminator among architectural choices. For example, SPGM reported 20.4 dB SI-SDRi on Libri2Mix, exceeding SepFormer by 0.3 dB while matching more parameter-heavy models on this benchmark (Yip et al., 2023). Second, the corpus is central to the evaluation of generative refiners and one-step generators: Fast-GeCo improved SepFormer from 10.58 dB to 12.98 dB SI-SNRi on the noisy test set (Wang et al., 2024), while AlphaFlowTSE reported 19.17 dB SI-SDR on clean Libri2Mix and 13.16 dB on noisy Libri2Mix together with PESQ 3.27 and 2.28 for clean and noisy conditions, respectively (Li et al., 11 Mar 2026).

Third, Libri2Mix is not only a signal-level benchmark. It is increasingly used to quantify downstream robustness. Earlier weakly supervised pre-training work reported 24.8%–24.9% WER for TS-HuBERT on noisy Libri2Mix test, improving over WavLM Base + cLN at 27.5% (Zhang et al., 2023). More recent TS-ASR adaptation with SQ-Whisper reduced noisy Libri2Mix test WER to 14.6% when train-360 and speed perturbation were added (Guo et al., 2024). This shift from SI-SDR-centric evaluation toward ASR, SV, and diarization endpoints is a notable feature of current Libri2Mix usage.

6. Limitations, extensions, and benchmark significance

Libri2Mix is synthetic, and that remains its defining limitation. Multiple papers note that it is built from LibriSpeech read speech and, in noisy variants, WHAM! noise, so it does not cover all real-world acoustic conditions, richer noise types, stronger reverberation, moving sources, or spontaneous conversational speech (Wang et al., 2024, Wang et al., 25 May 2025, Li et al., 11 Mar 2026). The benchmark’s statistical mismatch between train-100 and test length distributions adds a further caveat for studies of cropping, sampling, and efficiency (Ravenscroft et al., 2023).

At the same time, Libri2Mix has proved unusually extensible. MC-Libri2Mix adds multichannel reverberation and explicit spatial structure (Ge et al., 2022). Three-condition PSE-style tasks reinterpret it as mix_single, mix_clean, and mix_both, enabling universal TSE studies across one-speaker-plus-noise, two-speaker, and two-speaker-plus-noise conditions (Huang, 4 Dec 2025). Joint modeling work further embeds Libri2Mix within broader LibriMix settings that include 100% overlap and max mode for simultaneous diarization, separation, and ASR (Shakeel et al., 28 Aug 2025).

This suggests that Libri2Mix functions less as a single benchmark file tree than as a benchmark family: a controlled synthetic substrate on which different communities test local feature modeling, speaker conditioning, generative transport, chunk-wise inference, diarization heads, and downstream ASR/SV behavior. Its continuing value lies precisely in that role. It provides enough control for rigorous ablation, enough scale for modern pretrained models, and enough protocol diversity to expose when a method is exploiting a narrow configuration rather than learning a robust multi-speaker representation.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (16)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Libri2Mix.