Papers
Topics
Authors
Recent
Search
2000 character limit reached

RIR-Mega-Speech: A Reverberant Speech Corpus with Comprehensive Acoustic Metadata and Reproducible Evaluation

Published 25 Jan 2026 in eess.AS, cs.CL, cs.SD, and eess.SP | (2601.19949v1)

Abstract: Despite decades of research on reverberant speech, comparing methods remains difficult because most corpora lack per-file acoustic annotations or provide limited documentation for reproduction. We present RIR-Mega-Speech, a corpus of approximately 117.5 hours created by convolving LibriSpeech utterances with roughly 5,000 simulated room impulse responses from the RIR-Mega collection. Every file includes RT60, direct-to-reverberant ratio (DRR), and clarity index (C50C_{50}) computed from the source RIR using clearly defined, reproducible procedures. We also provide scripts to rebuild the dataset and reproduce all evaluation results. Using Whisper small on 1,500 paired utterances, we measure 5.20% WER (95% CI: 4.69--5.78) on clean speech and 7.70% (7.04--8.35) on reverberant versions, corresponding to a paired increase of 2.50 percentage points (2.06--2.98). This represents a 48% relative degradation. WER increases monotonically with RT60 and decreases with DRR, consistent with prior perceptual studies. While the core finding that reverberation harms recognition is well established, we aim to provide the community with a standardized resource where acoustic conditions are transparent and results can be verified independently. The repository includes one-command rebuild instructions for both Windows and Linux environments.

Authors (1)

Summary

  • The paper [provides RIR-Mega-Speech, a innovative reverberant speech corpus with over 117.5 hours of processed speech from LibriSpeech, offering detailed acoustic metadata] for each file, facilitating reproducible evaluations in reverberant automatic speech recognition (ASR) setting.
  • The corpus employs rigorous acoustic parameters like RT60, DRR, $C_{50}$ and incorporates per-utterance analyses to isolate the impact of reverberation, revealing a 48.2% relative increase in word error rate (WER) under reverberant conditions.
  • Detailed RT60 and DRR analyses highlight the robust degradation in ASR performance with increased reverberation and the marginal impact of loudness variation, positioning RIR-Mega-Speech as an essential tool for ASR researchers.

Motivation and contribution

Comparing methods for reverberant automatic speech recognition (ASR) is hampered by a persistent infrastructure problem: most existing corpora either lack per-file acoustic annotations, rely on non-redistributable or proprietary room impulse responses (RIRs), or provide insufficient documentation to reproduce published evaluations. The REVERB Challenge, CHiME-5/6, AISHELL-4, and AMI all contain reverberant speech but do not ship per-utterance RT60/DRR labels; RIR collections from the AEC and DNS challenges are not paired with transcribed speech; VCTK-based reverberant sets typically omit systematic acoustic annotation. The paper's stated contribution is deliberately not algorithmic novelty — the authors explicitly concede that the finding that reverberation degrades ASR is well established — but rather a standardized resource in which every file carries transparent acoustic metadata and every reported number can be independently regenerated.

The corpus, RIR-Mega-Speech, comprises approximately 117.5 hours of speech: roughly 5,200 LibriSpeech dev-clean/test-clean utterances convolved with about 5,000 simulated RIRs sampled from the RIR-Mega collection (Goswami, 21 Oct 2025), yielding 53,230 reverberant files. Each file is annotated with RT60, direct-to-reverberant ratio (DRR), clarity index (C50C_{50}), RMS loudness proxy, and duration, stored in a universal metadata CSV that permits filtering by acoustic condition without loading audio.

Corpus construction

Clean speech is drawn exclusively from LibriSpeech, chosen for reliable transcriptions and public availability; durations range from 1.5 to 36 seconds. Each clean utterance is convolved with up to ten randomly sampled RIRs via time-domain convolution y[n]=(x∗h)[n]y[n] = (x * h)[n], with exclusions for clipping or problematic metadata. Audio is stored as 16-bit PCM WAV at 16 kHz.

Acoustic parameters are computed from the source RIRs prior to convolution:

  • RT60 follows Schroeder backward integration with a line fit between −5 dB and −35 dB of the energy decay curve, extrapolated to −60 dB, per ISO 3382-1.
  • DRR uses an intentionally narrow 2.5 ms window centered on the first arrival, isolating only the true direct path rather than including early reflections up to 50 ms as is common in room acoustics.
  • C50C_{50} is the standard 50 ms clarity index.

The narrow DRR definition is a consequential design choice the authors flag explicitly: it produces extreme negative values (down to −141.96 dB) when the simulated direct peak is weak, values that may not align with perceptual relevance. Alternative definitions are deferred to future releases.

Splits are stratified by speaker (82.0% train / 8.7% dev / 9.3% test), preventing speaker leakage, but RIR selection is not stratified by acoustic parameters, so coverage across the RT60–DRR plane is uneven. The corpus statistics are summarized below:

Metric Mean Median Std Min Max
Duration (s) 7.96 6.52 4.94 1.52 36.07
RT60 (s) 0.44 0.36 0.25 0.09 1.51
DRR (dB) 3.32 6.58 22.11 −141.96 30.77

Coverage concentrates in RT60 of 0.2–0.6 s with DRR of 0–10 dB, reflecting office/classroom-dominated simulation; large halls appear in the tail but cathedrals, outdoor spaces, and vehicle interiors are absent. A useful property verified empirically is that utterance duration is uncorrelated with RT60, reducing confounding in WER-versus-acoustics analyses.

Evaluation methodology

The baseline evaluation uses Whisper small with default beam search (beam size 5), no hyperparameter tuning, and identical text normalization for references and hypotheses. The main comparison uses 1,500 paired utterances from the test split, where each utterance's clean version and one randomly selected reverberant variant are both decoded, enabling paired statistical tests that remove between-utterance variance. All uncertainty quantification uses nonparametric bootstrap at the utterance level (B=2000B = 2000 resamples) with 95% confidence intervals; trend analyses use fixed RT60 bins ([0.2, 0.4, 0.6, 0.8, 1.0, 1.2] s) and DRR bins ([−10, −5, 0, 5, 10, 15] dB). The authors note they do not adjust for multiple comparisons because the analyses are descriptive rather than confirmatory hypothesis tests.

Results

The headline result on the 1,500 paired utterances:

Condition WER (%)
Clean 5.20 (95% CI: 4.69–5.78)
Reverberant 7.70 (95% CI: 7.04–8.35)
Paired ΔWER +2.50 pp (95% CI: +2.06 to +2.98)
Relative increase +48.2%

Because the comparison is paired, the 2.50 percentage-point degradation is estimated with tight confidence intervals despite wide per-utterance scatter. Three trends emerge from the binned analyses:

  • RT60: WER rises monotonically from roughly 6% at RT60 = 0.2–0.4 s to about 10% at RT60 = 1.0–1.2 s, with non-overlapping CIs at the extremes.
  • DRR: WER decreases with increasing DRR, with the steepest effect below 0 dB and a plateau near clean-speech levels above 10 dB.
  • Duration: only a weak upward trend, indicating Whisper's chunked processing imposes little length penalty once acoustics are controlled.

The joint RT60–DRR heatmap shows the two dimensions interact: worst-case conditions combine long decay times with weak direct paths. These trends reproduce classical intelligibility findings (e.g., Nabelek and Pickett; Bradley et al.) but now quantified with confidence intervals on a rebuildable corpus.

Two ablations on a 500-utterance subset bound the effect. RMS loudness normalization to −20 dB yields 8.00% WER (CI: 6.57–9.58), statistically indistinguishable from the 7.70% baseline, so loudness variation is not a major difficulty factor here. Adding white noise at 10–15 dB SNR produces 30.95% WER (CI: 27.22–34.87) — a fourfold increase confirming that additive noise dominates over reverberation in this error budget, included as a sanity check rather than a substantive finding.

Error analysis of the 25 hardest utterances (WER above 50%) shows concentration at RT60 above 0.8 s and DRR below −5 dB, with errors dominated by phonetically similar consonant substitutions ("sit"→"zit", "bat"→"bad") and deletion of unstressed function words. Notably, two of eight manually inspected cases had reference transcriptions containing disfluencies, meaning some apparent model error is actually reference noise — a caveat on the absolute WER figures.

Limitations

The authors are candid about several constraints. All RIRs are physics-based simulations, which guarantee reproducible ground-truth parameters but may not capture diffraction around irregular objects, non-uniform surface scattering, furniture effects, HVAC noise, or time-varying conditions; validation against measured-RIR corpora such as REVERB or CHiME is recommended for generalization claims. Acoustic coverage is uneven because sampling was uniform rather than parameter-stratified. The 2.5 ms direct-only DRR definition is non-standard and produces physically implausible extreme values. Source speech is limited to read English audiobook material, excluding spontaneous speech, non-native accents, and other languages. Baseline coverage is limited to a single model (Whisper small); no dereverberation baselines or dev-set results are provided. A phoneme-level error analysis via forced alignment remains incomplete.

Conclusion

RIR-Mega-Speech provides what most reverberant corpora lack: per-file RT60/DRR/C50C_{50} metadata, speaker-stratified splits, bootstrap-based uncertainty reporting, and one-command rebuild scripts for Windows and Linux (2–3 h build time on 16 cores; 1–2 h evaluation on a single GPU). Its value lies in verifiability rather than novelty: the 48% relative WER degradation under reverberation and its monotonic dependence on RT60 and DRR are expected results, but they are now attached to reproducible ground truth. Planned extensions include alternative DRR definitions, STI metrics, more extreme acoustic coverage, additional model baselines, multilingual variants, and a lightweight subset for rapid benchmarking.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 2 tweets with 0 likes about this paper.