Clarity Prediction Challenge Corpus
- Clarity Prediction Challenge Corpus is a multimodal benchmark that combines acoustic signals, listener profiles, and simulation data to predict speech intelligibility.
- The corpus underpins a paired challenge design that separates signal enhancement from intelligibility prediction for hearing-device research.
- It enables both intrusive and non-intrusive modeling approaches by providing comprehensive datasets and standard evaluation metrics for hearing-aid performance.
The Clarity Prediction Challenge Corpus is the corpus and resource package that supports the prediction arm of the Clarity project, a five-year initiative to improve speech processing for hearing devices through public machine learning challenges (Graetzer et al., 2020). Within this framework, prediction systems do not generate enhanced audio; rather, they model human speech perception by estimating speech intelligibility and quality for hearing-impaired listeners from processed acoustic signals, listener characteristics, and associated references such as clean speech and transcripts (Graetzer et al., 2020). The corpus is therefore best understood as a multimodal benchmark for hearing-device speech perception modeling rather than as a collection of audio files alone.
1. Project context and paired-challenge design
The Clarity project is organized as three paired challenges over five years, with each pair containing an enhancement challenge and a prediction challenge (Graetzer et al., 2020). The enhancement challenge is focused on hearing-device signal processing, whereas the prediction challenge is focused on predicting speech intelligibility and quality for hearing-impaired listeners (Graetzer et al., 2020). This paired design establishes a pipeline in which simulated listener characteristics and binaural spatialized speech-in-noise signals are first processed by a baseline hearing-aid or enhancement model, and the result is then passed to a clarity prediction model (Graetzer et al., 2020).
The project is explicitly motivated by competition-driven progress in speech technology, including CHiME and similar ASR challenges, and it is designed to lower barriers for researchers from speech, machine learning, and hearing science to work on hearing impairment using open datasets, baseline models, and infrastructure (Graetzer et al., 2020). In the prediction pipeline described for the first round, the prediction model includes a hearing loss model and a speech intelligibility model, and it uses both the processed audio and the clean speech or transcripts to produce predicted speech intelligibility scores (Graetzer et al., 2020).
A recurrent misconception is that the prediction task is a variant of speech enhancement. The corpus design makes the opposite distinction: enhancement modifies the signal, whereas prediction estimates how understandable that signal will be to listeners with different hearing abilities (Graetzer et al., 2020).
2. Corpus composition and representational scope
The resource package provided for the prediction challenge comprises open-access datasets, models, and infrastructure (Graetzer et al., 2020). The project specifies five principal components: tools for generating realistic test and training materials for different listening scenarios, baseline models of hearing impairment, baseline models of hearing-device processing, baseline models of speech perception, and databases of speech perception in noise (Graetzer et al., 2020).
The speech-perception databases are defined to contain listening-test results from hearing-impaired listeners together with comprehensive characterization of each listener’s hearing ability (Graetzer et al., 2020). In practice, this means that the corpus combines several modalities: speech and acoustic scenes, listener hearing profiles, perception or response data, and reference transcripts for the speech content (Graetzer et al., 2020). The baseline schematic also states that the prediction model receives clean speech recorded in anechoic conditions and matching transcripts, while the figure caption adds the important clarification that “Transcripts are not used by the baseline but will be provided to entrants.” (Graetzer et al., 2020)
This composition is consequential for benchmark design. The corpus does not restrict prediction to signal-level degradation alone; it embeds intelligibility prediction in a listener-aware framework in which acoustic scene structure, hearing loss, and behavioral outcomes are jointly represented (Graetzer et al., 2020).
3. Round-one corpus design and listening-test protocol
For the first prediction scenario, the speech material consists of sentences from the British National Corpus (BNC), produced by British English speakers, recorded in anechoic conditions, and then presented as speech in noise or in a room environment (Graetzer et al., 2020). The acoustic scene is a single speech source in a cuboid “living room” that is moderately reverberant with low to moderate reverberation time (Graetzer et al., 2020). The receiver position is randomly selected but must be at least one meter from the source and the walls, and the receiver faces the source (Graetzer et al., 2020). Hearing aids have either two or three microphones per ear, and the scene includes multiple interferers, predominantly non-speech noise, such as television noise (Graetzer et al., 2020).
The listener panel for this round comprises fifty listeners, described as a mix of healthy and impaired hearing (Graetzer et al., 2020). Each listener undergoes a comprehensive hearing assessment in house, but the listening itself is carried out at home via a tablet and headphones (Graetzer et al., 2020). The protocol is sentence-based: listeners hear each sentence, repeat the sentence verbally, and speech intelligibility scores are computed from these responses (Graetzer et al., 2020).
The round-one design is controlled but intentionally ecologically relevant. It uses anechoic reference recordings, simulated domestic acoustics, hearing-aid microphone configurations, and home listening to connect laboratory-style benchmarking with conditions relevant to hearing-device deployment (Graetzer et al., 2020).
4. Prediction targets, evaluation, and challenge constraints
The explicit goal of the prediction challenge is to develop methods that can predict speech intelligibility and quality for hearing-impaired listeners (Graetzer et al., 2020). In the first-round description, the main target is the speech intelligibility score for each audio item after hearing-aid processing, and systems are evaluated by how close their predicted intelligibility is to intelligibility measured from actual listeners (Graetzer et al., 2020).
The stated evaluation metric is mean-squared error between predicted speech intelligibility and measured speech intelligibility from the listener panel (Graetzer et al., 2020). By contrast, the enhancement challenge ranks entries according to mean intelligibility score across all audio signals in the test set, and those rankings determine which systems are subsequently evaluated by the listener panel (Graetzer et al., 2020). This asymmetry underscores the distinct scientific roles of the paired tasks: one optimizes signal processing, the other models perceptual outcome.
The corpus is embedded in a challenge protocol with several benchmark constraints (Graetzer et al., 2020). Entries to the enhancement challenge must be causal, with no sample information more than 5 ms into the future (Graetzer et al., 2020). Prediction systems are scored on the test set using the listener-panel measurements (Graetzer et al., 2020). Only two entries from any one team can be evaluated by the listener panel, teams must provide a two-page technical document, and external data and pre-existing software must be disclosed (Graetzer et al., 2020). At the same time, there are no limits on computational cost, training data generation, or number of institutions, although teams must use the provided audio files and signal generation tool (Graetzer et al., 2020).
A plausible implication is that the corpus is designed simultaneously for algorithmic openness and evaluation discipline: entrants may innovate freely, but benchmark comparability is preserved through fixed signal generation, fixed audio inputs, and listener-panel scoring (Graetzer et al., 2020).
5. Later corpus instantiations: CPC2 and the 2023 dataset
Subsequent challenge editions instantiate the corpus under expanded conditions. The second Clarity Prediction Challenge (CPC2) provides a dataset of speech signals, processed binaural outputs, and measured listener intelligibility scores (Tu et al., 2023). Speech is simulated in domestic acoustic environments and mixed with noises, music, or competing speech, and the database contains speech–listener recognition performance pairs (Tu et al., 2023). CPC2 is divided into three partitions, each with a training set and an evaluation set, and performance is reported on the evaluation subsets using RMSE, NCC, and KT (Tu et al., 2023).
The Clarity Prediction Challenge 2023 dataset is described as a corpus of speech scenes associated with 6 talkers, 10 enhancement methods corresponding to 10 HA systems from the 2022 Clarity Enhancement Challenge, and 25 listeners who rated intelligibility scores (Zezario et al., 2023). It is organized into three tracks: Track 1: 2779 utterances, Track 2: 2796 utterances, and Track 3: 2772 utterances (Zezario et al., 2023). The corresponding test sets contain 305, 294, and 298 utterances, respectively, and the test condition includes unseen listeners and unseen HA systems (Zezario et al., 2023).
| Edition | Supplied data | Evaluation setting |
|---|---|---|
| Round one | BNC sentences, simulated living-room scenes, listener assessments, intelligibility responses, transcripts and clean references | Mean-squared error against listener-panel intelligibility (Graetzer et al., 2020) |
| CPC2 | Speech signals, processed binaural outputs, measured listener intelligibility scores; three partitions with train and evaluation subsets | RMSE, NCC, KT on three evaluation subsets (Tu et al., 2023) |
| CPC 2023 | Speech scenes with 6 talkers, 10 enhancement methods, 25 listeners; three tracks and held-out test sets | RMSE, LCC, SRCC; unseen listeners and unseen HA systems (Zezario et al., 2023) |
These later instantiations preserve the core objective of intelligibility prediction while broadening the generalization regime from the initial round-one configuration to multi-partition and multi-track evaluations involving new listeners and new hearing-aid systems (Tu et al., 2023).
6. Methodological uses and model classes supported by the corpus
The corpus has supported both intrusive and non-intrusive intelligibility prediction methods. On CPC2, one submission derives two predictors from a pretrained noise-robust ASR model: an intrusive system that uses decoder hidden representations from both reference and processed speech, and a non-intrusive system that relies on utterance-level recognition uncertainty, specifically negative entropy (Tu et al., 2023). The same work emphasizes that the ASR backbone is trained only on a simulated noisy speech corpus rather than on CPC2 data, and it reports accurate prediction performance on the CPC2 evaluation subsets, which the authors interpret as evidence of robustness to unseen scenarios (Tu et al., 2023).
On the 2023 dataset, MBI-Net+ extends the earlier MBI-Net system by replacing prior SSL features with Whisper embeddings, adding a system classifier over the ten different enhancement systems, and incorporating HASPI as a supplementary target in a multi-task objective (Zezario et al., 2023). The model is non-intrusive, operates on binaural hearing-aid input, and is reported to improve prediction performance by 7.11% RMSE relative to the original MBI-Net on the development set; on the challenge-wide comparison, it ranks third overall among non-intrusive systems, with RMSE 26.1 and LCC 0.76 (Zezario et al., 2023).
The corpus has also supported explicitly hearing-loss-aware intrusive modeling. A later study on the Clarity Prediction Challenge corpus simulates reduced frequency resolution via cochlear filter broadening and reduced temporal resolution via low-pass filtering of temporal envelopes, extracts spectro-temporal modulation (STM) representations, computes normalized cross-correlation (NCC) matrices between clean speech and speech in noise, and uses a Vision Transformer regressor to estimate intelligibility (Zhou et al., 30 Jul 2025). On that evaluation, the proposed method outperforms HASPI v2 with a 16.5% relative RMSE reduction for the mild hearing-loss group and a 6.1% reduction for the moderate-to-severe group (Zhou et al., 30 Jul 2025).
Taken together, these studies show that the corpus is sufficiently rich to support multiple modeling paradigms: ASR-derived transfer features, binaural non-intrusive deep architectures, and auditory-inspired intrusive predictors that explicitly simulate hearing-loss-related degradations (Tu et al., 2023).
7. Scientific role, scope, and conceptual boundaries
The corpus is intended to serve as a benchmark for speech perception modeling in the context of hearing aids and hearing loss (Graetzer et al., 2020). Its practical uses include modeling listener-specific hearing ability, predicting intelligibility for spatialized, reverberant, noisy speech, comparing algorithms against real human perceptual data, and advancing hearing-device design and evaluation beyond signal-based metrics alone (Graetzer et al., 2020).
Two conceptual boundaries are central to its interpretation. First, the prediction task is not an enhancement task: it estimates perceptual outcome after enhancement rather than producing the enhanced signal itself (Graetzer et al., 2020). Second, the corpus is not just audio: it combines speech scenes, binaural or hearing-aid processing outputs, listener hearing profiles or audiograms, transcripts or clean references, and human intelligibility measurements (Graetzer et al., 2020). Later editions reinforce this structure by explicitly testing generalization to unseen listeners and unseen HA systems (Zezario et al., 2023), while subsequent methods have used the corpus to stratify performance by hearing-loss severity and to analyze mild versus moderate-to-severe groups (Zhou et al., 30 Jul 2025).
A plausible implication is that the Clarity Prediction Challenge Corpus occupies a distinctive position among speech benchmarks. It is neither a conventional speech-enhancement corpus nor a purely signal-based intelligibility dataset; rather, it is a listener-centered benchmark in which hearing-device processing, acoustic scene simulation, and human perceptual outcomes are evaluated within a unified challenge framework (Graetzer et al., 2020).