---
title: Clarity Prediction Challenge Corpus
url: https://www.emergentmind.com/topics/clarity-prediction-challenge-corpus
type: topic
---

# Clarity Prediction Challenge Corpus

The Clarity Prediction Challenge Corpus is the corpus and resource package that supports the prediction arm of the Clarity project, a five-year initiative to improve speech processing for hearing devices through public machine learning challenges [2006.11140]. Within this framework, prediction systems do not generate enhanced audio; rather, they model human speech perception by estimating speech intelligibility and quality for hearing-impaired listeners from processed acoustic signals, listener characteristics, and associated references such as clean speech and transcripts [2006.11140]. The corpus is therefore best understood as a multimodal benchmark for hearing-device speech perception modeling rather than as a collection of audio files alone.

## 1. Project context and paired-challenge design

The Clarity project is organized as **three paired challenges over five years**, with each pair containing an **enhancement challenge** and a **prediction challenge** [2006.11140]. The enhancement challenge is focused on hearing-device signal processing, whereas the prediction challenge is focused on predicting speech intelligibility and quality for hearing-impaired listeners [2006.11140]. This paired design establishes a pipeline in which simulated listener characteristics and binaural spatialized speech-in-noise signals are first processed by a baseline hearing-aid or enhancement model, and the result is then passed to a clarity prediction model [2006.11140].

The project is explicitly motivated by competition-driven progress in speech technology, including **CHiME** and similar ASR challenges, and it is designed to lower barriers for researchers from speech, machine learning, and hearing science to work on hearing impairment using open datasets, baseline models, and infrastructure [2006.11140]. In the prediction pipeline described for the first round, the prediction model includes a **hearing loss model** and a **speech intelligibility model**, and it uses both the processed audio and the clean speech or transcripts to produce predicted speech intelligibility scores [2006.11140].

A recurrent misconception is that the prediction task is a variant of speech enhancement. The corpus design makes the opposite distinction: enhancement modifies the signal, whereas prediction estimates how understandable that signal will be to listeners with different hearing abilities [2006.11140].

## 2. Corpus composition and representational scope

The resource package provided for the prediction challenge comprises **open-access datasets, models, and infrastructure** [2006.11140]. The project specifies five principal components: **tools for generating realistic test and training materials for different listening scenarios**, **baseline models of hearing impairment**, **baseline models of hearing-device processing**, **baseline models of speech perception**, and **databases of speech perception in noise** [2006.11140].

The speech-perception databases are defined to contain **listening-test results from hearing-impaired listeners** together with **comprehensive characterization of each listener’s hearing ability** [2006.11140]. In practice, this means that the corpus combines several modalities: speech and acoustic scenes, listener hearing profiles, perception or response data, and reference transcripts for the speech content [2006.11140]. The baseline schematic also states that the prediction model receives **clean speech recorded in anechoic conditions** and matching transcripts, while the figure caption adds the important clarification that **“Transcripts are not used by the baseline but will be provided to entrants.”** [2006.11140]

This composition is consequential for benchmark design. The corpus does not restrict prediction to signal-level degradation alone; it embeds intelligibility prediction in a listener-aware framework in which acoustic scene structure, hearing loss, and behavioral outcomes are jointly represented [2006.11140].

## 3. Round-one corpus design and listening-test protocol

For the first prediction scenario, the speech material consists of sentences from the **British National Corpus (BNC)**, produced by **British English speakers**, recorded in **anechoic conditions**, and then presented as speech in noise or in a room environment [2006.11140]. The acoustic scene is a **single speech source** in a cuboid **“living room”** that is **moderately reverberant** with **low to moderate reverberation time** [2006.11140]. The receiver position is randomly selected but must be at least **one meter from the source and the walls**, and the receiver faces the source [2006.11140]. Hearing aids have either **two or three microphones per ear**, and the scene includes multiple interferers, predominantly **non-speech noise**, such as television noise [2006.11140].

The listener panel for this round comprises **fifty listeners**, described as a mix of **healthy and impaired hearing** [2006.11140]. Each listener undergoes a **comprehensive hearing assessment in house**, but the listening itself is carried out at home via a **tablet and headphones** [2006.11140]. The protocol is sentence-based: listeners hear each sentence, **repeat the sentence verbally**, and speech intelligibility scores are computed from these responses [2006.11140].

The round-one design is controlled but intentionally ecologically relevant. It uses anechoic reference recordings, simulated domestic acoustics, hearing-aid microphone configurations, and home listening to connect laboratory-style benchmarking with conditions relevant to hearing-device deployment [2006.11140].

## 4. Prediction targets, evaluation, and challenge constraints

The explicit goal of the prediction challenge is to develop methods that can **predict speech intelligibility and quality for hearing-impaired listeners** [2006.11140]. In the first-round description, the main target is the **speech intelligibility score** for each audio item after hearing-aid processing, and systems are evaluated by how close their predicted intelligibility is to intelligibility measured from actual listeners [2006.11140].

The stated evaluation metric is **mean-squared error** between **predicted speech intelligibility** and **measured speech intelligibility** from the listener panel [2006.11140]. By contrast, the enhancement challenge ranks entries according to **mean intelligibility score across all audio signals in the test set**, and those rankings determine which systems are subsequently evaluated by the listener panel [2006.11140]. This asymmetry underscores the distinct scientific roles of the paired tasks: one optimizes signal processing, the other models perceptual outcome.

The corpus is embedded in a challenge protocol with several benchmark constraints [2006.11140]. Entries to the **enhancement** challenge must be **causal**, with no sample information more than **5 ms into the future** [2006.11140]. Prediction systems are scored on the test set using the listener-panel measurements [2006.11140]. Only **two entries from any one team** can be evaluated by the listener panel, teams must provide a **two-page technical document**, and **external data and pre-existing software must be disclosed** [2006.11140]. At the same time, there are **no limits** on computational cost, training data generation, or number of institutions, although teams must use the provided **audio files** and **signal generation tool** [2006.11140].

A plausible implication is that the corpus is designed simultaneously for algorithmic openness and evaluation discipline: entrants may innovate freely, but benchmark comparability is preserved through fixed signal generation, fixed audio inputs, and listener-panel scoring [2006.11140].

## 5. Later corpus instantiations: CPC2 and the 2023 dataset

Subsequent challenge editions instantiate the corpus under expanded conditions. The **second Clarity Prediction Challenge (CPC2)** provides a dataset of **speech signals, processed binaural outputs, and measured listener intelligibility scores** [2310.19817]. Speech is simulated in domestic acoustic environments and mixed with **noises, music, or competing speech**, and the database contains **speech–listener recognition performance pairs** [2310.19817]. CPC2 is divided into **three partitions**, each with a **training set** and an **evaluation set**, and performance is reported on the evaluation subsets using **RMSE**, **NCC**, and **KT** [2310.19817].

The **Clarity Prediction Challenge 2023 dataset** is described as a corpus of speech scenes associated with **6 talkers**, **10 enhancement methods** corresponding to **10 HA systems** from the 2022 Clarity Enhancement Challenge, and **25 listeners** who rated intelligibility scores [2309.09548]. It is organized into **three tracks**: **Track 1: 2779 utterances**, **Track 2: 2796 utterances**, and **Track 3: 2772 utterances** [2309.09548]. The corresponding test sets contain **305**, **294**, and **298** utterances, respectively, and the test condition includes **unseen listeners and unseen HA systems** [2309.09548].

| Edition | Supplied data | Evaluation setting |
|---|---|---|
| Round one | BNC sentences, simulated living-room scenes, listener assessments, intelligibility responses, transcripts and clean references | Mean-squared error against listener-panel intelligibility [2006.11140] |
| CPC2 | Speech signals, processed binaural outputs, measured listener intelligibility scores; three partitions with train and evaluation subsets | RMSE, NCC, KT on three evaluation subsets [2310.19817] |
| CPC 2023 | Speech scenes with 6 talkers, 10 enhancement methods, 25 listeners; three tracks and held-out test sets | RMSE, LCC, SRCC; unseen listeners and unseen HA systems [2309.09548] |

These later instantiations preserve the core objective of intelligibility prediction while broadening the generalization regime from the initial round-one configuration to multi-partition and multi-track evaluations involving new listeners and new hearing-aid systems [2310.19817].

## 6. Methodological uses and model classes supported by the corpus

The corpus has supported both **intrusive** and **non-intrusive** intelligibility prediction methods. On CPC2, one submission derives two predictors from a pretrained noise-robust ASR model: an intrusive system that uses **decoder hidden representations** from both reference and processed speech, and a non-intrusive system that relies on **utterance-level recognition uncertainty**, specifically **negative entropy** [2310.19817]. The same work emphasizes that the ASR backbone is trained only on a simulated noisy speech corpus rather than on CPC2 data, and it reports accurate prediction performance on the CPC2 evaluation subsets, which the authors interpret as evidence of robustness to unseen scenarios [2310.19817].

On the 2023 dataset, **MBI-Net+** extends the earlier MBI-Net system by replacing prior SSL features with **Whisper embeddings**, adding a **system classifier** over the **ten different enhancement systems**, and incorporating **HASPI** as a supplementary target in a multi-task objective [2309.09548]. The model is non-intrusive, operates on binaural hearing-aid input, and is reported to improve prediction performance by **7.11% RMSE** relative to the original MBI-Net on the development set; on the challenge-wide comparison, it ranks **third overall among non-intrusive systems**, with **RMSE 26.1** and **LCC 0.76** [2309.09548].

The corpus has also supported explicitly hearing-loss-aware intrusive modeling. A later study on the Clarity Prediction Challenge corpus simulates **reduced frequency resolution** via **cochlear filter broadening** and **reduced temporal resolution** via **low-pass filtering of temporal envelopes**, extracts **spectro-temporal modulation (STM)** representations, computes **normalized cross-correlation (NCC)** matrices between clean speech and speech in noise, and uses a **Vision Transformer** regressor to estimate intelligibility [2507.22599]. On that evaluation, the proposed method outperforms **HASPI v2** with a **16.5% relative RMSE reduction** for the **mild hearing-loss group** and a **6.1% reduction** for the **moderate-to-severe group** [2507.22599].

Taken together, these studies show that the corpus is sufficiently rich to support multiple modeling paradigms: ASR-derived transfer features, binaural non-intrusive deep architectures, and auditory-inspired intrusive predictors that explicitly simulate hearing-loss-related degradations [2310.19817].

## 7. Scientific role, scope, and conceptual boundaries

The corpus is intended to serve as a benchmark for **speech perception modeling in the context of hearing aids and hearing loss** [2006.11140]. Its practical uses include modeling listener-specific hearing ability, predicting intelligibility for spatialized, reverberant, noisy speech, comparing algorithms against real human perceptual data, and advancing hearing-device design and evaluation beyond signal-based metrics alone [2006.11140].

Two conceptual boundaries are central to its interpretation. First, the prediction task is **not** an enhancement task: it estimates perceptual outcome after enhancement rather than producing the enhanced signal itself [2006.11140]. Second, the corpus is **not just audio**: it combines speech scenes, binaural or hearing-aid processing outputs, listener hearing profiles or audiograms, transcripts or clean references, and human intelligibility measurements [2006.11140]. Later editions reinforce this structure by explicitly testing generalization to **unseen listeners** and **unseen HA systems** [2309.09548], while subsequent methods have used the corpus to stratify performance by hearing-loss severity and to analyze mild versus moderate-to-severe groups [2507.22599].

A plausible implication is that the Clarity Prediction Challenge Corpus occupies a distinctive position among speech benchmarks. It is neither a conventional speech-enhancement corpus nor a purely signal-based intelligibility dataset; rather, it is a listener-centered benchmark in which hearing-device processing, acoustic scene simulation, and human perceptual outcomes are evaluated within a unified challenge framework [2006.11140].

Source: https://www.emergentmind.com/topics/clarity-prediction-challenge-corpus