---
title: ICASSP 2023 Clarity Challenge Overview
url: https://www.emergentmind.com/topics/icassp-2023-clarity-challenge
type: topic
---

# ICASSP 2023 Clarity Challenge Overview

Searching arXiv for the ICASSP 2023 Clarity Challenge overview and closely related Clarity papers.
The **ICASSP 2023 SP Clarity Challenge** was a machine-learning challenge on **speech enhancement for hearing aids** centered on a listener attending to a **target speaker in a noisy, domestic environment** with **multiple interferers** and **head rotation**. It formed part of the broader **Clarity** challenge program on hearing-device processing and speech perception modeling, and it extended the **Second Clarity Enhancement Challenge (CEC2)** by **fixing the amplification stage of the hearing aid**, evaluating systems with the **average of HASPI and HASQI**, and introducing **two evaluation sets**, one based on **simulation** and the other on **real-room measurements**. The principal outcome was a marked contrast between strong gains on the simulated set and much weaker gains on the measured set, foregrounding sim-to-real mismatch as a central research issue for hearing-aid enhancement [2311.14490].

## 1. Position within the Clarity program

The challenge belongs to the **Clarity** project, a five-year program organized around paired machine learning challenges for **hearing-device processing** (“enhancement”) and **speech perception / intelligibility prediction** (“prediction”). The project was designed to provide **open-access datasets**, **baseline models**, **tools and infrastructure**, and **public challenge tasks** for research on speech processing for hearing devices and on predicting intelligibility and quality for hearing-impaired listeners [2006.11140].

Within that program, the ICASSP 2023 challenge represents the enhancement side of the Clarity ecosystem in a specifically hearing-aid-oriented formulation. The underlying motivation is the longstanding problem of **speech-in-noise understanding for hearing-impaired people**, together with the observation that modern machine-learning methods from speech technology could be adapted to hearing-aid constraints if evaluated in an appropriate listening context [2006.11140].

A plausible implication is that the ICASSP 2023 challenge should be read not as an isolated benchmark, but as part of a broader effort to standardize hearing-aid enhancement research around common data, constraints, and evaluation procedures.

## 2. Task definition and acoustic scenario

The task was to process hearing-aid microphone inputs for a listener with hearing loss so as to improve both **speech intelligibility** and **speech quality**. The scenario was a listener attending to a target speaker in a **noisy domestic environment**, with **multiple interferers** and **head rotation** [2311.14490].

The input scenes were typical **living rooms** containing **one target talker**, **up to three competing interferers**, and listener motion. The challenge also provided **target talker enrollment sentences**, allowing systems to learn which speech to attend to. This makes the problem more specific than generic denoising: it is a **target-speaker-aware enhancement** task in a **binaural/spatial hearing-aid setting** [2311.14490].

The challenge extended CEC2 in three explicit ways. First, the **hearing-aid amplification stage was fixed**, so entrants could focus on the enhancement front-end. Second, the evaluation metric became the **average of HASPI and HASQI**, rather than HASPI alone. Third, a second evaluation set based on **real-room measurements** was introduced to test generalization beyond simulation [2311.14490].

This design places the challenge at the intersection of enhancement, spatial hearing, hearing-loss-aware processing, and robustness to distribution shift. A plausible implication is that success in this benchmark depends not only on suppression of interferers, but also on preserving cue structures compatible with downstream hearing-aid amplification and with objective intelligibility/quality models.

## 3. Data generation, evaluation sets, and measurement conditions

The challenge used two evaluation sets. **Evaluation set 1** was simulation-based and followed the CEC2-style generation pipeline in which **dry recordings of speech and noise were convolved with binaural room impulse responses from a geometric room acoustic model and applying HRTFs** [2311.14490].

**Evaluation set 2** was a measured set built from **real-room recordings**. The room was reported to have **mid-frequency reverberation time 0.27 s**, dimensions **6.6 × 5.8 × 2.8 m**, and **5.7 dBA background noise**. The measured dataset used **live actors in a listening room**, while still being post-processed into a format comparable to the challenge data [2311.14490].

The measured recording chain included a **Neumann KM184 cardioid microphone** at **50 cm** for close speech reference, and a **1st-order Sennheiser AMBEO VR Ambisonic mic** at the listener position. Interferers were later played and recorded using a **M-Audio BX8a loudspeaker**. Post-processing included **head rotations via spherical harmonic rotation**, **HRTFs to obtain hearing-aid microphone signals**, and mixing of target and interferers to the desired SNR [2311.14490].

The speech corpus for the measured set comprised **1,600 new sentences** selected from the **British National Corpus**, read by **5 male and 5 female actors**, ages **20 to 62**. Each actor recorded **160 unique sentences** in **10 talking positions**. The close cardioid recordings served as the reference speech for HASPI and HASQI [2311.14490].

The measured set was intended to be more realistic than the simulation, but it was not fully equivalent to real hearing-aid use, because **speech and noise were recorded separately** and **virtual head rotation** was used. This suggests that the challenge probed an intermediate regime between controlled simulation and uncontrolled field recording [2311.14490].

## 4. Evaluation criterion and scoring logic

The challenge evaluated systems using the average of two objective metrics: **HASPI** and **HASQI**. **HASPI** is the **Hearing Aid Speech Perception Index**, described as an objective estimate of intelligibility. **HASQI** is the **Hearing Aid Speech Quality Index**, described as an objective estimate of sound quality or naturalness [2311.14490].

The challenge metric was defined as:

$$
\text{Ave} = \frac{\text{HASPI} + \text{HASQI}}{2}
$$

This combined criterion was adopted because some systems in CEC2 had improved intelligibility at the expense of naturalness; adding HASQI was intended to discourage such trade-offs [2311.14490].

The overview paper reports that across the successful systems on Eval1, **HASPI and HASQI were highly correlated**, with correlation coefficient \( r = 0.943 \) using the best entry from each team. It also notes that systems often improved **HASPI** more than **HASQI**, and that quality gains tended to be about half the intelligibility gains [2311.14490].

This evaluation design is significant because it makes the challenge neither purely perceptual-quality-driven nor purely intelligibility-driven. A plausible implication is that entrants were incentivized to avoid aggressive enhancement strategies that raise intelligibility while introducing distortions detrimental to perceived quality.

## 5. Participation and official outcomes

The organizers report that **9 systems were submitted by 7 teams**. The result table lists **E02**, **E09**, **E14**, **E23**, **E28**, **E28d**, **E29**, **E29r**, **E30**, and the **E01 baseline**. The paper additionally notes that **E28d** used additional data, **E29r** used head rotation information, and **E01** was the baseline [2311.14490].

The principal results can be summarized as follows:

| Evaluation set | Baseline / best scores |
|---|---|
| **Eval1 (simulated)** | **E01 baseline: Ave = 0.197**; **best simulated score: E28d = 0.693** |
| **Eval2 (measured)** | **E01 baseline: Ave = 0.149**; **best measured score in table: E30 = 0.208** |

For Eval1, **five teams improved over baseline**, while **two produced worse scores**. Reported strong simulated-set entries included **E29r = 0.616**, **E14 = 0.606**, and **E30 = 0.522**. For Eval2, performance dropped sharply: alongside **E30 = 0.208**, the table reports **E14 = 0.201**, **E28d = 0.199**, **E29 = 0.180**, and **E29r = 0.180** [2311.14490].

The organizers summarize the outcome by stating that the more ecologically valid Eval2 set produced lower scores for both objective measures across all teams, and that although **four teams still managed to beat the baseline**, the improvement was much smaller than on Eval1 [2311.14490].

The overview paper characterizes the overall result in balanced terms: **modern speech enhancement methods can provide better signals for a simple hearing aid to amplify**, but the improvement is **marginal for the more ecologically-valid evaluation set based on real-room recording** [2311.14490].

## 6. Sim-to-real mismatch and its interpretation

The central scientific conclusion of the challenge was that a **mismatch between the simulated and measured data** harmed machine-learning-based speech enhancement [2311.14490]. This was not framed as a minor implementation issue, but as the defining limitation exposed by the 2023 benchmark.

The paper lists several concrete differences between the simulated and measured conditions. In simulation, talkers were **close-miked in a studio**, whereas in the measured set they were asked to speak to a distant microphone, altering speech style and prosody. The **real listening room’s acoustics** differed from the **geometric-model approximation**. Simulation used **omnidirectional interferers**, whereas the measured set used interferers reproduced through a **loudspeaker** with its own directivity. The measured Ambisonic microphone introduced **electronic/transducer noise** absent from simulation. Finally, simulation used **sixth-order Ambisonics**, whereas measurements used **first-order Ambisonics** [2311.14490].

Among these, the organizers single out three factors as most likely to explain the gap: **transducer noise**, **first-order Ambisonics**, and **differences between real and simulated room impulse responses**. Transducer noise was problematic because its onset timing differed from the interfering noises seen in training. Lower-order Ambisonics made it harder for systems to exploit **binaural cues** for noise suppression. Real-versus-simulated room impulse response differences introduced additional mismatch in the spatial-acoustic structure of the scenes [2311.14490].

This suggests that the ICASSP 2023 Clarity Challenge became, in effect, a benchmark for **sim-to-real generalization in hearing-aid enhancement**. A plausible implication is that progress on this challenge requires not only stronger network architectures, but tighter control of capture conditions, microphone characteristics, spatial representation, and acoustic realism across training and evaluation.

## 7. Subsequent systems and legacy

Later work treated the challenge as a platform for developing end-to-end low-latency hearing-aid enhancement systems. One example is a **multi-stage low-latency enhancement system** explicitly designed for the ICASSP 2023 Clarity Challenge. That system combined a **monaural denoising module**, a **neural beamforming module**, and a **post-processing module** for hearing-loss compensation; introduced an **asymmetric window pair** to satisfy the **5 ms latency requirement** while retaining **16 ms** spectral support; incorporated **head rotation information**; and used a **differentiable** post-processing stage compatible with the baseline **NALR fitting algorithm** [2508.04283].

That later paper reports challenge-dataset scale figures of **6000 scenes** for training, **2500 scenes** for validation, and **3000 scenes** for evaluation, with each scene containing **6-channel BTE device recordings** and a **head rotation signal**. It also reports that the submitted system achieved **HASPI 0.835**, **HASQI 0.393**, average **0.614**, and that a variant using head rotation achieved **HASPI 0.838**, **HASQI 0.393**, average **0.616** on the official evaluation set [2508.04283].

These later results should be interpreted carefully relative to the 2023 overview, because they come from a separate system paper rather than the challenge overview itself. Nonetheless, they show the kinds of architectural directions the challenge encouraged: explicit low-latency design, phase-aware multi-stage enhancement, spatial processing, hearing-loss-aware post-processing, and use of auxiliary motion cues [2508.04283].

The 2023 overview concludes that future enhancement challenges should move toward **measurements on real hearing aid microphones**. This reflects the principal lesson of the benchmark: simulation-only success is insufficient evidence of practical robustness for hearing-aid enhancement, and ecologically valid evaluation conditions are indispensable for progress [2311.14490].

Source: https://www.emergentmind.com/topics/icassp-2023-clarity-challenge