ICASSP 2023 Clarity Challenge Overview
- The paper highlights a major sim-to-real mismatch, where systems excel on simulated data but show limited gains on real-room recordings.
- ICASSP 2023 Clarity Challenge is a target-speaker-aware enhancement task in noisy domestic settings using both simulated and measured datasets.
- Evaluation incorporated fixed hearing aid amplification and dual metrics (HASPI and HASQI) to balance improvements in speech intelligibility and quality.
Searching arXiv for the ICASSP 2023 Clarity Challenge overview and closely related Clarity papers. The ICASSP 2023 SP Clarity Challenge was a machine-learning challenge on speech enhancement for hearing aids centered on a listener attending to a target speaker in a noisy, domestic environment with multiple interferers and head rotation. It formed part of the broader Clarity challenge program on hearing-device processing and speech perception modeling, and it extended the Second Clarity Enhancement Challenge (CEC2) by fixing the amplification stage of the hearing aid, evaluating systems with the average of HASPI and HASQI, and introducing two evaluation sets, one based on simulation and the other on real-room measurements. The principal outcome was a marked contrast between strong gains on the simulated set and much weaker gains on the measured set, foregrounding sim-to-real mismatch as a central research issue for hearing-aid enhancement (Cox et al., 2023).
1. Position within the Clarity program
The challenge belongs to the Clarity project, a five-year program organized around paired machine learning challenges for hearing-device processing (“enhancement”) and speech perception / intelligibility prediction (“prediction”). The project was designed to provide open-access datasets, baseline models, tools and infrastructure, and public challenge tasks for research on speech processing for hearing devices and on predicting intelligibility and quality for hearing-impaired listeners (Graetzer et al., 2020).
Within that program, the ICASSP 2023 challenge represents the enhancement side of the Clarity ecosystem in a specifically hearing-aid-oriented formulation. The underlying motivation is the longstanding problem of speech-in-noise understanding for hearing-impaired people, together with the observation that modern machine-learning methods from speech technology could be adapted to hearing-aid constraints if evaluated in an appropriate listening context (Graetzer et al., 2020).
A plausible implication is that the ICASSP 2023 challenge should be read not as an isolated benchmark, but as part of a broader effort to standardize hearing-aid enhancement research around common data, constraints, and evaluation procedures.
2. Task definition and acoustic scenario
The task was to process hearing-aid microphone inputs for a listener with hearing loss so as to improve both speech intelligibility and speech quality. The scenario was a listener attending to a target speaker in a noisy domestic environment, with multiple interferers and head rotation (Cox et al., 2023).
The input scenes were typical living rooms containing one target talker, up to three competing interferers, and listener motion. The challenge also provided target talker enrollment sentences, allowing systems to learn which speech to attend to. This makes the problem more specific than generic denoising: it is a target-speaker-aware enhancement task in a binaural/spatial hearing-aid setting (Cox et al., 2023).
The challenge extended CEC2 in three explicit ways. First, the hearing-aid amplification stage was fixed, so entrants could focus on the enhancement front-end. Second, the evaluation metric became the average of HASPI and HASQI, rather than HASPI alone. Third, a second evaluation set based on real-room measurements was introduced to test generalization beyond simulation (Cox et al., 2023).
This design places the challenge at the intersection of enhancement, spatial hearing, hearing-loss-aware processing, and robustness to distribution shift. A plausible implication is that success in this benchmark depends not only on suppression of interferers, but also on preserving cue structures compatible with downstream hearing-aid amplification and with objective intelligibility/quality models.
3. Data generation, evaluation sets, and measurement conditions
The challenge used two evaluation sets. Evaluation set 1 was simulation-based and followed the CEC2-style generation pipeline in which dry recordings of speech and noise were convolved with binaural room impulse responses from a geometric room acoustic model and applying HRTFs (Cox et al., 2023).
Evaluation set 2 was a measured set built from real-room recordings. The room was reported to have mid-frequency reverberation time 0.27 s, dimensions 6.6 × 5.8 × 2.8 m, and 5.7 dBA background noise. The measured dataset used live actors in a listening room, while still being post-processed into a format comparable to the challenge data (Cox et al., 2023).
The measured recording chain included a Neumann KM184 cardioid microphone at 50 cm for close speech reference, and a 1st-order Sennheiser AMBEO VR Ambisonic mic at the listener position. Interferers were later played and recorded using a M-Audio BX8a loudspeaker. Post-processing included head rotations via spherical harmonic rotation, HRTFs to obtain hearing-aid microphone signals, and mixing of target and interferers to the desired SNR (Cox et al., 2023).
The speech corpus for the measured set comprised 1,600 new sentences selected from the British National Corpus, read by 5 male and 5 female actors, ages 20 to 62. Each actor recorded 160 unique sentences in 10 talking positions. The close cardioid recordings served as the reference speech for HASPI and HASQI (Cox et al., 2023).
The measured set was intended to be more realistic than the simulation, but it was not fully equivalent to real hearing-aid use, because speech and noise were recorded separately and virtual head rotation was used. This suggests that the challenge probed an intermediate regime between controlled simulation and uncontrolled field recording (Cox et al., 2023).
4. Evaluation criterion and scoring logic
The challenge evaluated systems using the average of two objective metrics: HASPI and HASQI. HASPI is the Hearing Aid Speech Perception Index, described as an objective estimate of intelligibility. HASQI is the Hearing Aid Speech Quality Index, described as an objective estimate of sound quality or naturalness (Cox et al., 2023).
The challenge metric was defined as:
This combined criterion was adopted because some systems in CEC2 had improved intelligibility at the expense of naturalness; adding HASQI was intended to discourage such trade-offs (Cox et al., 2023).
The overview paper reports that across the successful systems on Eval1, HASPI and HASQI were highly correlated, with correlation coefficient using the best entry from each team. It also notes that systems often improved HASPI more than HASQI, and that quality gains tended to be about half the intelligibility gains (Cox et al., 2023).
This evaluation design is significant because it makes the challenge neither purely perceptual-quality-driven nor purely intelligibility-driven. A plausible implication is that entrants were incentivized to avoid aggressive enhancement strategies that raise intelligibility while introducing distortions detrimental to perceived quality.
5. Participation and official outcomes
The organizers report that 9 systems were submitted by 7 teams. The result table lists E02, E09, E14, E23, E28, E28d, E29, E29r, E30, and the E01 baseline. The paper additionally notes that E28d used additional data, E29r used head rotation information, and E01 was the baseline (Cox et al., 2023).
The principal results can be summarized as follows:
| Evaluation set | Baseline / best scores |
|---|---|
| Eval1 (simulated) | E01 baseline: Ave = 0.197; best simulated score: E28d = 0.693 |
| Eval2 (measured) | E01 baseline: Ave = 0.149; best measured score in table: E30 = 0.208 |
For Eval1, five teams improved over baseline, while two produced worse scores. Reported strong simulated-set entries included E29r = 0.616, E14 = 0.606, and E30 = 0.522. For Eval2, performance dropped sharply: alongside E30 = 0.208, the table reports E14 = 0.201, E28d = 0.199, E29 = 0.180, and E29r = 0.180 (Cox et al., 2023).
The organizers summarize the outcome by stating that the more ecologically valid Eval2 set produced lower scores for both objective measures across all teams, and that although four teams still managed to beat the baseline, the improvement was much smaller than on Eval1 (Cox et al., 2023).
The overview paper characterizes the overall result in balanced terms: modern speech enhancement methods can provide better signals for a simple hearing aid to amplify, but the improvement is marginal for the more ecologically-valid evaluation set based on real-room recording (Cox et al., 2023).
6. Sim-to-real mismatch and its interpretation
The central scientific conclusion of the challenge was that a mismatch between the simulated and measured data harmed machine-learning-based speech enhancement (Cox et al., 2023). This was not framed as a minor implementation issue, but as the defining limitation exposed by the 2023 benchmark.
The paper lists several concrete differences between the simulated and measured conditions. In simulation, talkers were close-miked in a studio, whereas in the measured set they were asked to speak to a distant microphone, altering speech style and prosody. The real listening room’s acoustics differed from the geometric-model approximation. Simulation used omnidirectional interferers, whereas the measured set used interferers reproduced through a loudspeaker with its own directivity. The measured Ambisonic microphone introduced electronic/transducer noise absent from simulation. Finally, simulation used sixth-order Ambisonics, whereas measurements used first-order Ambisonics (Cox et al., 2023).
Among these, the organizers single out three factors as most likely to explain the gap: transducer noise, first-order Ambisonics, and differences between real and simulated room impulse responses. Transducer noise was problematic because its onset timing differed from the interfering noises seen in training. Lower-order Ambisonics made it harder for systems to exploit binaural cues for noise suppression. Real-versus-simulated room impulse response differences introduced additional mismatch in the spatial-acoustic structure of the scenes (Cox et al., 2023).
This suggests that the ICASSP 2023 Clarity Challenge became, in effect, a benchmark for sim-to-real generalization in hearing-aid enhancement. A plausible implication is that progress on this challenge requires not only stronger network architectures, but tighter control of capture conditions, microphone characteristics, spatial representation, and acoustic realism across training and evaluation.
7. Subsequent systems and legacy
Later work treated the challenge as a platform for developing end-to-end low-latency hearing-aid enhancement systems. One example is a multi-stage low-latency enhancement system explicitly designed for the ICASSP 2023 Clarity Challenge. That system combined a monaural denoising module, a neural beamforming module, and a post-processing module for hearing-loss compensation; introduced an asymmetric window pair to satisfy the 5 ms latency requirement while retaining 16 ms spectral support; incorporated head rotation information; and used a differentiable post-processing stage compatible with the baseline NALR fitting algorithm (Ouyang et al., 6 Aug 2025).
That later paper reports challenge-dataset scale figures of 6000 scenes for training, 2500 scenes for validation, and 3000 scenes for evaluation, with each scene containing 6-channel BTE device recordings and a head rotation signal. It also reports that the submitted system achieved HASPI 0.835, HASQI 0.393, average 0.614, and that a variant using head rotation achieved HASPI 0.838, HASQI 0.393, average 0.616 on the official evaluation set (Ouyang et al., 6 Aug 2025).
These later results should be interpreted carefully relative to the 2023 overview, because they come from a separate system paper rather than the challenge overview itself. Nonetheless, they show the kinds of architectural directions the challenge encouraged: explicit low-latency design, phase-aware multi-stage enhancement, spatial processing, hearing-loss-aware post-processing, and use of auxiliary motion cues (Ouyang et al., 6 Aug 2025).
The 2023 overview concludes that future enhancement challenges should move toward measurements on real hearing aid microphones. This reflects the principal lesson of the benchmark: simulation-only success is insufficient evidence of practical robustness for hearing-aid enhancement, and ecologically valid evaluation conditions are indispensable for progress (Cox et al., 2023).