Significance of ITD-Based Equalisation-Cancellation Mechanisms in Music Scene Analysis

Determine the significance of interaural time difference (ITD)-based equalisation-cancellation mechanisms for spatial release from masking and musical scene analysis in music, including whether musical scene analysis prioritises simultaneous perception of multiple instruments rather than cancellation of competing sources.

Background

The paper evaluates spatial release from masking (SRM) in musical scene analysis (MSA) using auditory models and compares rendering conditions that provide interaural level differences (ILDs), interaural time differences (ITDs), and individually measured or augmented head-related transfer functions (HRTFs). For normal-hearing listeners, the models predict that adding ITDs to frequency-independent ILDs increases SRM, although the magnitude of this benefit differs between models.

The authors note that this model-based result differs from prior behavioural findings in speech-on-speech masking, where ITDs provided only a marginal additional benefit when ILDs were already present. They propose that music scene analysis may rely on different perceptual mechanisms from speech intelligibility, particularly because listeners may need to hear multiple instruments simultaneously rather than use binaural cancellation to suppress a competing talker. The behavioural and perceptual importance of ITD-based equalisation-cancellation mechanisms in music therefore remains unresolved.

References

However, the significance of ITD-based equalisation-cancellation mechanisms in music remains an open question; unlike BU speech models designed to cancel competing talkers, MSA may prioritise the ability to hear multiple instruments simultaneously.

Effects of HRTF Augmentation on Predicted Spatial Release from Masking in Music  (2608.28422 - Webb et al., 28 Aug 2026) in Section 4.1, “Spatial Cue Contributions to MSA” (page 4)