Interpretability of self-supervised audio embeddings for clinical speech analysis
Determine which aspects of speech signals are captured by self-supervised audio embeddings and how those aspects relate to underlying clinical constructs, in order to establish their clinical interpretability and suitability for health applications.
References
While recent advances in self-supervised learning and foundational models have produced powerful audio embeddings (i. vectors that capture salient acoustic information but are not directly explainable), these representations lack direct physiological or clinical interpretability, making it unclear which aspects of the signal they capture and how these relate to underlying clinical constructs.
— Towards clinical adoption of voice and speech as measures of health: the need for harmonization
(2609.28894 - Cummins et al., 24 Sep 2026) in Section 1, subsection “Spotlight on feature extraction”