Listenable Maps for Audio Classifiers (2403.13086v3)

Published 19 Mar 2024 in cs.SD, cs.LG, eess.AS, and eess.SP

Abstract: Despite the impressive performance of deep learning models across diverse tasks, their complexity poses challenges for interpretation. This challenge is particularly evident for audio signals, where conveying interpretations becomes inherently difficult. To address this issue, we introduce Listenable Maps for Audio Classifiers (L-MAC), a posthoc interpretation method that generates faithful and listenable interpretations. L-MAC utilizes a decoder on top of a pretrained classifier to generate binary masks that highlight relevant portions of the input audio. We train the decoder with a loss function that maximizes the confidence of the classifier decision on the masked-in portion of the audio while minimizing the probability of model output for the masked-out portion. Quantitative evaluations on both in-domain and out-of-domain data demonstrate that L-MAC consistently produces more faithful interpretations than several gradient and masking-based methodologies. Furthermore, a user study confirms that, on average, users prefer the interpretations generated by the proposed technique.

References (40)

Citations (6)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Tweets

https://twitter.com/fpaissan_/status/1810672931090956544

https://twitter.com/mirco_ravanelli/status/1810674109459677435

https://twitter.com/AudioAndSpeech/status/1770710232588591551

Listenable Maps for Audio Classifiers (2403.13086v3)

Summary

Related Papers

Tweets