Unified representation of source content, transformation, and transformed audio
Develop a single audio representation model that simultaneously captures source content, the content-dependent transformation induced by audio processing, and the resulting transformed audio.
References
A complete representation of audio processing involves three aspects: the source content, the transformation, and the transformed audio. We addressed the latter two with distinct embeddings, and demonstrated that such embeddings should be not only informative but also geometrically organized for the downstream task. Extending the framework to capture all three aspects in a single representation model remains an open direction.
— Exploring the Design Space of Representation Learning for Audio Transformations
(2608.28127 - Lee et al., 28 Aug 2026) in Section 5, Discussion, page 5