Unified representation of source content, transformation, and transformed audio

Develop a single audio representation model that simultaneously captures source content, the content-dependent transformation induced by audio processing, and the resulting transformed audio.

Background

The paper distinguishes between a transformation embedding, intended to organize the processing applied to an audio signal, and a processed-audio embedding, intended to retain both source and processing information. Because audio transformations are often content-dependent, the two aspects cannot be completely separated in the proposed framework. The authors therefore identify the joint modeling of source content, transformation, and transformed audio within one representation model as an unresolved extension of their work.

References

A complete representation of audio processing involves three aspects: the source content, the transformation, and the transformed audio. We addressed the latter two with distinct embeddings, and demonstrated that such embeddings should be not only informative but also geometrically organized for the downstream task. Extending the framework to capture all three aspects in a single representation model remains an open direction.

Exploring the Design Space of Representation Learning for Audio Transformations  (2608.28127 - Lee et al., 28 Aug 2026) in Section 5, Discussion, page 5