Determine how to integrate acoustic features into Layer 1

Determine whether and how to integrate MERT-derived music-audio representations into the proposed Layer 1 song-embedding architecture as an auxiliary input, soft target-label source, or parallel pathway.

Background

The current Layer 1 proposal uses canonical track identifiers and does not incorporate audio-derived features. MERT is identified as a possible source of acoustic representations, but the architecture could use those representations in several different ways. The paper explicitly states that the design choice remains unresolved.

References

Incorporating MERT, a pretrained self-supervised model for music-audio representation, is under consideration as either an auxiliary input channel, a source of soft target labels, or a separate parallel pathway. No decision has been made among these options.

Project Qualia: Recovering Experiential Music Structure from Session Co-occurrence Data  (2609.10862 - Mohammed et al., 9 Sep 2026) in Section 6.4, “Unresolved Design Parameters”