Determine how to integrate acoustic features into Layer 1
Determine whether and how to integrate MERT-derived music-audio representations into the proposed Layer 1 song-embedding architecture as an auxiliary input, soft target-label source, or parallel pathway.
References
Incorporating MERT, a pretrained self-supervised model for music-audio representation, is under consideration as either an auxiliary input channel, a source of soft target labels, or a separate parallel pathway. No decision has been made among these options.
— Project Qualia: Recovering Experiential Music Structure from Session Co-occurrence Data
(2609.10862 - Mohammed et al., 9 Sep 2026) in Section 6.4, “Unresolved Design Parameters”