Distinguish experiential similarity from correlated acoustic or contextual properties
Determine whether the genre- and era-coherent structure recovered from artist-residual song embeddings reflects listeners’ shared experiential sense of musical feel rather than production era, instrumentation, tempo, or another correlated acoustic or contextual property.
References
But genre coherence and experiential similarity are distinct claims: a cluster could form because tracks share production era, instrumentation, or tempo, properties that overlap with genre by construction without being identical to it. The residual analysis demonstrates the signal is not reducible to artist identity; distinguishing whether it reflects a listener's sense of shared ``feel,'' specifically, versus a narrower or different kind of acoustic or contextual similarity, was outside the scope of the current analysis and remains open.
Several questions remain open as of this report: whether a JEPA-style objective recovers experiential structure more effectively than the residual-corrected Song2Vec embeddings already do; whether Layer 1 representations, once trained, are suitable state representations for a Layer 2 transition model and Layer 3 policy (a use motivated by the JEPA and world-model literature cited above but not evaluated on this dataset); and whether the cross-genre clustering in Section 4.3 reflects experiential interchangeability specifically, as distinct from a correlated confound, such as production era or energy level, that the residual-subtraction method does not disambiguate.