Validate Layer 1 representations for downstream world modeling
Determine whether representations learned by the proposed Layer 1 song-embedding model are suitable state representations for a Layer 2 transition model and a Layer 3 session-level policy.
References
Several questions remain open as of this report: whether a JEPA-style objective recovers experiential structure more effectively than the residual-corrected Song2Vec embeddings already do; whether Layer 1 representations, once trained, are suitable state representations for a Layer 2 transition model and Layer 3 policy (a use motivated by the JEPA and world-model literature cited above but not evaluated on this dataset); and whether the cross-genre clustering in Section 4.3 reflects experiential interchangeability specifically, as distinct from a correlated confound, such as production era or energy level, that the residual-subtraction method does not disambiguate.