Validate Layer 1 representations for downstream world modeling

Determine whether representations learned by the proposed Layer 1 song-embedding model are suitable state representations for a Layer 2 transition model and a Layer 3 session-level policy.

Background

Project Qualia proposes a hierarchy in which Layer 1 produces experiential song representations, Layer 2 predicts state transitions, and Layer 3 uses those transitions for session-level planning. The current dataset and experiments do not evaluate whether the proposed Layer 1 representations support these downstream uses. This suitability remains an explicit open question.

References

Several questions remain open as of this report: whether a JEPA-style objective recovers experiential structure more effectively than the residual-corrected Song2Vec embeddings already do; whether Layer 1 representations, once trained, are suitable state representations for a Layer 2 transition model and Layer 3 policy (a use motivated by the JEPA and world-model literature cited above but not evaluated on this dataset); and whether the cross-genre clustering in Section 4.3 reflects experiential interchangeability specifically, as distinct from a correlated confound, such as production era or energy level, that the residual-subtraction method does not disambiguate.

Project Qualia: Recovering Experiential Music Structure from Session Co-occurrence Data  (2609.10862 - Mohammed et al., 9 Sep 2026) in Section 6.6, “Scope of Present Claims”