Evaluate cross-architecture applicability and develop generalizable methods
Determine how well existing interpretability methods (including sparse dictionary learning and circuit analysis) apply to architectures such as diffusion models, vision transformers, RWKV, and state space models, and develop techniques that generalize effectively across architectures.
References
Assessing how well interpretability methods apply to architectures beyond those for which they were developed, and whether we can develop techniques that generalize effectively across architectures remain open questions.
We do not examine whether similar cancellations arise in other architectures, such as diffusion transformers or state-space models, where residual streams and update branches may interact differently, and note that strong residual cancellation may have broader implications for training dynamics and generalization.
Current limitations include the absence of spectral localization, mechanistic evaluation restricted to AMP tasks with ESM2-8M, and unresolved interpretations of non-DC components. Local time-frequency or adaptive spectral methods [51, 52], controlled sequence interventions, and structure-aware validation may clarify these components and test whether the observed mechanisms generalize across modalities.