Theoretical explanation of the mirror-shadow crossover
Characterize the excision loss landscape and determine why entanglement governs the relative prevalence of the mirror and shadow basins, including whether the proposed learning-rate-ratio prediction holds independently of optimizer-specific normalization.
References
A complete account would characterize the excision loss landscape's local curvature around plausible decoder configurations as a function of entanglement directly, rather than reasoning informally about relative step costs as we do above. We view this, together with the empirical check proposed in \S\ref{app:basin_open_question}, as the most direct open theoretical question raised by this work, and one we think is better pursued as its own dedicated analysis than compressed into an appendix of an empirically focused paper.
Our excision objective is a specific, bespoke choice, reconstruction loss on the retain set plus a squared-activation penalty on $B$, and it remains open whether the same bifurcation appears under objectives drawn directly from the unlearning literature, such as representation misdirection \citep{li2024wmdp} or a student-teacher distillation objective.