Theoretical explanation of the mirror-shadow crossover

Characterize the excision loss landscape and determine why entanglement governs the relative prevalence of the mirror and shadow basins, including whether the proposed learning-rate-ratio prediction holds independently of optimizer-specific normalization.

Background

The experiments establish that the relative frequency of mirror and shadow outcomes changes sharply with feature entanglement, but they do not derive the mechanism producing this dependence. The appendix proposes that entanglement changes the relative local cost of decoder rotation and output-bias suppression.

That explanation is explicitly exploratory and untested. The authors identify a complete characterization of local curvature as the central theoretical problem, alongside an empirical test involving mismatched decoder-to-bias learning rates and comparisons between AdamW and plain stochastic gradient descent.

References

A complete account would characterize the excision loss landscape's local curvature around plausible decoder configurations as a function of entanglement directly, rather than reasoning informally about relative step costs as we do above. We view this, together with the empirical check proposed in \S\ref{app:basin_open_question}, as the most direct open theoretical question raised by this work, and one we think is better pursued as its own dedicated analysis than compressed into an appendix of an empirically focused paper.

— Hidden not Deleted: How Networks Suppress Entangled Features  (2609.27593 - Samanta et al., 23 Sep 2026) in Appendix, Section “Toward a Theoretical Account of the Mirror/Shadow Bifurcation,” subsection “Directions for future theoretical work”

Our excision objective is a specific, bespoke choice, reconstruction loss on the retain set plus a squared-activation penalty on $B$, and it remains open whether the same bifurcation appears under objectives drawn directly from the unlearning literature, such as representation misdirection \citep{li2024wmdp} or a student-teacher distillation objective.

— Hidden not Deleted: How Networks Suppress Entangled Features  (2609.27593 - Samanta et al., 23 Sep 2026) in Appendix, Section “Directions for Empirical Validation,” paragraph “Alternative unlearning objectives”