Optimal preservation of entangled features under whitened linear erasure

Determine whether optimally whitened linear operators can achieve quantitatively improved preservation of feature A while erasing feature B at intermediate entanglement, compared with the unwhitened rank-1 projection studied in the Toy Model of Superposition.

Background

The paper studies erasure of one feature from an antipodal or partially entangled pair while preserving the reconstruction of the other feature. Its baseline rank-1 nullspace projection successfully suppresses feature B, but preservation of feature A degrades as the encoder directions become more entangled.

The authors note that the degradation at exact antipodality is unavoidable for any linear operator, because the two features occupy the same one-dimensional subspace. At intermediate entanglement, however, the reported degradation is specific to the unwhitened projection used in the experiments, leaving unresolved whether covariance-whitened linear methods could produce better preservation curves.

References

At intermediate $\rho$ (Table~\ref{tab:linear_baseline}), the degradation is specific to our unwhitened rank-1 projection; optimally whitened linear operators might yield quantitatively different preservation curves, which we leave to future work.

— Hidden not Deleted: How Networks Suppress Entangled Features  (2609.27593 - Samanta et al., 23 Sep 2026) in Section 3.3, subsection “Linear erasure fails”