Develop a theory of middle-layer vulnerability and principled sparsity allocation
Develop a formal theoretical account of how internal representations evolve with depth, explain why middle layers exhibit heightened sensitivity to pruning, and use that account to derive principled rather than heuristic layer-wise sparsity allocations.
References
Finally, while the middle-layer sensitivity finding is empirically robust, our framework does not yet explain it theoretically. A formal account of how the structure of internal representations evolves with depth, and how this motivates principled rather than heuristic sparsity allocation, remains an open problem.
— When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs
(2608.25941 - Gupte et al., 26 Aug 2026) in Section 5, “Limitations and Future Work” (Section \ref{sec:limitations})