Develop a theory of middle-layer vulnerability and principled sparsity allocation

Develop a formal theoretical account of how internal representations evolve with depth, explain why middle layers exhibit heightened sensitivity to pruning, and use that account to derive principled rather than heuristic layer-wise sparsity allocations.

Background

The experiments find that middle layers are more vulnerable to pruning than early or late layers, while the proposed layer-wise sparsity schedule is presented only as a preliminary heuristic. The paper hypothesizes that accumulated upstream perturbations propagated through residual connections contribute to this pattern, but does not provide a formal explanation.

The unresolved problem is therefore twofold: characterize the depth-dependent evolution of internal representations sufficiently to explain middle-layer sensitivity, and translate that characterization into a theoretically justified sparsity-allocation rule rather than an empirically motivated schedule.

References

Finally, while the middle-layer sensitivity finding is empirically robust, our framework does not yet explain it theoretically. A formal account of how the structure of internal representations evolves with depth, and how this motivates principled rather than heuristic sparsity allocation, remains an open problem.

When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs  (2608.25941 - Gupte et al., 26 Aug 2026) in Section 5, “Limitations and Future Work” (Section \ref{sec:limitations})