Relationship between effective depth and pruning tolerance
Determine whether model-level effective depth, particularly normalized effective depth, reliably predicts the number or fraction of layers that a decoder-only language model can remove using BI-ranked pruning while keeping perplexity within a specified tolerance.
References
A separate, exploratory question is whether model-level $/L$ correlates with how many BI-ranked layers a model can lose before its perplexity degrades substantially.
— The Residual Stream's Effective Depth
(2609.31098 - Gahtan et al., 25 Sep 2026) in Appendix S19, “Pruning Boundary Checks”