Timescale separation in ReLU networks from small initialization
Establish whether two-layer ReLU networks trained with gradient flow from small initialization exhibit a timescale separation between weight directions analogous to linear networks, and characterize the mechanism and conditions under which this directional separation arises.
References
Because the ReLU activation function is piece-wise linear, we conjecture that ReLU networks trained from small initialization have a timescale separation between different directions, similar to the mechanism in linear networks.
— Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
(2512.20607 - Zhang et al., 23 Dec 2025) in Appendix C — Additional Discussion (ReLU activation function)
Several questions remain open. The most immediate is whether the technical modification to gradient flow can be removed, allowing us to analyze vanilla gradient flow directly.
— Learning Orthogonal Multi-Index Models Beyond Small Initialization: Incremental Learning, Competitive Dynamics and Symmetry
(2609.10879 - Zhou et al., 9 Sep 2026) in Section Conclusion