Conditions for using large learning rates in stochastic optimization
Determine the necessary and sufficient conditions under which large fixed learning rates (for example, step size γ on the order of D/G) can be safely used in stochastic optimization problems while maintaining optimal convergence behavior, beyond currently known special cases such as quadratic losses.
References
More generally, the full conditions under which large learning rates can be used are not yet fully understood for stochastic problems.
— The Road Less Scheduled
(2405.15682 - Defazio et al., 2024) in Subsection "On Large Learning Rates"
More broadly, it remains to understand whether analogous hierarchical mechanisms govern large-stepsize transitions beyond separable logistic regression.
— Tight Transition Time Bounds for Separable Logistic Regression at the Edge of Stability
(2610.01459 - Wen et al., 1 Oct 2026) in Conclusion