Convergence under non-vanishing gradient noise
Extend the convergence analysis of the OSGM-SGD framework to establish convergence guarantees when the stochastic gradient noise does not vanish.
References
Two questions remain open: does {purple} admit convergence guarantees when the stepsize is updated at each inner iteration rather than only in the outer loop? Can the analysis of {purple} be extended to show convergence even with non-vanishing gradient noise?
— Stochastic Gradient Methods with Online Scaling
(2609.11751 - Zhang et al., 10 Sep 2026) in Section Conclusion