Convergence under non-vanishing gradient noise

Extend the convergence analysis of the OSGM-SGD framework to establish convergence guarantees when the stochastic gradient noise does not vanish.

Background

The paper proves high-probability convergence for its OSGM-SGD instantiation under a relative gradient-noise condition and sufficiently large batch sizes. These assumptions effectively require the stochastic gradient noise to be controlled relative to the true gradient, and the stated results do not resolve the regime of persistent, non-vanishing gradient noise.

The conclusion explicitly asks whether the analysis can be extended to prove convergence in that broader stochastic setting.

References

Two questions remain open: does {purple} admit convergence guarantees when the stepsize is updated at each inner iteration rather than only in the outer loop? Can the analysis of {purple} be extended to show convergence even with non-vanishing gradient noise?

Stochastic Gradient Methods with Online Scaling  (2609.11751 - Zhang et al., 10 Sep 2026) in Section Conclusion