Papers
Topics
Authors
Recent
Search
2000 character limit reached

Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency

Published 2 Sep 2026 in stat.ML, cs.LG, and math.OC | (2609.02728v1)

Abstract: We study when and how momentum improves large-batch training in the one-pass regime, using power-law kernel regression as a tractable setting. We first characterize risk stability through the critical learning rate, defined as the largest learning rate for stable training, and obtain η<em>SGD<sup>crit</sup>1η<em>{\mathrm{SGD}}<sup>{\mathrm{crit}}\eqsim</sup> 1, η</em>Polyak<sup>crit</sup>min1,B(1ρ)η</em>{\mathrm{Polyak}}<sup>{\mathrm{crit}}\eqsim</sup> \min{1,B(1-ρ)}, and ηNesterov<sup>crit</sup>min1,B<sup>β(1ρ)η_{\mathrm{Nesterov}}<sup>{\mathrm{crit}}\eqsim</sup> \min{1,B<sup>β(1-ρ)}, where BB is the batch size, ρρ is the momentum factor, and $β&gt;1$ is the capacity exponent. Within this admissible region, we derive scaling laws for the full risk dynamics, capturing the progression from an early transient, through power-law decay, to a noise floor. We then minimize the final-step risk over the admissible learning rates and momentum factors under a fixed data budget, yielding a three-regime batch-size phase diagram that reveals how the role of momentum changes with batch size. Notably, Polyak enlarges the critical batch size, the largest batch size preserving the best small-batch data-scaling exponent, thereby enabling greater parallelism without sacrificing data efficiency. In contrast, Nesterov achieves better data efficiency in the large-batch regime because its look-ahead mechanism suppresses noise accumulation. Numerical experiments validate the predicted stability boundaries, risk dynamics, and batch-size phase diagram.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.