Matching upper bounds for stochastic momentum contraction rates

Establish matching upper bounds on the contraction rates for stochastic momentum methods in finite-dimensional linear regression, complementing the lower bounds derived by Morwani et al. (2026) and determining whether heavy-ball momentum retains data efficiency over a larger batch-size range.

Background

The paper discusses concurrent work by Morwani et al. (2026), which studies batch-size tradeoffs for stochastic momentum in finite-dimensional linear regression and derives lower bounds on contraction rates. Those lower bounds suggest that the heavy-ball method may preserve data efficiency across a larger range of batch sizes, but the corresponding upper bounds needed to match the lower-bound characterization have not been established.

The unresolved problem is to derive such matching upper bounds, thereby determining whether the suggested heavy-ball data-efficiency behavior is sharp. This issue is distinct from the present paper’s analysis, which studies sharp risk dynamics and batch-size scaling in infinite-dimensional power-law kernel regression for canonical Polyak and Nesterov methods.

References

These bounds suggest that heavy-ball may retain data efficiency over a larger batch-size range, but matching upper bounds are left open.

Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency  (2609.02728 - Wang et al., 2 Sep 2026) in Section 2, Related work, paragraph “Concurrent work”