Polyak Heavy-Ball Momentum
- Polyak Heavy-Ball Momentum is a momentum-based first-order optimization method defined by a two-term recurrence that incorporates past gradient information.
- It can be equivalently reformulated as a conjugate-gradient scheme, analyzed via spectral methods, control theory, and continuous-time dynamics.
- Recent research extends HB with adaptive step sizes and stochastic variants, revealing its regime-dependent behavior and convergence guarantees across various models.
Searching arXiv for recent and foundational papers on Polyak Heavy-Ball momentum. Polyak Heavy-Ball Momentum (HB), also called Polyak momentum or the heavy-ball method, is a momentum-based first-order optimization method introduced by Polyak in the 1960s and typically written in discrete time as
with step size and momentum parameter in the classical regime. Across quadratic optimization, stochastic approximation, composite convex minimization, non-convex problems satisfying the Polyak–Łojasiewicz (PL) inequality, and more recent settings such as actor-critic reinforcement learning and min-max games, HB appears as a common inertial template whose behavior depends strongly on geometry, parameterization, and the interaction between momentum and noise. On convex quadratics it admits especially sharp characterizations: it can be expressed as a two-term recurrence, as a conjugate-gradient-type method in certain adaptive variants, as a linear dynamical system amenable to spectral or control analysis, and as a discrete approximation to damped second-order dynamics (Goujaud et al., 2022, Ugrinovskii et al., 2023). More recent work has also clarified that HB can be interpreted through modified losses, invariant manifolds, Hamiltonian formalisms, and last-iterate convergence analyses, while simultaneously revealing important limits: several acceleration guarantees remain local, quadratic, or otherwise structurally restricted (Cattaneo et al., 10 Sep 2025, Kassing et al., 2024).
1. Canonical iteration and equivalent formulations
The classical heavy-ball recursion is
or, in the fixed-parameter form emphasized in several works,
with the step size and the momentum coefficient (Goujaud et al., 2022, Ugrinovskii et al., 2023). An equivalent momentum-buffer formulation writes
and both formulations generate the same iterate sequence under compatible initialization (Wang et al., 2022, Wang et al., 2020).
A recurrent theme in the literature is that the apparent two-step nature of HB hides simpler internal structure. One line of analysis rewrites HB through an auxiliary moving-average variable. In stochastic heavy ball, an “iterate moving average” representation has the form
with parameter matching
This reformulation is central to almost sure convergence analyses and to momentum-corrected Polyak step-size rules for stochastic HB (Sebbouh et al., 2020, Oikonomou et al., 2024).
Another line emphasizes continuous-time structure. The classical heavy-ball ODE is
0
or equivalently
1
which motivates the interpretation of HB as a damped inertial system rather than merely a gradient method with an ad hoc memory term (Aujol et al., 2024, Kassing et al., 2024). A Hamiltonian formulation further views HB as a special case of a broader time-varying Hamiltonian family with 2, distinguishing it from Nesterov-type dynamics corresponding to different kinetic scaling (Diakonikolas et al., 2019).
These equivalent forms matter because distinct proofs exploit different parameterizations. Spectral-radius analyses prefer the two-step linear recurrence on quadratics; stochastic last-iterate theory prefers the moving-average representation; modified-loss analyses exploit exponentially attractive invariant manifolds; and differential-geometric PL analyses proceed through the second-order ODE picture (Ugrinovskii et al., 2023, Cattaneo et al., 10 Sep 2025, Kassing et al., 2024).
2. Quadratic optimization as the canonical model problem
The most explicit theory concerns convex quadratics
3
with 4 symmetric positive semidefinite (Goujaud et al., 2022). In this regime the gradient is affine in the error, 5, so HB becomes a linear recurrence whose behavior is determined by the spectrum of the Hessian.
For strongly convex quadratics with eigenvalues in 6, the classical tuned heavy-ball parameters are
7
and the associated accelerated contraction factor is
8
(Can et al., 2019). A robust-control analysis of fixed-parameter first-order methods shows that, over the finite-dimensional quadratic class 9, no convergent method in a broad multistep class can achieve a worst-case asymptotic root-convergence factor smaller than
0
and HB attains this benchmark, establishing worst-case asymptotic optimality within that class (Ugrinovskii et al., 2023).
Quadratics also expose the relation between HB and average-case optimality. For random quadratic objectives with Hessian spectrum supported on 1, optimal average-case first-order methods have coefficients 2 that vary with time and depend on the full spectral density. Yet under mild positivity conditions on the spectral density, these coefficients converge asymptotically to the Polyak momentum coefficients,
3
so Polyak momentum emerges as the universal asymptotic limit of optimal average-case methods (2002.04664). This suggests that HB is not merely an extremal worst-case design but also the asymptotic attractor of average-case optimal recurrences on quadratics.
A separate development addresses convex quadratics without strong convexity. Using the Störmer–Verlet discretization of a harmonic oscillator, a modified HB4 scheme is shown to admit a Lyapunov function that decreases monotonically under the step-size condition
5
yielding
6
for convex quadratics (Orvieto, 2023). The construction is notable because it avoids standard eigenvalue arguments and instead derives the Lyapunov function from a modified Hamiltonian containing an essential cross term. This does not settle global acceleration of HB for general convex smooth objectives, but it shows that convex-quadratic acceleration can be certified even in the absence of strong convexity.
3. Adaptive Polyak step-sizes and conjugate-gradient structure
A major recent development concerns the combination of Polyak step-sizes with momentum. For quadratic minimization, an adaptive heavy-ball scheme is defined by
7
with
8
and, for 9,
0
Here the “natural” step-size 1 is twice the classical Polyak step-size (Goujaud et al., 2022).
The central result is that this adaptive HB method is exactly equivalent to the projection problem
2
and more generally to a family of 3-minimization problems over the gradient span (Goujaud et al., 2022). In the special case 4, the framework recovers classical conjugate gradient; in the special case 5, it yields the Polyak-step HB method. The main structural implication is that Polyak HB on convex quadratics is not merely a heuristic momentum scheme but an exact conjugate-gradient-type method expressed as a two-term recurrence.
This equivalence immediately yields three strong properties on 6-dimensional quadratic problems. First, finite termination: 7 Second, instance optimality over the span of past gradients. Third, the worst-case convergence behavior inherited from conjugate-gradient-type methods on quadratics (Goujaud et al., 2022). The result gives a definitive positive answer, within the quadratic setting, to the question of whether Polyak step-sizes can be combined with momentum in a principled and provably stronger way than plain gradient descent with Polyak step-size.
A broader generalized Polyak step-size framework extends this logic to momentum methods of the form
8
Minimizing the post-update squared distance to the optimizer yields
9
For heavy ball, taking 0 and 1 leads to
2
and, under 3-smoothness,
4
This framework is presented as a principled generalization of Polyak step-size to momentum methods, with truncation added in least-squares settings to maintain nonnegativity and stability (Wang et al., 2023).
4. Deterministic convergence beyond strongly convex quadratics
Outside the quadratic model, the most developed positive theory concerns structured regimes rather than arbitrary smooth objectives. One such regime is quadratic growth for composite optimization. For objectives 5, where 6 is convex and differentiable with 7-Lipschitz gradient and 8 is proper, lower semicontinuous, convex, and proximable, a V-FISTA-type inertial iteration
9
with constant 0, is studied under the quadratic growth condition
1
For appropriately chosen 2, the method achieves exponential decay of the form
3
without assuming uniqueness of the minimizer (Aujol et al., 2024). This extends heavy-ball-type acceleration beyond strong convexity to composite problems such as LASSO and provides robustness when the conditioning parameter is not exactly known.
A different extension concerns PL geometry. Under the PL inequality
4
continuous-time HB is shown to converge globally from any initialization with objective gap decay
5
under the optimal friction choice 6 (Kassing et al., 2024). In discrete time, the guarantee is local: if the iterates enter a neighborhood of the minimizer set, then the objective decays geometrically with rate governed by 7, and the optimally tuned parameters
8
recover the classical accelerated dependence on 9 (Kassing et al., 2024). The proof uses a differential-geometric perspective on PL landscapes, decomposing motion into normal and tangential directions relative to the minimizer manifold.
There are also results for non-convex objectives where acceleration is proved only after structural restrictions or after a transient phase. For a class of PL problems satisfying an “average-out” curvature condition based on the average Hessian
0
HB can achieve a local contraction factor
1
after a burn-in period 2, provided an additional co-diagonalization condition holds (Wang et al., 2022). The erratum explicitly restricts this acceleration result: it naturally holds in one dimension and when the Hessian is diagonal, but more general high-dimensional cases remain open (Wang et al., 2022).
A complementary deterministic non-convex perspective studies the role of HB in reaching a benign region. In phase retrieval and cubic-regularized minimization, heavy ball is proved to enter a region where standard local convergence theory applies faster than gradient descent, with iteration complexity improvements that scale roughly as 3 for small step sizes (Wang et al., 2020). This suggests that, in some non-convex problems, the principal benefit of HB is not a globally accelerated rate statement but faster transit into a locally favorable regime.
5. Stochastic heavy ball, last-iterate theory, and Polyak-type stochastic step-sizes
In stochastic optimization, HB is typically written as
4
or with additive oracle noise
5
where 6 and 7 in the standard unbiased bounded-variance model (Gadat et al., 2016, Can et al., 2019).
A basic distinction in the stochastic literature is between averaged-iterate guarantees and last-iterate guarantees. A sharp analysis of adaptive Polyak heavy-ball methods in constrained convex optimization shows that, with time-varying momentum
8
the projected HB recursion can be rewritten through
9
as
0
This transformation yields an optimal last-iterate convergence rate
1
for nonsmooth constrained convex problems, improving on the known last-iterate rate 2 for SGD (Tao et al., 2021). The same paper extends the result to an adaptive preconditioned setting with diagonal second-moment estimates and the same time-varying momentum schedule 3, again obtaining
4
(Tao et al., 2021). Constant momentum, by contrast, suffices for optimal averaged convergence but not for the same last-iterate guarantee.
A different almost sure analysis shows that stochastic heavy ball converges to a minimizer almost surely under smooth convexity, and that the last iterate satisfies
5
Under constant stepsize in the deterministic case this gives the sharper
6
improving the previously known 7 bound (Sebbouh et al., 2020). This work stresses that, unlike SGD, HB’s internal averaging makes the last iterate the natural object of asymptotic analysis.
The stochastic quadratic setting admits yet another perspective. For strongly convex quadratics with i.i.d. noise, the lifted state 8 forms a Markov chain with a unique invariant distribution 9, and stochastic HB converges to this invariant law linearly in 0 with accelerated factor 1: 2 where 3 grows at most linearly in 4 (Can et al., 2019). Persistent noise prevents pointwise convergence to the minimizer, but the transient accelerated factor remains the same as in the noiseless quadratic theorem, while the stationary spread is determined by a discrete Lyapunov equation for the limiting covariance.
Recent work has also developed Polyak-type step-size rules specifically for stochastic heavy ball. Using the IMA representation, the key observation is that the heavy-ball step-size should be the IMA step-size multiplied by 5. This yields momentum-corrected rules such as
6
for MomSPS7, as well as decreasing variants MomDecSPS and MomAdaSPS (Oikonomou et al., 2024). For convex smooth problems, MomSPS8 gives 9 convergence of the averaged iterate to a neighborhood in the non-interpolated case and exact 0 convergence under interpolation; MomDecSPS gives exact 1 convergence under bounded iterates; and MomAdaSPS adapts between 2 under interpolation and 3 otherwise (Oikonomou et al., 2024). The paper explicitly shows that the naive Polyak step-size inserted directly into SHB can diverge for moderate or large 4, so the momentum correction is not a cosmetic modification.
More general stochastic HB theory allows biased gradients, random stepsizes, block updates, and functions satisfying PL or KL conditions. Under stochastic analogs of Robbins–Monro conditions, almost-sure convergence is established even when conditional variance grows over time or finite-difference gradient approximations are used (Tadipatri et al., 2023). This broadens the classical unbiased bounded-variance framework and shows that the HB template remains analyzable under substantially weaker oracle assumptions.
6. Interpretive frameworks, extensions, and regime-dependent limitations
Several recent works reinterpret HB in ways that sharpen both its meaning and its limitations. One modified-loss perspective studies full-batch and mini-batch HB with fixed 5 and shows that, on an exponentially attractive invariant manifold,
6
Thus, restricted to the manifold, HB is exactly gradient descent on a modified loss, although 7 is defined only implicitly (Cattaneo et al., 10 Sep 2025). The same work derives memoryless approximations of arbitrary order, identifies tree-indexed coefficient systems containing Eulerian and Narayana polynomials, and extends the approximation theory to mini-batch HB with an additional 8-type regularization term after averaging over permutations (Cattaneo et al., 10 Sep 2025). A plausible implication is that HB’s memory effects can be understood as structured objective deformation rather than solely as transient filtering.
A Hamiltonian perspective places HB within a broader continuum of momentum methods generated by a time-varying Hamiltonian
9
with HB corresponding to 00 and Nesterov-type methods to 01 (Diakonikolas et al., 2019). In this view, HB-like methods yield 02 function-value convergence for smooth convex optimization and 03 rates for the minimum gradient norm in convex and nonconvex settings, while more aggressive kinetic scaling recovers accelerated Nesterov-type rates (Diakonikolas et al., 2019). This suggests that HB and Nesterov acceleration are not isolated designs but points on a parameterized dynamical spectrum.
Extensions beyond classical minimization reveal that HB is highly regime-dependent. In actor-critic reinforcement learning with Markovian noise, heavy-ball momentum inserted only into the critic recursion,
04
yields an HB-A2C method that finds an 05-approximate stationary point with 06 iterations under linear value approximation and mixing assumptions (Dong et al., 2024). The role of momentum here is to balance initialization error and stochastic approximation error in the critic dynamics rather than to provide the classical quadratic acceleration story.
In min-max games, the behavior is markedly different from minimization. Continuous-time analysis for simultaneous and alternating HB updates finds that smaller momentum can enlarge the local stability region and bias trajectories toward shallower slope regions, especially under alternating updates (Feng et al., 26 May 2025). This contrasts sharply with minimization, where larger positive momentum is typically associated with a larger stable step-size range. The result indicates that “HB momentum” is not a monolithic mechanism whose effects transfer unchanged across optimization regimes.
A recurring misconception is that heavy-ball momentum is either universally accelerated or merely heuristic. The literature supports neither extreme. On quadratics, there are precise optimality, finite-time, and conjugate-gradient-equivalence theorems (Goujaud et al., 2022, Ugrinovskii et al., 2023). Under PL, quadratic growth, or related structure, there are strong local or global rates in carefully delimited settings (Aujol et al., 2024, Kassing et al., 2024). In stochastic convex optimization, last-iterate improvements over SGD can be rigorously established with suitable momentum scheduling (Tao et al., 2021). At the same time, several positive results require exact knowledge of 07, quadratic structure, local neighborhoods, diagonal Hessians, or other restrictive conditions, and naive combinations of momentum with adaptive step-size rules can diverge (Goujaud et al., 2022, Wang et al., 2022, Oikonomou et al., 2024).
For that reason, Polyak Heavy-Ball Momentum is best understood not as a single theorem but as a family of inertial recursions whose properties depend on representation, geometry, and noise model. Its enduring importance stems from the fact that it occupies a central intersection of optimization, numerical integration, approximation theory, stochastic approximation, and dynamical systems: it is simultaneously a classical linear multistep method, a conjugate-gradient-type scheme in adaptive quadratic form, a damped second-order flow, a last-iterate averaging mechanism, and, under suitable reductions, gradient descent on a modified loss (Goujaud et al., 2022, Cattaneo et al., 10 Sep 2025).