---
title: Polyak Step-Size Method
url: https://www.emergentmind.com/topics/polyak-step-size
type: topic
---

# Polyak Step-Size Method

Polyak step-size is an adaptive learning-rate rule that selects the step from the current objective gap divided by a squared gradient norm. For a differentiable convex objective with known optimum value $f^\star$, the classical rule is
\[
\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2},
\qquad
x_{k+1}=x_k-\alpha_k\nabla f(x_k).
\]
It was derived by minimizing a quadratic upper bound on the next squared distance to an optimizer. The method does not require an externally supplied smoothness constant, but its classical form requires the optimal objective value, which is generally unavailable. Subsequent work has extended the construction to stochastic optimization, momentum, proximal and mirror descent, constrained problems, preconditioned updates, high-dimensional sparse estimation, decentralized optimization, reinforcement learning, and objectives whose curvature degenerates near the optimum.

## 1. Classical rule and geometric derivation

Consider the convex subgradient iteration
\[
x^{k+1}=x^k-\gamma_k g^k,
\qquad
g^k\in\partial f(x^k).
\]
For an optimizer $x^\star$ and optimal value $f^\star=f(x^\star)$, convexity gives
\[
\langle g^k,x^k-x^\star\rangle\geq f(x^k)-f^\star.
\]
Consequently,
\[
\|x^{k+1}-x^\star\|^2
\leq
\|x^k-x^\star\|^2
-2\gamma_k\bigl(f(x^k)-f^\star\bigr)
+\gamma_k^2\|g^k\|^2.
\]
Minimizing the right-hand side as a quadratic function of $\gamma_k$ yields
\[
\boxed{
\gamma_k^{\mathrm{Polyak}}
=
\frac{f(x^k)-f^\star}{\|g^k\|^2}
}.
\]
For smooth objectives, $g^k=\nabla f(x^k)$.

The numerator is the current objective gap and the denominator normalizes the step by the local gradient scale. Thus the rule tends to take larger steps far from the optimum and smaller steps near it. If $f^\star$ is known exactly, no estimate of a global Lipschitz or smoothness constant is required. The same construction can be interpreted geometrically: the affine subgradient minorant
\[
f(x^k)+\langle g^k,x-x^k\rangle
\]
is intersected with the optimal level $f^\star$, and the update is the point on that hyperplane closest to $x^k$. This projection interpretation is the basis of the Polyak Minorant Method, which replaces affine subgradient models with arbitrary convex lower models and incorporates constraints [2310.07922].

A scaled form is frequently used:
\[
x^{k+1}=x^k-\gamma
\frac{f(x^k)-f^\star}{\|\nabla f(x^k)\|^2}
\nabla f(x^k),
\qquad \gamma\in(0,2].
\]
The case $\gamma=1$ is the usual Polyak method. For $\gamma\in(0,2)$, the elementary distance recursion contains the positive descent coefficient $\gamma(2-\gamma)$.

The requirement that $f^\star$ be known is both the principal strength and the principal limitation of the rule. The method is adaptive once the target value is available, but is not parameter-free in an absolute sense. In nonsmooth convex optimization, exact last-iterate behavior is substantially weaker than many best-iterate or averaged-iterate guarantees: with uniformly bounded subgradients, the standard Polyak method has an exact worst-case last-iterate rate of order
\[
O\!\left(\frac{BR}{N^{1/4}}\right),
\]
where $B$ bounds subgradients and $R$ bounds the initial distance to a solution [2407.15195].

## 2. Stochastic Polyak step-size

For a finite-sum objective
\[
f(x)=\frac1n\sum_{i=1}^n f_i(x),
\]
stochastic gradient descent samples $i_k$ and updates
\[
x^{k+1}=x^k-\gamma_k\nabla f_{i_k}(x^k).
\]
Applying the Polyak construction to the sampled component gives the stochastic Polyak step-size (SPS):
\[
\boxed{
\gamma_k^{\mathrm{SPS}}
=
\frac{f_{i_k}(x^k)-f_{i_k}^\ast}
{c\|\nabla f_{i_k}(x^k)\|^2}
},
\]
where
\[
f_i^\ast=\inf_x f_i(x),
\qquad c>0.
\]
A bounded version is
\[
\boxed{
\gamma_k^{\mathrm{SPS}_{\max}}
=
\min\left\{
\frac{f_{i_k}(x^k)-f_{i_k}^\ast}
{c\|\nabla f_{i_k}(x^k)\|^2},
\gamma_{\max}
\right\}.
}
\]
The cap prevents very large steps when a sampled gradient is small. The loss and gradient used by SPS are already computed in an SGD iteration, so unbounded SPS requires no full-objective or full-gradient evaluation and has essentially no additional per-iteration computational cost [2002.10542].

The component optimum $f_i^\ast$ is weaker information than the global optimum $f^\star$. For many nonnegative unregularized losses, $f_i^\ast=0$ is available or is the relevant infimum. Examples include squared loss in realizable regression, logistic loss in separable classification, exponential loss, and squared-hinge loss in linearly separable classification. For unregularized logistic loss, the infimum can be zero without being attained by any finite parameter vector. With regularization, individual minima are generally nonzero, although they can sometimes be computed analytically or cheaply offline.

The parameter $c$ controls the scale of the update. The cited SPS analysis recommends $c=1/2$ in strongly convex and interpolation analyses and $c=1$ for a convenient convex-objective bound. A central inequality is
\[
\gamma_k^2\|\nabla f_{i_k}(x^k)\|^2
\leq
\frac{\gamma_k}{c}
\bigl(f_{i_k}(x^k)-f_{i_k}^\ast\bigr),
\]
which is an equality for unbounded SPS. This permits convergence analysis without bounded stochastic gradients or a standard bounded-variance assumption in several settings.

Define the stochastic inconsistency quantity
\[
\sigma^2
=
\mathbb E_i\bigl[f_i(x^\star)-f_i^\ast\bigr].
\]
Under interpolation,
\[
f_i(x^\star)=f_i^\ast\quad\forall i,
\qquad \sigma=0.
\]
For convex smooth components and a strongly convex aggregate objective, capped SPS converges linearly in expected squared distance to a neighborhood whose size depends on $\sigma^2$ and $\gamma_{\max}$. Under interpolation, ordinary SPS with $c=1/2$ converges linearly to the true solution. For convex non-strongly-convex objectives, the averaged iterate has an $O(1/K)$ objective-gap rate up to a residual neighborhood; under interpolation, the residual vanishes. For smooth PL objectives, SPS yields linear objective convergence under the stated parameter restrictions. For general smooth nonconvex objectives, weak growth yields an $O(1/K)$ stationarity rate up to a neighborhood governed by the growth residual.

The same analysis gives constant-step SGD bounds. If the cap is sufficiently small, capped SPS is effectively constant-step SGD. In the strongly convex case,
\[
\mathbb E\|x^k-x^\star\|^2
\leq
(1-\mu\gamma)^k\|x^0-x^\star\|^2
+\frac{2\sigma^2}{\mu}.
\]
Under interpolation, the residual term disappears. The neighborhood is expressed in terms of the optimal objective difference $\sigma^2$, rather than solely in terms of the conventional gradient-noise quantity
\[
z^2=\mathbb E\|\nabla f_i(x^\star)\|^2.
\]

## 3. Interpolation, non-interpolation, and optimal-value estimation

Interpolation means that a common point minimizes every component:
\[
f_i(x^\star)=f_i^\ast
\qquad\text{for every }i.
\]
It occurs in realizable least-squares problems, consistent linear systems, separable classification, sufficiently expressive kernel models, and over-parameterized neural networks. Under interpolation, stochastic gradients vanish at a common solution and constant-step stochastic methods need not settle in a persistent noise neighborhood.

For consistent linear systems, the interpolation interpretation is especially direct. A stochastic quadratic reformulation can satisfy
\[
f_S(x)-f_S^\ast=\frac12\|\nabla f_S(x)\|^2.
\]
With $c=1/2$, SPS therefore produces a unit step and recovers the theoretically optimal step for randomized sketch-and-project methods, including randomized Kaczmarz. The associated distance guarantee depends on the smallest nonzero eigenvalue of the expected system matrix [2002.10542].

Outside interpolation, replacing $f_i^\ast$ by zero can produce systematically excessive steps. The residual $\sigma^2$ then determines the limiting neighborhood in many SPS bounds. Several methods address this limitation by estimating individual or global target values.

**MOTAPS**, or Moving Targetted Polyak Stepsize, maintains one auxiliary scalar for every data point. These variables track individual losses and permit a moving estimate of the global target. TAPS assumes that the global optimum value is known; MOTAPS also estimates that value and therefore does not require interpolation. Its auxiliary objective contains a dampening parameter $\lambda$. Fixed $\lambda>0$ produces convergence to a neighborhood, with an error term proportional to
\[
\frac{\lambda f(w^\star)^2}{\mu(n+1)}
\]
or
\[
\frac{\lambda f(w^\star)^2}{2(n+1)}.
\]
Eventually decreasing the step-size can remove the fixed-$\lambda$ error floor asymptotically. The convergence theory is conditional on star-convexity or strong star-convexity of an auxiliary objective, which does not automatically follow from convexity of the original component losses [2106.11851].

**Twin-Polyak** maintains two iterates and updates the higher-valued one using the lower-valued objective as a proxy for $f^\star$. If $f(x^k)>f(y^k)$,
\[
x^{k+1}
=
x^k
-
2\frac{f(x^k)-f(y^k)}
{\|\nabla f(x^k)\|^2}
\nabla f(x^k),
\qquad
y^{k+1}=y^k.
\]
The lower-valued sequence supplies an observed target rather than a known optimum. The method is invariant to positive objective scaling and objective translation. Under deterministic twin-gap assumptions, it has linear convergence for strongly convex smooth objectives and an $O(k^{-1/2})$ best-iterate rate for convex bounded-gradient objectives. Its stochastic variants are primarily empirical because the stepsize depends on the same random minibatch as the gradient [2509.14854].

**PSADLA**, or Polyak Stepsize with Accelerated Decision-Guided Level Adjustment, maintains a lower level $\bar f_k<f^\star$ and uses
\[
s^k=\gamma\frac{f(x^k)-\bar f_k}{\|g^k\|^2}.
\]
A Polyak Stepsize Violation Detector (PSVD) tests the feasibility of a block of linear inequalities encoding whether recent updates remain compatible with a possible optimizer. If the system becomes infeasible, the level is raised by a convex combination of the previous level and the best objective value observed in the block. The method proves convergence of the level values and running-best objective values to $f^\star$, including when approximate or surrogate subgradients satisfy the stated conditions [2311.18255].

## 4. Extensions to composite, momentum, mirror, and constrained optimization

For composite objectives
\[
\psi(x)=f(x)+\phi(x),
\]
ordinary SPS applied to the complete sample objective requires a lower bound on $f_i+\phi$. Such a bound may be loose when $\phi$ is a regularizer. ProxSPS applies Polyak adaptation only to the stochastic loss and treats $\phi$ through an exact proximal step:
\[
x^{k+1}
=
\arg\min_y
\left\{
\max\bigl(f(x^k;S_k)+\langle g_k,y-x^k\rangle,C(S_k)\bigr)
+\phi(y)
+\frac{1}{2\alpha_k}\|y-x^k\|^2
\right\}.
\]
It requires only
\[
C(s)\leq\inf_x f(x;s),
\]
not a lower bound on the regularized loss. For squared $\ell_2$ regularization,
\[
x^{k+1}
=
\frac{1}{1+\alpha_k\lambda}
\left(x^k-\tau_k^+g_k\right).
\]
ProxSPS has convergence results for nonsmooth weakly convex, smooth convex, strongly convex, and smooth weakly convex settings. Its nonconvex stationarity guarantee is expressed through the Moreau envelope and has an approximately $\widetilde O(1/\sqrt K)$ rate; strongly convex settings yield linear convergence up to a noise neighborhood [2301.04935].

Momentum requires modifying the Polyak derivation because the update contains an inherited displacement. For heavy-ball momentum,
\[
x_{k+1}=x_k-\eta_k\nabla f(x_k)+\beta(x_k-x_{k-1}),
\]
the generalized adaptive learning rate includes a momentum-alignment term:
\[
\eta_k
=
\frac{f(x_k)-f^\ast}{\|\nabla f(x_k)\|^2}
+
\beta
\frac{\langle\nabla f(x_k),x_k-x_{k-1}\rangle}
{\|\nabla f(x_k)\|^2}.
\]
ALR-HB therefore reduces the gradient step when momentum opposes the current gradient. For moving averaged gradients,
\[
d_k=\beta d_{k-1}+\nabla f(x_k),
\]
ALR-MAG uses
\[
\eta_k=\frac{f(x_k)-f^\ast}{\|d_k\|^2}.
\]
Its deterministic analysis establishes monotone distance decrease and global linear convergence under smooth semi-strong convexity. Stochastic extensions include ALR-SMAG and ALR-SHB. A later IMA-based analysis proposes MomSPS$_{\max}$, MomDecSPS, and MomAdaSPS. MomSPS$_{\max}$ scales SPS by $(1-\beta)$ and converges to a neighborhood without interpolation, while interpolation gives an $O(1/T)$ rate. MomDecSPS and MomAdaSPS converge to the exact minimizer in non-interpolated convex problems, with rates of order $O(1/\sqrt T)$ and $O(1/T+\sigma/\sqrt T)$, respectively [2305.12939; 2406.04142].

Mirror-descent variants replace Euclidean geometry with a Bregman divergence. For a mirror map $h$, the update is
\[
x_{k+1}
\in
\arg\min_{y\in\mathcal X}
\left\{
\langle g(x_k),y-x_k\rangle
+\frac{1}{\eta_k}D_h(y,x_k)
\right\}.
\]
Two Polyak-type rules replace $f^\star$ by adaptive target levels. The first maintains
\[
\widehat f_k=\min_{\kappa\leq k}f(x_\kappa)-\delta_k
\]
and guarantees an objective value within a prescribed tolerance $\delta$ of optimality. The second uses level reductions driven by accumulated mirror movement and satisfies
\[
\inf_k f(x_k)=f^\star.
\]
Both results concern function values rather than convergence of the complete iterate sequence and apply to convex locally Lipschitz objectives, including objectives that are not globally Lipschitz or smooth [2210.01532].

The Polyak Minorant Method generalizes the affine projection interpretation to constrained convex optimization. It constructs convex minorants $\hat f_i^k$ of the objective and constraints, defines a model-based target set, and projects onto it:
\[
x^{k+1}
=
\Pi_{\mathcal X^k}(x^k).
\]
With one affine unconstrained objective minorant, this is exactly the classical Polyak subgradient update. With multiple stored minorants, the method becomes structurally related to cutting-plane and bundle methods. Fejér monotonicity gives asymptotic convergence of objective and constraint violations; a sharpness condition gives linear distance convergence [2310.07922].

## 5. Geometry, preconditioning, sparsity, and higher-order variants

The scalar Polyak rule adapts step length but does not by itself correct anisotropic curvature. A preconditioned extension uses
\[
w_{t+1}=w_t-\gamma_t M_t m_t,
\]
where $M_t\succ0$ is a preconditioner. For a weighted norm induced by $B_t\succ0$, the weighted projection derivation gives
\[
\boxed{
w_{t+1}
=
w_t
-
\frac{f_i(w_t)-f_i^\ast}
{\|\nabla f_i(w_t)\|_{B_t^{-1}}^2}
B_t^{-1}\nabla f_i(w_t).
}
\]
The preconditioner therefore affects both the direction and the denominator. PSPS instantiations use Hutchinson Hessian-diagonal estimates, AdaGrad accumulations, or Adam-style first- and second-moment scaling. Slack variants PSPSL1 and PSPSL2 address non-interpolating problems. These methods are empirically useful on badly scaled classification data, but the cited work does not provide convergence rates for the proposed preconditioned algorithms or for the Hutchinson estimation error [2310.07922].

SP2 extends the interpolation interpretation to second-order local models. For a sampled component, let
\[
c_t=f_i(w^t),\qquad
g_t=\nabla f_i(w^t),\qquad
H_t=\nabla^2 f_i(w^t).
\]
The quadratic interpolation model is
\[
q_t(w)
=
c_t+\langle g_t,w-w^t\rangle
+\frac12\langle H_t(w-w^t),w-w^t\rangle.
\]
SP2 projects onto $q_t(w)=0$. Unlike Newton's method, it seeks a root of the local quadratic rather than its minimizer, so positive-definite Hessians and convexity are not required. Negative curvature can make a positive local model cross zero and therefore facilitate the interpolation step. SP2$^+$ approximates the projection using an ordinary SP step followed by a curvature correction involving only the Hessian-vector product $H_tg_t$. The cleanest convergence theory is for sums of interpolated positive-semidefinite quadratics, where exact SP2 acts through null-space projections and can reach the solution immediately when a sampled component matrix is invertible [2207.08171].

For high-dimensional sparse M-estimation, the full gradient can be dense even when the iterate is sparse. Sparse Polyak replaces the full-gradient denominator by a hard-thresholded gradient norm:
\[
\boxed{
\gamma_t
=
\frac{\max\{f(\theta_t)-\widehat f,0\}}
{5\|\operatorname{HT}_s(\nabla f(\theta_t))\|^2}.
}
\]
The update is
\[
\theta_{t+1}
=
\operatorname{HT}_s\bigl(\theta_t-\gamma_t\nabla f(\theta_t)\bigr).
\]
Under restricted strong convexity and restricted smoothness, the resulting step remains bounded below by a quantity of order $1/\bar L$ until statistical precision is reached. The method has linear convergence governed by the restricted condition number $\bar\kappa$, rather than by the ambient dimension, under the stated sparsity requirement [2509.09802].

## 6. Rates, asymptotic behavior, and applications

For smooth strongly convex functions, Polyak gradient descent has a worst-case linear rate of order
\[
O\!\left(\left(1-\frac1\kappa\right)^K\right),
\qquad
\kappa=\frac{L}{\mu}.
\]
An analytic worst-case analysis establishes the sharp contraction factor
\[
q_1=
\left(\frac{L-\mu}{L+\mu}\right)^2
\]
for the standard choice on the relevant squared-distance measure. This matches the worst-case contraction factor of exact line search, although exact line search contracts objective error while Polyak descent contracts squared distance. On quadratics, the method asymptotically concentrates in the two-dimensional subspace spanned by the eigenvectors associated with the smallest and largest eigenvalues. The limiting stepsize approaches $2/(L+\mu)$, producing opposite-sign contraction factors in the two extreme eigendirections and an alternating zigzag trajectory [2407.04914].

For smooth convex functions, Polyak gradient descent has an $O(1/K)$ best-iterate rate, and explicit quadratic constructions show that this dependence is tight. Under Hölder smoothness with exponent $\nu\in(0,1]$, the best-iterate rate becomes
\[
O\!\left(K^{-(\nu+1)/2}\right),
\]
with matching lower bounds for the corresponding class. Hölder growth conditions produce geometry-dependent rates, including linear convergence when the growth exponent matches the smoothness structure. These results establish a form of universality: the same update adapts to unknown smoothness exponents, growth exponents, condition numbers, and curvature constants once $f^\star$ is available [2512.06231].

Statistical analyses distinguish population contraction from empirical perturbation. Under generalized smoothness, generalized Łojasiewicz conditions, and empirical-to-population gradient stability, the population Polyak operator contracts geometrically even in singular models. The empirical iterates reach a statistical radius
\[
r_n
=
O\!\left(
\varepsilon(n,\delta)^{1/(\alpha+1-\gamma)}
\right)
\]
after
\[
O\!\left(\log\frac1{\varepsilon(n,\delta)}\right)
\]
iterations, whereas fixed-step gradient descent can require polynomially many iterations when curvature degenerates at the optimum. The theory applies to generalized linear models, Gaussian mixtures, and mixed linear regression, including low-signal regimes with singular Fisher information [2110.07810].

Applications and extensions include over-parameterized classification and matrix factorization, composite learning with weight decay, nonsmooth constrained optimization, decentralized diffusion, and reinforcement learning. In policy-gradient reinforcement learning, Polyak adaptation is applied to return maximization by replacing the unknown optimal return with the higher return of two nearby policies. Entropy regularization and upper clipping are used to prevent softmax-gradient collapse and explosive updates. Experiments on Acrobot, CartPole, and LunarLander report faster convergence and more stable final policies than fixed-learning-rate Adam, but no formal convergence theorem is provided for the complete stochastic RL method [2404.07525].

A decentralized graph-aware rule derives local learning rates by minimizing a network-wide one-step distance to the optimizer. The resulting coupled linear system depends on the two-hop matrix $A^2$, neighboring objective values, and gradient correlations. It reduces to the classical Polyak rule in the single-agent limit, but the proposed decentralized stochastic recursion has no complete convergence-rate theorem [2509.14854]. For randomized methods solving generalized absolute value equations, the objective changes with the current iterate:
\[
f_S^k(x)
=
\frac12
\left\|
S^\top(Ax-B|x^k|-b)
\right\|^2.
\]
A Polyak ratio based on the sampled residual yields linear convergence in expectation under the stated sketching and spectral conditions [2608.00952].

The principal limitations of Polyak step-size methods are consistent across these settings. Exact or reliable target values are often unavailable; invalid lower bounds can cause overly aggressive or negative steps; small gradient norms create denominator instability; clipping and damping introduce parameters; stochastic objectives and gradients can destroy deterministic unbiasedness arguments; and nonconvex guarantees generally require interpolation, growth, PL, quasar-convexity, star-convexity, or other structural assumptions. Polyak adaptation therefore should not be identified with a universally parameter-free or universally accelerated method. Its central contribution is adaptive scale selection from objective-gap information, with the strongest behavior appearing when the target value is known or estimable and when the objective possesses interpolation, restricted geometry, or a suitable growth structure.

Source: https://www.emergentmind.com/topics/polyak-step-size