Papers
Topics
Authors
Recent
Search
2000 character limit reached

Polyak Step-Size Method

Updated 23 August 2026
  • Polyak step-size is an adaptive learning-rate method that uses the current objective gap and gradient norm to determine step size, adapting to the local curvature and proximity to the optimum. It does not require a globally smoothness constant, but needs the optimal objective value which is often unavailable in real-world scenarios. Examples of applications include stochastic optimization, momentum, proximal and mirror descent, constrained problems, high-dimensional sparse estimation, decentralized optimization, and reinforcement learning.
  • The stochastic Polyak step-size (SPS) for finite-sum objectives updates step size using a sampled component, similarly modifying updates to be robust to varying gradients. Methods like MOTAPS, Twin-Polyak, and PSADLA extend SPS for unknown or unknown components, targeting global optimums with adaptive scaling, to projecting methods that handle varying constraints by adapting its objective levels.
  • Polyak step-size enables faster convergence of iterative optimization methods in objective evaluations, addressing the overfitting effect when varying scale. With precision-targeted and projectional strategies pushed in non-convex landscapes, Polyak step-size has a significant role as a targeted convergence entity.

Polyak step-size is an adaptive learning-rate rule that selects the step from the current objective gap divided by a squared gradient norm. For a differentiable convex objective with known optimum value f⋆f^\star, the classical rule is

αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).

It was derived by minimizing a quadratic upper bound on the next squared distance to an optimizer. The method does not require an externally supplied smoothness constant, but its classical form requires the optimal objective value, which is generally unavailable. Subsequent work has extended the construction to stochastic optimization, momentum, proximal and mirror descent, constrained problems, preconditioned updates, high-dimensional sparse estimation, decentralized optimization, reinforcement learning, and objectives whose curvature degenerates near the optimum.

1. Classical rule and geometric derivation

Consider the convex subgradient iteration

xk+1=xk−γkgk,gk∈∂f(xk).x^{k+1}=x^k-\gamma_k g^k, \qquad g^k\in\partial f(x^k).

For an optimizer x⋆x^\star and optimal value f⋆=f(x⋆)f^\star=f(x^\star), convexity gives

⟨gk,xk−x⋆⟩≥f(xk)−f⋆.\langle g^k,x^k-x^\star\rangle\geq f(x^k)-f^\star.

Consequently,

∥xk+1−x⋆∥2≤∥xk−x⋆∥2−2γk(f(xk)−f⋆)+γk2∥gk∥2.\|x^{k+1}-x^\star\|^2 \leq \|x^k-x^\star\|^2 -2\gamma_k\bigl(f(x^k)-f^\star\bigr) +\gamma_k^2\|g^k\|^2.

Minimizing the right-hand side as a quadratic function of γk\gamma_k yields

γkPolyak=f(xk)−f⋆∥gk∥2.\boxed{ \gamma_k^{\mathrm{Polyak}} = \frac{f(x^k)-f^\star}{\|g^k\|^2} }.

For smooth objectives, gk=∇f(xk)g^k=\nabla f(x^k).

The numerator is the current objective gap and the denominator normalizes the step by the local gradient scale. Thus the rule tends to take larger steps far from the optimum and smaller steps near it. If αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).0 is known exactly, no estimate of a global Lipschitz or smoothness constant is required. The same construction can be interpreted geometrically: the affine subgradient minorant

αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).1

is intersected with the optimal level αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).2, and the update is the point on that hyperplane closest to αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).3. This projection interpretation is the basis of the Polyak Minorant Method, which replaces affine subgradient models with arbitrary convex lower models and incorporates constraints (Devanathan et al., 2023).

A scaled form is frequently used: αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).4 The case αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).5 is the usual Polyak method. For αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).6, the elementary distance recursion contains the positive descent coefficient αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).7.

The requirement that αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).8 be known is both the principal strength and the principal limitation of the rule. The method is adaptive once the target value is available, but is not parameter-free in an absolute sense. In nonsmooth convex optimization, exact last-iterate behavior is substantially weaker than many best-iterate or averaged-iterate guarantees: with uniformly bounded subgradients, the standard Polyak method has an exact worst-case last-iterate rate of order

αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).9

where xk+1=xk−γkgk,gk∈∂f(xk).x^{k+1}=x^k-\gamma_k g^k, \qquad g^k\in\partial f(x^k).0 bounds subgradients and xk+1=xk−γkgk,gk∈∂f(xk).x^{k+1}=x^k-\gamma_k g^k, \qquad g^k\in\partial f(x^k).1 bounds the initial distance to a solution (Zamani et al., 2024).

2. Stochastic Polyak step-size

For a finite-sum objective

xk+1=xk−γkgk,gk∈∂f(xk).x^{k+1}=x^k-\gamma_k g^k, \qquad g^k\in\partial f(x^k).2

stochastic gradient descent samples xk+1=xk−γkgk,gk∈∂f(xk).x^{k+1}=x^k-\gamma_k g^k, \qquad g^k\in\partial f(x^k).3 and updates

xk+1=xk−γkgk,gk∈∂f(xk).x^{k+1}=x^k-\gamma_k g^k, \qquad g^k\in\partial f(x^k).4

Applying the Polyak construction to the sampled component gives the stochastic Polyak step-size (SPS): xk+1=xk−γkgk,gk∈∂f(xk).x^{k+1}=x^k-\gamma_k g^k, \qquad g^k\in\partial f(x^k).5 where

xk+1=xk−γkgk,gk∈∂f(xk).x^{k+1}=x^k-\gamma_k g^k, \qquad g^k\in\partial f(x^k).6

A bounded version is

xk+1=xk−γkgk,gk∈∂f(xk).x^{k+1}=x^k-\gamma_k g^k, \qquad g^k\in\partial f(x^k).7

The cap prevents very large steps when a sampled gradient is small. The loss and gradient used by SPS are already computed in an SGD iteration, so unbounded SPS requires no full-objective or full-gradient evaluation and has essentially no additional per-iteration computational cost (Loizou et al., 2020).

The component optimum xk+1=xk−γkgk,gk∈∂f(xk).x^{k+1}=x^k-\gamma_k g^k, \qquad g^k\in\partial f(x^k).8 is weaker information than the global optimum xk+1=xk−γkgk,gk∈∂f(xk).x^{k+1}=x^k-\gamma_k g^k, \qquad g^k\in\partial f(x^k).9. For many nonnegative unregularized losses, x⋆x^\star0 is available or is the relevant infimum. Examples include squared loss in realizable regression, logistic loss in separable classification, exponential loss, and squared-hinge loss in linearly separable classification. For unregularized logistic loss, the infimum can be zero without being attained by any finite parameter vector. With regularization, individual minima are generally nonzero, although they can sometimes be computed analytically or cheaply offline.

The parameter x⋆x^\star1 controls the scale of the update. The cited SPS analysis recommends x⋆x^\star2 in strongly convex and interpolation analyses and x⋆x^\star3 for a convenient convex-objective bound. A central inequality is

x⋆x^\star4

which is an equality for unbounded SPS. This permits convergence analysis without bounded stochastic gradients or a standard bounded-variance assumption in several settings.

Define the stochastic inconsistency quantity

x⋆x^\star5

Under interpolation,

x⋆x^\star6

For convex smooth components and a strongly convex aggregate objective, capped SPS converges linearly in expected squared distance to a neighborhood whose size depends on x⋆x^\star7 and x⋆x^\star8. Under interpolation, ordinary SPS with x⋆x^\star9 converges linearly to the true solution. For convex non-strongly-convex objectives, the averaged iterate has an f⋆=f(x⋆)f^\star=f(x^\star)0 objective-gap rate up to a residual neighborhood; under interpolation, the residual vanishes. For smooth PL objectives, SPS yields linear objective convergence under the stated parameter restrictions. For general smooth nonconvex objectives, weak growth yields an f⋆=f(x⋆)f^\star=f(x^\star)1 stationarity rate up to a neighborhood governed by the growth residual.

The same analysis gives constant-step SGD bounds. If the cap is sufficiently small, capped SPS is effectively constant-step SGD. In the strongly convex case,

f⋆=f(x⋆)f^\star=f(x^\star)2

Under interpolation, the residual term disappears. The neighborhood is expressed in terms of the optimal objective difference f⋆=f(x⋆)f^\star=f(x^\star)3, rather than solely in terms of the conventional gradient-noise quantity

f⋆=f(x⋆)f^\star=f(x^\star)4

3. Interpolation, non-interpolation, and optimal-value estimation

Interpolation means that a common point minimizes every component: f⋆=f(x⋆)f^\star=f(x^\star)5 It occurs in realizable least-squares problems, consistent linear systems, separable classification, sufficiently expressive kernel models, and over-parameterized neural networks. Under interpolation, stochastic gradients vanish at a common solution and constant-step stochastic methods need not settle in a persistent noise neighborhood.

For consistent linear systems, the interpolation interpretation is especially direct. A stochastic quadratic reformulation can satisfy

f⋆=f(x⋆)f^\star=f(x^\star)6

With f⋆=f(x⋆)f^\star=f(x^\star)7, SPS therefore produces a unit step and recovers the theoretically optimal step for randomized sketch-and-project methods, including randomized Kaczmarz. The associated distance guarantee depends on the smallest nonzero eigenvalue of the expected system matrix (Loizou et al., 2020).

Outside interpolation, replacing f⋆=f(x⋆)f^\star=f(x^\star)8 by zero can produce systematically excessive steps. The residual f⋆=f(x⋆)f^\star=f(x^\star)9 then determines the limiting neighborhood in many SPS bounds. Several methods address this limitation by estimating individual or global target values.

MOTAPS, or Moving Targetted Polyak Stepsize, maintains one auxiliary scalar for every data point. These variables track individual losses and permit a moving estimate of the global target. TAPS assumes that the global optimum value is known; MOTAPS also estimates that value and therefore does not require interpolation. Its auxiliary objective contains a dampening parameter ⟨gk,xk−x⋆⟩≥f(xk)−f⋆.\langle g^k,x^k-x^\star\rangle\geq f(x^k)-f^\star.0. Fixed ⟨gk,xk−x⋆⟩≥f(xk)−f⋆.\langle g^k,x^k-x^\star\rangle\geq f(x^k)-f^\star.1 produces convergence to a neighborhood, with an error term proportional to

⟨gk,xk−x⋆⟩≥f(xk)−f⋆.\langle g^k,x^k-x^\star\rangle\geq f(x^k)-f^\star.2

or

⟨gk,xk−x⋆⟩≥f(xk)−f⋆.\langle g^k,x^k-x^\star\rangle\geq f(x^k)-f^\star.3

Eventually decreasing the step-size can remove the fixed-⟨gk,xk−x⋆⟩≥f(xk)−f⋆.\langle g^k,x^k-x^\star\rangle\geq f(x^k)-f^\star.4 error floor asymptotically. The convergence theory is conditional on star-convexity or strong star-convexity of an auxiliary objective, which does not automatically follow from convexity of the original component losses (Gower et al., 2021).

Twin-Polyak maintains two iterates and updates the higher-valued one using the lower-valued objective as a proxy for ⟨gk,xk−x⋆⟩≥f(xk)−f⋆.\langle g^k,x^k-x^\star\rangle\geq f(x^k)-f^\star.5. If ⟨gk,xk−x⋆⟩≥f(xk)−f⋆.\langle g^k,x^k-x^\star\rangle\geq f(x^k)-f^\star.6,

⟨gk,xk−x⋆⟩≥f(xk)−f⋆.\langle g^k,x^k-x^\star\rangle\geq f(x^k)-f^\star.7

The lower-valued sequence supplies an observed target rather than a known optimum. The method is invariant to positive objective scaling and objective translation. Under deterministic twin-gap assumptions, it has linear convergence for strongly convex smooth objectives and an ⟨gk,xk−x⋆⟩≥f(xk)−f⋆.\langle g^k,x^k-x^\star\rangle\geq f(x^k)-f^\star.8 best-iterate rate for convex bounded-gradient objectives. Its stochastic variants are primarily empirical because the stepsize depends on the same random minibatch as the gradient (Fainman et al., 18 Sep 2025).

PSADLA, or Polyak Stepsize with Accelerated Decision-Guided Level Adjustment, maintains a lower level ⟨gk,xk−x⋆⟩≥f(xk)−f⋆.\langle g^k,x^k-x^\star\rangle\geq f(x^k)-f^\star.9 and uses

∥xk+1−x⋆∥2≤∥xk−x⋆∥2−2γk(f(xk)−f⋆)+γk2∥gk∥2.\|x^{k+1}-x^\star\|^2 \leq \|x^k-x^\star\|^2 -2\gamma_k\bigl(f(x^k)-f^\star\bigr) +\gamma_k^2\|g^k\|^2.0

A Polyak Stepsize Violation Detector (PSVD) tests the feasibility of a block of linear inequalities encoding whether recent updates remain compatible with a possible optimizer. If the system becomes infeasible, the level is raised by a convex combination of the previous level and the best objective value observed in the block. The method proves convergence of the level values and running-best objective values to ∥xk+1−x⋆∥2≤∥xk−x⋆∥2−2γk(f(xk)−f⋆)+γk2∥gk∥2.\|x^{k+1}-x^\star\|^2 \leq \|x^k-x^\star\|^2 -2\gamma_k\bigl(f(x^k)-f^\star\bigr) +\gamma_k^2\|g^k\|^2.1, including when approximate or surrogate subgradients satisfy the stated conditions (Liu et al., 2023).

4. Extensions to composite, momentum, mirror, and constrained optimization

For composite objectives

∥xk+1−x⋆∥2≤∥xk−x⋆∥2−2γk(f(xk)−f⋆)+γk2∥gk∥2.\|x^{k+1}-x^\star\|^2 \leq \|x^k-x^\star\|^2 -2\gamma_k\bigl(f(x^k)-f^\star\bigr) +\gamma_k^2\|g^k\|^2.2

ordinary SPS applied to the complete sample objective requires a lower bound on ∥xk+1−x⋆∥2≤∥xk−x⋆∥2−2γk(f(xk)−f⋆)+γk2∥gk∥2.\|x^{k+1}-x^\star\|^2 \leq \|x^k-x^\star\|^2 -2\gamma_k\bigl(f(x^k)-f^\star\bigr) +\gamma_k^2\|g^k\|^2.3. Such a bound may be loose when ∥xk+1−x⋆∥2≤∥xk−x⋆∥2−2γk(f(xk)−f⋆)+γk2∥gk∥2.\|x^{k+1}-x^\star\|^2 \leq \|x^k-x^\star\|^2 -2\gamma_k\bigl(f(x^k)-f^\star\bigr) +\gamma_k^2\|g^k\|^2.4 is a regularizer. ProxSPS applies Polyak adaptation only to the stochastic loss and treats ∥xk+1−x⋆∥2≤∥xk−x⋆∥2−2γk(f(xk)−f⋆)+γk2∥gk∥2.\|x^{k+1}-x^\star\|^2 \leq \|x^k-x^\star\|^2 -2\gamma_k\bigl(f(x^k)-f^\star\bigr) +\gamma_k^2\|g^k\|^2.5 through an exact proximal step: ∥xk+1−x⋆∥2≤∥xk−x⋆∥2−2γk(f(xk)−f⋆)+γk2∥gk∥2.\|x^{k+1}-x^\star\|^2 \leq \|x^k-x^\star\|^2 -2\gamma_k\bigl(f(x^k)-f^\star\bigr) +\gamma_k^2\|g^k\|^2.6 It requires only

∥xk+1−x⋆∥2≤∥xk−x⋆∥2−2γk(f(xk)−f⋆)+γk2∥gk∥2.\|x^{k+1}-x^\star\|^2 \leq \|x^k-x^\star\|^2 -2\gamma_k\bigl(f(x^k)-f^\star\bigr) +\gamma_k^2\|g^k\|^2.7

not a lower bound on the regularized loss. For squared ∥xk+1−x⋆∥2≤∥xk−x⋆∥2−2γk(f(xk)−f⋆)+γk2∥gk∥2.\|x^{k+1}-x^\star\|^2 \leq \|x^k-x^\star\|^2 -2\gamma_k\bigl(f(x^k)-f^\star\bigr) +\gamma_k^2\|g^k\|^2.8 regularization,

∥xk+1−x⋆∥2≤∥xk−x⋆∥2−2γk(f(xk)−f⋆)+γk2∥gk∥2.\|x^{k+1}-x^\star\|^2 \leq \|x^k-x^\star\|^2 -2\gamma_k\bigl(f(x^k)-f^\star\bigr) +\gamma_k^2\|g^k\|^2.9

ProxSPS has convergence results for nonsmooth weakly convex, smooth convex, strongly convex, and smooth weakly convex settings. Its nonconvex stationarity guarantee is expressed through the Moreau envelope and has an approximately γk\gamma_k0 rate; strongly convex settings yield linear convergence up to a noise neighborhood (Schaipp et al., 2023).

Momentum requires modifying the Polyak derivation because the update contains an inherited displacement. For heavy-ball momentum,

γk\gamma_k1

the generalized adaptive learning rate includes a momentum-alignment term: γk\gamma_k2 ALR-HB therefore reduces the gradient step when momentum opposes the current gradient. For moving averaged gradients,

γk\gamma_k3

ALR-MAG uses

γk\gamma_k4

Its deterministic analysis establishes monotone distance decrease and global linear convergence under smooth semi-strong convexity. Stochastic extensions include ALR-SMAG and ALR-SHB. A later IMA-based analysis proposes MomSPSγk\gamma_k5, MomDecSPS, and MomAdaSPS. MomSPSγk\gamma_k6 scales SPS by γk\gamma_k7 and converges to a neighborhood without interpolation, while interpolation gives an γk\gamma_k8 rate. MomDecSPS and MomAdaSPS converge to the exact minimizer in non-interpolated convex problems, with rates of order γk\gamma_k9 and γkPolyak=f(xk)−f⋆∥gk∥2.\boxed{ \gamma_k^{\mathrm{Polyak}} = \frac{f(x^k)-f^\star}{\|g^k\|^2} }.0, respectively (Wang et al., 2023, Oikonomou et al., 2024).

Mirror-descent variants replace Euclidean geometry with a Bregman divergence. For a mirror map γkPolyak=f(xk)−f⋆∥gk∥2.\boxed{ \gamma_k^{\mathrm{Polyak}} = \frac{f(x^k)-f^\star}{\|g^k\|^2} }.1, the update is

γkPolyak=f(xk)−f⋆∥gk∥2.\boxed{ \gamma_k^{\mathrm{Polyak}} = \frac{f(x^k)-f^\star}{\|g^k\|^2} }.2

Two Polyak-type rules replace γkPolyak=f(xk)−f⋆∥gk∥2.\boxed{ \gamma_k^{\mathrm{Polyak}} = \frac{f(x^k)-f^\star}{\|g^k\|^2} }.3 by adaptive target levels. The first maintains

γkPolyak=f(xk)−f⋆∥gk∥2.\boxed{ \gamma_k^{\mathrm{Polyak}} = \frac{f(x^k)-f^\star}{\|g^k\|^2} }.4

and guarantees an objective value within a prescribed tolerance γkPolyak=f(xk)−f⋆∥gk∥2.\boxed{ \gamma_k^{\mathrm{Polyak}} = \frac{f(x^k)-f^\star}{\|g^k\|^2} }.5 of optimality. The second uses level reductions driven by accumulated mirror movement and satisfies

γkPolyak=f(xk)−f⋆∥gk∥2.\boxed{ \gamma_k^{\mathrm{Polyak}} = \frac{f(x^k)-f^\star}{\|g^k\|^2} }.6

Both results concern function values rather than convergence of the complete iterate sequence and apply to convex locally Lipschitz objectives, including objectives that are not globally Lipschitz or smooth (You et al., 2022).

The Polyak Minorant Method generalizes the affine projection interpretation to constrained convex optimization. It constructs convex minorants γkPolyak=f(xk)−f⋆∥gk∥2.\boxed{ \gamma_k^{\mathrm{Polyak}} = \frac{f(x^k)-f^\star}{\|g^k\|^2} }.7 of the objective and constraints, defines a model-based target set, and projects onto it: γkPolyak=f(xk)−f⋆∥gk∥2.\boxed{ \gamma_k^{\mathrm{Polyak}} = \frac{f(x^k)-f^\star}{\|g^k\|^2} }.8 With one affine unconstrained objective minorant, this is exactly the classical Polyak subgradient update. With multiple stored minorants, the method becomes structurally related to cutting-plane and bundle methods. Fejér monotonicity gives asymptotic convergence of objective and constraint violations; a sharpness condition gives linear distance convergence (Devanathan et al., 2023).

5. Geometry, preconditioning, sparsity, and higher-order variants

The scalar Polyak rule adapts step length but does not by itself correct anisotropic curvature. A preconditioned extension uses

γkPolyak=f(xk)−f⋆∥gk∥2.\boxed{ \gamma_k^{\mathrm{Polyak}} = \frac{f(x^k)-f^\star}{\|g^k\|^2} }.9

where gk=∇f(xk)g^k=\nabla f(x^k)0 is a preconditioner. For a weighted norm induced by gk=∇f(xk)g^k=\nabla f(x^k)1, the weighted projection derivation gives

gk=∇f(xk)g^k=\nabla f(x^k)2

The preconditioner therefore affects both the direction and the denominator. PSPS instantiations use Hutchinson Hessian-diagonal estimates, AdaGrad accumulations, or Adam-style first- and second-moment scaling. Slack variants PSPSL1 and PSPSL2 address non-interpolating problems. These methods are empirically useful on badly scaled classification data, but the cited work does not provide convergence rates for the proposed preconditioned algorithms or for the Hutchinson estimation error (Devanathan et al., 2023).

SP2 extends the interpolation interpretation to second-order local models. For a sampled component, let

gk=∇f(xk)g^k=\nabla f(x^k)3

The quadratic interpolation model is

gk=∇f(xk)g^k=\nabla f(x^k)4

SP2 projects onto gk=∇f(xk)g^k=\nabla f(x^k)5. Unlike Newton's method, it seeks a root of the local quadratic rather than its minimizer, so positive-definite Hessians and convexity are not required. Negative curvature can make a positive local model cross zero and therefore facilitate the interpolation step. SP2gk=∇f(xk)g^k=\nabla f(x^k)6 approximates the projection using an ordinary SP step followed by a curvature correction involving only the Hessian-vector product gk=∇f(xk)g^k=\nabla f(x^k)7. The cleanest convergence theory is for sums of interpolated positive-semidefinite quadratics, where exact SP2 acts through null-space projections and can reach the solution immediately when a sampled component matrix is invertible (Li et al., 2022).

For high-dimensional sparse M-estimation, the full gradient can be dense even when the iterate is sparse. Sparse Polyak replaces the full-gradient denominator by a hard-thresholded gradient norm: gk=∇f(xk)g^k=\nabla f(x^k)8 The update is

gk=∇f(xk)g^k=\nabla f(x^k)9

Under restricted strong convexity and restricted smoothness, the resulting step remains bounded below by a quantity of order αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).00 until statistical precision is reached. The method has linear convergence governed by the restricted condition number αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).01, rather than by the ambient dimension, under the stated sparsity requirement (Qiao et al., 11 Sep 2025).

6. Rates, asymptotic behavior, and applications

For smooth strongly convex functions, Polyak gradient descent has a worst-case linear rate of order

αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).02

An analytic worst-case analysis establishes the sharp contraction factor

αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).03

for the standard choice on the relevant squared-distance measure. This matches the worst-case contraction factor of exact line search, although exact line search contracts objective error while Polyak descent contracts squared distance. On quadratics, the method asymptotically concentrates in the two-dimensional subspace spanned by the eigenvectors associated with the smallest and largest eigenvalues. The limiting stepsize approaches αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).04, producing opposite-sign contraction factors in the two extreme eigendirections and an alternating zigzag trajectory (Huang et al., 2024).

For smooth convex functions, Polyak gradient descent has an αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).05 best-iterate rate, and explicit quadratic constructions show that this dependence is tight. Under Hölder smoothness with exponent αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).06, the best-iterate rate becomes

αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).07

with matching lower bounds for the corresponding class. Hölder growth conditions produce geometry-dependent rates, including linear convergence when the growth exponent matches the smoothness structure. These results establish a form of universality: the same update adapts to unknown smoothness exponents, growth exponents, condition numbers, and curvature constants once αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).08 is available (He et al., 6 Dec 2025).

Statistical analyses distinguish population contraction from empirical perturbation. Under generalized smoothness, generalized Łojasiewicz conditions, and empirical-to-population gradient stability, the population Polyak operator contracts geometrically even in singular models. The empirical iterates reach a statistical radius

αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).09

after

αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).10

iterations, whereas fixed-step gradient descent can require polynomially many iterations when curvature degenerates at the optimum. The theory applies to generalized linear models, Gaussian mixtures, and mixed linear regression, including low-signal regimes with singular Fisher information (Ren et al., 2021).

Applications and extensions include over-parameterized classification and matrix factorization, composite learning with weight decay, nonsmooth constrained optimization, decentralized diffusion, and reinforcement learning. In policy-gradient reinforcement learning, Polyak adaptation is applied to return maximization by replacing the unknown optimal return with the higher return of two nearby policies. Entropy regularization and upper clipping are used to prevent softmax-gradient collapse and explosive updates. Experiments on Acrobot, CartPole, and LunarLander report faster convergence and more stable final policies than fixed-learning-rate Adam, but no formal convergence theorem is provided for the complete stochastic RL method (Li et al., 2024).

A decentralized graph-aware rule derives local learning rates by minimizing a network-wide one-step distance to the optimizer. The resulting coupled linear system depends on the two-hop matrix αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).11, neighboring objective values, and gradient correlations. It reduces to the classical Polyak rule in the single-agent limit, but the proposed decentralized stochastic recursion has no complete convergence-rate theorem (Fainman et al., 18 Sep 2025). For randomized methods solving generalized absolute value equations, the objective changes with the current iterate: αk=f(xk)−f⋆∥∇f(xk)∥2,xk+1=xk−αk∇f(xk).\alpha_k=\frac{f(x_k)-f^\star}{\|\nabla f(x_k)\|^2}, \qquad x_{k+1}=x_k-\alpha_k\nabla f(x_k).12 A Polyak ratio based on the sampled residual yields linear convergence in expectation under the stated sketching and spectral conditions (Chen et al., 2 Aug 2026).

The principal limitations of Polyak step-size methods are consistent across these settings. Exact or reliable target values are often unavailable; invalid lower bounds can cause overly aggressive or negative steps; small gradient norms create denominator instability; clipping and damping introduce parameters; stochastic objectives and gradients can destroy deterministic unbiasedness arguments; and nonconvex guarantees generally require interpolation, growth, PL, quasar-convexity, star-convexity, or other structural assumptions. Polyak adaptation therefore should not be identified with a universally parameter-free or universally accelerated method. Its central contribution is adaptive scale selection from objective-gap information, with the strongest behavior appearing when the target value is known or estimable and when the objective possesses interpolation, restricted geometry, or a suitable growth structure.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Polyak Step-Size.