Polyak Step-Size Method
- Polyak step-size is an adaptive learning-rate method that uses the current objective gap and gradient norm to determine step size, adapting to the local curvature and proximity to the optimum. It does not require a globally smoothness constant, but needs the optimal objective value which is often unavailable in real-world scenarios. Examples of applications include stochastic optimization, momentum, proximal and mirror descent, constrained problems, high-dimensional sparse estimation, decentralized optimization, and reinforcement learning.
- The stochastic Polyak step-size (SPS) for finite-sum objectives updates step size using a sampled component, similarly modifying updates to be robust to varying gradients. Methods like MOTAPS, Twin-Polyak, and PSADLA extend SPS for unknown or unknown components, targeting global optimums with adaptive scaling, to projecting methods that handle varying constraints by adapting its objective levels.
- Polyak step-size enables faster convergence of iterative optimization methods in objective evaluations, addressing the overfitting effect when varying scale. With precision-targeted and projectional strategies pushed in non-convex landscapes, Polyak step-size has a significant role as a targeted convergence entity.
Polyak step-size is an adaptive learning-rate rule that selects the step from the current objective gap divided by a squared gradient norm. For a differentiable convex objective with known optimum value , the classical rule is
It was derived by minimizing a quadratic upper bound on the next squared distance to an optimizer. The method does not require an externally supplied smoothness constant, but its classical form requires the optimal objective value, which is generally unavailable. Subsequent work has extended the construction to stochastic optimization, momentum, proximal and mirror descent, constrained problems, preconditioned updates, high-dimensional sparse estimation, decentralized optimization, reinforcement learning, and objectives whose curvature degenerates near the optimum.
1. Classical rule and geometric derivation
Consider the convex subgradient iteration
For an optimizer and optimal value , convexity gives
Consequently,
Minimizing the right-hand side as a quadratic function of yields
For smooth objectives, .
The numerator is the current objective gap and the denominator normalizes the step by the local gradient scale. Thus the rule tends to take larger steps far from the optimum and smaller steps near it. If 0 is known exactly, no estimate of a global Lipschitz or smoothness constant is required. The same construction can be interpreted geometrically: the affine subgradient minorant
1
is intersected with the optimal level 2, and the update is the point on that hyperplane closest to 3. This projection interpretation is the basis of the Polyak Minorant Method, which replaces affine subgradient models with arbitrary convex lower models and incorporates constraints (Devanathan et al., 2023).
A scaled form is frequently used: 4 The case 5 is the usual Polyak method. For 6, the elementary distance recursion contains the positive descent coefficient 7.
The requirement that 8 be known is both the principal strength and the principal limitation of the rule. The method is adaptive once the target value is available, but is not parameter-free in an absolute sense. In nonsmooth convex optimization, exact last-iterate behavior is substantially weaker than many best-iterate or averaged-iterate guarantees: with uniformly bounded subgradients, the standard Polyak method has an exact worst-case last-iterate rate of order
9
where 0 bounds subgradients and 1 bounds the initial distance to a solution (Zamani et al., 2024).
2. Stochastic Polyak step-size
For a finite-sum objective
2
stochastic gradient descent samples 3 and updates
4
Applying the Polyak construction to the sampled component gives the stochastic Polyak step-size (SPS): 5 where
6
A bounded version is
7
The cap prevents very large steps when a sampled gradient is small. The loss and gradient used by SPS are already computed in an SGD iteration, so unbounded SPS requires no full-objective or full-gradient evaluation and has essentially no additional per-iteration computational cost (Loizou et al., 2020).
The component optimum 8 is weaker information than the global optimum 9. For many nonnegative unregularized losses, 0 is available or is the relevant infimum. Examples include squared loss in realizable regression, logistic loss in separable classification, exponential loss, and squared-hinge loss in linearly separable classification. For unregularized logistic loss, the infimum can be zero without being attained by any finite parameter vector. With regularization, individual minima are generally nonzero, although they can sometimes be computed analytically or cheaply offline.
The parameter 1 controls the scale of the update. The cited SPS analysis recommends 2 in strongly convex and interpolation analyses and 3 for a convenient convex-objective bound. A central inequality is
4
which is an equality for unbounded SPS. This permits convergence analysis without bounded stochastic gradients or a standard bounded-variance assumption in several settings.
Define the stochastic inconsistency quantity
5
Under interpolation,
6
For convex smooth components and a strongly convex aggregate objective, capped SPS converges linearly in expected squared distance to a neighborhood whose size depends on 7 and 8. Under interpolation, ordinary SPS with 9 converges linearly to the true solution. For convex non-strongly-convex objectives, the averaged iterate has an 0 objective-gap rate up to a residual neighborhood; under interpolation, the residual vanishes. For smooth PL objectives, SPS yields linear objective convergence under the stated parameter restrictions. For general smooth nonconvex objectives, weak growth yields an 1 stationarity rate up to a neighborhood governed by the growth residual.
The same analysis gives constant-step SGD bounds. If the cap is sufficiently small, capped SPS is effectively constant-step SGD. In the strongly convex case,
2
Under interpolation, the residual term disappears. The neighborhood is expressed in terms of the optimal objective difference 3, rather than solely in terms of the conventional gradient-noise quantity
4
3. Interpolation, non-interpolation, and optimal-value estimation
Interpolation means that a common point minimizes every component: 5 It occurs in realizable least-squares problems, consistent linear systems, separable classification, sufficiently expressive kernel models, and over-parameterized neural networks. Under interpolation, stochastic gradients vanish at a common solution and constant-step stochastic methods need not settle in a persistent noise neighborhood.
For consistent linear systems, the interpolation interpretation is especially direct. A stochastic quadratic reformulation can satisfy
6
With 7, SPS therefore produces a unit step and recovers the theoretically optimal step for randomized sketch-and-project methods, including randomized Kaczmarz. The associated distance guarantee depends on the smallest nonzero eigenvalue of the expected system matrix (Loizou et al., 2020).
Outside interpolation, replacing 8 by zero can produce systematically excessive steps. The residual 9 then determines the limiting neighborhood in many SPS bounds. Several methods address this limitation by estimating individual or global target values.
MOTAPS, or Moving Targetted Polyak Stepsize, maintains one auxiliary scalar for every data point. These variables track individual losses and permit a moving estimate of the global target. TAPS assumes that the global optimum value is known; MOTAPS also estimates that value and therefore does not require interpolation. Its auxiliary objective contains a dampening parameter 0. Fixed 1 produces convergence to a neighborhood, with an error term proportional to
2
or
3
Eventually decreasing the step-size can remove the fixed-4 error floor asymptotically. The convergence theory is conditional on star-convexity or strong star-convexity of an auxiliary objective, which does not automatically follow from convexity of the original component losses (Gower et al., 2021).
Twin-Polyak maintains two iterates and updates the higher-valued one using the lower-valued objective as a proxy for 5. If 6,
7
The lower-valued sequence supplies an observed target rather than a known optimum. The method is invariant to positive objective scaling and objective translation. Under deterministic twin-gap assumptions, it has linear convergence for strongly convex smooth objectives and an 8 best-iterate rate for convex bounded-gradient objectives. Its stochastic variants are primarily empirical because the stepsize depends on the same random minibatch as the gradient (Fainman et al., 18 Sep 2025).
PSADLA, or Polyak Stepsize with Accelerated Decision-Guided Level Adjustment, maintains a lower level 9 and uses
0
A Polyak Stepsize Violation Detector (PSVD) tests the feasibility of a block of linear inequalities encoding whether recent updates remain compatible with a possible optimizer. If the system becomes infeasible, the level is raised by a convex combination of the previous level and the best objective value observed in the block. The method proves convergence of the level values and running-best objective values to 1, including when approximate or surrogate subgradients satisfy the stated conditions (Liu et al., 2023).
4. Extensions to composite, momentum, mirror, and constrained optimization
For composite objectives
2
ordinary SPS applied to the complete sample objective requires a lower bound on 3. Such a bound may be loose when 4 is a regularizer. ProxSPS applies Polyak adaptation only to the stochastic loss and treats 5 through an exact proximal step: 6 It requires only
7
not a lower bound on the regularized loss. For squared 8 regularization,
9
ProxSPS has convergence results for nonsmooth weakly convex, smooth convex, strongly convex, and smooth weakly convex settings. Its nonconvex stationarity guarantee is expressed through the Moreau envelope and has an approximately 0 rate; strongly convex settings yield linear convergence up to a noise neighborhood (Schaipp et al., 2023).
Momentum requires modifying the Polyak derivation because the update contains an inherited displacement. For heavy-ball momentum,
1
the generalized adaptive learning rate includes a momentum-alignment term: 2 ALR-HB therefore reduces the gradient step when momentum opposes the current gradient. For moving averaged gradients,
3
ALR-MAG uses
4
Its deterministic analysis establishes monotone distance decrease and global linear convergence under smooth semi-strong convexity. Stochastic extensions include ALR-SMAG and ALR-SHB. A later IMA-based analysis proposes MomSPS5, MomDecSPS, and MomAdaSPS. MomSPS6 scales SPS by 7 and converges to a neighborhood without interpolation, while interpolation gives an 8 rate. MomDecSPS and MomAdaSPS converge to the exact minimizer in non-interpolated convex problems, with rates of order 9 and 0, respectively (Wang et al., 2023, Oikonomou et al., 2024).
Mirror-descent variants replace Euclidean geometry with a Bregman divergence. For a mirror map 1, the update is
2
Two Polyak-type rules replace 3 by adaptive target levels. The first maintains
4
and guarantees an objective value within a prescribed tolerance 5 of optimality. The second uses level reductions driven by accumulated mirror movement and satisfies
6
Both results concern function values rather than convergence of the complete iterate sequence and apply to convex locally Lipschitz objectives, including objectives that are not globally Lipschitz or smooth (You et al., 2022).
The Polyak Minorant Method generalizes the affine projection interpretation to constrained convex optimization. It constructs convex minorants 7 of the objective and constraints, defines a model-based target set, and projects onto it: 8 With one affine unconstrained objective minorant, this is exactly the classical Polyak subgradient update. With multiple stored minorants, the method becomes structurally related to cutting-plane and bundle methods. Fejér monotonicity gives asymptotic convergence of objective and constraint violations; a sharpness condition gives linear distance convergence (Devanathan et al., 2023).
5. Geometry, preconditioning, sparsity, and higher-order variants
The scalar Polyak rule adapts step length but does not by itself correct anisotropic curvature. A preconditioned extension uses
9
where 0 is a preconditioner. For a weighted norm induced by 1, the weighted projection derivation gives
2
The preconditioner therefore affects both the direction and the denominator. PSPS instantiations use Hutchinson Hessian-diagonal estimates, AdaGrad accumulations, or Adam-style first- and second-moment scaling. Slack variants PSPSL1 and PSPSL2 address non-interpolating problems. These methods are empirically useful on badly scaled classification data, but the cited work does not provide convergence rates for the proposed preconditioned algorithms or for the Hutchinson estimation error (Devanathan et al., 2023).
SP2 extends the interpolation interpretation to second-order local models. For a sampled component, let
3
The quadratic interpolation model is
4
SP2 projects onto 5. Unlike Newton's method, it seeks a root of the local quadratic rather than its minimizer, so positive-definite Hessians and convexity are not required. Negative curvature can make a positive local model cross zero and therefore facilitate the interpolation step. SP26 approximates the projection using an ordinary SP step followed by a curvature correction involving only the Hessian-vector product 7. The cleanest convergence theory is for sums of interpolated positive-semidefinite quadratics, where exact SP2 acts through null-space projections and can reach the solution immediately when a sampled component matrix is invertible (Li et al., 2022).
For high-dimensional sparse M-estimation, the full gradient can be dense even when the iterate is sparse. Sparse Polyak replaces the full-gradient denominator by a hard-thresholded gradient norm: 8 The update is
9
Under restricted strong convexity and restricted smoothness, the resulting step remains bounded below by a quantity of order 00 until statistical precision is reached. The method has linear convergence governed by the restricted condition number 01, rather than by the ambient dimension, under the stated sparsity requirement (Qiao et al., 11 Sep 2025).
6. Rates, asymptotic behavior, and applications
For smooth strongly convex functions, Polyak gradient descent has a worst-case linear rate of order
02
An analytic worst-case analysis establishes the sharp contraction factor
03
for the standard choice on the relevant squared-distance measure. This matches the worst-case contraction factor of exact line search, although exact line search contracts objective error while Polyak descent contracts squared distance. On quadratics, the method asymptotically concentrates in the two-dimensional subspace spanned by the eigenvectors associated with the smallest and largest eigenvalues. The limiting stepsize approaches 04, producing opposite-sign contraction factors in the two extreme eigendirections and an alternating zigzag trajectory (Huang et al., 2024).
For smooth convex functions, Polyak gradient descent has an 05 best-iterate rate, and explicit quadratic constructions show that this dependence is tight. Under Hölder smoothness with exponent 06, the best-iterate rate becomes
07
with matching lower bounds for the corresponding class. Hölder growth conditions produce geometry-dependent rates, including linear convergence when the growth exponent matches the smoothness structure. These results establish a form of universality: the same update adapts to unknown smoothness exponents, growth exponents, condition numbers, and curvature constants once 08 is available (He et al., 6 Dec 2025).
Statistical analyses distinguish population contraction from empirical perturbation. Under generalized smoothness, generalized Łojasiewicz conditions, and empirical-to-population gradient stability, the population Polyak operator contracts geometrically even in singular models. The empirical iterates reach a statistical radius
09
after
10
iterations, whereas fixed-step gradient descent can require polynomially many iterations when curvature degenerates at the optimum. The theory applies to generalized linear models, Gaussian mixtures, and mixed linear regression, including low-signal regimes with singular Fisher information (Ren et al., 2021).
Applications and extensions include over-parameterized classification and matrix factorization, composite learning with weight decay, nonsmooth constrained optimization, decentralized diffusion, and reinforcement learning. In policy-gradient reinforcement learning, Polyak adaptation is applied to return maximization by replacing the unknown optimal return with the higher return of two nearby policies. Entropy regularization and upper clipping are used to prevent softmax-gradient collapse and explosive updates. Experiments on Acrobot, CartPole, and LunarLander report faster convergence and more stable final policies than fixed-learning-rate Adam, but no formal convergence theorem is provided for the complete stochastic RL method (Li et al., 2024).
A decentralized graph-aware rule derives local learning rates by minimizing a network-wide one-step distance to the optimizer. The resulting coupled linear system depends on the two-hop matrix 11, neighboring objective values, and gradient correlations. It reduces to the classical Polyak rule in the single-agent limit, but the proposed decentralized stochastic recursion has no complete convergence-rate theorem (Fainman et al., 18 Sep 2025). For randomized methods solving generalized absolute value equations, the objective changes with the current iterate: 12 A Polyak ratio based on the sampled residual yields linear convergence in expectation under the stated sketching and spectral conditions (Chen et al., 2 Aug 2026).
The principal limitations of Polyak step-size methods are consistent across these settings. Exact or reliable target values are often unavailable; invalid lower bounds can cause overly aggressive or negative steps; small gradient norms create denominator instability; clipping and damping introduce parameters; stochastic objectives and gradients can destroy deterministic unbiasedness arguments; and nonconvex guarantees generally require interpolation, growth, PL, quasar-convexity, star-convexity, or other structural assumptions. Polyak adaptation therefore should not be identified with a universally parameter-free or universally accelerated method. Its central contribution is adaptive scale selection from objective-gap information, with the strongest behavior appearing when the target value is known or estimable and when the objective possesses interpolation, restricted geometry, or a suitable growth structure.