Papers
Topics
Authors
Recent
Search
2000 character limit reached

Mirror Polyak: Non-Euclidean Optimization

Updated 23 August 2026
  • Mirror Polyak extends Polyak-type optimization, utilizing Bregman projections and mirror descent in Legendre-type geometry, allowing for updated stepsizes derived from objective gaps and geometry. This typically involves a convex domain $\mathcal X$, a convex reference function $h$, and dual coordinates $\hat x=\nabla h(x)$.
  • In applications such as stochastic convex feasibility and quantum optimization, these methods leverage entropy and quantum divergences to optimize non-Euclidean metrics while preserving interiority and iterated solutions.
  • Popally stochastic methods adapted for stochastic and quantum optimization obtain $O(1/\sqrt{t})$ rates for smooth objectives, $O(1/ \sqrt{T})$ iteratively, and bounded stationary results, often with a favorable adaptivity rate in larger settings

“Mirror Polyak” is a name used for several related but non-identical extensions of Polyak-type optimization to non-Euclidean geometry. Its central formulation replaces Euclidean projection or squared-gradient normalization with a Bregman projection or mirror-descent update, while selecting the stepsize from an objective gap and a geometry-dependent quantity. The term also appears descriptively for inertial mirror descent, mirror stochastic Polyak stepsizes, entropic methods for quantum divergences and linear systems, stochastic block Bregman projection, and PL-based adaptive mirror methods. The most direct modern formulation defines Mirror Polyak as a Bregman projection of the current iterate onto the affine hyperplane given by the first-order model at the optimal value (Kunstner et al., 18 Aug 2026).

1. Terminology and conceptual scope

Polyak’s classical subgradient stepsize for minimizing a proper convex function ff is

αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),

with update

xt+1=xt−αtgt.x_{t+1}=x_t-\alpha_tg_t.

Its defining feature is adaptivity to the current objective gap rather than dependence on a prescribed smoothness, Lipschitz, or strong-convexity constant. Under suitable regularity, the rule yields O(L/T)O(L/T) for convex LL-smooth functions, O(G/T)O(G/\sqrt{T}) for convex GG-Lipschitz functions, geometric convergence for smooth strongly convex functions, and O(G2/(μT))O(G^2/(\mu T)) for Lipschitz strongly convex functions (Kunstner et al., 18 Aug 2026).

Mirror Polyak transports this principle from Euclidean geometry to a Legendre-type mirror geometry. Let X\mathcal X be a convex domain, let hh be a convex reference function, and define

αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),0

The divergence αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),1 is nonnegative but generally asymmetric and is not a metric. Its dual coordinate is αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),2. Standard mirror descent updates these coordinates through

αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),3

In the strictest and most recent usage, Mirror Polyak is the Bregman-projection rule

αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),4

The constraint is the affine first-order model of αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),5 evaluated at the optimal value. Thus, the method preserves Polyak’s projection interpretation while replacing Euclidean distance by a Bregman divergence.

The phrase is not used uniformly across the literature. In particular:

  • “Mirror Polyak” denotes the Bregman-projection method in “Mirror Polyak and a Primal-Dual Lifting” (Kunstner et al., 18 Aug 2026).
  • It denotes mirror descent equipped with a Polyak stepsize in quantum Rényi-divergence minimization (You et al., 2021).
  • It describes mirror stochastic Polyak stepsizes such as mSPS and DecmSPS (D'Orazio et al., 2021, Zhang et al., 31 Mar 2026).
  • It is used conceptually for inertial mirror descent, whose Euclidean specialization yields Polyak’s heavy-ball dynamics (Nazin, 2017).
  • It is not a formal optimization term in the geometric theorem on reflections of convex bodies; there, “mirror” means Euclidean reflection, and “Polyak” has no technical role (Morales-Amaya, 2024).

2. Bregman-projection formulation

For an unconstrained domain, the Bregman-projection formulation is equivalent to a mirror-descent step with an implicitly determined stepsize. Writing αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),6,

αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),7

The Polyak condition becomes

αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),8

Since αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),9, this condition is equivalent to

xt+1=xt−αtgt.x_{t+1}=x_t-\alpha_tg_t.0

The stepsize is therefore generally implicit. It can be obtained by solving a one-dimensional convex root-finding or minimization problem. The relevant upper bound is

xt+1=xt−αtgt.x_{t+1}=x_t-\alpha_tg_t.1

where xt+1=xt−αtgt.x_{t+1}=x_t-\alpha_tg_t.2. The final term is convex in xt+1=xt−αtgt.x_{t+1}=x_t-\alpha_tg_t.3 because

xt+1=xt−αtgt.x_{t+1}=x_t-\alpha_tg_t.4

For quadratic xt+1=xt−αtgt.x_{t+1}=x_t-\alpha_tg_t.5, the implicit rule becomes explicit. If

xt+1=xt−αtgt.x_{t+1}=x_t-\alpha_tg_t.6

then xt+1=xt−αtgt.x_{t+1}=x_t-\alpha_tg_t.7 and the method reduces to classical Polyak subgradient descent.

If xt+1=xt−αtgt.x_{t+1}=x_t-\alpha_tg_t.8 is xt+1=xt−αtgt.x_{t+1}=x_t-\alpha_tg_t.9-strongly convex with respect to a norm O(L/T)O(L/T)0, then

O(L/T)O(L/T)1

Consequently,

O(L/T)O(L/T)2

Thus the Bregman-projection version takes at least as large a step as the corresponding norm-based mirror Polyak rule under this normalization.

For constrained optimization over O(L/T)O(L/T)3, the update is

O(L/T)O(L/T)4

subject to

O(L/T)O(L/T)5

Under Legendre, interiority, and projection-constraint qualifications, the resulting iterate satisfies

O(L/T)O(L/T)6

This Bregman contraction is the fundamental stationarity and convergence inequality for the method. The interiority requirement is substantive: for entropy-like mirror maps, if the optimum lies on the boundary, the projection may be attained only in the limit as O(L/T)O(L/T)7.

3. Relative convergence guarantees

The principal theoretical advantage of Mirror Polyak is that its guarantees are expressed in relative geometry rather than through a reference norm.

A function O(L/T)O(L/T)8 is O(L/T)O(L/T)9-smooth relative to LL0 if

LL1

and is LL2-strongly convex relative to LL3 if

LL4

Relative Lipschitz continuity is defined by

LL5

These conditions reduce to the standard Euclidean notions when LL6.

For convex objectives that are LL7-smooth relative to LL8, Mirror Polyak satisfies

LL9

The same bound holds for the best iterate. The proof combines Bregman contraction with the relative-smoothness inequality

O(G/T)O(G/\sqrt{T})0

For objectives that are both relatively smooth and relatively strongly convex,

O(G/T)O(G/\sqrt{T})1

This is a geometric rate expressed through the relative condition number O(G/T)O(G/\sqrt{T})2. Unlike norm-based arguments, the result does not require a uniform lower bound such as O(G/T)O(G/\sqrt{T})3; arbitrarily small Mirror Polyak steps can occur in the relative setting.

For convex objectives that are relatively O(G/T)O(G/\sqrt{T})4-Lipschitz,

O(G/T)O(G/\sqrt{T})5

For relatively Lipschitz and relatively strongly convex objectives, the guarantee becomes

O(G/T)O(G/\sqrt{T})6

Hence Mirror Polyak preserves the usual Euclidean orders for relative smoothness and relative Lipschitzness, with a logarithmic deterioration in the relatively Lipschitz and strongly convex regime. The proof uses two complementary contraction mechanisms: contraction of Bregman distance and contraction of the objective gap (Kunstner et al., 18 Aug 2026).

4. Variants of Mirror Polyak stepsizes

Several methods use explicit or estimated objective gaps rather than the implicit Bregman-projection stepsize.

Mirror stochastic Polyak stepsize

For stochastic mirror descent,

O(G/T)O(G/\sqrt{T})7

the mirror stochastic Polyak stepsize is

O(G/T)O(G/\sqrt{T})8

The safeguarded version is

O(G/T)O(G/\sqrt{T})9

The Euclidean specialization recovers the stochastic Polyak stepsize for SGD or stochastic projected gradient descent. The mirror generalization replaces the Euclidean gradient norm by the dual norm and replaces Euclidean projection by Bregman projection (D'Orazio et al., 2021).

Under smoothness, the sampled objective satisfies the self-bounding inequality

GG0

Consequently, the adaptive stepsize has a lower bound

GG1

before capping. Under strong convexity it also satisfies

GG2

The resulting convergence results avoid bounded-gradient and bounded-variance assumptions in the smooth convex setting. Instead, the stochastic neighborhood is governed by the finite optimal objective difference

GG3

or, in constrained problems,

GG4

Under interpolation, the relevant quantity vanishes. In the relatively strongly convex case, mSPS with a cap gives

GG5

where GG6. For smooth convex objectives, the averaged iterate satisfies an GG7 bound to a stochastic neighborhood, which vanishes under interpolation (D'Orazio et al., 2021).

Target-estimation rules without GG8

The two Polyak-type stepsizes of (You et al., 2022) remove the requirement that the exact optimal value be known. Both use

GG9

but construct O(G2/(μT))O(G^2/(\mu T))0 differently.

The first uses

O(G2/(μT))O(G^2/(\mu T))1

where O(G2/(μT))O(G^2/(\mu T))2 is increased after a successful crossing of the estimated level and otherwise decreases toward a prescribed floor O(G2/(μT))O(G^2/(\mu T))3. Its guarantee is

O(G2/(μT))O(G^2/(\mu T))4

The second is an adaptive level method. It maintains levels O(G2/(μT))O(G^2/(\mu T))5, a running minimum, and a travel budget O(G2/(μT))O(G^2/(\mu T))6. The level is halved when the cumulative normalized travel exceeds O(G2/(μT))O(G^2/(\mu T))7. It guarantees

O(G2/(μT))O(G^2/(\mu T))8

These are asymptotic best-objective guarantees. They do not establish convergence of the full iterate sequence or finite-time complexity rates. Their assumptions include a Legendre mirror map, strong convexity of the mirror map, domain compatibility, and local boundedness of subgradients (You et al., 2022).

Entropic Mirror Polyak

For the negative entropy mirror map

O(G2/(μT))O(G^2/(\mu T))9

the update on the positive orthant is

X\mathcal X0

For nonnegative linear systems with

X\mathcal X1

the proposed stepsize is

X\mathcal X2

where

X\mathcal X3

The safeguard ensures X\mathcal X4, permitting the exponential remainder bound X\mathcal X5 for X\mathcal X6. The method achieves an X\mathcal X7 best-iterate rate and convergence of the full sequence under the stated assumptions (Malitsky et al., 5 May 2025).

For an unrestricted linear system, the positive and negative parts are represented as X\mathcal X8, and the EGX\mathcal X9 method uses multiplicative updates for hh0 and hh1. Its invariant

hh2

supports a linear-rate result under strict positivity of the limiting lifted solution. The method also establishes an entropy-Bregman implicit bias: the limit is the Bregman projection of the initialization onto the solution set.

5. Inertia, PL geometry, and other interpretations

Inertial mirror descent

In “Algorithms of Inertial Mirror Descent in Convex Problems of Stochastic Optimization,” inertial mirror descent modifies the continuous-time mirror relation from

hh3

to

hh4

with accumulated-gradient dynamics

hh5

Setting hh6 recovers continuous-time mirror descent. Choosing hh7 yields the pointwise continuous-time bound

hh8

For the Euclidean potential

hh9

constant inertia αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),00 gives

αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),01

which is Polyak’s continuous-time heavy-ball equation. In this usage, Mirror Polyak means a mirror-geometric generalization of inertial or heavy-ball dynamics. The discrete stochastic theorem, however, gives an expected αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),02 rate for nonsmooth stochastic convex optimization, not a Nesterov-style accelerated rate (Nazin, 2017).

Generalized mirror descent under PL

Generalized mirror descent uses an invertible, possibly nonlinear and time-dependent map αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),03:

αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),04

When αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),05, this is standard mirror descent; when αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),06 is linear, it becomes preconditioned gradient descent; Adagrad is obtained through a time-dependent diagonal map.

If αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),07 is αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),08-smooth and satisfies the Polyak–Łojasiewicz inequality

αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),09

strong monotonicity and Lipschitzness of αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),10 yield linear convergence with generalized condition number

αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),11

The analysis does not require convexity of αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),12. For stochastic nonlinear mirrors, Taylor expansions of αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),13 yield expected linear convergence under analyticity, derivative bounds, PLαt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),14, interpolation, and sufficiently small adaptive stepsizes (Radhakrishnan et al., 2020).

This is conceptually distinct from Bregman-projection Mirror Polyak. Here, “Polyak” refers to PL-based linear-convergence analysis, whereas the stepsize is not necessarily selected by an objective-gap-over-gradient rule.

Bilevel optimization

Adaptive mirror-descent methods for nonconvex bilevel optimization use a lower-level PL condition to control the lower-level residual and the hypergradient error. If

αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),15

then lower-level objective error is controlled by the squared lower-level gradient. This residual controls tracking error, which in turn controls the hypergradient approximation error.

AdaPAG and AdaVSPAG use adaptive quadratic Bregman geometries, proximal upper-level steps, lower-level adaptive steps, and clipped Hessian and cross-Hessian quantities. The deterministic method obtains an average squared stationarity rate αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),16 and gradient complexity αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),17. The variance-reduced stochastic method obtains oracle complexity αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),18 for an αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),19-stationary solution (Huang, 2023).

These methods are best regarded as another conceptual use of “Mirror Polyak”: mirror geometry supplies adaptive proximal updates, while the PL condition supplies error control. The paper does not formally name a separate method Mirror Polyak.

6. Applications, comparisons, and limitations

Mirror Polyak is applicable when the geometry of the feasible set or objective is poorly represented by Euclidean distance. Entropy geometry is suited to probability simplices, positive variables, and quantum states; matrix geometries include quantum relative entropy and log-determinant structures; quadratic adaptive mirrors support preconditioning and coordinate scaling.

In quantum optimization, entropic mirror descent with an adjusted Polyak stepsize is used to minimize Petz and sandwiched Rényi divergences, Petz–Augustin information, sandwiched Augustin information, conditional sandwiched Rényi entropy, and sandwiched Rényi information. The update preserves full-rank density matrices:

αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),20

The method avoids the boundary singularities that affect Euclidean projected methods. Its formal guarantee is an approximate function-value result,

αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),21

rather than an explicit convergence rate for every iterate (You et al., 2021).

In stochastic convex feasibility, stochastic block Bregman projection combines block sampling with a decreasing mirror stochastic Polyak stepsize. For possibly inconsistent systems, the inner objective minimizes aggregate constraint violation, and the Polyak-like rule uses a lower bound on each sampled block optimum. Under a Bregman distance growth condition, the method obtains ergodic αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),22 convergence under interpolation and αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),23 convergence in the non-interpolated case. It also obtains linear convergence in expectation to the inner minimizer set under the growth condition (Zhang et al., 31 Mar 2026).

The main limitations are structural:

  • Knowledge of the optimum: the exact Bregman-projection rule requires αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),24. Explicit stochastic rules require sampled optimal values or lower bounds. Level-estimation methods avoid exact knowledge but provide weaker or asymptotic guarantees.
  • Interiority: Legendre maps, entropy maps, and logarithmic coordinates may require iterates and comparison solutions to lie in the interior.
  • Mirror-map regularity: convergence depends on strong convexity, relative smoothness, local boundedness, analyticity, derivative bounds, or Bregman growth conditions, depending on the variant.
  • Stochastic noise: mSPS results typically converge to a neighborhood determined by αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),25 or αt=f(xt)−f⋆∥gt∥22,gt∈∂f(xt),\alpha_t=\frac{f(x_t)-f_\star}{\|g_t\|_2^2}, \qquad g_t\in\partial f(x_t),26, unless interpolation eliminates the term.
  • No universal rate: the different methods have guarantees ranging from exact relative geometric rates to best-iterate sublinear bounds and asymptotic infimum results.
  • Iterate versus objective convergence: several results control only the best objective value or an averaged iterate, rather than convergence of the entire sequence.
  • Geometry-dependent conditioning: the mirror map affects stepsizes, Bregman distances, dual norms, implicit bias, and the effective condition number.
  • Terminological ambiguity: “Mirror Polyak” may denote a Bregman-projection stepsize, a norm-based mirror stochastic Polyak rule, inertial mirror descent, or a PL-based mirror method. It does not refer to the geometric “mirror” in the theorem on reflections of convex bodies (Morales-Amaya, 2024).

The common principle across the optimization uses is the transport of Polyak’s objective-gap adaptation into a non-Euclidean update geometry. In the Bregman-projection formulation, this transport is exact: the affine first-order model reaches the optimal value, while the update minimizes Bregman displacement. In explicit variants, the same principle is approximated through dual norms, sampled objective gaps, lower-level residuals, adaptive levels, or primal-dual lifting. The resulting methods retain Polyak-type adaptivity while exploiting feasibility preservation, relative regularity, preconditioning, entropy geometry, and structural implicit bias.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Mirror Polyak.