Mirror Polyak: Non-Euclidean Optimization
- Mirror Polyak extends Polyak-type optimization, utilizing Bregman projections and mirror descent in Legendre-type geometry, allowing for updated stepsizes derived from objective gaps and geometry. This typically involves a convex domain $\mathcal X$, a convex reference function $h$, and dual coordinates $\hat x=\nabla h(x)$.
- In applications such as stochastic convex feasibility and quantum optimization, these methods leverage entropy and quantum divergences to optimize non-Euclidean metrics while preserving interiority and iterated solutions.
- Popally stochastic methods adapted for stochastic and quantum optimization obtain $O(1/\sqrt{t})$ rates for smooth objectives, $O(1/ \sqrt{T})$ iteratively, and bounded stationary results, often with a favorable adaptivity rate in larger settings
“Mirror Polyak” is a name used for several related but non-identical extensions of Polyak-type optimization to non-Euclidean geometry. Its central formulation replaces Euclidean projection or squared-gradient normalization with a Bregman projection or mirror-descent update, while selecting the stepsize from an objective gap and a geometry-dependent quantity. The term also appears descriptively for inertial mirror descent, mirror stochastic Polyak stepsizes, entropic methods for quantum divergences and linear systems, stochastic block Bregman projection, and PL-based adaptive mirror methods. The most direct modern formulation defines Mirror Polyak as a Bregman projection of the current iterate onto the affine hyperplane given by the first-order model at the optimal value (Kunstner et al., 18 Aug 2026).
1. Terminology and conceptual scope
Polyak’s classical subgradient stepsize for minimizing a proper convex function is
with update
Its defining feature is adaptivity to the current objective gap rather than dependence on a prescribed smoothness, Lipschitz, or strong-convexity constant. Under suitable regularity, the rule yields for convex -smooth functions, for convex -Lipschitz functions, geometric convergence for smooth strongly convex functions, and for Lipschitz strongly convex functions (Kunstner et al., 18 Aug 2026).
Mirror Polyak transports this principle from Euclidean geometry to a Legendre-type mirror geometry. Let be a convex domain, let be a convex reference function, and define
0
The divergence 1 is nonnegative but generally asymmetric and is not a metric. Its dual coordinate is 2. Standard mirror descent updates these coordinates through
3
In the strictest and most recent usage, Mirror Polyak is the Bregman-projection rule
4
The constraint is the affine first-order model of 5 evaluated at the optimal value. Thus, the method preserves Polyak’s projection interpretation while replacing Euclidean distance by a Bregman divergence.
The phrase is not used uniformly across the literature. In particular:
- “Mirror Polyak” denotes the Bregman-projection method in “Mirror Polyak and a Primal-Dual Lifting” (Kunstner et al., 18 Aug 2026).
- It denotes mirror descent equipped with a Polyak stepsize in quantum Rényi-divergence minimization (You et al., 2021).
- It describes mirror stochastic Polyak stepsizes such as mSPS and DecmSPS (D'Orazio et al., 2021, Zhang et al., 31 Mar 2026).
- It is used conceptually for inertial mirror descent, whose Euclidean specialization yields Polyak’s heavy-ball dynamics (Nazin, 2017).
- It is not a formal optimization term in the geometric theorem on reflections of convex bodies; there, “mirror” means Euclidean reflection, and “Polyak” has no technical role (Morales-Amaya, 2024).
2. Bregman-projection formulation
For an unconstrained domain, the Bregman-projection formulation is equivalent to a mirror-descent step with an implicitly determined stepsize. Writing 6,
7
The Polyak condition becomes
8
Since 9, this condition is equivalent to
0
The stepsize is therefore generally implicit. It can be obtained by solving a one-dimensional convex root-finding or minimization problem. The relevant upper bound is
1
where 2. The final term is convex in 3 because
4
For quadratic 5, the implicit rule becomes explicit. If
6
then 7 and the method reduces to classical Polyak subgradient descent.
If 8 is 9-strongly convex with respect to a norm 0, then
1
Consequently,
2
Thus the Bregman-projection version takes at least as large a step as the corresponding norm-based mirror Polyak rule under this normalization.
For constrained optimization over 3, the update is
4
subject to
5
Under Legendre, interiority, and projection-constraint qualifications, the resulting iterate satisfies
6
This Bregman contraction is the fundamental stationarity and convergence inequality for the method. The interiority requirement is substantive: for entropy-like mirror maps, if the optimum lies on the boundary, the projection may be attained only in the limit as 7.
3. Relative convergence guarantees
The principal theoretical advantage of Mirror Polyak is that its guarantees are expressed in relative geometry rather than through a reference norm.
A function 8 is 9-smooth relative to 0 if
1
and is 2-strongly convex relative to 3 if
4
Relative Lipschitz continuity is defined by
5
These conditions reduce to the standard Euclidean notions when 6.
For convex objectives that are 7-smooth relative to 8, Mirror Polyak satisfies
9
The same bound holds for the best iterate. The proof combines Bregman contraction with the relative-smoothness inequality
0
For objectives that are both relatively smooth and relatively strongly convex,
1
This is a geometric rate expressed through the relative condition number 2. Unlike norm-based arguments, the result does not require a uniform lower bound such as 3; arbitrarily small Mirror Polyak steps can occur in the relative setting.
For convex objectives that are relatively 4-Lipschitz,
5
For relatively Lipschitz and relatively strongly convex objectives, the guarantee becomes
6
Hence Mirror Polyak preserves the usual Euclidean orders for relative smoothness and relative Lipschitzness, with a logarithmic deterioration in the relatively Lipschitz and strongly convex regime. The proof uses two complementary contraction mechanisms: contraction of Bregman distance and contraction of the objective gap (Kunstner et al., 18 Aug 2026).
4. Variants of Mirror Polyak stepsizes
Several methods use explicit or estimated objective gaps rather than the implicit Bregman-projection stepsize.
Mirror stochastic Polyak stepsize
For stochastic mirror descent,
7
the mirror stochastic Polyak stepsize is
8
The safeguarded version is
9
The Euclidean specialization recovers the stochastic Polyak stepsize for SGD or stochastic projected gradient descent. The mirror generalization replaces the Euclidean gradient norm by the dual norm and replaces Euclidean projection by Bregman projection (D'Orazio et al., 2021).
Under smoothness, the sampled objective satisfies the self-bounding inequality
0
Consequently, the adaptive stepsize has a lower bound
1
before capping. Under strong convexity it also satisfies
2
The resulting convergence results avoid bounded-gradient and bounded-variance assumptions in the smooth convex setting. Instead, the stochastic neighborhood is governed by the finite optimal objective difference
3
or, in constrained problems,
4
Under interpolation, the relevant quantity vanishes. In the relatively strongly convex case, mSPS with a cap gives
5
where 6. For smooth convex objectives, the averaged iterate satisfies an 7 bound to a stochastic neighborhood, which vanishes under interpolation (D'Orazio et al., 2021).
Target-estimation rules without 8
The two Polyak-type stepsizes of (You et al., 2022) remove the requirement that the exact optimal value be known. Both use
9
but construct 0 differently.
The first uses
1
where 2 is increased after a successful crossing of the estimated level and otherwise decreases toward a prescribed floor 3. Its guarantee is
4
The second is an adaptive level method. It maintains levels 5, a running minimum, and a travel budget 6. The level is halved when the cumulative normalized travel exceeds 7. It guarantees
8
These are asymptotic best-objective guarantees. They do not establish convergence of the full iterate sequence or finite-time complexity rates. Their assumptions include a Legendre mirror map, strong convexity of the mirror map, domain compatibility, and local boundedness of subgradients (You et al., 2022).
Entropic Mirror Polyak
For the negative entropy mirror map
9
the update on the positive orthant is
0
For nonnegative linear systems with
1
the proposed stepsize is
2
where
3
The safeguard ensures 4, permitting the exponential remainder bound 5 for 6. The method achieves an 7 best-iterate rate and convergence of the full sequence under the stated assumptions (Malitsky et al., 5 May 2025).
For an unrestricted linear system, the positive and negative parts are represented as 8, and the EG9 method uses multiplicative updates for 0 and 1. Its invariant
2
supports a linear-rate result under strict positivity of the limiting lifted solution. The method also establishes an entropy-Bregman implicit bias: the limit is the Bregman projection of the initialization onto the solution set.
5. Inertia, PL geometry, and other interpretations
Inertial mirror descent
In “Algorithms of Inertial Mirror Descent in Convex Problems of Stochastic Optimization,” inertial mirror descent modifies the continuous-time mirror relation from
3
to
4
with accumulated-gradient dynamics
5
Setting 6 recovers continuous-time mirror descent. Choosing 7 yields the pointwise continuous-time bound
8
For the Euclidean potential
9
constant inertia 00 gives
01
which is Polyak’s continuous-time heavy-ball equation. In this usage, Mirror Polyak means a mirror-geometric generalization of inertial or heavy-ball dynamics. The discrete stochastic theorem, however, gives an expected 02 rate for nonsmooth stochastic convex optimization, not a Nesterov-style accelerated rate (Nazin, 2017).
Generalized mirror descent under PL
Generalized mirror descent uses an invertible, possibly nonlinear and time-dependent map 03:
04
When 05, this is standard mirror descent; when 06 is linear, it becomes preconditioned gradient descent; Adagrad is obtained through a time-dependent diagonal map.
If 07 is 08-smooth and satisfies the Polyak–Łojasiewicz inequality
09
strong monotonicity and Lipschitzness of 10 yield linear convergence with generalized condition number
11
The analysis does not require convexity of 12. For stochastic nonlinear mirrors, Taylor expansions of 13 yield expected linear convergence under analyticity, derivative bounds, PL14, interpolation, and sufficiently small adaptive stepsizes (Radhakrishnan et al., 2020).
This is conceptually distinct from Bregman-projection Mirror Polyak. Here, “Polyak” refers to PL-based linear-convergence analysis, whereas the stepsize is not necessarily selected by an objective-gap-over-gradient rule.
Bilevel optimization
Adaptive mirror-descent methods for nonconvex bilevel optimization use a lower-level PL condition to control the lower-level residual and the hypergradient error. If
15
then lower-level objective error is controlled by the squared lower-level gradient. This residual controls tracking error, which in turn controls the hypergradient approximation error.
AdaPAG and AdaVSPAG use adaptive quadratic Bregman geometries, proximal upper-level steps, lower-level adaptive steps, and clipped Hessian and cross-Hessian quantities. The deterministic method obtains an average squared stationarity rate 16 and gradient complexity 17. The variance-reduced stochastic method obtains oracle complexity 18 for an 19-stationary solution (Huang, 2023).
These methods are best regarded as another conceptual use of “Mirror Polyak”: mirror geometry supplies adaptive proximal updates, while the PL condition supplies error control. The paper does not formally name a separate method Mirror Polyak.
6. Applications, comparisons, and limitations
Mirror Polyak is applicable when the geometry of the feasible set or objective is poorly represented by Euclidean distance. Entropy geometry is suited to probability simplices, positive variables, and quantum states; matrix geometries include quantum relative entropy and log-determinant structures; quadratic adaptive mirrors support preconditioning and coordinate scaling.
In quantum optimization, entropic mirror descent with an adjusted Polyak stepsize is used to minimize Petz and sandwiched Rényi divergences, Petz–Augustin information, sandwiched Augustin information, conditional sandwiched Rényi entropy, and sandwiched Rényi information. The update preserves full-rank density matrices:
20
The method avoids the boundary singularities that affect Euclidean projected methods. Its formal guarantee is an approximate function-value result,
21
rather than an explicit convergence rate for every iterate (You et al., 2021).
In stochastic convex feasibility, stochastic block Bregman projection combines block sampling with a decreasing mirror stochastic Polyak stepsize. For possibly inconsistent systems, the inner objective minimizes aggregate constraint violation, and the Polyak-like rule uses a lower bound on each sampled block optimum. Under a Bregman distance growth condition, the method obtains ergodic 22 convergence under interpolation and 23 convergence in the non-interpolated case. It also obtains linear convergence in expectation to the inner minimizer set under the growth condition (Zhang et al., 31 Mar 2026).
The main limitations are structural:
- Knowledge of the optimum: the exact Bregman-projection rule requires 24. Explicit stochastic rules require sampled optimal values or lower bounds. Level-estimation methods avoid exact knowledge but provide weaker or asymptotic guarantees.
- Interiority: Legendre maps, entropy maps, and logarithmic coordinates may require iterates and comparison solutions to lie in the interior.
- Mirror-map regularity: convergence depends on strong convexity, relative smoothness, local boundedness, analyticity, derivative bounds, or Bregman growth conditions, depending on the variant.
- Stochastic noise: mSPS results typically converge to a neighborhood determined by 25 or 26, unless interpolation eliminates the term.
- No universal rate: the different methods have guarantees ranging from exact relative geometric rates to best-iterate sublinear bounds and asymptotic infimum results.
- Iterate versus objective convergence: several results control only the best objective value or an averaged iterate, rather than convergence of the entire sequence.
- Geometry-dependent conditioning: the mirror map affects stepsizes, Bregman distances, dual norms, implicit bias, and the effective condition number.
- Terminological ambiguity: “Mirror Polyak” may denote a Bregman-projection stepsize, a norm-based mirror stochastic Polyak rule, inertial mirror descent, or a PL-based mirror method. It does not refer to the geometric “mirror” in the theorem on reflections of convex bodies (Morales-Amaya, 2024).
The common principle across the optimization uses is the transport of Polyak’s objective-gap adaptation into a non-Euclidean update geometry. In the Bregman-projection formulation, this transport is exact: the affine first-order model reaches the optimal value, while the update minimizes Bregman displacement. In explicit variants, the same principle is approximated through dual norms, sampled objective gaps, lower-level residuals, adaptive levels, or primal-dual lifting. The resulting methods retain Polyak-type adaptivity while exploiting feasibility preservation, relative regularity, preconditioning, entropy geometry, and structural implicit bias.