Papers
Topics
Authors
Recent
Search
2000 character limit reached

Mirror Descent with Bregman Divergence

Updated 1 January 2026
  • Mirror descent with Bregman divergence is a geometry-aware optimization framework that leverages strictly convex mirror maps to respect problem structures like sparsity and manifold constraints.
  • The method utilizes dual gradient steps and Bregman projections to ensure robust convergence, achieving sublinear to linear rates depending on convexity of the objective.
  • It underpins a wide range of applications, from machine learning and statistics to control and reinforcement learning, and supports distributed and primal-dual optimization settings.

Mirror descent with Bregman divergence is a general first-order optimization framework that extends classical gradient descent by leveraging non-Euclidean geometries, as defined by strictly convex "mirror maps." The essence of mirror descent is the utilization of Bregman divergence—generated by a mirror map—to measure proximity and dictate update steps, enabling algorithms to respect problem structure such as sparsity, simplex constraints, and manifold geometry. This approach yields robust convergence guarantees over a wide family of domains, supports primal-dual and distributed settings, and underpins a vast array of applications in statistics, machine learning, control, and reinforcement learning.

1. Definition and General Framework

Let h:dom(h)→Rh: \mathsf{dom}(h) \to \mathbb{R} be a strictly convex, differentiable function with open domain, referred to as the "mirror map." Given an optimization problem min⁡x∈Xf(x)\min_{x \in X} f(x) over a closed convex set X⊂RnX \subset \mathbb{R}^n, the Bregman divergence associated with hh is defined as

Dh(p,q)=h(p)−h(q)−⟨∇h(q),p−q⟩D_h(p, q) = h(p) - h(q) - \langle \nabla h(q), p - q \rangle

where p,q∈X∩dom(h)p, q \in X \cap \mathrm{dom}(h). DhD_h is nonnegative, convex in its first argument, and recovers squared Euclidean distance for h(x)=12∥x∥22h(x) = \frac{1}{2}\|x\|_2^2. The mirror descent update is specified by:

  • Dual step: ∇h(yt+1)=∇h(xt)−ηt∇f(xt)\nabla h(y^{t+1}) = \nabla h(x^t) - \eta_t \nabla f(x^t)
  • Primal projection: xt+1=arg⁡min⁡x∈X{Dh(x,yt+1)}x^{t+1} = \arg\min_{x \in X}\{ D_h(x, y^{t+1}) \} or, equivalently,

min⁡x∈Xf(x)\min_{x \in X} f(x)0

This framework respects the geometry of min⁡x∈Xf(x)\min_{x \in X} f(x)1 by encoding it into the choice of min⁡x∈Xf(x)\min_{x \in X} f(x)2, and supports a variety of optimization landscapes (Raskutti et al., 2013).

2. Specialization to Probability Simplex: Entropic Mirror Descent

On the probability simplex min⁡x∈Xf(x)\min_{x \in X} f(x)3, the canonical mirror map is the negative entropy min⁡x∈Xf(x)\min_{x \in X} f(x)4. For this choice,

min⁡x∈Xf(x)\min_{x \in X} f(x)5

which is simply the Kullback-Leibler divergence when min⁡x∈Xf(x)\min_{x \in X} f(x)6. The Bregman projection onto the simplex consists of normalization: min⁡x∈Xf(x)\min_{x \in X} f(x)7 The mirror descent update becomes: min⁡x∈Xf(x)\min_{x \in X} f(x)8 This entropic form arises naturally in learning probability distributions, portfolio optimization, and boosting (Halder, 2018).

3. Variational Principle and Fixed Point Structure

Mirror descent is variationally equivalent to minimizing a composite objective of the form: min⁡x∈Xf(x)\min_{x \in X} f(x)9 where X⊂RnX \subset \mathbb{R}^n0 is the Kullback-Leibler divergence to a reference vector X⊂RnX \subset \mathbb{R}^n1 ("influence"), and X⊂RnX \subset \mathbb{R}^n2 is the "extropy" (entropy of the complement). The fixed point X⊂RnX \subset \mathbb{R}^n3 of mirror descent (e.g., the DeGroot–Friedkin map) solves

X⊂RnX \subset \mathbb{R}^n4

Strict convexity ensures existence and uniqueness of X⊂RnX \subset \mathbb{R}^n5, and standard Lyapunov arguments establish convergence (Halder, 2018).

Mirror descent can be viewed as gradient descent in the dual Riemannian geometry, with the metric tensor given by the Hessian X⊂RnX \subset \mathbb{R}^n6. The Legendre transform X⊂RnX \subset \mathbb{R}^n7 defines the dual geometry, and the update in dual coordinates corresponds to natural gradient descent: X⊂RnX \subset \mathbb{R}^n8 which, via the chain rule, becomes steepest descent on the dual manifold (Raskutti et al., 2013). In exponential families, mirror descent with negative entropy achieves asymptotic statistical efficiency, attaining the Cramér–Rao lower bound for parameter estimation.

5. Convergence Analysis and Lyapunov Perspective

Mirror descent admits rigorous convergence guarantees:

  • Sublinear X⊂RnX \subset \mathbb{R}^n9 convergence for convex objectives with constant step size
  • Linear (geometric) rate hh0 for strongly convex objectives, where hh1 depends on the strong convexity of hh2 and hh3
  • For mirror descent steps, the Bregman divergence hh4 serves as a Lyapunov function
  • Integral Quadratic Constraint (IQC) analyses show that the Bregman Lyapunov is a special case of Popov-criterion storage functions, enabling tight rates via matrix inequalities (Li et al., 2022, Li et al., 2023).

6. Practical Applications and Specialized Algorithms

Mirror descent with Bregman divergence is foundational in:

  • Composite, distributed, and online optimization, where geometry-aware updates outperform Euclidean approaches (Yuan et al., 2020, Chen et al., 2021)
  • Policy optimization in reinforcement learning, where PMD-style updates guarantee finite-step optimality and adapt to geometry-inducing divergences (Lin et al., 2022)
  • Stochastic control, both with vector-valued and measure-valued actions. Relative smoothness and strong convexity with respect to hh5 provide linear or exponential rates, depending on regularization (Sethi et al., 3 Jun 2025, Kerimkulov et al., 2024)
  • Implicit regularization in separable data: choice of the mirror map directly affects margin bounds and learning behavior (Li et al., 2021)
  • Optimization over curved manifolds and norm-constrained sets: dual-norm mirror descent and generalized logarithmic mirrors extend the method to non-Euclidean settings, often yielding closed-form projection-free updates (Nock et al., 2016, Cichocki, 8 Jun 2025)
  • Statistical learning in exponential families, phase retrieval, optimal transport (Sinkhorn algorithm as a mirror descent with KL divergence) (Godeme et al., 2022, 2002.03758)

7. Algorithmic Templates and Implementation Considerations

The prototypical mirror descent algorithm is:

Dh(p,q)=h(p)−h(q)−⟨∇h(q),p−q⟩D_h(p, q) = h(p) - h(q) - \langle \nabla h(q), p - q \rangle8

  • The efficiency of mirror descent depends on the choice of hh6 and the tractability of inverting hh7.
  • For simplex domains, exponentiated-gradient and its generalizations (Tempesta logarithms, Tsallis/Kaniadakis mirrors) provide closed-form updates and enable domain adaptation via hyperparameters (Cichocki, 8 Jun 2025).
  • For distributed and non-smooth optimization, Bregman damping and ergodic gap analysis yield hh8 rates (for saddle point/constrained problems) (Chen et al., 2021).

8. Theoretical and Empirical Insights

Mirror descent unifies proximal, primal-dual, and natural gradient frameworks. The geometry is entirely governed by the mirror map and its Bregman divergence, providing both interpretability and flexibility. Analysis via IQC and Lyapunov methods confirms the tightness of classical rates and allows systematic extension to advanced settings—stochastic, distributed, measure-valued, nonconvex—maintaining robust guarantees (Li et al., 2023, Fatkhullin et al., 2024). Proper tuning of the underlying mirror geometry yields optimal statistical efficiency and domain-adaptive regularization.

Summary Table: Core Components

Component Definition / Role Classical Case
Mirror Map hh9 Strictly convex, differentiable potential encoding geometry of Dh(p,q)=h(p)−h(q)−⟨∇h(q),p−q⟩D_h(p, q) = h(p) - h(q) - \langle \nabla h(q), p - q \rangle0 Dh(p,q)=h(p)−h(q)−⟨∇h(q),p−q⟩D_h(p, q) = h(p) - h(q) - \langle \nabla h(q), p - q \rangle1
Bregman Divergence Dh(p,q)=h(p)−h(q)−⟨∇h(q),p−q⟩D_h(p, q) = h(p) - h(q) - \langle \nabla h(q), p - q \rangle2 Squared Euclidean distance
Dual Step Dh(p,q)=h(p)−h(q)−⟨∇h(q),p−q⟩D_h(p, q) = h(p) - h(q) - \langle \nabla h(q), p - q \rangle3 Additive (Euclidean) update
Primal Projection Dh(p,q)=h(p)−h(q)−⟨∇h(q),p−q⟩D_h(p, q) = h(p) - h(q) - \langle \nabla h(q), p - q \rangle4 Standard Euclidean projection
Typical Geometry Simplex (entropy), orthant, manifold, dual-norm, measure space Dh(p,q)=h(p)−h(q)−⟨∇h(q),p−q⟩D_h(p, q) = h(p) - h(q) - \langle \nabla h(q), p - q \rangle5
Convergence Rate Dh(p,q)=h(p)−h(q)−⟨∇h(q),p−q⟩D_h(p, q) = h(p) - h(q) - \langle \nabla h(q), p - q \rangle6 (convex); Dh(p,q)=h(p)−h(q)−⟨∇h(q),p−q⟩D_h(p, q) = h(p) - h(q) - \langle \nabla h(q), p - q \rangle7 (strongly convex); exponential (strong regularizer) Same under Euclidean geometry

The generality, geometry-awareness, and provable efficiency of mirror descent with Bregman divergence position it as a cornerstone method in modern convex, stochastic, distributed, and nonconvex optimization.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Mirror Descent with Bregman Divergence.