Papers
Topics
Authors
Recent
Search
2000 character limit reached

Rényi Divergence Gradient: Theory & Applications

Updated 1 December 2025
  • Rényi divergence gradients are explicit derivative formulations quantifying differences between probability distributions using tilted and weighted expectations.
  • They underpin optimization techniques across classical statistics, variational inference, quantum information, and control through structured gradient flows and geometric insights.
  • Their advanced formulation extends conventional divergences like the KL divergence, offering robust convergence properties in applications from deep learning to risk-sensitive policy updates.

Rényi divergence gradients underpin a family of information-theoretic optimization methods used in classical statistics, information geometry, quantum information theory, variational inference, machine learning, and control. The explicit forms, computational characteristics, and geometric structures of these gradients are central to theoretical analysis and algorithmic development across these fields.

1. Definition and General Formulas

The Rényi divergence of order α1\alpha\neq 1, between probability measures pp and qq on a common domain, is

Dα(pq)=1α1log(p(x)q(x))αq(x)dxD_\alpha(p \| q) = \frac{1}{\alpha - 1} \log \int \left( \frac{p(x)}{q(x)} \right)^\alpha q(x) dx

In the limit α1\alpha\to 1, this reduces to the Kullback-Leibler divergence. The gradient of Dα(pθq)D_\alpha(p_\theta\|q) with respect to parameters θ\theta of a parametric distribution pθ(x)p_\theta(x) takes the generic form

θDα(pθq)=αα1Erθ[θlogpθ(x)]\nabla_\theta D_\alpha(p_\theta\|q) = \frac{\alpha}{\alpha-1} \mathbb{E}_{r_\theta}\left[\nabla_\theta \log p_\theta(x)\right]

where rθ(x)r_\theta(x) is the "tilted" distribution proportional to pp0, i.e.,

pp1

An equivalent representation brings the expectation under pp2 with non-linear weights

pp3

so that

pp4

These formulas provide the foundation for Rényi-based optimization in both continuous and discrete domains (Ito et al., 2024).

2. Discrete and Geometric Interpretations

For discrete probability vectors pp5, the gradient with respect to pp6 is

pp7

where pp8 (Wong, 2017). In vector notation, this is

pp9

This view aligns with the geometry of statistical manifolds: Rényi divergence gradients are associated with canonical divergences on dually flat or constant curvature manifolds, connecting to optimal transport and Bregman geometry (Wong, 2017). Flows induced by these gradients follow geodesics defined by the underlying Riemannian metric.

3. Fokker-Planck Equations and Gradient-Flow Structure

On qq0, for the Fokker-Planck PDE with drift qq1 and strictly convex potential qq2, the time-derivative (gradient flow) of Rényi divergence qq3 along the solution qq4 has the form

qq5

where qq6, qq7, and qq8 is the "escort" distribution. qq9 is the relative Dα(pq)=1α1log(p(x)q(x))αq(x)dxD_\alpha(p \| q) = \frac{1}{\alpha - 1} \log \int \left( \frac{p(x)}{q(x)} \right)^\alpha q(x) dx0-Fisher information. This framework enables exponential decay rates for the Rényi divergence, extending log-Sobolev inequalities for Dα(pq)=1α1log(p(x)q(x))αq(x)dxD_\alpha(p \| q) = \frac{1}{\alpha - 1} \log \int \left( \frac{p(x)}{q(x)} \right)^\alpha q(x) dx1 to arbitrary Dα(pq)=1α1log(p(x)q(x))αq(x)dxD_\alpha(p \| q) = \frac{1}{\alpha - 1} \log \int \left( \frac{p(x)}{q(x)} \right)^\alpha q(x) dx2 (Cao et al., 2018).

4. Quantum Generalizations: Sandwiched Rényi Gradients

For quantum states (faithful density operators) Dα(pq)=1α1log(p(x)q(x))αq(x)dxD_\alpha(p \| q) = \frac{1}{\alpha - 1} \log \int \left( \frac{p(x)}{q(x)} \right)^\alpha q(x) dx3 and Dα(pq)=1α1log(p(x)q(x))αq(x)dxD_\alpha(p \| q) = \frac{1}{\alpha - 1} \log \int \left( \frac{p(x)}{q(x)} \right)^\alpha q(x) dx4 in a finite-dimensional Dα(pq)=1α1log(p(x)q(x))αq(x)dxD_\alpha(p \| q) = \frac{1}{\alpha - 1} \log \int \left( \frac{p(x)}{q(x)} \right)^\alpha q(x) dx5, the "sandwiched" Rényi Dα(pq)=1α1log(p(x)q(x))αq(x)dxD_\alpha(p \| q) = \frac{1}{\alpha - 1} \log \int \left( \frac{p(x)}{q(x)} \right)^\alpha q(x) dx6-divergence is

Dα(pq)=1α1log(p(x)q(x))αq(x)dxD_\alpha(p \| q) = \frac{1}{\alpha - 1} \log \int \left( \frac{p(x)}{q(x)} \right)^\alpha q(x) dx7

The gradient with respect to Dα(pq)=1α1log(p(x)q(x))αq(x)dxD_\alpha(p \| q) = \frac{1}{\alpha - 1} \log \int \left( \frac{p(x)}{q(x)} \right)^\alpha q(x) dx8 is

Dα(pq)=1α1log(p(x)q(x))αq(x)dxD_\alpha(p \| q) = \frac{1}{\alpha - 1} \log \int \left( \frac{p(x)}{q(x)} \right)^\alpha q(x) dx9

where α1\alpha\to 10, α1\alpha\to 11 (Takahashi et al., 2016). For α1\alpha\to 12, this recovers the quantum relative entropy gradient α1\alpha\to 13. Analogous results hold for gradient flows of sandwiched Rényi divergence under GNS-detailed-balance Lindblad semigroups, where the Lindblad equation is proven to be the gradient flow of the divergence with respect to a non-commutative Otto-Wasserstein-like metric (Cao et al., 2018).

5. Applications in Variational Inference and Learning

In exponential families, the gradient of the Rényi divergence with respect to the natural parameter α1\alpha\to 14 is

α1\alpha\to 15

where α1\alpha\to 16 is the expectation of the sufficient statistics under α1\alpha\to 17, and α1\alpha\to 18 is the α1\alpha\to 19-geometric average density (Guilmeau et al., 2022). This is used in Rényi-divergence minimization via Bregman proximal gradient algorithms, which interpolate between standard moment-matching and "geometric" averages. This relaxed update enables robust inference and provable convergence rates.

In deep learning, e.g., Deep Mutual Learning (DML) with Rényi divergence (Huang et al., 2022), the gradient w.r.t. network parameters Dα(pθq)D_\alpha(p_\theta\|q)0 is

Dα(pθq)D_\alpha(p_\theta\|q)1

with Dα(pθq)D_\alpha(p_\theta\|q)2 the tilted "peer-regularization" distribution. This structure generalizes cross-entropy and is tunable via Dα(pθq)D_\alpha(p_\theta\|q)3, with limiting behavior reducing to KL-based mutual learning as Dα(pθq)D_\alpha(p_\theta\|q)4.

6. Quantum Machine Learning and Barren Plateau Avoidance

The maximal Rényi divergence of order two for quantum states,

Dα(pθq)D_\alpha(p_\theta\|q)5

exhibits gradients

Dα(pθq)D_\alpha(p_\theta\|q)6

This unboundedness circumvents gradient vanishing ("barren plateau") phenomena prevalent with bounded, linear cost functions in QNN training (Kieferova et al., 2021).

7. Control and Reinforcement Learning

In risk-sensitive control viewed as inference, Rényi divergence gradients drive new classes of policy-gradient methods. For variational objectives parametrized by risk-sensitivity parameter Dα(pθq)D_\alpha(p_\theta\|q)7 (with Dα(pθq)D_\alpha(p_\theta\|q)8), the policy gradients are

Dα(pθq)D_\alpha(p_\theta\|q)9

with θ\theta0, and for policy update steps, weights θ\theta1 interpolate between mass-covering (θ\theta2) and zero-forcing (θ\theta3) behaviors. Actor-critic updates involve similarly explicit Rényi-gradient terms, interpolating MaxEnt SAC and its risk-sensitive generalizations (Ito et al., 2024).


References

  • "Exponential decay of Rényi divergence under Fokker-Planck equations" (Cao et al., 2018)
  • "Information geometry of sandwiched Rényi θ\theta4-divergence" (Takahashi et al., 2016)
  • "Logarithmic divergences from optimal transport and Rényi geometry" (Wong, 2017)
  • "Rényi Divergence Deep Mutual Learning" (Huang et al., 2022)
  • "Quantum Generative Training Using Rényi Divergences" (Kieferova et al., 2021)
  • "Gradient flow structure and exponential decay of the sandwiched Rényi divergence for primitive Lindblad equations with GNS-detailed balance" (Cao et al., 2018)
  • "Regularized Rényi divergence minimization through Bregman proximal gradient algorithms" (Guilmeau et al., 2022)
  • "Risk-sensitive control as inference with Rényi divergence" (Ito et al., 2024)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Rényi Divergence Gradient.