---
title: Donsker–Varadhan Variational Formula
url: https://www.emergentmind.com/topics/donsker-varadhan-variational-formula
type: topic
---

# Donsker–Varadhan Variational Formula

The Donsker–Varadhan variational formula is a foundational result in probability theory, statistical mechanics, large deviations, spectral theory, risk-sensitive control, and modern machine learning. It provides a dual representation of quantities such as the Kullback–Leibler divergence, principal eigenvalues of positive operators, and large deviation rate functions for empirical measures of Markov processes. The formula and its extensions unify concepts from convex analysis, information theory, stochastic control, and statistical inference, with modern applications ranging from reinforcement learning policy optimization to neural network-based divergence estimation.

## 1. Classical Variational Principle

The classical Donsker–Varadhan (DV) variational formula gives a dual representation of the Kullback–Leibler divergence (relative entropy) between two probability measures. For probability measures $Q$ and $P$ on a measurable space $(\Omega, \mathcal{F})$ with $Q \ll P$, the DV formula is:
\[
D_\mathrm{KL}(Q\|P) = \sup_{f\in\mathcal{M}_b(\Omega)} \left\{ \mathbb{E}_Q[f] - \log \mathbb{E}_P[e^{f}] \right\}.
\]
The supremum is attained at $f^*(x) = \log\frac{dQ}{dP}(x)$. This representation applies to general measurable spaces, with the admissible test functions extendable to bounded continuous or Lipschitz functions where appropriate. For Markov processes, the associated rate function for empirical measures admits an equivalent formula using the generator $L$:
\[
I(\mu) = -\inf_{u>0\in\operatorname{Dom}(L)} \int \frac{L u}{u}\, d\mu,
\]
where the infimum is over strictly positive functions in the domain of $L$ [2007.03814][1302.6647].

## 2. Extensions: Rényi and Generalized Divergences

The DV formula generalizes to the Rényi α-divergences. For $\alpha\neq 0,1$, the order-$\alpha$ Rényi divergence $R_\alpha(Q\|P)$ admits the following variational form:
\[
R_\alpha(Q\|P) = \sup_{g \in \Gamma} \left\{ \frac{1}{\alpha-1} \log \int e^{(\alpha-1)g} dQ - \frac{1}{\alpha} \log \int e^{\alpha g} dP \right\},
\]
for any function class $\Gamma$ containing the bounded measurable functions. At optimality, $g^*(x) = \log\frac{dQ}{dP}(x)$ in the absolutely continuous case, and in the limit $\alpha\to1$, the formula recovers the classical DV variational representation for Kullback–Leibler divergence [2007.03814].

A maxitive analogue of the Donsker–Varadhan formula exists in possibility theory, substituting integrals with suprema and KL divergence with the max-relative entropy, formalizing a max-plus (tropical) duality essential for possibilistic variational inference [2511.21223].

## 3. Spectral and Large Deviations Theory

The DV formula is the cornerstone for analyzing large deviations of empirical measures of Markov processes:

- **Diffusions**: For a reversible diffusion $X_t$ with infinitesimal generator
  $
  Lf(x) = \frac{1}{2}\operatorname{Tr}(a(x) D^2 f(x)) + b(x) \cdot \nabla f(x),
  $
  the rate function for occupation-time large deviations is
  \[
  I(\mu) = \sup_{f \in C_b^2} \left\{ \int Lf\, d\mu - \int e^{-f} (L e^{f})\, d\mu \right\}.
  \]
  This formula connects to the Freidlin–Wentzell action in small noise limits [2211.02593].

- **Jump Markov processes**: For pure jump Markov processes on a compact Polish space with generator
  $
  Lf(x) = q(x) \int [f(y)-f(x)] \alpha(x,dy),
  $
  and reversible invariant measure $\pi$, the explicit DV rate function is:
  \[
  I(\eta) =
  \begin{cases}
    \displaystyle
    \int q(x) \eta(dx) - \int \sqrt{\theta(x)\theta(y)} q(x)\alpha(x,dy)\pi(dx), & \eta \ll \pi,\, \theta = \frac{d\eta}{d\pi}, \\
    +\infty, & \text{otherwise},
  \end{cases}
  \]
  equivalently expressed via the Dirichlet form $E_\pi(g,g)$ for $g = \sqrt{\theta}$ [1302.6647].

- **Spectral theory**: The DV formula arises in characterizing principal eigenvalues of positive (linear or nonlinear) operators via
  \[
  \lambda = \sup_{\mu \in \mathcal{P}(Q)} \inf_{f>0} \int \frac{Gf(x)}{f(x)} \mu(dx),
  \]
  crucial for risk-sensitive control and spectral theory of elliptic operators [1312.5834][1501.00676][1903.10714].

## 4. Applications in Modern Machine Learning

The Donsker–Varadhan variational principle underlies numerous algorithms and estimator designs in contemporary machine learning:

- **Reinforcement learning**: The DV formula expresses the equivalence between entropy-regularized policy gradients and soft Q-learning, with the optimal policy taking the Gibbs (Boltzmann) form:
  \[
  \pi^*(a|s) \propto \exp \left( \frac{r(s,a) + \gamma V(s')}{\beta} \right).
  \]
  The Bellman operator incorporates a log-sum-exp (softmax) and the Legendre transform with respect to entropy regularization [1712.08650].

- **Density estimation by deep networks**: Plugging deep neural networks into the DV representation enables principled estimation of data densities. For example, with $P$ empirical and $Q$ uniform, training $T_\theta$ to maximize
  \[
  \mathbb{E}_P[T_\theta] - \log \mathbb{E}_Q[e^{T_\theta}]
  \]
  produces $T^*$ approximating $\log$-density up to an additive constant. This approach is competitive for high-dimensional density estimation, outperforms kernel density estimation, and improves downstream tasks such as anomaly detection and classification [2104.06612].

- **Variational estimation of Rényi divergences**: The Rényi-DV formula provides the foundation for statistically consistent neural network estimators of $R_\alpha(Q\|P)$. Universal approximation of the test function by neural nets is sufficient for consistency, and rates scale efficiently with sample size even in thousands of dimensions [2007.03814].

- **Possibilistic inference**: Replacement of integrals by max operations and KL by max-entropy allows application in imprecise probability and robust inference [2511.21223].

## 5. Risk-Sensitive Control, Dynamic Programming, and Collatz–Wielandt Duality

The variational structure of the DV formula generalizes to the principal eigenvalue of controlled positive operators, yielding risk-sensitive reward rates:
\[
\lambda^* = \sup_{n \in \mathcal{M}} \left\{ \int r(x,u,y) n(dx,du,dy) - \int D(\eta_2(\cdot|x,u)\|p(\cdot|x,u)) n_0(dx)\eta_1(du|x) \right\},
\]
where $n$ runs over ergodic occupation measures induced by admissible policies $\eta$, and $D$ is the Kullback–Leibler divergence; $r$ denotes per-stage reward [1903.10714][1501.00676].

The abstract Collatz–Wielandt duality, extended to nonlinear and controlled contexts, provides the min-max structure underlying the DV formula and its generalizations [1312.5834].

## 6. Further Structural and Computational Developments

The DV variational principle also:

- Underlies large deviation analysis for Markov processes with degenerate rates (e.g., absorbing states) and self-interacting pure jump processes, yielding explicit and non-convex large deviation rate functions in the latter case [1310.5829][2510.14659].

- Admits max-plus (tropical) analogues for possibility theory, crucial for robust variational inference frameworks [2511.21223].

- Provides lower bounds for principal eigenvalues of elliptic operators in terms of exit times, offering both classical moment-based and quantile-based (probabilistic) bounds with implications for metastability and Monte Carlo spectral estimation [1611.09294].

- Facilitates explicit computation in Piecewise Deterministic Markov Processes (PDMP) such as the zig-zag sampler, where the DV rate function quantifies empirical measure convergence and elucidates the effect of refreshment rates [1912.06635].

## 7. Summary Table: Donsker–Varadhan Formula—Contexts and Variational Forms

| Context                                | Variational Form                                                                   | Reference            |
|-----------------------------------------|------------------------------------------------------------------------------------|----------------------|
| KL divergence ($Q \ll P$)               | $\sup_f \{ \mathbb{E}_Q[f] - \log \mathbb{E}_P[e^{f}] \}$                          | [2007.03814][2104.06612]       |
| Rényi divergence                        | $\sup_g \{ \frac{1}{\alpha-1}\log\mathbb{E}_Q[e^{(\alpha-1)g}] -\frac{1}{\alpha}\log\mathbb{E}_P[e^{\alpha g}] \}$ | [2007.03814]         |
| Empirical measure LDP (diffusion)       | $-\inf_{u>0} \int \frac{Lu}{u}\, d\mu$                                             | [1302.6647][2211.02593]  |
| Controlled Markov chain (reward rate)   | $\sup_n \{ \int r\, dn - \int D(\cdot)\, dn \}$                                    | [1903.10714][1501.00676] |
| Possibility theory (maxitive)           | $\sup_g \inf_\theta \{-\ell(\theta)-\log \frac{g(\theta)}{\pi(\theta)}\}$          | [2511.21223]          |
| Zig-zag process PDMP                    | $-\inf_{u>0} \int \frac{Lu}{u}\, d\mu$, explicit minimization in $\eta$            | [1912.06635]          |

The DV variational framework thus serves as a unifying principle across stochastic processes, divergence estimation, spectral theory, control, and contemporary machine learning, underpinning both theoretical advances and practical algorithmic developments.

Source: https://www.emergentmind.com/topics/donsker-varadhan-variational-formula