---
title: Donsker–Varadhan Representation
url: https://www.emergentmind.com/topics/donsker-varadhan-representation
type: topic
---

# Donsker–Varadhan Representation

The Donsker–Varadhan representation is a set of fundamental variational formulas encoding the rate function for occupation-time large deviations of Markov processes, both in finite and infinite dimensions, discrete and continuous time, and for both equilibrium and nonequilibrium systems. These representations are central in the mathematical theory of large deviations, ergodic properties of Markov processes, risk-sensitive control, information-theoretic mutual information estimation, and modern statistical mechanics. The universal ingredient in all Donsker–Varadhan formulas is a supremum (or infimum) over suitable test functions, potentials, or path measures, connected to spectral properties of Feynman–Kac semigroups and the structure of the Markov generator.

## 1. Core Variational Formulation

The prototypical setting is an irreducible continuous-time Markov process $(X_t)_{t\ge0}$ on a finite or compact state space $K$, with generator $L$ and unique stationary law $\rho$. The empirical occupation measure over $[0,T]$ is
\[
\mu_T(x)=\frac1T\int_0^T \mathbf{1}_{\{X_t=x\}}\,dt,
\]
which satisfies a large-deviation principle as $T\to\infty$:
\[
\mathbb{P}_\rho[\mu_T\approx \mu] \asymp \exp\{-T\,\mathcal{I}(\mu)\}.
\]
The Donsker–Varadhan rate functional admits several equivalent forms, e.g.:
\[
\mathcal{I}(\mu)
= \sup_{g>0} \left\{ -\sum_{x} \frac{\mu(x)}{g(x)}\sum_{y}k(x,y)[g(y)-g(x)] \right\}
\]
or, under spectral representation,
\[
\mathcal{I}(\mu)=\sup_{f} \left\{\int f\,d\mu - \Lambda(f)\right\},\quad
\Lambda(f)=\lim_{T\to\infty}\frac{1}{T}\log\mathbb{E}_x\big[e^{\int_0^T f(X_s)\,ds}\big],
\]
with the latter limit given by the principal eigenvalue of $L+M_f$ (where $M_f$ is the multiplication operator by $f$) [1102.2690], [2211.02593], [1310.5829], [1602.02545].

Under detailed balance, the rate functional reduces to a Dirichlet form (quadratic in $\sqrt{\mu/\rho}$) or relative entropy up to a constant, but out of equilibrium the Donsker–Varadhan functional encodes excess dynamical activity not captured by entropy alone [1102.2690], [1602.02545].

## 2. Generalizations: Infinite Dimensions and Path Space

For Markov models on compact metric spaces, processes with degeneracies, or infinite-dimensional systems (e.g., SPDEs, cellular automata), the Donsker–Varadhan variational principle extends structurally:

- **Abstract Feller/compact-space setting**:
  \[
  I(\mu) = \sup_{\phi\in C_b(H)}\left\{\int_H \phi\,d\mu - \log\lambda_\phi\right\},
  \]
  where $\lambda_\phi$ is the principal eigenvalue of the Feynman–Kac semigroup with test potential $\phi$ [2510.24119].

- **Empirical flow and joint LDPs**: For jump processes, one obtains a joint rate function for empirical measure and empirical flow, with Donsker–Varadhan representation contracted to the occupation component [1310.5829].

- **SPDEs and stochastic PDEs**: For white-forced Navier–Stokes or nonlinear Schrödinger systems, level-2 and level-3 DV-type rate functionals are constructed using principal eigenvalues of Markov semigroups on path space, controlled by Feynman–Kac asymptotics and detailed regularity properties [2506.14119], [2510.24119].

- **Infinite-dimensional PCA**: The Donsker–Varadhan action functional is defined for empirical measures of probabilistic cellular automata, with finiteness of the DV rate functional only on measures stationary on the spatial tail σ-algebra [2509.01745].

## 3. Information-Theoretic Extensions: KL, Mutual Information, Deep Learning

The Donsker–Varadhan representation forms the backbone of several information-theoretic quantities:

- **Kullback–Leibler Divergence**:
  \[
  D_\mathrm{KL}(P\Vert Q) = \sup_T \mathbb{E}_{P}[T] - \log\mathbb{E}_{Q}[e^T],
  \]
  with $T$ ranging over measurable functions. The optimal $T^*$ is (up to a constant) the log-density ratio. Approximations of $T^*$ using neural networks underlie techniques for deep data density estimation and mutual information estimation in high-dimensional or continuous settings [2104.06612, 2506.22789].

- **Mutual Information**: Since MI is a KL between a joint law and product of marginals,
  \[
  I(X;Y) = \sup_T \mathbb{E}_{P_{XY}}[T(x,y)] - \log \mathbb{E}_{P_XP_Y}[e^{T(x,y)}],
  \]
  this form is directly used for variational MI estimators, privacy-leakage control, and fairness-aware representation learning (e.g., in speech embeddings) via critic neural networks parameterizing $T$ [2506.22789].

- **Risk-sensitive control**: Optimal asymptotic growth rates for controlled Markov processes are given by controlled Donsker–Varadhan formulas, maximizing reward minus relative entropy over valid ergodic occupation measures (admissible policies), with explicit finite-state LP/DP reductions [1903.10714].

## 4. Analytical and Physical Interpretation

The Donsker–Varadhan rate functional has multiple interpretations and analytic features:

- **Spectral character**: The rate function is the Legendre transform of the asymptotic cumulant generating function governed by the principal eigenvalue of the Feynman–Kac semigroup with added potential [1102.2690], [2211.02593].

- **Lyapunov and monotonicity**: Under a sector condition (normal linear-response), the DV functional is monotone decreasing along Markov evolution, functioning as a Lyapunov functional for convergence to stationary distributions, even out of equilibrium [1102.2690].

- **Dynamical activity**: For Markov jump processes, the DV functional measures the excess in expected escape rate (activity) required to hold an atypical occupation profile, quantifying the "cost" in dynamical terms [1102.2690].

- **Thermodynamic and entropy relations**: For diffusions with detailed balance, the DV functional corresponds (up to scaling) to nonadiabatic entropy production rates. In nonequilibrium steady state, the functional generalizes to incorporate both deviations in empirical current and density, reflecting the structure of entropy production in stochastic thermodynamics [1602.02545].

- **Bridge and path-space representations**: Recent work has shown the DV rate for occupational measures can alternatively be formulated as a two-stage infimum over joint laws (e.g., pairwise transition frequencies and within-block fluctuations), linking classical variational and probabilistic bridge-based perspectives [2407.00216].

## 5. Applications in Modern Probability and Mathematical Physics

Donsker–Varadhan representations appear across diverse domains:

- **Large deviations theory**: Classical and hydrodynamic limits for exclusion processes, analysis of dynamical phase transitions, and fluctuation theorems leverage DV formulations at both microscopic and macroscopic levels [2111.05892].
- **Infinite-dimensional systems**: Extensions to SPDEs and deterministic or stochastic PDEs rely on Feynman–Kac principal eigenvalue methods, often requiring uniform Feller properties, exponential mixing, and sophisticated coupling arguments [2510.24119], [2506.14119].
- **Statistical mechanics**: Nonequilibrium large-deviation rate functions, contractions to currents or observables, and entropy relations are derived from or reduced to DV formulas [1602.02545], [1102.2690].
- **Machine learning and information theory**: DV-based neural estimation underpins recent advances in data density estimation, privacy-presevering representation learning, and variational MI bounds in high-dimensional and continuous data [2104.06612], [2506.22789].

## 6. Extensions, Controlled Formulations, and Limitations

- **Controlled Markov processes**: For models with controls or actions, the DV representation is extended to ergodic occupation measures of both state and control, introducing an additional entropy penalty for deviation from the passive transition kernel. Explicit LP/DP programs exist in finite state-action spaces, connecting to the Collatz–Wielandt formula for positive operators [1903.10714].

- **Criteria for LDP applicability**: Establishment of the full LDP and the validity of the DV representation, particularly in infinite-dimensional or degenerate settings (absorbing states, slow mixing), generally requires irreducibility, uniform Feller properties, and sometimes coupling/squeezing conditions. Special attention is needed in models with degeneracies or non-ergodicity, where the zero-level set of the DV functional may not be a singleton and finiteness can restrict to stationary (or tail-stationary) measures [1310.5829], [2509.01745].

- **Alternative representations**: Bridge- and flow-based decompositions provide alternative perspectives, often more amenable to certain generalizations, path-level large deviations, or explicit computations for occupation-flow LDPs [2407.00216], [1310.5829].

---

The Donsker–Varadhan representation and its generalizations thus form a unifying backbone across large-deviation theory, spectral analysis of stochastic processes, statistical physics, controlled Markov dynamics, information-theoretic estimation, and contemporary computational methods. Their wide impact derives from their variational structure, spectral dualities, and operational connections to dynamical and information-theoretic quantities.

Source: https://www.emergentmind.com/topics/donsker-varadhan-representation