---
title: Transport-Information Inequalities
url: https://www.emergentmind.com/topics/transport-information-inequalities-ti
type: topic
---

# Transport-Information Inequalities

Transport-information inequalities (TI), also termed transportation cost-information inequalities, are a central topic at the interface of probability, analysis, and geometry, quantifying how the cost of transporting mass between probability distributions is controlled by information-theoretic divergences such as entropy or Fisher information. These inequalities generalize and connect classical functional inequalities (logarithmic Sobolev, Poincaré), concentration of measure, PDE regularity, and stochastic process theory. Applications span Euclidean spaces, path spaces (diffusions and interacting particles), discrete Markov chains, trees, and even quantum systems.

## 1. General Formulations and Key Definitions

A transport-information inequality typically bounds an optimal transport cost $T_c(\nu,\mu)$ (for example, a $p$-Wasserstein metric) between probability measures $\nu$ and reference measure $\mu$ in terms of an "information" divergence such as entropy $H(\nu\mid\mu)$ or Fisher information $I(\nu\mid\mu)$:

- **Transport-Entropy (TE) Inequality:** 
  $$
  \alpha(T_c(\nu,\mu)) \leq H(\nu\mid\mu)
  $$
  for convex increasing $\alpha$, transport cost $c$, and relative entropy $H(\nu\mid\mu) = \int \log \frac{d\nu}{d\mu} d\nu$ [1003.3852].

- **Transport-Information (TI) Inequality:** 
  $$
  \alpha(T_c(\nu,\mu)) \leq I(\nu\mid\mu)
  $$
  with $I(\nu\mid\mu)$ the Donsker–Varadhan (or relative Fisher) information associated to a symmetric Markov generator $L$:
  $$
  I(\nu\mid\mu) = \mathcal{E}(\sqrt{f},\sqrt{f}),
  $$
  for $f = d\nu/d\mu$ and Dirichlet form $\mathcal{E}$ [1003.3852, 2603.13135].

The canonical examples are the $L^p$-Wasserstein distances $W_p$, with $\alpha(r) = r^p/C$ and $c(x,y)=d^p(x,y)$:
- **$T_2$ (Talagrand's):** $W_2^2(\nu, \mu)\leq C H(\nu|\mu)$
- **$W_2I$ (quadratic TI):** $W_2^2(\nu,\mu)\leq C I(\nu\mid \mu)$

Related divergences include Rényi, $f$-divergences, and their convex-analytic variants, all of which can induce transport-type inequalities [2311.14520].

## 2. Classical Models: Euclidean, Markov Semigroups, and Function Spaces

### 2.1. Euclidean Diffusions and Semigroup Methods

For $\mu$ on $\mathbb{R}^d$ or Riemannian manifolds, TI inequalities can be derived under convexity or curvature-dimension assumptions. A paradigmatic result for $\mu = N(0,I)$ (standard Gaussian) is:
$$
W_2^2(\nu,\gamma) \leq 2 H(\nu\mid\gamma)
$$
Building from this, Kolesnikov [1007.1103] showed (for uniformly convex potentials) that the Fisher information of $\mu$ controls the $W^{2,2}$ Sobolev regularity of the transport map, leading to:
$$
I(\mu) \geq K \int |D^2\Phi|^2_{HS}\,d\mu \implies W_2^2(\mu,\nu) \leq (2/K) H(\mu\mid \nu)
$$

Semigroup (Ornstein–Uhlenbeck) interpolation provides De Bruijn-type formulas:
$$
\frac{d}{dt} H(\nu^t\mid\gamma) = - I(\nu^t\mid\gamma), \quad \nu^t = P_t\nu
$$
and leads to strengthened inequalities involving the Stein kernel—most notably the Ledoux–Nourdin–Peccati HSI and WSH inequalities, which interpolate and strictly strengthen the log-Sobolev and Talagrand's inequalities when Stein discrepancy is small [1403.5855].

### 2.2. Lyapunov and Variational Perspectives

TI inequalities can be equivalently characterized by Lyapunov drift conditions: existence of $W > 0$ such that
$$
L W(x) \leq [ -c\,d(x,x_0)^2 + b ]\, W(x)
$$
implies $W_2(\nu,\mu) \leq \sqrt{K H(\nu|\mu)}$, i.e., $T_2$-type inequality [1001.1822, 1506.02489]. The approach is robust under perturbations and admits generalization to non-convex and degenerate settings.

Variational characterizations interpret $T_2$ inequalities as the absence of nontrivial critical points of entropy–transport functionals $G_a(\nu) = a H(\nu|\mu) - W_2^2(\nu,\mu)$ [1508.07642].

### 2.3. Markov Process and Product Structure

The $W_2I$ inequality for Markov processes is equivalent to dimension-free Gaussian concentration for empirical averages of i.i.d. copies:
$$
\mathbb{P}_\nu\Big(\frac{1}{t}\int_0^t f(X_s)ds - \mathbb{E}_\mu f \geq r\Big) \leq \left\|\frac{d\nu}{d\mu^{\otimes n}}\right\|_{L^2} \exp\left(-\frac{t r^2}{C}\right)
$$
for all $n \in \mathbb{N}$, $1$-Lipschitz $f$ [2012.02304]. Such characterizations connect functional, semigroup, and large deviations formulations.

## 3. Transport-Information in Discrete and Structured Spaces

### 3.1. Discrete Markov Chains and Graphs

Transport-information inequalities extend to discrete settings (finite Markov chains, graphs), with the discrete gradient operator (carré du champ). Under discrete curvature-dimension (Bakry–Émery, Ollivier), one obtains:

- **Bakry–Émery CD$(\kappa,\infty)$:** 
  $$
  W_{1,d_\Gamma}(f\pi, \pi) \leq \frac{2}{\kappa} \sqrt{I_\pi(f)}
  $$
- **Ollivier coarse Ricci $\kappa_c$:** 
  $$
  W_{1,d_g}(f\pi, \pi) \leq \frac{2\sqrt{J}}{\kappa_c} \sqrt{I_\pi(f)}
  $$
where $I_\pi(f):=4\int \Gamma(\sqrt f) d\pi$ [1509.07160].

On finite graphs, Ma–Wang–Wu develop explicit path-based estimates for the constant $C_G$ in the discrete $W_1$–information inequality:
$$
W_{1,\rho}(\nu, \mu)^2 \leq 2 C_G I(\nu|\mu)
$$
where $C_G$ is controlled via a sum over network paths and edges [1505.04552].

### 3.2. Trees and Bifurcating Markov Chains

For processes on trees, e.g., bifurcating Markov chains modeling cell division, the law of the process up to generation $n$ satisfies
$$
W_p^d(\nu, \mu) \leq \sqrt{2C_N H(\nu|\mu)}
$$
with $C_N$ scaling appropriately with tree depth and structure; this enables non-asymptotic concentration bounds for empirical means [1501.06693].

### 3.3. Point Processes: Poisson, Binomial

Transport-information inequalities lift via tensorization and symmetrization to Poisson and mixed binomial point processes. In particular, if a base measure $\mu$ satisfies $T_2$, the Poisson process law $\Pi_\nu$ on the configuration space satisfies
$$
W_2^2(\Pi_1, \Pi_2) \leq a\left[H(\Pi_1|\Pi_\nu) + H(\Pi_2|\Pi_\nu)\right]
$$
with preservation of constants [2002.04923]. Universal Marton-type ($M^2$) inequalities imply robust concentration properties for convex functionals of the random point fields.

## 4. Functional and Dimensional Extensions

### 4.1. Reweighted and Mixture Models

For measures $\pi$ decomposed as a mixture of components $(\pi_i)$ each admitting TI inequalities with profile $\alpha$, Niles-Weed [2603.13135] shows that proximity in Fisher information to $\pi$ guarantees proximity in transport (with the same profile) to some reweighted mixture $\tilde\pi$ in the same family. This "reweighted TI" mechanism allows effective treatment of non-log-concave or multimodal targets in sampling and metastability analysis, and yields structurally optimal convergence rates.

### 4.2. Quantum and Infinite-Dimensional Models

Quantum extensions of TI are established via quantum Wasserstein distances (quantum Ricci curvature, Petz maps), and show that high-temperature Gibbs states satisfy
$$
\|\rho-\omega\|_{W_1} \leq \sqrt{C S(\rho\|\omega)}
$$
with $C=O(|V|)$, yielding quantum Gaussian concentration, equivalence of ensembles, and exponential improvement over previous versions of the Eigenstate Thermalization Hypothesis [2106.15819].

For infinite-dimensional path spaces (diffusions, SPDEs), TI inequalities control the law of the process, e.g., for solutions of stochastic wave equations, where the constant depends on the time horizon, the "size" of the fundamental solution, and the Lipschitz constant of the drift [1811.06385].

## 5. Methods of Proof and Structural Results

### 5.1. Semigroup Interpolation and $\Gamma$-Calculus

Heat-flow (e.g., Ornstein–Uhlenbeck) interpolation and $\Gamma$-calculus provide both sharp estimates and dynamical proofs, yielding exponential convergence in entropy, moment bounds, and direct connections between log-Sobolev, Talagrand, and TI inequalities [1403.5855, 1001.1822].

### 5.2. Variational and Duality Techniques

Bobkov–Götze duality, Kantorovich duality, and convex-analytic tensorization allow characterization of TI inequalities in terms of optimization problems on probability measures and test functions, leading to unification with large deviations and Sanov's theorem [1508.07642, 2012.02304].

### 5.3. Curvature and Path Methods

Finite and discrete models leverage curvature conditions (Bakry–Émery, Ollivier) and path decompositions (random-path methods), giving explicit dimension and geometry dependence of TI constants [1509.07160, 1505.04552].

### 5.4. Lyapunov Function and Drift

Lyapunov drift conditions provide robust sufficient (and under suitable regularity, necessary) criteria for TI inequalities, extending to general state spaces and degenerate situations [1001.1822, 1506.02489].

## 6. Applications: Concentration, Stability, and Extensions

- **Concentration of measure:** TI inequalities imply sub-Gaussian (or better) deviation bounds for Lipschitz observables and empirical functions, uniformly over system size and complexity [1003.3852, 1501.06693, 1505.04552].
- **Stability:** TI properties are stable under bounded density perturbations (Holley–Stroock), tensorization, and push-forward through measurable maps [1506.02489, 2002.04923, 2603.13135].
- **Functional inequalities:** TI implies and is implied by stronger/related inequalities: log-Sobolev, Poincaré, weak/super-Poincaré, and dimension-free concentration [1506.02489, 1508.07642, 1001.1822].
- **Sampling and Markov processes:** Reweighted TI enables new guarantees for convergence and sampling in non-log-concave, multimodal, or high-dimensional pathological regimes [2603.13135].
- **Quantum processes and random matrices:** Quantum TCIs provide concentration and mixing results extending the classical theory [2106.15819].
- **Interacting particle systems and SPDEs:** TI controls concentration, hydrodynamic limits, and deviation principles for high-dimensional, non-product processes [1811.06385, 1808.02164].

## Table: Canonical TI Inequalities and Settings

| Setting                      | TI Formulation                                     | Key Constant/Assumptions                                            |
|------------------------------|----------------------------------------------------|---------------------------------------------------------------------|
| Gaussian measure $\gamma$    | $W_2^2(\nu,\gamma)\leq 2 H(\nu|\gamma)$           | Convexity, semigroup smoothing [1003.3852, 1403.5855]               |
| Uniformly convex log-density | $W_2^2(\mu,\nu)\leq (2/K) D(\mu||\nu)$            | $D^2W\geq K\operatorname{Id}$ [1007.1103]                           |
| Discrete Markov chain        | $W_1(\nu,\pi)\leq (2/\kappa)\sqrt{I_\pi(f)}$      | CD$(\kappa,\infty)$ curvature [1509.07160]                          |
| Point processes              | $W_2^2(\Pi_1,\Pi_2)\leq a [H(\Pi_1|\Pi_\nu)+H(\Pi_2|\Pi_\nu)]$ | Base $T_2$ condition, tensorization [2002.04923]         |
| Mixtures                     | $\inf_{\tilde\pi\in C}W_2^2(\mu,\tilde\pi)\leq C^2 FI(\mu\|\pi)$ | Reweighted over mixture components [2603.13135]          |
| Quantum system               | $\|\rho-\omega\|_{W_1}\leq \sqrt{C S(\rho\|\omega)}$ | Curvature, mLSI, or local indistinguishability [2106.15819]         |
| Path space/SPDE              | $W_2^2(Q,P^\nu)\leq 2C\,H(Q|P^\nu)$               | Lipschitz drift, Gaussian/correlated noise [1811.06385, 1808.02164] |

## 7. Perspectives and Current Directions

Transport-information inequalities continue to drive new developments in:
- Sampling theory for non-log-concave and multimodal distributions, via modular control of Fisher information and proximity to mixtures [2603.13135].
- Quantum information geometry and concentration, leveraging non-commutative analogues of curvature and optimal transport [2106.15819].
- Infinite-dimensional and interacting systems, including reflected diffusions and stochastic PDEs [1811.06385, 1808.02164].
- Integration with functional analysis, large deviations, and spectral theory via convex-analytic duality and variational frameworks [1508.07642, 2012.02304].
- Discrete and combinatorial models, where curvature, path, and concentration phenomena are realized in non-Euclidean or random environments [1509.07160, 1505.04552, 1501.06693].
- Explicit computation and optimization of constants and profiles, especially via geometric or path-based methods.
- Robustness, stability, and transference under perturbation, tensorization, and map-based contraction [1506.02489, 2603.13135].

These advances underline the central role of transport-information inequalities as fundamental, structurally robust bridges between geometry, probability, and analysis.

Source: https://www.emergentmind.com/topics/transport-information-inequalities-ti