---
title: 'Statistical Ricci Curvature: Transport & Geometry'
url: https://www.emergentmind.com/topics/statistical-ricci-curvature
type: topic
---

# Statistical Ricci Curvature: Transport & Geometry

Searching arXiv for recent and foundational papers on statistical Ricci curvature, optimal-transport curvature, and discrete/simplicial variants.
Statistical Ricci curvature denotes a family of Ricci-type constructions in which curvature is encoded through probability measures, Markov kernels, entropy, diffusion semigroups, or dual affine connections, rather than only through the classical smooth Ricci tensor. In the literature, this includes Ollivier’s coarse Ricci curvature on metric-measure spaces, entropy-convexity curvature for finite Markov chains, Wasserstein-statistical curvature for parametrized models, and the statistical Ricci tensor obtained from dual connections on statistical manifolds [1111.2687] [1807.07095] [1902.09298]. This suggests that the phrase functions as an umbrella term for several closely related programs whose common aim is to recover Ricci-type control from transport, information, or stochastic structure.

## 1. Conceptual scope

Several non-equivalent definitions appear under the same broad theme. In finite-state transport geometry, one studies probability densities \(\rho\) on a Markov chain and asks whether the Boltzmann–Shannon entropy is geodesically convex in a discrete transport metric \(W\). In parametric statistics, one studies a family \(p(\theta)\) of positive probability vectors and imposes convexity of the Kullback–Leibler divergence on a parameter manifold endowed with a Wasserstein metric tensor. In information geometry with dual affine connections, one contracts a statistical curvature tensor to obtain a statistical Ricci tensor [1111.2687] [1807.07095] [1902.09298].

| Setting | Basic object | Ricci-type formulation |
|---|---|---|
| Finite Markov chains | Entropy \(H(\rho)\) on \(\mathscr P(X)\) with transport metric \(W\) | \(H(\rho_t)\) is \(\kappa\)-convex along \(W\)-geodesics |
| Parametrized statistics | \(D_{\mathrm{KL}}\bigl(p(\theta)\|q\bigr)\) on \((\Theta,G_W)\) | \(\mathrm{Hess}_W D_{\mathrm{KL}} \succeq \kappa\,G_W(\theta)\) |
| Statistical manifolds | Statistical curvature \(S=(R+R^*)/2\) | \(\mathrm{Ric}^{\nabla,\nabla^*}(X,Y)=\sum_i g(S(e_i,X)Y,e_i)\) |

The unifying pattern is that curvature lower bounds are phrased as contraction, convexity, or Bochner-type inequalities. A plausible implication is that “statistical” does not single out one canonical object; rather, it identifies a methodological shift from tensorial pointwise definitions to probabilistic, transport, or dual-connection formulations.

## 2. Optimal transport, Markov chains, and coarse curvature

A standard starting point is Ollivier’s coarse Ricci curvature. On a metric space \((X,d)\) with probability measures \(\{m_x\}_{x\in X}\), the \(1\)-Wasserstein distance is
\[
W(m_x,m_y)=\inf_{\pi:\pi_x=m_x,\;\pi_y=m_y}\sum_{u,v}\pi(u,v)\,d(u,v),
\]
and the curvature along \((x,y)\) is
\[
\kappa(x,y)=1-\frac{W(m_x,m_y)}{d(x,y)}.
\]
On graphs, \(m_x\) is often the one-step random-walk measure; on simplicial complexes and directed graphs this same formula is adapted to higher-order or asymmetric neighborhoods [1906.07404].

For finite reversible Markov chains, an exact discrete analogue of the Lott–Sturm–Villani picture replaces the classical \(W_2\)-metric by a discrete transport metric \(W\) under which the heat semigroup becomes the gradient flow of the entropy
\[
H(\rho)=\sum_{x\in X}\pi(x)\,\rho(x)\,\log\rho(x).
\]
A Markov kernel \(K\) is said to satisfy \(\mathrm{Ric}(K)\ge\kappa\) when, for every \(W\)-geodesic \((\rho_t)\),
\[
H(\rho_t)\le (1-t)\,H(\rho_0)+t\,H(\rho_1)-\frac{\kappa}{2}\,t(1-t)\,W(\rho_0,\rho_1)^2.
\]
This is equivalent to an Evolution Variational Inequality, to a lower bound on the Hessian of \(H\) in the \(W\)-Riemannian sense, and to a discrete Bochner-type inequality \(B(\rho,\psi)\ge\kappa\,A(\rho,\psi)\) [1111.2687].

The same framework yields the usual analytic consequences. If \(\mathrm{Ric}(K)\ge\kappa\), then one obtains a modified logarithmic Sobolev inequality, a modified Talagrand inequality, a Poincaré inequality, and exponential contractivity of the heat flow:
\[
W(P_t\rho,P_t\sigma)\le e^{-\kappa t}\,W(\rho,\sigma).
\]
Tensorisation is preserved, and for the simple random walk on the discrete hypercube \(Q^n\) the sharp lower bound \(\mathrm{Ric}(K_n)\ge 2/n\) is obtained [1111.2687].

A related probabilistic development considers multi-step coarse Ricci curvature
\[
\kappa_k(x,y)=1-\frac{W_1(P_x^k,P_y^k)}{d(x,y)}.
\]
This allows one-step negative curvature to be averaged out over longer time scales. The resulting bounds imply mixing-time estimates, spectral-gap bounds such as
\[
\gamma^*\ge 1-(1-\kappa_k)^{1/k}\ge \kappa_k/k,
\]
as well as concentration inequalities and nonasymptotic MCMC bias and variance bounds [1404.2802].

## 3. Parametric statistics and the Wasserstein statistical manifold

In parametric statistics, the state space is finite or graph-based, but the primary geometry lives on the parameter domain. One considers a family of positive probability vectors
\[
p(\theta)=\bigl(p_i(\theta)\bigr)_{i=1}^n,\qquad \theta\in\Theta\subset\mathbb R^d,
\]
and endows \(\mathcal P_+(I)\) with an \(L^2\)-Wasserstein Riemannian metric
\[
g_p(\sigma,\sigma)=\sigma^{\mathsf T}L(p)^\dagger \sigma.
\]
Pulling this metric back by \(\theta\mapsto p(\theta)\) yields the tensor
\[
G_W(\theta)=J_\theta p(\theta)^{\mathsf T}\,L\bigl(p(\theta)\bigr)\,J_\theta p(\theta),
\]
and \((\Theta,G_W)\) is called a Wasserstein statistical manifold [1807.07095].

The principal functional is the Kullback–Leibler divergence relative to a reference measure \(q\),
\[
D_{\mathrm{KL}\bigl(p(\theta)\,\|\,q\bigr)}=\sum_{i=1}^n p_i(\theta)\,\log\frac{p_i(\theta)}{q_i}.
\]
Its Wasserstein gradient flow is
\[
\dot\theta=-\,\nabla_W D_{\mathrm{KL}}
=-\,G_W(\theta)^{-1}\,\nabla_\theta D_{\mathrm{KL}\bigl(p(\theta)\,\|\,q\bigr)},
\]
which drives \(\theta(t)\) toward the minimizer \(\theta^*\) of the divergence. A Ricci-curvature lower bound \(\kappa\) is defined by demanding that \(D_{\mathrm{KL}}\) be \(\kappa\)-convex along every constant-speed geodesic in \((\Theta,G_W)\), equivalently
\[
\mathrm{Hess}_W\,D_{\mathrm{KL}\bigl(p(\theta)\|q\bigr)}\succeq \kappa\,G_W(\theta).
\]

A central identity expresses the Wasserstein Hessian in mixed information-geometric form:
\[
\mathrm{Hess}_W\,D_{\mathrm{KL}}
=
G_F(\theta)
+\sum_{i=1}^n d_{\theta\theta}p_i(\theta)\,\log\frac{p_i(\theta)}{q_i}
-\sum_{k=1}^d\Gamma^{W,k}(\theta)\,\partial_{\theta_k}D_{\mathrm{KL}},
\]
and the matrix inequality
\[
G_F(\theta)
+\sum_{i=1}^n d_{\theta\theta}p_i(\theta)\,\log\frac{p_i(\theta)}{q_i}
-\sum_{k=1}^d\Gamma^{W,k}(\theta)\,\partial_{\theta_k}D_{\mathrm{KL}}
\succeq
\kappa\,G_W(\theta)
\]
is termed the RIW (Ricci–Information–Wasserstein) condition [1807.07095].

When \(\kappa>0\), the KL divergence becomes strongly convex in Riemannian sense, and the gradient flow converges exponentially fast:
\[
D_{\mathrm{KL}\bigl(p(\theta_t)\|q\bigr)}-D_{\mathrm{KL}\bigl(p(\theta^*)\|q\bigr)}
\le
\exp(-2\kappa t)\,
\Bigl[D_{\mathrm{KL}\bigl(p(\theta_0)\|q\bigr)}-D_{\mathrm{KL}\bigl(p(\theta^*)\|q\bigr)\Bigr].
\]
The same curvature bound yields a logarithmic-Sobolev inequality, a Talagrand-transport–entropy inequality, and an HWI inequality on parameter space. The paper’s examples on one-dimensional exponential families on a three-state simplex show numerically that the minimal eigenvalue \(\kappa\) of \(\mathrm{Hess}_W D_{\mathrm{KL}}\) bounds the empirical contraction constant \(K\) from below, and that on small parameter domains the two coincide almost exactly [1807.07095].

## 4. Statistical manifolds with dual affine connections

In a distinct but classical information-geometric sense, a statistical manifold is a Riemannian manifold \((B,g)\) endowed with a pair of torsion-free affine connections \(\nabla\) and \(\nabla^*\) satisfying
\[
X[g(Y,Z)]=g(\nabla_XY,Z)+g(Y,\nabla_X^*Z).
\]
Equivalently,
\[
\nabla=\nabla^g+K,\qquad \nabla^*=\nabla^g-K,
\]
where \(K\) is symmetric in its two lower arguments and totally symmetric when raised by \(g\) [1902.09298].

The relevant curvature tensor is the statistical curvature
\[
S(X,Y)Z=R(X,Y)Z+\bigl[K_X,K_Y\bigr]Z,
\qquad\text{equivalently}\qquad
R+R^*=2S.
\]
The sectional curvature of an orthonormal pair \(\{E,F\}\) is
\[
K(E\wedge F)=g(S(E,F)F,E),
\]
and the statistical Ricci tensor is obtained by contraction:
\[
\mathrm{Ric}^{\nabla,\nabla^*}(X,Y)
=
\sum_{i=1}^n g\bigl(S(e_i,X)Y,e_i\bigr).
\]
In this terminology, the statistical Ricci curvature in the direction \(X\) is \(\mathrm{Ric}^{\nabla,\nabla^*}(X)=\mathrm{Ric}^{\nabla,\nabla^*}(X,X)\) [1902.09298].

The Kenmotsu setting provides a concrete example. For Kenmotsu statistical manifolds of constant \(\phi\)-sectional curvature, the statistical Ricci tensor has explicit \(g\)–\(\eta\otimes\eta\) form, and a non-trivial Kenmotsu statistical manifold of odd dimension is never Ricci-flat. More generally, the paper states that a Kenmotsu statistical manifold of constant \(\phi\)-sectional curvature is never Ricci-flat, apart from the trivial even-dimensional construction with \(K\equiv 0\) and \(c=-1\) [1902.09298].

For statistical submanifolds, one also obtains a Chen–Ricci inequality. The statistical Ricci curvature of the submanifold is bounded below by twice its classical Ricci curvature minus explicit quadratic expressions in the mean curvature vectors \(H,H^*\) and the ambient \(\phi\)-geometry. This suggests a second major use of the adjective “statistical”: here it refers not to transport on spaces of measures, but to dual affine connections and Amari–Chentsov curvature [1902.09298].

## 5. Discrete higher-order, directed, and network formulations

Transport-based curvature extends beyond undirected graphs. For a simple, locally finite, strongly connected directed graph, the \(\alpha\)-lazy random-walk measure is
\[
m_x^\alpha(v)=
\begin{cases}
\alpha,& v=x,\\
(1-\alpha)/d_{\mathrm{out}}(x),& (x,v)\in E,\\
0,& \text{otherwise},
\end{cases}
\]
the directed distance \(d(x,y)\) is the length of a shortest directed path, and
\[
\kappa_\alpha(x,y)=1-\frac{W(m_x^\alpha,m_y^\alpha)}{d(x,y)}.
\]
The renormalized limit
\[
\kappa(x,y):=\lim_{\alpha\to1}\frac{\kappa_\alpha(x,y)}{1-\alpha}
\]
exists and defines the Lin–Lu–Yau curvature in the directed setting. Exact criteria are given for Ricci-flat regular directed graphs, and the Cartesian product satisfies a tensorization formula in which curvature is weighted by \(d_G/(d_G+d_H)\) or \(d_H/(d_G+d_H)\) [1602.07779].

On simplicial complexes, one replaces vertices by \(i\)-faces. If \(F,F'\in S_i(K)\) are adjacent whenever they lie in a common \((i+1)\)-face, then the induced graph on \(S_i(K)\) carries a graph distance \(d(F,F')\). A random-walk measure \(m_F\) is placed on neighboring \(i\)-faces, and the curvature is
\[
\kappa(F,F')=1-\frac{W(m_F,m_{F'})}{d(F,F')}.
\]
In the normalized case \(w(\bar F)=1\), Yamada proves upper and lower bounds in terms of \(\deg(F)\), \(\deg(F')\), and common neighbors \(\Gamma(F,F')\), and also derives eigenvalue estimates for the Horak–Jost up-Laplacian:
\[
(i+1)\,k-i\le \lambda\le (i+2)-(i+1)\,k
\]
for every non-trivial eigenvalue \(\lambda\) whenever \(\kappa(F,F')\ge k\) on all adjacent \(i\)-faces [1906.07404].

A semigroup-based alternative is the large scale Ricci curvature on graphs. For a weighted graph \((V,w,m)\), the \(R\)-gradient is
\[
|\nabla_R f|(x)=\max_{y:d(x,y)\le R}|f(y)-f(x)|,
\]
and the Gradient-Ollivier curvature \(K_R(x,y)\) is characterized by the heat-semigroup estimate
\[
|\nabla_R P_t f|\le e^{-Kt}\,P_t|\nabla_R f|.
\]
This hybrid of Ollivier and Bakry–Émery curvature yields Bonnet–Myers diameter bounds, Lichnerowicz eigenvalue estimates, Harnack inequalities, and Buser inequalities, and the hexagonal lattice satisfies \(K_2\ge0\) [1906.06222].

Network science has developed further statistical uses of these ideas. Bakry–Émery–Ricci curvature on vertices is defined as the largest \(K\) such that
\[
\Gamma_2(f)(x)\ge \frac1n\,(\Delta f(x))^2+K\,\Gamma(f)(x)
\]
for all \(f\). Empirically, most vertices in model and real-world networks have negative curvature, Bakry–Émery–Ricci curvature has high positive correlation with both Forman-Ricci and Ollivier-Ricci curvature, it exhibits a high negative correlation with vertex centrality measure and degree, and it does not correlate with the clustering coefficient [2402.06616].

A recent transport-based variant is Sobolev–Ricci Curvature (SRC), defined on a graph by
\[
\kappa_{\mathrm{S}_p}(x,y)=1-\frac{S_p(\mu_x,\mu_y)}{D_p(x,y)},
\]
where \(S_p\) is a Sobolev transport distance on neighborhood measures and \(D_p(x,y)=S_p(\delta_x,\delta_y)\). On trees with length measure and \(p=1\), SRC recovers Ollivier-Ricci curvature, and in the Dirac limit \(\kappa_{\mathrm{S}_p}(x,y)\to0\). SRC is used in Sobolev–Ricci Flow and in curvature-guided edge pruning aimed at preserving manifold structure [2603.12652].

## 6. Measure-valued, local, and semigroup characterizations

Another line of work replaces pointwise tensors by measure-valued curvature. For certain singular torsion-free connections on a \(C^{1,1}\)-manifold, one defines a quadratic form
\[
Q(V,W)=\int_M\bigl[(\mathrm{div}\,V)(\mathrm{div}\,W)-\langle\nabla V,\nabla W\rangle\bigr],
\]
and, under a lower bound on \(Q(V,V)\), obtains a unique measure-valued tensor \(R\) such that
\[
Q(V,W)=\int_M \langle V,RW\rangle.
\]
This \(R\) is the Ricci measure. The framework recovers the smooth Ricci tensor, remains stable under \(L^1\)-perturbations, and supports a weak notion of Ricci flow through an integral Bochner identity [1503.04725].

In nonsmooth metric-measure geometry, Gigli’s measure-valued Ricci tensor is used to characterize \(\mathrm{RCD}(K,\infty)\) spaces locally. Writing
\[
\mathrm{Ricci}(V_f,V_f)=\Gamma_2(f)-|Hess[f]|_{HS}^2\,m,
\]
one decomposes \(\mathrm{Ricci}(V_f,V_f)=\mathrm{Ricci}^{ac}(V_f,V_f)\,m+\mathrm{Ricci}^{sing}(V_f,V_f)\). The conditions
\[
\mathrm{Ricci}^{ac}(V_f,V_f)\ge K\,|Df|^2\quad\text{and}\quad \mathrm{Ricci}^{sing}(V_f,V_f)\ge0
\]
for every \(f\in \mathrm{TestF}\) are equivalent to \(\mathrm{RCD}(K,\infty)\), and the scalar measure
\[
\mathrm{Ricci}_{loc}(f):=|Df|^2\,\mathrm{Ricci}^{ac}(V_f,V_f)\,m
\]
encodes the lower bound locally. The same paper proves that an \(L^p\)-gradient estimate for the heat flow, for some \(p>1\), already forces the \(L^1\)-gradient estimate and hence the synthetic Ricci bound [1702.00740].

A related smooth construction defines a coarse Ricci curvature on \(M\times M\) directly from a diffusion operator \(L\). With
\[
f_{x,y}(z)=\tfrac12\bigl(d^2(x,y)-d^2(y,z)+d^2(z,x)\bigr),
\qquad
\mathrm{Ric}_L(x,y)=\Gamma_2\bigl(L;f_{x,y},f_{x,y}\bigr)(x),
\]
one recovers the classical Ricci tensor via
\[
\mathrm{Ric}\bigl(\gamma'(0),\gamma'(0)\bigr)
=
\frac12\frac{d^2}{ds^2}\Bigl[\mathrm{Ric}_{\Delta_g}\bigl(x,\gamma(s)\bigr)\Bigr]_{s=0}.
\]
Lower bounds \(\mathrm{Ric}_{\Delta_g}(x,y)\ge K\,d^2(x,y)\) are equivalent to \(\mathrm{Ric}\ge Kg\) [1505.04166].

Across these formulations, positive lower bounds control spectral gaps, convergence rates of gradient flows, functional inequalities, and heat-semigroup contraction; negative curvature identifies tree-like or bottleneck structure in graphs and networks [1807.07095] [1602.07779]. A common misconception is that these notions are interchangeable. The literature instead indicates that each construction is tied to a specific geometric mechanism—transport, entropy, diffusion, higher-order adjacency, or dual connections—and its theorems depend on that mechanism’s own metric, Laplacian, or curvature tensor.

Source: https://www.emergentmind.com/topics/statistical-ricci-curvature