---
title: Curvature-Based Covariance Expansion
url: https://www.emergentmind.com/topics/curvature-based-covariance-expansion
type: topic
---

# Curvature-Based Covariance Expansion

Curvature-based covariance expansion denotes a family of constructions in which a covariance, mean-squared error, or covariance-derived spectral object is written as a leading flat or first-order term plus curvature-dependent corrections. In the cited literature, this idea appears in several technically distinct settings: second-order asymptotics of efficient estimators on statistical manifolds, finite-sample expansions for empirical Fréchet means on Riemannian and affine manifolds, local covariance analysis of embedded submanifolds, superlinear relations between SGD noise covariance and loss curvature, and covariance-domain parameterizations of near-field array manifolds through Fresnel curvature [2604.12725][1906.07418][1804.10425][2602.05600][2603.28918]. Taken together, these formulations suggest that the term is best understood as an umbrella for methods that encode intrinsic, extrinsic, or model-specific curvature information directly into covariance structure.

## 1. Second-order covariance corrections in regular statistical models

In regular parametric inference, the classical first-order asymptotic statement is
\[
\Cov_\theta(\hat\theta_n)=\frac1n\,I(\theta)^{-1}+o(n^{-1}),
\]
equivalently
\[
\sqrt n(\hat\theta_n-\theta)\dto \Normal\bigl(0,I(\theta)^{-1}\bigr).
\]
The curvature-based refinement developed in the information-geometric setting assumes a score-root, first-order efficient estimator with stochastic expansion
\[
\hat\theta_n-\theta =\frac1{\sqrt n}A(\theta) +\frac1n\,B(\theta) +o_p(n^{-1}),
\]
and yields
\[
\Cov_\theta(\hat\theta_n)=\frac1n\,I(\theta)^{-1}+\frac1{n^2}\,C(\theta)+o(n^{-2}),\qquad
C(\theta)=I(\theta)^{-1}P(\theta)I(\theta)^{-1}.
\]
The tensor \(P_{ij}\) is the second-order correction tensor; in arbitrary coordinates it is given by the expectation-based normal-coordinate formula built from score moments \(s_i=\partial_i\log p_\theta(X)\), the \(e\)-connection coefficients \(\Gamma^{(e)}_{ijk}=\E[\partial_i\partial_j\log p\;\partial_k\log p]\), and the third-order log-density moment \(\kappa_{ijk}=\E[\partial_i\partial_j\partial_k\log p]\) [2604.12725].

A central structural result is the canonical splitting
\[
P_{ij}=\tfrac12\,R^\sharp_{ij}+S^\sharp_{ij}+D_{ij}.
\]
Here \(R^\sharp_{ij}=g^{k\ell}R_{ikj\ell}\) is an intrinsic Ricci-type contraction of the Fisher–Rao curvature tensor, \(S^\sharp_{ij}=g^{k\ell}\langle \II_{ik},\II_{j\ell}\rangle_{L^2}\) is an extrinsic Gram-type contraction of the second fundamental form of the square-root density immersion, and \(D_{ij}\) is a Hellinger discrepancy tensor collecting higher-order probabilistic information not determined by the immersion geometry alone. The extrinsic contribution is positive semidefinite, since
\[
v^i v^jS^\sharp_{ij}=\sum_k\|v^i\II_{ik}\|_{L^2}^2\ge 0.
\]

The same framework proves that \(R^\sharp\), \(S^\sharp\), and \(D\) are coordinate-invariant \((0,2)\)-tensors, so the full \(n^{-2}\) correction \(I^{-1}PI^{-1}\) is invariant under smooth reparameterization. In a full exponential family, \(R\equiv 0\) and the probabilistic discrepancy \(D\) exactly cancels \(S^\sharp\), so \(P_{ij}\equiv 0\). In that case the curvature-based correction vanishes identically and the covariance attains the classical Cramér–Rao behavior up to \(o(n^{-2})\) [2604.12725].

## 2. Resolved-space expansions in singular models

When Fisher information degenerates, the regular expansion above no longer applies directly. Under the additive normal-crossing assumption, a resolution of singularities produces a real-analytic proper map
\[
\pi:\widetilde\Theta\to\Theta
\]
such that in local resolved coordinates \(u=(u_1,\dots,u_r)\) near a singular point \(\theta_0\),
\[
K(\pi(u))=\sum_{j=1}^r c_j\,u_j^{2k_j},\qquad c_j>0,\ k_j\in\mathbb N,
\]
and
\[
|\det D\pi(u)|=\prod_{j=1}^r |u_j|^{h_j}\,\varphi(u),\qquad \varphi(0)\neq 0.
\]
Pulling back the KL-Hessian yields the diagonal resolved metric
\[
\widetilde g_{jj}(u)=2k_j(2k_j-1)\,c_j\,u_j^{2k_j-2},
\]
which is degenerate on \(\{u_j=0\}\) when \(k_j\ge 2\) [2604.12725].

The real log canonical threshold is
\[
\lambda=\sum_{j=1}^r \frac{h_j+1}{2k_j},
\]
and it governs the marginal-likelihood asymptotics
\[
\int e^{-nK(\theta)}\varphi(\theta)\,d\theta \sim C\,n^{-\lambda}.
\]
Under the same additive form, the posterior mean-squared error decays like
\[
\E_{\rm post}\|\theta-\theta_0\|^2\sim C'\,n^{-\min_j(1/k_j)}.
\]
On the regular part of the resolved space, the curvature-corrected covariance expansion becomes
\[
\Cov_{\theta_0}(\hat\theta_n)\sim n^{-2\lambda}
\Bigl[\widetilde g^{-1}+n^{-\mu}\Bigl(\tfrac12\,\widetilde R^\sharp+\widetilde S^\sharp\Bigr)\Bigr],
\qquad
\mu=\min_j \frac1{k_j}.
\]
This recovers the regular theory when \(k_j=1\) and \(h_j=0\) for all \(j\) [2604.12725].

The same paper interprets these resolved-space terms diagnostically. Directions in which \(g(\theta)\) is small or vanishing are first-order degeneracies, whereas directions in which \(R^\sharp\) or \(S^\sharp\) is large indicate rapid accumulation of curvature-induced error at second order. It further proposes curvature-aware regularization through penalties such as \(\operatorname{tr}(I^{-1}PI^{-1})\), \(\operatorname{tr}(R^\sharp)\), \(\operatorname{tr}(S^\sharp)\), and \(\|\II\|^2\), and a formal curvature-refined natural-gradient preconditioner \(I^{-1}(I+n^{-1}P)^{-1}\) [2604.12725].

## 3. Small-sample covariance of empirical Fréchet means

In Riemannian and torsion-free affine manifolds, the empirical Fréchet mean admits a non-asymptotic high-concentration expansion that makes curvature effects explicit at finite sample size. For a probability measure \(\mu\) supported in a convex neighborhood, the population mean \(\bar x\) is defined by
\[
\int_M \log_{\bar x}(y)\,\mu(dy)=0.
\]
A key input is the Taylor expansion of the neighboring log
\[
l_x(v,w)=\Pi_{x_v\to x}\log_{x_v}(x_w),
\]
namely
\[
\begin{aligned}
l_x(v,w) &= w-v+\tfrac16\,R(w,v)(v-2w) \\
&\quad +\tfrac1{24}\,(\nabla_vR)(w,v)(2v-3w)
+\tfrac1{24}\,(\nabla_wR)(w,v)(v-2w)+O(\varepsilon^5),
\end{aligned}
\]
which makes the curvature and curvature-gradient contributions explicit in the log-map under a shift of foot-point [1906.07418].

From this expansion, the first-moment bias and the second-moment correction of the empirical mean \(\bar x_n\) follow. The bias is
\[
\mathbb E[\log_{\bar x}(\bar x_n)]
=\frac{1}{6n}\Bigl(1-\frac1n\Bigr)\bigl[\,{}_2\nabla R(\cdot,\cdot)\,\bigr]\,{}_2+O(\varepsilon^5),
\]
or in coordinates
\[
\mathrm{Bias}^a
=\frac{1}{6n}\Bigl(1-\frac1n\Bigr)\nabla_bR^a{}_{cde}\,\Sigma^{cd}\,\Sigma^{be}+O(\varepsilon^5).
\]
The covariance expansion is
\[
\Cov\bigl(\log_{\bar x}(\bar x_n)\bigr)
=\frac1n\Bigl(
\Sigma-\tfrac13\Bigl(1-\tfrac1n\Bigr)
\bigl[\Sigma^{cd}\bigl(\Sigma^{ae}R^b{}_{cde}+\Sigma^{be}R^a{}_{cde}\bigr)\bigr]_{ab}
\Bigr)+O(\varepsilon^5).
\]
Thus the \(1/n\) Euclidean scaling persists, but curvature modulates the prefactor through contractions of the population covariance \(\Sigma\) with the curvature tensor [1906.07418].

For constant sectional curvature \(\kappa\), \(\nabla R\equiv 0\), so the \(1/n\)-bias vanishes to order four. In the isotropic case \(\Sigma=\frac{\sigma^2}{d}g^{ab}\), the variance becomes
\[
\mathrm{Var}(\bar x_n)
=\frac{\sigma^2}{n}
\Bigl[
1+\frac{2\kappa\sigma^2}{3}\Bigl(1-\frac1d\Bigr)\Bigl(1-\frac1n\Bigr)
\Bigr]
+O(\varepsilon^5).
\]
The sign of \(\kappa\) determines the qualitative effect: negative curvature accelerates convergence of the mean, while positive curvature retards it. The expansion is consistent with the Bhattacharya–Patrangenaru central limit theorem through the expansion of the expected Hessian of the squared distance, and its validity requires the small-diameter regime together with KKC conditions in the Riemannian case or ALC conditions in the affine case [1906.07418].

## 4. Local covariance analysis of embedded submanifolds

A different but closely related use of curvature-based covariance expansion arises in geometric inference on embedded submanifolds \(M^n\subset \mathbb R^{n+k}\). The construction begins with local domains around a point \(p\in M\): the spherical cap
\[
D_p(\varepsilon)=\{q\in M:\|\exp_p^{-1}(q)\|\le \varepsilon\}
\]
and the tangent-cylinder
\[
\Cyl_p(\varepsilon)=\{q\in M:\|\proj_{T_p}(\exp_p^{-1}(q))\|\le \varepsilon\}.
\]
Their covariance matrices are computed from the embedded coordinates \(X=\exp_p^{-1}(q)\), either about the barycenter for \(D_p(\varepsilon)\) or about the base point for \(\Cyl_p(\varepsilon)\). The associated volume asymptotics are
\[
\Vol(\Cyl_p(\varepsilon))
=V_n(\varepsilon)\Bigl[1+\frac{\varepsilon^2}{2(n+2)}\,\tr+O(\varepsilon^4)\Bigr],
\]
\[
\Vol(D_p(\varepsilon))
=V_n(\varepsilon)\Bigl[1+\frac{\varepsilon^2}{8(n+2)}\bigl(2\,\tr-\|H\|^2\bigr)+O(\varepsilon^3)\Bigr],
\]
with \(\tr=\|H\|^2-R\) in Euclidean ambient space, where \(H\) is the mean-curvature vector and \(R\) the scalar curvature [1804.10425].

The curvature tensors entering these expansions are organized through a generalized third fundamental form \(\III\), defined for general codimension by
\[
\langle \III(x,y)n,m\rangle = \langle {}_m x,\ {}_n y\rangle,
\]
where \({}_n:T_p\to T_p\) is the Weingarten map associated with the normal vector \(n\). Its two natural traces are \(\operatorname{tr}_\perp\III\in (T_p^*)^2\) and \(\operatorname{tr}_\parallel\III\in \operatorname{End}(N_p)\), and one has
\[
\operatorname{tr}\III=\|H\|^2-R
\qquad
(\text{in }\mathbb R^{n+k}\text{ ambient}).
\]
The covariance eigenvalues then admit asymptotic expansions whose scale separation encodes tangential and normal curvature information [1804.10425].

In the cylinder case, the first \(n\) eigenvalues scale like \(\varepsilon^2\) and split at order \(\varepsilon^4\), while the last \(k\) eigenvalues scale like \(\varepsilon^4\). In the sphere case, the tangential eigenvalues involve the Weingarten map in the direction \(H\), and the normal eigenvalues involve \(\operatorname{tr}_\parallel\III-\frac1{n+2}H\otimes H\). As \(\varepsilon\to 0\), the corresponding eigenvectors converge to the principal directions of these operators [1804.10425].

For hypersurfaces \(M^n\subset \mathbb R^{n+1}\), the construction recovers the principal curvatures \(\kappa_1,\dots,\kappa_n\) and principal directions \(e_1,\dots,e_n\). The paper gives explicit curvature-recovery formulas at scale \(\varepsilon\) for \(H_{\rm est}^2\) and \(\kappa_{\mu,\rm est}^2\) in terms of \(\Vol\), \(\lambda_{n+1}\), and \(\lambda_\mu\), both for spherical descriptors and cylindrical descriptors. In this setting, covariance analysis serves directly as a local estimator of the second fundamental form and, consequently, of the Riemann tensor of general submanifolds [1804.10425].

## 5. Noise–curvature covariance relations in stochastic optimization

In deep learning, curvature-based covariance expansion appears in the analysis of SGD noise. Let \(\ell_p(\mathbf w)\) be the per-sample loss, \(\mathbf h_p(\mathbf w)=\nabla^2\ell_p(\mathbf w)\) the per-sample Hessian, and \(\mathbf H(\mathbf w)=\mathbb E_p[\mathbf h_p(\mathbf w)]\) the average Hessian. For batch size \(B\), the SGD noise covariance is
\[
\mathbf C(\mathbf w)
=\frac1B\Bigl(\mathbb E_p[\nabla \ell_p\,\nabla \ell_p^\top]-\nabla\mathcal L\,\nabla\mathcal L^\top\Bigr)
\approx \frac1B\,\mathbb E_p[\nabla \ell_p\,\nabla \ell_p^\top]
\qquad (\|\nabla\mathcal L\|\approx 0).
\]
Using Activity–Weight Duality, a matched-pair construction between consecutive minibatches, and a Taylor expansion of gradients, the leading-order result is
\[
\mathbf C \approx \frac{\sigma_w^2}{2B}\,\mathbb E_p[\mathbf h_p^{\,2}],
\]
under the empirical isotropic-perturbation assumption for the local covariance of the minimal dual-weight perturbations \(\Delta \mathbf w_p\) in each sample’s local Hessian eigenbasis [2602.05600].

This formula replaces the common identification \(\mathbf C\propto \mathbf H\) by a more general second-moment relation in the per-sample Hessians. In the eigenbasis of the full-batch Hessian \(\mathbf H\mathbf v_i=H_{ii}\mathbf v_i\), the covariance matrix is approximately diagonal, because the off-diagonal terms \(C_{ij}\) with \(i\neq j\) vanish in high dimension under the approximate independence of the projections \(\mathbf u_m^{(p)}\cdot \mathbf v_i\) and \(\mathbf u_m^{(p)}\cdot \mathbf v_j\). Consequently,
\[
[\mathbf C,\mathbf H]\approx 0.
\]
The theory further yields an approximate power law
\[
C_{ii}\propto H_{ii}^{\,\gamma},
\qquad 1\le \gamma\le 2,
\]
with \(\gamma\) determined by the joint spectrum of the per-sample Hessian eigenvalues and their alignments with the global Hessian eigenvectors [2602.05600].

The experimental evidence reported in the paper is consistent with this picture. Projection of the empirical SGD covariance onto the top \(100\) eigenvectors of the full-batch Hessian produces an “arrowhead” pattern, and after normalization by the diagonal nearly all off-diagonals vanish. Log–log plots of \(C_{ii}\) against \(H_{ii}\) for the top \(\sim 10^3\) directions show a straight-line relation. Measured exponents \(\gamma_{\mathrm{emp}}\) range approximately \(1.3\text{–}1.5\) for cross-entropy and approximately \(1.0\) for MSE, and suppression experiments attribute \(\gamma>1\) to positive coupling between large per-sample curvatures and their alignment with global eigenvectors [2602.05600].

## 6. Covariance-domain curvature parameterization in hybrid near-field arrays

In hybrid near-field MIMO channel estimation, “curvature” refers not to Riemannian curvature but to the quadratic phase term of the Fresnel steering law. For a path at \((\theta,r)\), the near-field steering vector under the Fresnel approximation is
\[
a_{\rm NF}(\theta,r)_m \simeq \exp\bigl[j\,\omega(\theta)\,\bar m-j\,\kappa(\theta,r)\,\bar m^2\bigr],
\]
with
\[
\omega(\theta)=\frac{2\pi}{\lambda}d_{\rm ant}\cos\theta,
\qquad
\kappa(\theta,r)=\frac{\pi d_{\rm ant}^2}{\lambda r}\sin^2\theta.
\]
Thus \(\theta\) controls the linear phase slope \(\omega\) and \(r\) controls curvature \(\kappa\). After hybrid combining by \(W\in \mathbb C^{M\times N_{\rm RF}}\), the natural sufficient statistic is the compressed sample covariance
\[
\widehat R_y=\frac1N\sum_{n=1}^N y(n)y(n)^H\in \mathbb C^{N_{\rm RF}\times N_{\rm RF}},
\]
and the modeled compressed covariance is
\[
R_y=\sum_{\ell=1}^d p_\ell\,d(\theta_\ell,r_\ell)d(\theta_\ell,r_\ell)^H+N_0W^H W,
\qquad
d(\theta,r)=W^Ha_{\rm NF}(\theta,r)
\]
[2603.28918].

The Curvature-Learning KL estimator grids only the angle dimension and learns the per-angle inverse range \(u_i=1/r_i\) directly from the compressed covariance. For an angle grid \(\Theta=\{\theta_1,\dots,\theta_{Q_\theta}\}\) and \(c_i=(\pi d_{\rm ant}^2/\lambda)\sin^2\theta_i\), the dictionary atoms are
\[
a_i(u_i)_m=\exp[j\,\omega(\theta_i)\,\bar m-j\,c_i u_i \bar m^2],
\qquad
d_i(u_i)=W^H a_i(u_i).
\]
The KL-based objective is
\[
\mathcal L=\log\det R_y+\operatorname{tr}(R_y^{-1}\widehat R_y)+\lambda\|p\|_1,
\]
with optimization variables \(p\ge 0\), \(u\in [u_{\min},u_{\max}]^{Q_\theta}\), and \(N_0\ge 0\). The reduction from a \(Q_\theta Q_r\)-atom polar dictionary to a \(Q_\theta\)-element dictionary eliminates the range-dimension dictionary coherence that plagues polar codebooks in the strong near-field regime [2603.28918].

The implemented CL-KL loop freezes the noise floor
\[
\widehat N_0=\text{mean of the smallest }(N_{\rm RF}-d)\text{ eigenvalues of }(\widehat R_y+\widehat R_y^H)/2,
\]
uses three warm starts, updates the powers by projected gradient with Armijo backtracking, and then performs a four-pass global matched-filter scan alternating between fine angular refinement and \(u\)-refinement. Each Phase 1 iteration costs
\[
O(N_{\rm RF}^3+Q_\theta N_{\rm RF}^2),
\]
with the dominant operation an \(N_{\rm RF}\times N_{\rm RF}\) inversion. The measured runtime is approximately \(70\) ms per trial for \(M=64\), \(N_{\rm RF}=8\), \(N=64\), and \(d=3\), and it remains nearly constant across \(M\in\{32,64,128,256\}\) because the dominant cost is \(N_{\rm RF}^3\), not \(M\) [2603.28918].

At \(N_{\rm MC}=400\) with \(f_c=28\) GHz, \(M=64\), \(N_{\rm RF}=8\), \(N=64\), \(d=3\), and \(r\in [0.05,1.0]\,r_{\rm RD}\), CL-KL achieves the lowest channel NMSE among all six evaluated methods at \(\mathrm{SNR}\in\{-5,0,+5,+10\}\) dB, including four full-array baselines using \(64\times\) more data. At \(\mathrm{SNR}=+10\) dB, the reported channel NMSE is approximately \(-5.16\) dB for CL-KL, approximately \(-4.06\) dB for P-SOMP, and approximately \(-4.80\) dB for full-array DL-OMP. The method is further validated against a compressed-domain Cramér–Rao bound and is reported robust to non-Gaussian QPSK sources with a maximum NMSE gap below \(0.6\) dB [2603.28918].

Across these settings, curvature-based covariance expansion has no single universal formula. In regular information geometry it is an \(n^{-2}\) refinement of Fisher asymptotics; in singular learning it is a resolved-space correction controlled by the RLCT; in manifold statistics it is a finite-sample modulation of empirical-mean covariance by \(R\) and \(\nabla R\); in submanifold geometry it is an asymptotic expansion of local covariance eigenstructure that recovers principal curvatures and directions; in SGD theory it is a second-moment relation \(\mathbf C\propto \mathbb E[\mathbf h_p^2]\); and in near-field signal processing it is a covariance-domain expansion in Fresnel curvature. The common theme is that covariance is not treated as a purely second-order Euclidean object, but as a carrier of curvature information intrinsic to the model, embedding, or propagation law.

Source: https://www.emergentmind.com/topics/curvature-based-covariance-expansion