---
title: 'Lindeberg Replacement Method: Comparison Technique'
url: https://www.emergentmind.com/topics/lindeberg-replacement-method
type: topic
---

# Lindeberg Replacement Method: Comparison Technique

The Lindeberg replacement method is a comparison technique for proving distributional approximation and universality by successively replacing individual coordinates, summands, layers, or columns of a random object with better understood counterparts and controlling the cumulative error through a telescoping decomposition. In the formulations developed for smooth observables, the basic mechanism is a coordinatewise Taylor expansion; in other settings it is coupled with Kolmogorov forward equations, Stein kernels, Gaussian smoothing, or representation-theoretic decompositions. The method appears in generalized form for smooth functions of independent coordinates [1004.0557], in stable approximation without Fourier inversion [1809.10864], in total-variation versions of the central limit theorem [2511.02391], in quantitative universality for deep neural networks [2605.02771], and in column-swapping arguments for random matrices over finite local rings [2601.10821].

## 1. General scheme and abstract formulation

A standard formulation begins with two random vectors \(U=(U_1,\dots,U_n)\) and \(V=(V_1,\dots,V_n)\) in \(\mathbb{R}^n\), each with independent components, and a smooth test function \(f:\mathbb{R}^n\to\mathbb{R}\). In the generalized principle stated in "Applications of Lindeberg Principle in Communications and Statistical Learning" [1004.0557], one defines
\[
a_i=\bigl|\mathbb{E}[U_i]-\mathbb{E}[V_i]\bigr|,\qquad
b_i=\bigl|\mathbb{E}[U_i^2]-\mathbb{E}[V_i^2]\bigr|,
\]
assumes \(\max_i\{\mathbb{E}|U_i|^3+\mathbb{E}|V_i|^3\}\le M_3<\infty\), and requires bounded first, second, and third coordinate derivatives:
\[
\bigl|\partial_i^r f(x)\bigr|\le L_r(f)\qquad \forall x\in\mathbb{R}^n,\;i=1,\dots,n,\;r=1,2,3.
\]
Under these hypotheses,
\[
\bigl|\mathbb{E}[f(U)]-\mathbb{E}[f(V)]\bigr|
\le
\sum_{i=1}^n\Bigl(a_iL_1(f)+\tfrac12 b_iL_2(f)\Bigr)
+\tfrac16 nL_3(f)M_3.
\]

The corresponding “local” version replaces global derivative bounds by expectations of derivatives evaluated along mixed vectors \((U_{1:i-1},s,V_{i+1:n})\), together with integral remainders involving \(\partial_i^3 f\). This variant is designed for high-dimensional applications in which uniform control of all \(\partial_i^r f\) over \(\mathbb{R}^n\) is unavailable [1004.0557].

The core identity is the telescoping sum
\[
\mathbb{E}[f(U)]-\mathbb{E}[f(V)]
=
\sum_{i=1}^n
\bigl[
\mathbb{E}\,f(U_1,\dots,U_i,V_{i+1},\dots,V_n)
-
\mathbb{E}\,f(U_1,\dots,U_{i-1},V_i,\dots,V_n)
\bigr].
\]
Each summand isolates a one-coordinate replacement. A Taylor expansion in the \(i\)-th coordinate then separates the effect of mean mismatch, variance mismatch, and the third-order remainder. This suggests that the replacement method is best understood not as a single theorem but as a modular proof pattern: telescoping plus a one-step comparison estimate.

## 2. Taylor expansion, moment matching, and cancellation

In the smooth setting, the method operates by expanding the observable in the coordinate being replaced. The proof sketch in [1004.0557] gives the basic form:
\[
f(W_{i-1}\cup\{s=U_i\})
=
f(W_i|_{s=0})
+ U_i\,\partial_i f|_{s=0}
+ \tfrac12 U_i^2\,\partial_i^2 f|_{s=0}
+
\int_0^{U_i}\tfrac{(U_i-s)^2}{2}\,\partial_i^3 f|_s\,ds,
\]
and similarly for \(V_i\). Because \(W_i|_{s=0}\) is independent of \(\{U_i,V_i\}\), expectations factorize, producing the explicit contributions of the first two moments and a remainder term controlled by third derivatives and third moments.

When the first two moments match, the leading terms cancel and only the third-order remainder remains. This is the mechanism emphasized in several later applications. In the deep-network setting, for example, the layerwise weights in the original and Gaussian networks have matching first two moments, so the first two Taylor terms cancel and the remainder is \(O(n_\ell^{-3/2})\), which after summing over \(n_\ell\) terms gives \(O(n_\ell^{-1/2})\) per layer [2605.02771]. In that sense, moment matching converts a qualitative exchange argument into a quantitative one.

The same structural idea persists when the one-step estimate is not written as an ordinary Euclidean Taylor formula. In stable approximation, the expansion is “Taylor-like” and the compensating term is expressed through the generator \(\mathcal L^{\alpha,\beta}\) of the stable Lévy process rather than the Gaussian second derivative [1809.10864]. In the total-variation central limit theorem, the object being replaced is a Stein kernel contribution \(\tau_k(X_k)\) by its mean \(\sigma_k^2\), and the telescoping is carried out inside Stein’s identity [2511.02391]. A plausible implication is that the decisive feature is not the literal form of the expansion, but the existence of a one-step approximation whose leading part is additive and whose error is summable.

## 3. Stable-law approximation and the generator method

"Approximation to the stable law by Lindeberg principle" develops the method for one-dimensional possibly asymmetric \(\alpha\)-stable laws with \(\alpha\in(0,2)\) [1809.10864]. For fixed \(\alpha\in(0,2)\), \(\beta\in[-1,1]\), and when \(\alpha=1\) with \(\beta=0\), a real random variable \(Y\) has the law \(Sa(1,\beta)\) if its characteristic function is
\[
\forall t\in\mathbb{R},\quad
\mathbb{E}e^{itY}
=
\exp\!\Big\{-|t|^\alpha\big(1-i\beta\,\sgn(t)\tan\tfrac{\pi\alpha}{2}\big)\Big\}
\quad(\alpha\neq1),
\]
with the usual log-correction form when \(\alpha=1,\beta=0\). Equivalently, \(Y\) is the one-dimensional Lévy process at time one with generator
\[
\mathcal{L}^{\alpha,\beta}f(x)
=
d_\alpha\int_{\mathbb{R}}\Big[f(x+y)-f(x)-y\,\mathbf1_{|y|<1}f'(x)\Big]
\frac{(1+\beta)\mathbf1_{y>0}+(1-\beta)\mathbf1_{y<0}}{2|y|^{1+\alpha}}\,dy.
\]

The comparison is performed in the smooth Wasserstein distance. For \(k\in\mathbb{N}\),
\[
H_k
=
\Big\{h\in C^k(\mathbb{R})\;:\;\|h^{(j)}\|_\infty\le1,\;j=0,1,\dots,k\Big\},
\]
and
\[
d_{W_k}(\mu,\nu)
=
\sup_{h\in H_k}\Big|\int h\,d\mu-\int h\,d\nu\Big|.
\]
The random variables \(X_i\) are assumed to lie in the domain of normal attraction described by
\[
F_X(x)
=
\begin{cases}
1-\displaystyle\frac{A+\varepsilon(x)}{x^\alpha}(1+\beta),&x>0,\\[1ex]
\displaystyle\frac{A+\varepsilon(x)}{|x|^\alpha}(1-\beta),&x<0,
\end{cases}
\]
with \(\varepsilon\) bounded and vanishing at infinity. The normalized sums are
\[
S_n=n^{-1/\alpha}\Big(X_1+\cdots+X_n-c_n\Big),
\]
where \(c_n=n\,\mathbb{E}[X\mathbf1_{|X|<n^{1/\alpha}}]\) if \(\alpha>1\), and \(c_n=0\) otherwise.

The replacement framework interpolates between the normalized sum of the \(X\)'s and an i.i.d. sum of \(Y\)'s. With suitable \(Z_j\),
\[
\mathbb{E}[f(S_n)]-\mathbb{E}[f(Y)]
=
\sum_{j=1}^n
\Big\{
\mathbb{E}\bigl[f\bigl(n^{-1/\alpha}X_j+Z_j\bigr)\bigr]
-
\mathbb{E}\bigl[f\bigl(n^{-1/\alpha}Y_j+Z_j\bigr)\bigr]
\Big\}.
\]
For the \(X_j\)-term, Lemma 3.9 gives, when \(1<\alpha<2\),
\[
\mathbb{E}[f(Z+aX)]-\mathbb{E}[f(Z)]-a\,\mathbb{E}[X]\,\mathbb{E}[f'(Z)]
-2A\alpha\,a^{\alpha-1}\,\mathbb{E}\bigl[\mathcal L^{\alpha,\beta}f(Z)\bigr]
=:R_1,
\]
with a remainder controlled by \(a^2\), a tail integral involving \(|\varepsilon(x)|\), and \(\sup_{|x|\ge a^{-1}}|\varepsilon(x)|\). For the \(Y_j\)-term, the Kolmogorov forward equation
\[
\mathbb{E}\bigl[f(Z+Y_t)-f(Z)\bigr]
=
\int_0^t\mathbb{E}\bigl[\mathcal L^{\alpha,\beta}f(Z+Y_s)\bigr]\,ds
\]
produces compensating generator terms. The paper states that in the telescoping sum the forward-equation terms exactly cancel the compensating \(\mathbb{E}[\mathcal L f(Z)]\)-terms from the Taylor-like expansion [1809.10864].

The resulting rates are explicit. Under \(|\varepsilon(x)|\le K\) for large \(|x|\), there is \(C=C(\alpha,A,K)\) such that for every \(f\in C^3(\mathbb{R})\),
\[
\bigl|\mathbb{E}[f(S_n)]-\mathbb{E}[f(Y)]\bigr|
\le
C\bigl(\|f\|_\infty+\|f'\|_\infty+\|f''\|_\infty+\|f^{(3)}\|_\infty\bigr)\times R_n.
\]
In particular,
\[
d_{W_3}(\mathcal L(S_n),Sa(1,\beta))=O\bigl(n^{-(2-\alpha)/\alpha}\bigr)
\quad(1<\alpha<2),
\]
with the paper also giving \(O(n^{-1}\log n)\) for \(\alpha=1\) and \(O(n^{-1})\) for \(0<\alpha<1\). The paper further states that it is the first time that the general stable central limit theorem is proved by the Lindeberg principle, and that the theorem with \(\alpha\in(0,1]\) is proved by a new method other than Fourier analysis [1809.10864].

## 4. Central-limit theorems in total variation

"Total variation bounds in the Lindeberg central limit theorem" reworks the replacement idea in a non-i.i.d. Gaussian approximation problem, but with total-variation distance rather than a smooth test-function metric [2511.02391]. Let
\[
S_n=\frac1{b_n}\sum_{k=1}^n X_k,\qquad
b_n^2=\sum_{k=1}^n \sigma_k^2,\qquad
\sigma_k^2=\mathbb{E}[X_k^2],
\]
where the \(X_k\) are independent, centered, and absolutely continuous. The total-variation distance from \(N\sim N(0,1)\) is
\[
d_{TV}(S_n,N)
:=
\tfrac12\sup_{\|h\|_\infty\le1}\bigl|\mathbb{E}[h(S_n)]-\mathbb{E}[h(N)]\bigr|.
\]

The theorem assumes finite standardised Fisher information
\[
J_{st}(X_k)=\sigma_k^2\int_{-\infty}^\infty \frac{p_k'(x)^2}{p_k(x)}\,dx-1,
\]
sets \(J(X_k)=1+J_{st}(X_k)\), \(J_{\max,n}=\max_{1\le k\le n}J(X_k)\), and \(\delta_n=\max_{1\le k\le n}\sigma_k^2/b_n^2\). Under these hypotheses,
\[
d_{TV}(S_n,N)
\le
2\,
\exp\!\Bigl\{
\tfrac12\,
\frac{\sum_{k=1}^n \mathbb{E}\bigl[|X_k|^2(b_n\wedge|X_k|)\bigr]}{b_n^3}
\log\!\Bigl(\frac{8\pi\,J_{\max,n}}{1-\delta_n}\Bigr)
\Bigr\}.
\]
Equivalently,
\[
d_{TV}(S_n,N)
\le
\Biggl(
\frac{8\pi\,J_{\max,n}}{1-\delta_n}
\Biggr)^{\frac12\,
\frac{\sum_{k=1}^n \mathbb{E}[|X_k|^2(b_n\wedge|X_k|)]}{b_n^3}}.
\]

The proof combines Stein’s equation with a Lindeberg-style telescoping. For each bounded \(h\), \(f\) solves
\[
f'(x)-x\,f(x)=h(x)-\mathbb{E}[h(N)],
\]
and Lemma 2.1 yields \(\|f\|_\infty\le\sqrt{2\pi}\) and \(\|f'\|_\infty\le4\). Each \(X_k\) is equipped with a Stein kernel \(\tau_k(X_k)\), and the replacement step consists of replacing \(\tau_k(X_k)\) by its mean \(\sigma_k^2\). Inserting this into Stein’s identity produces
\[
\mathbb{E}[h(S_n)]-\mathbb{E}[h(N)]
=
\frac1{b_n^2}\sum_{k=1}^n
\mathbb{E}\bigl[f'(S_n)(\sigma_k^2-\tau_k(X_k))\bigr].
\]
The paper explicitly describes this as the classical telescoping argument of Lindeberg, but carried out in the language of Stein kernels [2511.02391].

A central consequence is an exact equivalence statement under bounded Fisher information: if \(\sup_k J(X_k)<\infty\), then the usual Lindeberg condition
\[
\frac1{b_n^2}\sum_{k=1}^n
\mathbb{E}\bigl[X_k^2\mathbf1\{|X_k|>\varepsilon b_n\}\bigr]\to0
\qquad \forall \varepsilon>0,
\]
together with the Feller condition \(\max_k \sigma_k^2/b_n^2\to0\), is equivalent to \(d_{TV}(S_n,N)\to0\) [2511.02391]. The paper therefore states that, under suitable assumptions, Lindeberg’s condition is sufficient and necessary for convergence in total variation.

## 5. Layerwise exchange in deep neural networks

In "Universality in Deep Neural Networks: An approach via the Lindeberg exchange principle," the method is adapted to the infinite-width limit of fully connected deep networks with general weights [2605.02771]. The setting fixes depth \(L\) and widths \(n_0,\dots,n_{L+1}\), and defines switched-layers networks \(z^{(L+1;L-k)}(x)\) whose first \(L-k\) layers use the original weights \(W^{(\ell)}\) and whose last \(k\) layers use independent Gaussian matrices \(G^{(\ell)}\) with the same variance \(C_W/n_{\ell-1}\).

The core weak comparison theorem states that if
\[
\sigma\in C_b^{3\cdot 2^{L-1}}(\mathbb{R}),
\]
and if \(z^{(L+1)}\) and \(\tilde z^{(L+1)}\) have the same architecture, same biases, and same activation \(\sigma\), but the hidden weights in \(z^{(L+1)}\) satisfy Assumption 2.1 while those in \(\tilde z^{(L+1)}\) are Gaussian with matching first two moments, then for every
\[
F\in C_b^{3\cdot2^{L-1}}(\mathbb{R}^{n_{L+1}})
\]
one has
\[
\bigl|\mathbb{E}[F(z^{(L+1)}(x))]-\mathbb{E}[F(\tilde z^{(L+1)}(x))]\bigr|
\le
C\sum_{\ell=1}^L \frac1{\sqrt{n_\ell}}.
\]
If instead one assumes Gaussian biases and only \(\sigma\in C_b^3(\mathbb{R})\), then the same bound holds for any bounded \(F\), or for \(F\in C_b^1\) with error controlled by \(\|\nabla F\|_\infty\) [2605.02771].

The telescoping decomposition is layerwise:
\[
\Delta_k
=
\mathbb{E}\bigl[F(z^{(L+1;L-k)}(x))\bigr]
-
\mathbb{E}\bigl[F(z^{(L+1;L-k-1)}(x))\bigr],
\qquad
\mathbb{E}[F(z^{(L+1)}(x))]-\mathbb{E}[F(\tilde z^{(L+1)}(x))]
=
\sum_{k=1}^L \Delta_k.
\]
Conditioning on the pre-activations of the relevant layer, one views the next layer as a sum of independent random variables and applies a conditional Taylor expansion up to third order plus moment matching to obtain
\[
|\Delta_k|\lesssim \frac1{\sqrt{n_{L-k+1}}}.
\]
The paper describes each layer \(\ell\) as contributing a “replacement error”
\[
R_\ell=\frac1{\sqrt{n_\ell}},
\]
and these add up linearly [2605.02771].

After all layers are replaced, one compares the fully Gaussian network to the Gaussian-process limit in \(W_2\). Under Assumptions 2.1, the stated activation regularity, and invertibility of the limiting covariances \(K^{(\ell)}(x)\),
\[
W_2\bigl(z^{(L+1)}(x),\mathcal Z^{(L+1)}(x)\bigr)
\le
C\Bigl(n_L^{-1/2}+\bigl(\sum_{k=1}^{L-1} n_k^{-1/2}\bigr)^{1/4}\Bigr).
\]
If all hidden layers have the same width \(n\), then
\[
W_2\bigl(z^{(L+1)}(x),\mathcal Z^{(L+1)}(x)\bigr)=O(n^{-1/4}),
\]
uniformly in \(L\) and \(x\) up to the constant \(C\) [2605.02771]. The role of Gaussian biases is explicitly that their smoothing effect reduces the derivative requirement on \(\sigma\) from \(3\cdot2^{L-1}\) bounded derivatives to only \(3\).

## 6. Universality, sparse limits, and algebraic variants

The method is used in [1004.0557] to prove universality properties in communications, statistical learning, and random matrix theory, and also to show that dense systems can be viewed as limits of properly defined sparse systems. The paper states universality and sparse–dense equivalence results for CDMA capacity, LASSO, Wishart spectra, and MIMO capacity. In each case the proof strategy is to express the quantity of interest as \(\mathbb{E}f(A)\) for a smooth \(f\), verify bounds on its third derivatives, and apply the modified Lindeberg theorem. For example, for standard-type spreading matrices \(A,B\),
\[
\lim_{n\to\infty}
\Bigl[
\mathbb{E}\,C_n(n^{-1/2}A)-\mathbb{E}\,C_n(n^{-1/2}B)
\Bigr]=0,
\]
and analogous convergence statements are given for LASSO and for the Stieltjes transform of the Wishart ensemble [1004.0557].

The same paper formulates a sparse-limit viewpoint by defining a \(\gamma\)-sparsification \(A^\gamma\) through
\[
A^\gamma_{ij}=
\begin{cases}
A_{ij},&\text{w.p. }\gamma/n,\\
0,&\text{w.p. }1-\gamma/n,
\end{cases}
\quad
\text{then rescale by }1/\sqrt{\gamma}.
\]
It then states that if the third partial derivatives of \(f(A)\) are \(O(1/n)\), replacing a nonzero by an independent Gaussian or another i.i.d. entry induces an error of order \(O(\gamma^{-3/2}n^{-1/2})\), so over \(O(\gamma n)\) nonzeros the total error is \(O(\gamma^{-1/2})\) [1004.0557]. This suggests a form of universality obtained by first passing to analytically simpler sparse models and then returning to dense ones.

"Universality results for random matrices over finite local rings" shows that the replacement method also has a discrete and algebraic form [2601.10821]. Here the random object is a matrix \(M_{n,n+u}\) over a finite local ring \(R\), and the comparison matrix \(U_{n,n+u}\) has independent Haar-uniform entries. The replacement is not coordinatewise in \(\mathbb{R}^n\) but columnwise. For a fixed column \(j\), if \(A_1\) uses the original \(j\)-th column \(v_1\) and \(A_0\) uses an independent Haar-uniform column \(v_0\), then
\[
d_{TV}\bigl(\coker(A_1),\coker(A_0)\bigr)
=
\big\|\proj_{\im(M)}(\nu_1-\nu_0)\big\|_1,
\]
where \(\nu_1,\nu_0\) are the laws of \(v_1,v_0\). A central inequality bounds, for any signed measure \(\nu\) on \(R^n\),
\[
\sum_{M'}\big\|\proj_{\im(M')}\nu\big\|_1
\le
\sum_{I\subset R} O\!\bigl((1+\varepsilon)^n\|\nu\bmod I\|_2\bigr)
+
O\!\bigl((1+\varepsilon)^n\Delta^n\bigr).
\]
Applying this with \(\nu=\nu_1-\nu_0\) gives a cost \(C(R,u,\varepsilon)\theta^n\) for a single-column swap, and summing over all columns yields
\[
d_{TV}\bigl(\coker(M_{n,n+u}),\coker(U_{n,n+u})\bigr)\le C\,\theta^n.
\]
The same argument extends to the span of the column-space and to the joint law of \((\coker(M_{n,n}),\det M_{n,n})\) [2601.10821].

Taken together, these examples show that the Lindeberg replacement method is not tied to a single ambient space, metric, or analytic technology. In the cited works, the exchange step is implemented by ordinary Taylor expansion, a Taylor-like generator expansion, a Stein-kernel substitution, Gaussian smoothing, or a column-swapping argument with Fourier-isotypic decomposition. What remains invariant is the telescoping architecture: a complex comparison is reduced to a sequence of one-step replacements whose errors can be summed.

Source: https://www.emergentmind.com/topics/lindeberg-replacement-method