The Lindeberg replacement method is a comparison technique that sequentially replaces parts of a random object to prove distributional approximations and universality.
It employs a telescoping sum and Taylor expansions to isolate the effects of mean and variance mismatches while controlling cumulative errors.
Its applications span central limit theorems, deep neural network universality, stable-law approximations, and random matrix theory, demonstrating practical versatility.
The Lindeberg replacement method is a comparison technique for proving distributional approximation and universality by successively replacing individual coordinates, summands, layers, or columns of a random object with better understood counterparts and controlling the cumulative error through a telescoping decomposition. In the formulations developed for smooth observables, the basic mechanism is a coordinatewise Taylor expansion; in other settings it is coupled with Kolmogorov forward equations, Stein kernels, Gaussian smoothing, or representation-theoretic decompositions. The method appears in generalized form for smooth functions of independent coordinates (Korada et al., 2010), in stable approximation without Fourier inversion (Chen et al., 2018), in total-variation versions of the central limit theorem (Dung et al., 4 Nov 2025), in quantitative universality for deep neural networks (Giovagnini et al., 4 May 2026), and in column-swapping arguments for random matrices over finite local rings (Lvov, 15 Jan 2026).
1. General scheme and abstract formulation
A standard formulation begins with two random vectors U=(U1,…,Un) and V=(V1,…,Vn) in Rn, each with independent components, and a smooth test function f:Rn→R. In the generalized principle stated in "Applications of Lindeberg Principle in Communications and Statistical Learning" (Korada et al., 2010), one defines
ai=E[Ui]−E[Vi],bi=E[Ui2]−E[Vi2],
assumes imax{E∣Ui∣3+E∣Vi∣3}≤M3<∞, and requires bounded first, second, and third coordinate derivatives: ∂irf(x)≤Lr(f)∀x∈Rn,i=1,…,n,r=1,2,3.
Under these hypotheses,
The corresponding “local” version replaces global derivative bounds by expectations of derivatives evaluated along mixed vectors (U1:i−1,s,Vi+1:n), together with integral remainders involving ∂i3f. This variant is designed for high-dimensional applications in which uniform control of all V=(V1,…,Vn)0 over V=(V1,…,Vn)1 is unavailable (Korada et al., 2010).
The core identity is the telescoping sum
V=(V1,…,Vn)2
Each summand isolates a one-coordinate replacement. A Taylor expansion in the V=(V1,…,Vn)3-th coordinate then separates the effect of mean mismatch, variance mismatch, and the third-order remainder. This suggests that the replacement method is best understood not as a single theorem but as a modular proof pattern: telescoping plus a one-step comparison estimate.
2. Taylor expansion, moment matching, and cancellation
In the smooth setting, the method operates by expanding the observable in the coordinate being replaced. The proof sketch in (Korada et al., 2010) gives the basic form: V=(V1,…,Vn)4
and similarly for V=(V1,…,Vn)5. Because V=(V1,…,Vn)6 is independent of V=(V1,…,Vn)7, expectations factorize, producing the explicit contributions of the first two moments and a remainder term controlled by third derivatives and third moments.
When the first two moments match, the leading terms cancel and only the third-order remainder remains. This is the mechanism emphasized in several later applications. In the deep-network setting, for example, the layerwise weights in the original and Gaussian networks have matching first two moments, so the first two Taylor terms cancel and the remainder is V=(V1,…,Vn)8, which after summing over V=(V1,…,Vn)9 terms gives Rn0 per layer (Giovagnini et al., 4 May 2026). In that sense, moment matching converts a qualitative exchange argument into a quantitative one.
The same structural idea persists when the one-step estimate is not written as an ordinary Euclidean Taylor formula. In stable approximation, the expansion is “Taylor-like” and the compensating term is expressed through the generator Rn1 of the stable Lévy process rather than the Gaussian second derivative (Chen et al., 2018). In the total-variation central limit theorem, the object being replaced is a Stein kernel contribution Rn2 by its mean Rn3, and the telescoping is carried out inside Stein’s identity (Dung et al., 4 Nov 2025). A plausible implication is that the decisive feature is not the literal form of the expansion, but the existence of a one-step approximation whose leading part is additive and whose error is summable.
3. Stable-law approximation and the generator method
"Approximation to the stable law by Lindeberg principle" develops the method for one-dimensional possibly asymmetric Rn4-stable laws with Rn5 (Chen et al., 2018). For fixed Rn6, Rn7, and when Rn8 with Rn9, a real random variable f:Rn→R0 has the law f:Rn→R1 if its characteristic function is
f:Rn→R2
with the usual log-correction form when f:Rn→R3. Equivalently, f:Rn→R4 is the one-dimensional Lévy process at time one with generator
The random variables f:Rn→R9 are assumed to lie in the domain of normal attraction described by
ai=E[Ui]−E[Vi],bi=E[Ui2]−E[Vi2],0
with ai=E[Ui]−E[Vi],bi=E[Ui2]−E[Vi2],1 bounded and vanishing at infinity. The normalized sums are
ai=E[Ui]−E[Vi],bi=E[Ui2]−E[Vi2],2
where ai=E[Ui]−E[Vi],bi=E[Ui2]−E[Vi2],3 if ai=E[Ui]−E[Vi],bi=E[Ui2]−E[Vi2],4, and ai=E[Ui]−E[Vi],bi=E[Ui2]−E[Vi2],5 otherwise.
The replacement framework interpolates between the normalized sum of the ai=E[Ui]−E[Vi],bi=E[Ui2]−E[Vi2],6's and an i.i.d. sum of ai=E[Ui]−E[Vi],bi=E[Ui2]−E[Vi2],7's. With suitable ai=E[Ui]−E[Vi],bi=E[Ui2]−E[Vi2],8,
ai=E[Ui]−E[Vi],bi=E[Ui2]−E[Vi2],9
For the imax{E∣Ui∣3+E∣Vi∣3}≤M3<∞0-term, Lemma 3.9 gives, when imax{E∣Ui∣3+E∣Vi∣3}≤M3<∞1,
imax{E∣Ui∣3+E∣Vi∣3}≤M3<∞2
with a remainder controlled by imax{E∣Ui∣3+E∣Vi∣3}≤M3<∞3, a tail integral involving imax{E∣Ui∣3+E∣Vi∣3}≤M3<∞4, and imax{E∣Ui∣3+E∣Vi∣3}≤M3<∞5. For the imax{E∣Ui∣3+E∣Vi∣3}≤M3<∞6-term, the Kolmogorov forward equation
imax{E∣Ui∣3+E∣Vi∣3}≤M3<∞7
produces compensating generator terms. The paper states that in the telescoping sum the forward-equation terms exactly cancel the compensating imax{E∣Ui∣3+E∣Vi∣3}≤M3<∞8-terms from the Taylor-like expansion (Chen et al., 2018).
The resulting rates are explicit. Under imax{E∣Ui∣3+E∣Vi∣3}≤M3<∞9 for large ∂irf(x)≤Lr(f)∀x∈Rn,i=1,…,n,r=1,2,3.0, there is ∂irf(x)≤Lr(f)∀x∈Rn,i=1,…,n,r=1,2,3.1 such that for every ∂irf(x)≤Lr(f)∀x∈Rn,i=1,…,n,r=1,2,3.2,
∂irf(x)≤Lr(f)∀x∈Rn,i=1,…,n,r=1,2,3.3
In particular,
∂irf(x)≤Lr(f)∀x∈Rn,i=1,…,n,r=1,2,3.4
with the paper also giving ∂irf(x)≤Lr(f)∀x∈Rn,i=1,…,n,r=1,2,3.5 for ∂irf(x)≤Lr(f)∀x∈Rn,i=1,…,n,r=1,2,3.6 and ∂irf(x)≤Lr(f)∀x∈Rn,i=1,…,n,r=1,2,3.7 for ∂irf(x)≤Lr(f)∀x∈Rn,i=1,…,n,r=1,2,3.8. The paper further states that it is the first time that the general stable central limit theorem is proved by the Lindeberg principle, and that the theorem with ∂irf(x)≤Lr(f)∀x∈Rn,i=1,…,n,r=1,2,3.9 is proved by a new method other than Fourier analysis (Chen et al., 2018).
4. Central-limit theorems in total variation
"Total variation bounds in the Lindeberg central limit theorem" reworks the replacement idea in a non-i.i.d. Gaussian approximation problem, but with total-variation distance rather than a smooth test-function metric (Dung et al., 4 Nov 2025). Let
where the E[f(U)]−E[f(V)]≤i=1∑n(aiL1(f)+21biL2(f))+61nL3(f)M3.1 are independent, centered, and absolutely continuous. The total-variation distance from E[f(U)]−E[f(V)]≤i=1∑n(aiL1(f)+21biL2(f))+61nL3(f)M3.2 is
sets E[f(U)]−E[f(V)]≤i=1∑n(aiL1(f)+21biL2(f))+61nL3(f)M3.5, E[f(U)]−E[f(V)]≤i=1∑n(aiL1(f)+21biL2(f))+61nL3(f)M3.6, and E[f(U)]−E[f(V)]≤i=1∑n(aiL1(f)+21biL2(f))+61nL3(f)M3.7. Under these hypotheses,
The proof combines Stein’s equation with a Lindeberg-style telescoping. For each bounded (U1:i−1,s,Vi+1:n)0, (U1:i−1,s,Vi+1:n)1 solves
(U1:i−1,s,Vi+1:n)2
and Lemma 2.1 yields (U1:i−1,s,Vi+1:n)3 and (U1:i−1,s,Vi+1:n)4. Each (U1:i−1,s,Vi+1:n)5 is equipped with a Stein kernel (U1:i−1,s,Vi+1:n)6, and the replacement step consists of replacing (U1:i−1,s,Vi+1:n)7 by its mean (U1:i−1,s,Vi+1:n)8. Inserting this into Stein’s identity produces
(U1:i−1,s,Vi+1:n)9
The paper explicitly describes this as the classical telescoping argument of Lindeberg, but carried out in the language of Stein kernels (Dung et al., 4 Nov 2025).
A central consequence is an exact equivalence statement under bounded Fisher information: if ∂i3f0, then the usual Lindeberg condition
∂i3f1
together with the Feller condition ∂i3f2, is equivalent to ∂i3f3 (Dung et al., 4 Nov 2025). The paper therefore states that, under suitable assumptions, Lindeberg’s condition is sufficient and necessary for convergence in total variation.
5. Layerwise exchange in deep neural networks
In "Universality in Deep Neural Networks: An approach via the Lindeberg exchange principle," the method is adapted to the infinite-width limit of fully connected deep networks with general weights (Giovagnini et al., 4 May 2026). The setting fixes depth ∂i3f4 and widths ∂i3f5, and defines switched-layers networks ∂i3f6 whose first ∂i3f7 layers use the original weights ∂i3f8 and whose last ∂i3f9 layers use independent Gaussian matrices V=(V1,…,Vn)00 with the same variance V=(V1,…,Vn)01.
The core weak comparison theorem states that if
V=(V1,…,Vn)02
and if V=(V1,…,Vn)03 and V=(V1,…,Vn)04 have the same architecture, same biases, and same activation V=(V1,…,Vn)05, but the hidden weights in V=(V1,…,Vn)06 satisfy Assumption 2.1 while those in V=(V1,…,Vn)07 are Gaussian with matching first two moments, then for every
V=(V1,…,Vn)08
one has
V=(V1,…,Vn)09
If instead one assumes Gaussian biases and only V=(V1,…,Vn)10, then the same bound holds for any bounded V=(V1,…,Vn)11, or for V=(V1,…,Vn)12 with error controlled by V=(V1,…,Vn)13 (Giovagnini et al., 4 May 2026).
The telescoping decomposition is layerwise: V=(V1,…,Vn)14
Conditioning on the pre-activations of the relevant layer, one views the next layer as a sum of independent random variables and applies a conditional Taylor expansion up to third order plus moment matching to obtain
V=(V1,…,Vn)15
The paper describes each layer V=(V1,…,Vn)16 as contributing a “replacement error”
After all layers are replaced, one compares the fully Gaussian network to the Gaussian-process limit in V=(V1,…,Vn)18. Under Assumptions 2.1, the stated activation regularity, and invertibility of the limiting covariances V=(V1,…,Vn)19,
V=(V1,…,Vn)20
If all hidden layers have the same width V=(V1,…,Vn)21, then
V=(V1,…,Vn)22
uniformly in V=(V1,…,Vn)23 and V=(V1,…,Vn)24 up to the constant V=(V1,…,Vn)25 (Giovagnini et al., 4 May 2026). The role of Gaussian biases is explicitly that their smoothing effect reduces the derivative requirement on V=(V1,…,Vn)26 from V=(V1,…,Vn)27 bounded derivatives to only V=(V1,…,Vn)28.
6. Universality, sparse limits, and algebraic variants
The method is used in (Korada et al., 2010) to prove universality properties in communications, statistical learning, and random matrix theory, and also to show that dense systems can be viewed as limits of properly defined sparse systems. The paper states universality and sparse–dense equivalence results for CDMA capacity, LASSO, Wishart spectra, and MIMO capacity. In each case the proof strategy is to express the quantity of interest as V=(V1,…,Vn)29 for a smooth V=(V1,…,Vn)30, verify bounds on its third derivatives, and apply the modified Lindeberg theorem. For example, for standard-type spreading matrices V=(V1,…,Vn)31,
V=(V1,…,Vn)32
and analogous convergence statements are given for LASSO and for the Stieltjes transform of the Wishart ensemble (Korada et al., 2010).
The same paper formulates a sparse-limit viewpoint by defining a V=(V1,…,Vn)33-sparsification V=(V1,…,Vn)34 through
V=(V1,…,Vn)35
It then states that if the third partial derivatives of V=(V1,…,Vn)36 are V=(V1,…,Vn)37, replacing a nonzero by an independent Gaussian or another i.i.d. entry induces an error of order V=(V1,…,Vn)38, so over V=(V1,…,Vn)39 nonzeros the total error is V=(V1,…,Vn)40 (Korada et al., 2010). This suggests a form of universality obtained by first passing to analytically simpler sparse models and then returning to dense ones.
"Universality results for random matrices over finite local rings" shows that the replacement method also has a discrete and algebraic form (Lvov, 15 Jan 2026). Here the random object is a matrix V=(V1,…,Vn)41 over a finite local ring V=(V1,…,Vn)42, and the comparison matrix V=(V1,…,Vn)43 has independent Haar-uniform entries. The replacement is not coordinatewise in V=(V1,…,Vn)44 but columnwise. For a fixed column V=(V1,…,Vn)45, if V=(V1,…,Vn)46 uses the original V=(V1,…,Vn)47-th column V=(V1,…,Vn)48 and V=(V1,…,Vn)49 uses an independent Haar-uniform column V=(V1,…,Vn)50, then
V=(V1,…,Vn)51
where V=(V1,…,Vn)52 are the laws of V=(V1,…,Vn)53. A central inequality bounds, for any signed measure V=(V1,…,Vn)54 on V=(V1,…,Vn)55,
V=(V1,…,Vn)56
Applying this with V=(V1,…,Vn)57 gives a cost V=(V1,…,Vn)58 for a single-column swap, and summing over all columns yields
V=(V1,…,Vn)59
The same argument extends to the span of the column-space and to the joint law of V=(V1,…,Vn)60 (Lvov, 15 Jan 2026).
Taken together, these examples show that the Lindeberg replacement method is not tied to a single ambient space, metric, or analytic technology. In the cited works, the exchange step is implemented by ordinary Taylor expansion, a Taylor-like generator expansion, a Stein-kernel substitution, Gaussian smoothing, or a column-swapping argument with Fourier-isotypic decomposition. What remains invariant is the telescoping architecture: a complex comparison is reduced to a sequence of one-step replacements whose errors can be summed.