---
title: Mean-Field Limits of Neural Networks
url: https://www.emergentmind.com/topics/mean-field-limits-of-neural-networks
type: topic
---

# Mean-Field Limits of Neural Networks

Mean-field limits of neural networks are asymptotic descriptions in which a network with width tending to infinity, together with a compatible scaling of optimization, is replaced by a deterministic evolution of a probability law over parameters, neurons, functions, or paths. In this regime, empirical averages over neurons converge to continuum objects, stochastic gradient dynamics are recast as continuity equations, integro-differential systems, or McKean–Vlasov equations, and the resulting limit is generally nonlinear and feature-learning rather than a fixed-kernel linearization [1903.04440], [2003.05884], [1906.00193].

## 1. Scaling regimes and state variables

The basic two-layer mean-field model writes
\[
f_m(x;\theta_1,\dots,\theta_m)=m^{-1}\sum_{i=1}^m a_i\,\phi(w_i^T x),
\qquad \theta_i=(a_i,w_i)\in\mathbb R^{1+d_0},
\]
with empirical measure
\[
\rho_k^m=\frac1m\sum_{i=1}^m \delta_{\theta_i^{(k)}}.
\]
The population loss is
\[
R(\rho)=\mathbb E_{x,y}[\ell(y,f[\rho;x])],
\qquad
f[\rho;x]=\int a\,\phi(w^T x)\,\rho(da,dw).
\]
Under mean-field scaling, initialization is \(O(1)\), learning rates are width-dependent, and parameter increments are \(O(1/m)\), so that finite parameter changes persist in the infinite-width limit [2003.05884].

For deep fully connected networks, the scaling is layer dependent. In the formulation of Sirignano and Spiliopoulos, an \(L\)-hidden-layer network has forward normalization by \(1/N_{\ell-1}\) at each hidden layer and \(1/N_L\) at the output. For the \(L\)-layer case, the learning-rate scaling is
\[
\alpha_C=\frac{N_L}{N_1},\qquad
\alpha_{W,1}=1,\qquad
\alpha_{W,\ell}=\frac{N_\ell}{N_1}\quad (\ell=2,\dots,L).
\]
This scaling “balances contributions of each layer in the infinite-width limit”; with a constant rate, the network “freezes” [1903.04440].

Several mathematically distinct state descriptions have been developed for these limits.

| Setting | State variable | Limit evolution |
|---|---|---|
| Two-layer mean-field training | \(\rho_t\in \mathcal P(\mathbb R^{1+d_0})\) | discrete map or parameter-space PDE |
| Deep sequential limit | empirical layerwise parameter distributions | deterministic integro-differential system / Liouville PDE |
| Three-layer neuronal embedding | \(w_i(t,\cdot)\) on a fixed probability space | deterministic ODE system |
| Deep path-space limit | law \(\mu_t\) of input-output paths | McKean–Vlasov ODE |
| Functional-space three-layer limit | \(\mu_t\in P(\mathbb R\times U)\) | kernel gradient flow with time-varying kernel |

The deep sequential formulation uses empirical parameter distributions over products of layerwise parameters and output weights, seeded by an initial law \(\mu_0\) with compact support and continuous densities [1903.04440]. The three-layer neuronal-embedding framework places finite-width networks and the infinite-width limit on a single probability space \((\Omega,\mathcal F,P)\), with deterministic anchor functions \(w_1^0,w_2^0,w_3^0\), so that initialization for all widths is coupled through sampled abstract neurons [2105.05228]. A different deep construction treats each input-to-output parameter path as a particle and studies the empirical distribution of such paths, yielding a path-space McKean–Vlasov limit for deep networks with fixed random features near the input and output [1906.00193]. For partially trained three-layer models with a fixed random first layer, the relevant state is a measure on a functional space, \(\mu_t\in P(\mathbb R\times U)\), where neurons are represented by output weights and pre-activation functions [2210.16286].

These formulations are not interchangeable. A plausible implication is that “the” mean-field limit of a neural network is not a single universal object, but a family of asymptotic descriptions indexed by architecture, training rule, and parametrization.

## 2. Limiting equations and derivation methods

In the two-layer setting, discrete-time gradient descent on the empirical measure is expressed by a deterministic Markov operator:
\[
\rho_{k+1}^m=T(\rho_k^m),
\qquad
\rho_{k+1}^\infty=T(\rho_k^\infty),
\]
where \(T\) pushes each atom by the parameter update \(\theta\mapsto \theta-\eta \nabla_\theta R(\rho)\). The corresponding continuous-time limit, obtained by rescaling time by \(t=k/m\), is the parameter-space PDE
\[
\partial_t \rho_t(\theta)=\nabla_\theta\cdot\bigl[\rho_t(\theta)\nabla_\theta R(\rho_t)\bigr].
\]
The induced network evolution is likewise closed at the level of the output function \(f_t\) through a kernel built from \(\nabla_\theta \phi\) and the current measure \(\rho_t\) [2003.05884].

For deep networks, Sirignano and Spiliopoulos obtain a limit output
\[
g_t(x)=\int_{\mathbb R} C(t;c)\,\sigma(Z(t;c,x))\,\mu_C(dc),
\]
where the parameter trajectories satisfy coupled deterministic ODEs indexed by initialization. In the two-hidden-layer case these equations govern \(\tilde C_t^c\), \(\tilde W_t^{1,w}\), and \(\tilde W_t^{2,c,w,u}\), while in measure form each layer obeys a continuity equation
\[
\partial_t\rho_t^{(\ell)}(w)+\nabla_w\cdot\bigl(V^{(\ell)}[\rho_t](w)\,\rho_t^{(\ell)}(w)\bigr)=0.
\]
The velocity fields are explicit nonlocal functionals of \(\sigma\), \(\sigma'\), and the current network output [1903.04440].

The deep path-space theory replaces neuronwise empirical measures by a law \(\mu_t\) over parameter paths. A typical particle satisfies a self-consistent ODE
\[
\Theta(t)=\Theta(0)-\int_0^t \alpha(s)\,\mathbb E_{(X,Y)}[\bar g(X,Y,\mu_s,\Theta(s))]\,ds,
\]
with drift obtained by replacing finite-width backpropagation averages by integrals under \(\mu_s\). This yields existence, uniqueness, and propagation of chaos for the simultaneous width-\(\infty\) limit of all hidden layers [1906.00193].

The proofs of these limits rely on different technical packages. The sequential deep limit uses tightness in Skorokhod space, uniform moment bounds via Gronwall, martingale-vanishing arguments, identification by test-function calculus, and uniqueness via coupling or fixed-point arguments [1903.04440]. The discrete-time two-layer theory proves convergence by showing that the operator \(T\) is Lipschitz-continuous in \(W_2\) and then inducting on the iteration index [2003.05884]. The deep McKean–Vlasov path-space theory constructs a closed subset of measures on path space on which the McKean map is well defined and eventually contracting under Wasserstein distance [1906.00193]. The neuronal-embedding framework avoids measure-over-measure closure issues by coupling all widths to abstract neurons on a fixed probability space [2105.05228].

A recurrent point of methodology is that mean-field derivations are not purely formal law-of-large-numbers arguments. In deep models, the main obstruction is not only width but also the nested dependence created by backpropagation through multiple layers.

## 3. Optimization, expressivity, and global convergence

Under suitable assumptions, the limiting dynamics can inherit strong optimization properties. In the deep analysis of Sirignano and Spiliopoulos, assuming \(\sigma\in C_b^2(\mathbb R)\) and, for global convergence, that \(\sigma\) is bounded, non-constant, monotone, and hence discriminatory, together with full support assumptions on the data measure and the initialization, the limit system is a gradient flow in probability space with Lyapunov function
\[
\mathcal L([\rho_t])=\frac12\int (f(x)-g_t(x))^2\,\pi(dx).
\]
They prove \(d\mathcal L/dt\le 0\) and that any stationary measure yields zero loss; any limit point \(\rho_*\) as \(t\to\infty\) satisfies \(g_*(x)=f(x)\) for all \(x\) [1903.04440].

For unregularized three-layer networks, Nguyen and Pham prove a global convergence theorem in the mean-field regime under bounded-Lipschitz assumptions on \(\varphi_1,\varphi_2,\varphi_3'\) and \(\partial_2\mathcal L\), full support of the first-layer initialization on \(\mathbb R^d\), convergence of the mean-field parameters \(W(t)\), and density of \(\{\varphi_1(\langle u,\cdot\rangle):u\in\mathbb R^d\}\) in \(L^2(P_X)\). If \(\mathcal L(y,\cdot)\) is convex, then
\[
\lim_{t\to\infty}\mathcal L(W(t))
=
\inf_f \mathbb E[\mathcal L(Y,f(X))].
\]
More generally, when \(\partial_2\mathcal L(y,\hat y)=0\implies \mathcal L(y,\hat y)=0\) and \(Y\) is a deterministic function of \(X\), then \(\mathcal L(W(t))\to 0\). A central ingredient is a universal-approximation property at any finite training time, obtained through an algebraic topology argument showing that the first-layer support remains all of \(\mathbb R^d\) [2105.05228].

The functional-space mean-field theory of partially trained three-layer networks yields a different optimization picture. There the limit output
\[
f_t(x)=\int a\,h(x)\,\mu_t(da,dh)
\]
obeys a kernel gradient flow with a time-varying kernel
\[
K_t(x,x')=K_{a,t}(x,x')+G(x,x')\,Q_t(x,x').
\]
Because \(K_t\) is symmetric and, under mild assumptions, remains strictly positive-definite for all \(t\), the empirical \(L_2\)-loss satisfies a linear-rate decay estimate of the form
\[
\mathcal L_t\le \mathcal L_0 e^{-2\lambda_* t/n}
\]
whenever \(\lambda_{\min}(K_t)\ge \lambda_*>0\) uniformly in time [2210.16286].

These results clarify a common misconception: nonconvexity of the finite-width parameterization does not preclude strong convergence statements for the mean-field limit. The available theorems are conditional on activation regularity, support conditions, convergence modes, or positivity of time-varying kernels, but they establish that multilayer mean-field dynamics can be globally optimizing rather than merely descriptive.

## 4. Fluctuations, finite-width corrections, and trajectorial stability

The law-of-large-numbers limit is only the first term in the large-width expansion. For a single hidden layer, the centered fluctuation
\[
\eta_t^N=\sqrt N\,(\mu_t^N-\mu_t)
\]
converges in a dual Sobolev space \(W^{-J,2}\) to a Gaussian process solving a linear stochastic partial differential equation
\[
d\eta_t=\mathcal L_t[\eta_t]\,dt+dM_t,
\]
or equivalently
\[
d\eta_t=\mathcal L_t[\eta_t]\,dt+\Sigma_t^{1/2}\,dW_t.
\]
The proof uses weak convergence, relative compactness, martingale problems, and uniqueness in a suitable Sobolev space [1808.09372].

For multilayer networks, Nguyen and Pham derive a second-order mean-field limit through the neuronal-embedding framework. The rescaled parameter deviations
\[
R_i^{(N)}(t,j_{i-1},j_i)
=
\sqrt N\Bigl[w_i^{(N)}(t,j_{i-1},j_i)-w_i(t,C_{i-1}(j_{i-1}),C_i(j_i))\Bigr]
\]
are approximated in \(L^2\) by a linear ODE system driven by a Gaussian sampling fluctuation process. They further obtain a central-limit theorem for the output fluctuation \(\hat G^{(N)}(t,x)\), with convergence in finite-dimensional moments at rate \(N^{-1/8+o(1)}\). Under additional assumptions, the width-scaled asymptotic variance
\[
V^*(t)=\mathbb E_X[\hat G(t,X)^2]
\]
is non-increasing and tends to zero as \(t\to\infty\), so gradient descent in the mean-field regime progressively biases training toward solutions with “minimal fluctuation” in the learned output function [2110.15954].

Mean-field theory is also used as a quantitative approximation to finite-width training trajectories. Sirignano and Spiliopoulos state that finite networks, “even moderately wide,” follow the mean-field curves closely and refer to a CIFAR10 example [1903.04440]. Nguyen’s nonrigorous multilayer formalism reports that as the uniform width \(n\) grows, the full training-loss and test-loss curves “lock in” onto a limiting trajectory; for depths \(3,4,5,8\), curves for \(n=400,800,\dots\) “essentially coincide” [1902.02880]. In the discrete-time two-layer theory, the mean-field limit is shown to approximate finite-width networks better than the NTK limit when learning rates are not very small, precisely because it retains the \(f_{aw}\) interaction term associated with feature learning [2003.05884].

This finite-width perspective changes the role of the mean-field limit. It is not only an asymptotic simplification; it is also a controlled surrogate for wide-but-finite dynamics, with explicit next-order corrections in shallow networks and second-order fluctuation systems in multilayer settings.

## 5. Relation to NTK, \(μ\)P, and control-theoretic extensions

Mean-field and NTK limits arise from different scalings. Under NTK scaling,
\[
f_m(x)=m^{-1/2}\sum a_i\phi(w_i^T x),
\]
with \(O(1)\) learning rate and \(O(m^{-1/2})\) parameter movement, the network admits a first-order Taylor linearization and the kernel converges to a fixed \(\Theta_\infty\). By contrast, the mean-field scaling preserves parameter movement of order \(O(1/m)\) over \(O(m)\) iterations and retains the nonlinear \(f_{aw}\) term. The discrete-time comparison shows that if \(\eta=O(1)\), then \(f_{aw}\) matters at finite width and the mean-field limit tracks finite-width networks better; if \(\eta\ll 1\), \(f_{aw}\) becomes negligible and the NTK approximation suffices [2003.05884].

The same framework also identifies an “intermediate” lazy limit that is neither NTK nor mean-field and shows a depth-dependent optimizer effect: for \(L\ge 3\), plain gradient descent with MF-type hyper-scaling has no non-trivial discrete-time mean-field limit, while RMSProp does. Under RMSProp, normalization removes the layer-wise width dependence, all layers move \(O(m^{-1})\), and one recovers a non-trivial discrete-time mean-field limit for any depth [2003.05884]. This is a precise statement about discrete-time scaling rather than a universal impossibility result for deep mean-field analysis.

The \(μ\)P literature extends the scope of mean-field analysis beyond classical \(1/n\) parametrizations. For noisy gradient descent with entropic regularization in wide two-layer networks under \(μ\)P, Acciaio, Heiss, Pammer, and Yan formulate the dynamics as a Fokker–Planck PDE
\[
\partial_t\mu_t
=
\nabla\cdot\Bigl(\mu_t\nabla_\theta \frac{\delta F_\lambda}{\delta\mu}(\mu_t)\Bigr)
+\lambda \Delta\mu_t,
\]
prove global existence and uniqueness in a maximal weighted-moment class \(M_{w^*}\), obtain a uniform-in-time squared-Wasserstein propagation rate \(O(N^{-1})\), characterize identifiability modulo finite-rank realization symmetry, and derive a sparse-dictionary decomposition of the long-time limit under a Barron–Hermite target condition [2605.24710].

A complementary control-theoretic extension interprets the mean-field gradient flow of a two-layer network as a McKean–Vlasov stochastic-control problem. The measure-valued continuity equation
\[
\partial_t \mu_t+\nabla_\theta\cdot(\mu_t v_t)=0,
\qquad
v_t(\theta)=-\nabla_\theta \frac{\delta L}{\delta\mu}(\theta;\mu_t),
\]
is linked to a Hamilton–Jacobi–Bellman equation on \(\mathcal P_2(\Theta)\) and a Dynamic Programming Principle. This yields a Finsler-type metric \(d_G\) on probability measures and the variational characterization
\[
\min_{\nu\in\mathcal P_2(\Theta)}
\left[
L(\nu)+\frac1{2T}d_G(\mu_0,\nu)^2
\right]
\]
for early stopping, up to a controllable error. The long-time limit selects, among global minimizers, one of minimal \(d_G(\mu_0,\cdot)\) [2603.20892].

Mean-field analysis has also been adapted to particle-based optimizers other than gradient descent. For two-layer networks trained by consensus-based optimization, the network-width limit lifts parameters to \(\mu\in\mathcal P_2(\Omega)\), the particle-ensemble limit lifts the ensemble to \(\rho\in\mathcal P(\mathcal P_2(\Omega))\), and the resulting dynamics become a gradient flow on the Wasserstein-over-Wasserstein space. In this setting the population variance contracts as
\[
\mathcal V(\rho^{k+1})\le (1-\Delta t)^2 \mathcal V(\rho^k)
\]
under the stated barycenter and absolute-continuity assumptions [2511.21466].

## 6. Broader meanings in neural-network research

The phrase “mean-field limit of neural networks” is not confined to supervised training of artificial feedforward models. In statistical mechanics and theoretical neuroscience, it refers to several neighboring but non-identical asymptotic programs.

For Hopfield-like and related rate networks, mean-field limits describe thermodynamic or stochastic population equations. One recent thermodynamic result proves that a measure-concentration assumption on order parameters suffices for existence of the asymptotic free energy of the Hopfield model and recovers the replica-symmetric free-energy formula through a decomposition into hard and soft spin-glass free energies [2409.10145]. A separate universality result establishes the mean-field equations for large networks of Hopfield-like neurons on \([0,T]\) without assuming i.i.d. zero-mean Gaussian synaptic weights; the limit is stochastic and characterized by a mean function \(m_Q\) and a correlation function \(K_Q\), with effective noise given through a Volterra equation [2408.14290]. For correlated Gaussian synaptic weights, Faugeras, Maclaurin, and Tanré prove an annealed large deviation principle and identify a unique Gaussian, non-Markovian minimizer, leading to an infinite countable family of linear non-Markovian SDEs in the limit [1901.10248].

For spiking or integrate-and-fire networks, mean-field limits are often PDE limits for empirical measures rather than parameter distributions. In dense stochastic integrate-and-fire networks with arbitrary synaptic weights satisfying \(O(1/N)\) scaling, the empirical measure converges, up to subsequence, to a spatially extended PDE indexed by a graphon variable \(\xi\in[0,1]\):
\[
\partial_t\mu(t,\xi,dx)
+\partial_x\bigl[(b(x)+h(t,\xi))\mu(t,\xi,dx)\bigr]
+f(x)\mu(t,\xi,dx)
-r(t,\xi)\delta_0(dx)=0,
\]
with \(h\) determined by a graphon kernel \(w(\xi,\zeta)\) [2409.06325]. For sparse integrate-and-fire networks, a tree-indexed extension of the BBGKY hierarchy yields convergence of one-particle observables to a non-exchangeable Vlasov equation under generalized mean-field scaling and non-vanishing diffusion [2309.04046]. Replica-mean-field theory for intensity-based spiking networks takes a different route: instead of letting interactions vanish as \(1/K\), it considers infinitely many interacting replicas of a fixed finite network and derives stationary ODEs and self-consistency equations via the Poisson Hypothesis, preserving finite-size effects such as saturation and sparse-induced metastability [1902.03504].

Even in rate-based noisy networks, the term can mean a macroscopic reduction of neuronal activity rather than a parameter-space transport equation. For a random network of noisy rate neurons on an Erdős–Rényi graph, a second-order stochastic mean-field model for the mean rate \(R(t)\) and variance \(S(t)\) distinguishes the effects of external and internal noise: in the thermodynamic limit, external noise reshapes the deterministic mean-field vector field, while internal noise affects only the variance equation [1512.03561].

These adjacent literatures show that “mean-field” is a family resemblance term. In machine learning, it usually denotes infinite-width training dynamics in parameter or function space; in statistical mechanics and neuroscience, it often denotes thermodynamic limits of interacting neuronal states, empirical membrane-potential laws, or free-energy formulas. The shared structure is passage from many-body randomness to a deterministic or self-consistent continuum description, but the state variables, limiting equations, and proof techniques differ substantially.

Source: https://www.emergentmind.com/topics/mean-field-limits-of-neural-networks