Papers
Topics
Authors
Recent
Search
2000 character limit reached

Information Spreading in Diffusion Models from Effective Field Theory

Published 14 Aug 2026 in hep-th and cond-mat.stat-mech | (2608.14308v1)

Abstract: We study score-matching diffusion models with a convolutional architecture. We argue that the inductive bias of locality means that the machinery of effective field theory from physics can be usefully applied to describe the denoising dynamics. We apply this formalism first to a simple toy example which permits an analytical description, and thereafter to MNIST, and show that in both cases, the mutual information between two points grows in a manner predicted by a simple effective field theory of Brownian motion.

Authors (2)

Summary

  • The paper shows that locality and translational equivariance reduce convolutional diffusion models near the pure-noise endpoint to a heat equation, with nonlinear terms suppressed as the signal vanishes.
  • The paper predicts that spatial mutual information follows diffusive scaling through the variable |x−y|²/ᾱt, a result supported by curve collapse in a ±1 grid model and MNIST experiments.
  • The paper identifies a smooth transition from diffusion to late-time patch snapping, showing that effective kernel size controls the saturated correlation length while the data enters through a small set of effective coefficients.

Overview

This paper, by Neogi and Iqbal (2608.14308), applies the machinery of effective field theory (EFT) to score-matching diffusion models with convolutional architectures. The central claim is that locality and translational equivariance of the network suffice to constrain the reverse-process dynamics near the pure-noise endpoint to a simple heat equation, whose predictions for the growth of spatial mutual information are verified empirically in a ±1\pm 1 grid toy model and on MNIST. The work builds on the Equivariant Local Score (ELS) machine of Kamb and Ganguli, which provides an exact microscopic description of the learned score function for ConvNets, and simplifies it into a small number of universal terms.

Effective field theory of the reverse process

The reverse process is taken in its deterministic probability-flow ODE form,

ϕ(t)t=tαˉtαˉt[ϕ(t)+π(ϕ(t),t)],\frac{\partial\phi(t)}{\partial t} = \frac{\partial_t \bar{\alpha}_t}{\bar{\alpha}_t}\left[\phi(t)+\pi(\phi(t),t)\right],

where π=xlogpt\pi = \nabla_x \log p_t is the score function. The authors expand the right-hand side in powers of the field and its spatial derivatives. Translational equivariance forbids explicit xx-dependence of coefficients; locality justifies truncation. Imposing rotational equivariance (vi=0{\bf v}_i = 0) and a Z2\mathbb{Z}_2 symmetry ϕϕ\phi \to -\phi, and introducing a scaling limit with field dimension ϕ~λd/2ϕ~\tilde{\phi} \to \lambda^{-d/2}\tilde{\phi} — justified both by Gaussian measure invariance and by the central limit theorem applied to box-averaged fields — only two terms survive in d=2d=2: a Laplacian term and a ϕ~3\tilde{\phi}^3 term.

The decisive further step uses the microscopic ELS: every nonlinearity ϕ(t)t=tαˉtαˉt[ϕ(t)+π(ϕ(t),t)],\frac{\partial\phi(t)}{\partial t} = \frac{\partial_t \bar{\alpha}_t}{\bar{\alpha}_t}\left[\phi(t)+\pi(\phi(t),t)\right],0 carries a factor ϕ(t)t=tαˉtαˉt[ϕ(t)+π(ϕ(t),t)],\frac{\partial\phi(t)}{\partial t} = \frac{\partial_t \bar{\alpha}_t}{\bar{\alpha}_t}\left[\phi(t)+\pi(\phi(t),t)\right],1, so all nonlinear terms vanish as ϕ(t)t=tαˉtαˉt[ϕ(t)+π(ϕ(t),t)],\frac{\partial\phi(t)}{\partial t} = \frac{\partial_t \bar{\alpha}_t}{\bar{\alpha}_t}\left[\phi(t)+\pi(\phi(t),t)\right],2. The dynamics near the noise endpoint therefore reduces exactly to the diffusion equation

ϕ(t)t=tαˉtαˉt[ϕ(t)+π(ϕ(t),t)],\frac{\partial\phi(t)}{\partial t} = \frac{\partial_t \bar{\alpha}_t}{\bar{\alpha}_t}\left[\phi(t)+\pi(\phi(t),t)\right],3

The training data enters only through a few Wilsonian coefficients such as ϕ(t)t=tαˉtαˉt[ϕ(t)+π(ϕ(t),t)],\frac{\partial\phi(t)}{\partial t} = \frac{\partial_t \bar{\alpha}_t}{\bar{\alpha}_t}\left[\phi(t)+\pi(\phi(t),t)\right],4. This is the paper's main theoretical result: the complicated patch-sum structure of the ELS collapses to Brownian motion at long distances and early reverse times.

Mutual information growth and diffusive scaling

Assuming Gaussianity — self-consistent since the initial condition is i.i.d. Gaussian noise and the evolution is linear — the mutual information between pixels is ϕ(t)t=tαˉtαˉt[ϕ(t)+π(ϕ(t),t)],\frac{\partial\phi(t)}{\partial t} = \frac{\partial_t \bar{\alpha}_t}{\bar{\alpha}_t}\left[\phi(t)+\pi(\phi(t),t)\right],5, where ϕ(t)t=tαˉtαˉt[ϕ(t)+π(ϕ(t),t)],\frac{\partial\phi(t)}{\partial t} = \frac{\partial_t \bar{\alpha}_t}{\bar{\alpha}_t}\left[\phi(t)+\pi(\phi(t),t)\right],6 is the correlation coefficient. Solving the diffusion equation for the connected two-point function gives

ϕ(t)t=tαˉtαˉt[ϕ(t)+π(ϕ(t),t)],\frac{\partial\phi(t)}{\partial t} = \frac{\partial_t \bar{\alpha}_t}{\bar{\alpha}_t}\left[\phi(t)+\pi(\phi(t),t)\right],7

so correlations depend on space and time only through the self-similar variable ϕ(t)t=tαˉtαˉt[ϕ(t)+π(ϕ(t),t)],\frac{\partial\phi(t)}{\partial t} = \frac{\partial_t \bar{\alpha}_t}{\bar{\alpha}_t}\left[\phi(t)+\pi(\phi(t),t)\right],8: curves at different times must collapse when plotted against this variable. This is a sharp, falsifiable prediction, and it also implies that the noise schedule plays the role of physical time while the schedule's detailed shape does not affect the spatial correlations built up.

Experiments

Two-sample model. For a dataset consisting of homogeneous ϕ(t)t=tαˉtαˉt[ϕ(t)+π(ϕ(t),t)],\frac{\partial\phi(t)}{\partial t} = \frac{\partial_t \bar{\alpha}_t}{\bar{\alpha}_t}\left[\phi(t)+\pi(\phi(t),t)\right],9 and π=xlogpt\pi = \nabla_x \log p_t0 grids, the ELS admits a closed-form evolution equation involving a π=xlogpt\pi = \nabla_x \log p_t1 of the local patch sum. Both the analytic ELS machine (with a π=xlogpt\pi = \nabla_x \log p_t2 kernel) and a trained π=xlogpt\pi = \nabla_x \log p_t3-equivariant ConvNet exhibit curve collapse against π=xlogpt\pi = \nabla_x \log p_t4, confirming diffusive scaling; very small separations (π=xlogpt\pi = \nabla_x \log p_t5 of a few pixels) are excluded because the trained model deviates from the pure ELS there.

Regime transition and correlation length. Feeding a linear interpolation test grid into the ELS score, the authors define the end of the diffusive regime via a factor-π=xlogpt\pi = \nabla_x \log p_t6 enhancement of the spatial gradient of the score, yielding π=xlogpt\pi = \nabla_x \log p_t7. A Taylor expansion of the patch sum gives π=xlogpt\pi = \nabla_x \log p_t8, so the saturated correlation length scales as π=xlogpt\pi = \nabla_x \log p_t9 — linearly in the effective kernel size. Empirically, fitting the Gaussian slope gives xx0, consistent with xx1 to integer precision despite the kernel size drifting over the reverse process. Notably, the transition is smooth: the authors explicitly state there is no discontinuity and that the behavior does not correspond to a conventional phase transition.

MNIST. Without the toy symmetries, additional linear terms appear, but a field redefinition absorbs them; the connected two-point function still obeys a diffusion equation in reparametrized time xx2, and mutual information remains controlled solely by the Laplacian coefficient. Curve collapse against the self-similar variable is again observed, though imperfectly at large distances, which the authors attribute to interactions of diffusive modes with boundary conditions (circular padding was used).

Snapping regime

Near the endpoint xx3, expanding the ELS in xx4 yields

xx5

an exponential relaxation of each patch toward its nearest training patch xx6, corresponding to the "memorisation" regime of Biroli et al. This complements the diffusive regime: spatial mutual information saturates at the value set by the architecture, after which patches decouple and snap individually.

Limitations and open questions

Several caveats are conceded directly. The theory assumes finite limits of the EFT coefficients as xx7 and treats them as constant; the marginal xx8 term would generate logarithmic corrections invisible at finite experimental resolution. The regime-transition criterion depends on the arbitrary factor xx9, so no sharp transition time exists. For general datasets, constructing interpolating test functions analytically is difficult because interpolated patches may lie closer to a third training patch. Large-distance collapse on MNIST degrades and is boundary-condition sensitive. The "fuzziness" of region boundaries lacks an analytical treatment. Finally, whether an EFT description extends to non-convolutional architectures — e.g., transformers, where locality may be inherited from data statistics rather than architecture — is left open.

Conclusion

The paper establishes that locality and equivariance reduce the early reverse dynamics of convolutional diffusion models to a heat equation, with the noise schedule acting as time and the dataset entering only through a few coefficients. Diffusive collapse of mutual-information curves is confirmed quantitatively in both an analytically tractable ELS model and trained networks on MNIST, and the late-time snapping dynamics is captured by a simple exponential relaxation equation. The framework provides a systematic inference method for the effective kernel size during generation and connects the dynamical regimes identified in prior statistical-mechanical analyses to concrete EFT equations.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 5 likes about this paper.