Papers
Topics
Authors
Recent
Search
2000 character limit reached

Deep neural networks as lattice gauge theories

Published 19 Aug 2026 in hep-th, cond-mat.dis-nn, and cs.LG | (2608.19331v1)

Abstract: We modify the NN/QFT duality [1] to incorporate the layerwise permutation symmetry of the network, resulting in a (0!+!1)(0!+!1)-dimensional lattice gauge theory, in which each layer of NN neurons acts as an NN-component lattice site, and the weight matrices play the role of gauge fields living on the links. In this framework, we compute the tree-level neuron-neuron propagator which describes the evolution of layer variance in the network, and develop the Feynman diagram machinery to compute interactions in the perturbative expansion in $1/N$. In particular, we obtain a recursive expression for all corrections to the exact propagator at O(1)O(1), representing statistical fluctuations in the ensemble of networks, including infinitely-many loop diagrams mediating the interactions from previous layers. We also present a preliminary analysis of neuron scattering amplitudes that contribute order-by-order in $1/N$, which provides a field-theoretic framework for studying higher-point correlations, and by extension information propagation, in deep networks. We remark on some interesting directions for future work at the intersection of neural networks and quantum field theory.

Summary

  • The paper recasts multilayer perceptrons at initialization as 0+1-dimensional lattice gauge theories with local permutation symmetry S_N, treating neurons as matter fields and weights as non-dynamical link variables.
  • The authors derive recursive bare propagators and a provably convergent resummation of all leading O(1) cactus corrections, capturing finite-width ensemble fluctuations beyond the Gaussian-process limit.
  • The theory predicts a hierarchy in which two-point functions survive at O(1), interlayer higher-point correlations emerge at O(1/N), and correlations spanning three or more layers first appear at O(1/N²).

From field theory to lattice gauge theory of deep networks

The NN/QFT correspondence developed by Grosvenor and Jefferson recasts fully-connected neural networks at initialization as statistical field theories in $0+1$ dimensions, with $1/N$ (where NN is the layer width) playing the role of a small parameter controlling finite-width corrections away from the Gaussian process limit (Grosvenor et al., 2021). The paper under review, "Deep neural networks as lattice gauge theories" (2608.19331), extends this framework in a specific and well-motivated direction: it incorporates the layerwise permutation symmetry of multilayer perceptrons into the dual description, promoting the network to a (0+1)(0+1)-dimensional lattice gauge theory with discrete gauge group SNS_N. Each layer of NN neurons becomes an NN-component lattice site carrying a fundamental-representation matter field, and each weight matrix WℓW_\ell acts as a link variable transforming in the adjoint, Wℓ↦PℓWℓPℓ−1TW_\ell \mapsto P_\ell W_\ell P_{\ell-1}^{\mathsf{T}}. This is, to the authors' knowledge, the first explicit accommodation of permutation symmetry within this class of dual descriptions.

The choice to work on the lattice is forced by the mathematics rather than merely preferred: because SNS_N is discrete, there is no principal-bundle connection or continuum covariant derivative available, so the continuum limit $1/N$0 taken in prior work must be abandoned. Instead, the network structure equation $1/N$1 is rewritten as a discretized covariant derivative $1/N$2, where $1/N$3 is the lattice spacing between layers. Two structural consequences follow immediately. First, in one dimension there are no plaquettes and hence no field strength term: the weight matrices are non-dynamical, fluctuating randomly at each link with no equations of motion. Second, only locally gauge-invariant quantities — such as $1/N$4 or interlayer products $1/N$5 — qualify as observables, which cleanly explains why interlayer correlators $1/N$6 vanish for $1/N$7 at tree level. The authors also note that for linear networks the symmetry elevates to continuous $1/N$8 rotations, restoring access to continuum methods; elementwise nonlinearities spoil this precisely because $1/N$9 holds only for permutations.

Construction of the action and bare propagators

The path integral follows the response-field formalism of (Grosvenor et al., 2021): delta functions enforcing the covariantized network dynamics are Fourier-transformed into complex response fields NN0, sources are added, and the Gaussian weight/bias distributions (NN1, biases NN2) are integrated out. A crucial simplification relative to the RNN case arises from treating NN3 as layer-dependent: after integrating out the link variables, the partition function becomes local in depth rather than bi-local. Introducing the permutation-invariant auxiliary field NN4 (with NN5 playing the 't Hooft coupling) decouples the theory into NN6 identical copies, and shifting NN7 by its nonzero vev NN8 expands around the true vacuum.

The resulting bare neuron-neuron propagator obeys the recursive relation

NN9

where (0+1)(0+1)0 are coefficients determined by the Taylor expansion of the activation function (for (0+1)(0+1)1, via Bernoulli numbers). All other propagators are trivial delta-function constraints on adjacent layers. The authors treat the activation nonlinearity as a formal power series retained to all orders, rather than truncating at quartic order as in much of the prior literature. They support this with a concrete estimate: for a 100-layer network initialized near the critical point ((0+1)(0+1)2, (0+1)(0+1)3), preactivations have standard deviation roughly (0+1)(0+1)4 against the convergence radius (0+1)(0+1)5 of (0+1)(0+1)6, so only about 5% of preactivations fall outside the analytic region. They candidly note that ReLU is excluded by the differentiability assumption, while observing that since the set of exactly-zero preactivations has measure zero, the practical stringency of this condition is unclear — possibly explaining why ReLU outperforms what such analyses predict.

An appendix clarifies that the conditional partition function (0+1)(0+1)7 is not canonically normalized as a function of the data (0+1)(0+1)8; expectation values are therefore functions of (0+1)(0+1)9, and the initial condition SNS_N0 must be supplied empirically (e.g., averaged over pixels of an input image).

Perturbative corrections: cactus towers at SNS_N1

The leading corrections to the bare propagator appear at SNS_N2 — not SNS_N3 — and represent statistical fluctuations across the ensemble of random initializations, persisting even at infinite width. Diagrammatically these are cactus-like structures reminiscent of vector models, but the lattice Feynman rules permit retaining arbitrarily many loops even at strong 't Hooft coupling SNS_N4. Each loop carries a factor of SNS_N5 from contracted neuron indices, offset by SNS_N6 factors from vertices; crucially, every diagram's correction to layer SNS_N7 depends on propagators at earlier layers (SNS_N8, SNS_N9, …), encoding the causal structure of information flow at initialization.

The central technical result is a recursive expression for the exact two-point function NN0 summing all NN1 diagrams:

NN2

with initial condition NN3. Here "recursion nodes" NN4 substitute for bare propagators inside loops, and a closed-form symmetry factor NN5 governs each NN6-loop, NN7-node flower diagram. The appendix proves convergence of this series by Gevrey-class analysis combined with Stirling estimates: the coefficients NN8 scale asymptotically as NN9 (Gevrey-0), and all three regimes of the inner sum decay rapidly enough to guarantee absolute convergence for fixed arguments. No closed form was found despite substantial effort — the authors state this plainly — though they obtain a striking reorganization: expressing the Bernoulli-number coefficients through Dirichlet series for the zeta function converts the sum over infinitely many loops and recursion nodes into sums over NN0 and error functions, evaluated via pole-subtracted cotangent partial-fraction identities, leaving a rapidly convergent sum over the spectral parameter NN1 plus an integral over NN2. The resulting pole structure invites comparison with Mellin-amplitude and heat-kernel spectral representations, though the authors caution that these remain suggestive reorganizations rather than identifications of a unique underlying operator.

A further point deserves emphasis: truncating the original sum over NN3 is not a consistent loop truncation, since even NN4 contains towers of up to NN5 loops via recursive insertion. Conversely, truncating the resummed sum over NN6 corresponds to approximating the Taylor coefficients of NN7 — a physically interpretable approximation scheme.

Subleading NN8 corrections arise from joining cactus stems at internal NN9 loops; exemplary diagrams scale as Wâ„“W_\ell0 and Wâ„“W_\ell1. Joining three or more stems is suppressed further, since each additional stem adds a Wâ„“W_\ell2 propagator without adding a compensating loop.

Neuron scattering amplitudes and the hierarchy of correlations

Because only gauge-invariant operators are physical, higher-point observables take the form Wâ„“W_\ell3 and products over more layers. A clean hierarchy emerges from the large-Wâ„“W_\ell4 counting:

Order Surviving correlators
Wâ„“W_\ell5 2-point functions only
Wâ„“W_\ell6 Wâ„“W_\ell7-point functions coupling exactly two distinct layers
Wâ„“W_\ell8 correlations among three or more layers

There are no WℓW_\ell9 scattering amplitudes, and no tree-level contributions to Wℓ↦PℓWℓPℓ−1TW_\ell \mapsto P_\ell W_\ell P_{\ell-1}^{\mathsf{T}}0 for Wℓ↦PℓWℓPℓ−1TW_\ell \mapsto P_\ell W_\ell P_{\ell-1}^{\mathsf{T}}1: the Wℓ↦PℓWℓPℓ−1TW_\ell \mapsto P_\ell W_\ell P_{\ell-1}^{\mathsf{T}}2 vertex necessarily shifts the layer index, so all tree-level 1PI diagrams couple adjacent layers. At Wℓ↦PℓWℓPℓ−1TW_\ell \mapsto P_\ell W_\ell P_{\ell-1}^{\mathsf{T}}3, two diagram classes contribute — those connecting distant layers through mixed Wℓ↦PℓWℓPℓ−1TW_\ell \mapsto P_\ell W_\ell P_{\ell-1}^{\mathsf{T}}4/Wℓ↦PℓWℓPℓ−1TW_\ell \mapsto P_\ell W_\ell P_{\ell-1}^{\mathsf{T}}5 loops (arbitrarily many allowed), and those containing a single internal Wℓ↦PℓWℓPℓ−1TW_\ell \mapsto P_\ell W_\ell P_{\ell-1}^{\mathsf{T}}6 joint restricted to 4-point same-layer insertions. Notably, amplitudes remain Wℓ↦PℓWℓPℓ−1TW_\ell \mapsto P_\ell W_\ell P_{\ell-1}^{\mathsf{T}}7 regardless of layer separation, since each additional distinct neuron loop brings its own cancelling Wℓ↦PℓWℓPℓ−1TW_\ell \mapsto P_\ell W_\ell P_{\ell-1}^{\mathsf{T}}8 factor. At Wℓ↦PℓWℓPℓ−1TW_\ell \mapsto P_\ell W_\ell P_{\ell-1}^{\mathsf{T}}9, three-layer correlators such as SNS_N0 first appear, scaling as SNS_N1.

This hierarchy provides a field-theoretic organization principle for information propagation in deep networks at initialization: increasingly long-range statistical correlations between layers are systematically suppressed in powers of SNS_N2. The analysis here is explicitly preliminary — the authors characterize their treatment of scattering amplitudes as qualitative and defer detailed computation.

Limitations and open questions

Several caveats bound the scope of the results. The framework applies strictly to random networks at initialization; training dynamics are absent, and the authors' proposed remedy — noting that SGD acts as an external driving term that would promote the theory to SNS_N3 dimensions and render the gauge fields dynamical — remains speculative future work. Empirical validation is deferred entirely to a companion paper, and two interpretive questions are left open: how the lattice spacing SNS_N4 should be fixed in practice, and whether the covariantized ODE faithfully represents an MLP away from the continuum limit or instead describes a ResNet-like architecture with residual connections controlled by SNS_N5. The activation function must be analytic at the origin (excluding ReLU), and although the exact-propagator series converges, combining it recursively with the Taylor-expanded bare propagator requires care when preactivations exceed the convergence radius. Finally, the SNS_N6 corrections appear to break down at strong coupling SNS_N7, which may connect to the known chaotic regime identified in earlier signal-propagation studies, but this consistency is asserted rather than demonstrated.

Conclusion

This paper reframes the NN/QFT duality for MLPs as a SNS_N8-dimensional lattice gauge theory with local SNS_N9 symmetry, in which weights are non-dynamical link variables and only permutation-invariant correlators are observable. Its concrete deliverables are the recursive bare propagator governing variance propagation, a provably convergent all-loops expression for the $1/N$00 ensemble fluctuations, and an order-by-order classification of higher-point "scattering" amplitudes revealing a $1/N$01 hierarchy in interlayer correlations. The most significant open problems are the incorporation of training dynamics (which would require a genuine plaquette term), empirical tests against trained networks, and extension to architectures such as CNNs and transformers, where the effective spacetime dimensionality of the dual theory would plausibly increase.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 2 likes about this paper.