- The paper introduces variance-exploding score-based diffusion models to overcome critical slowing down in lattice field theory sampling.
- It employs a fully convolutional U-Net architecture ensuring translation invariance and achieving high MALA acceptance rates above 65% across various lattice volumes.
- Validation against FA-HMC–Wolff ensembles shows the model accurately reproduces scalar observables and momentum-space propagators with near-unity effective sample sizes.
Diffusion Models for Sampling Near Criticality in Lattice Field Theories: Technical Overview
Motivation and Context
Critical slowing down is a longstanding computational bottleneck in Markov-chain Monte Carlo sampling of lattice field theories, especially near second-order phase transitions where the correlation length diverges and autocorrelation times scale as τint​∼ξz with z>1. Although algorithms like cluster updates and Fourier acceleration ameliorate this issue for special cases, generic interacting theories still require costly Markov-chain approaches, particularly near criticality. The real scalar ϕ4 theory provides a controlled benchmark for non-perturbative methods, featuring symmetric, near-critical, and broken phases under tunable parameters. This study targets the challenge of generating statistically independent field samples across this phase diagram without suffering from critical slowing down.
Methodology: Score-Based Diffusion Models for Lattice Sampling
This work implements variance-exploding (VE) score-based diffusion models for lattice sampling. The forward process is an N(0,Σ2(t)1) noise injection (zero drift), resulting in a monotonic variance schedule as configurations evolve from the target measure towards a broad Gaussian. The reverse process, initialized from this Gaussian, leverages a learned score network sθ​(ϕ,t) to iteratively denoise and map noise samples to configurations distributed according to the Boltzmann weight e−S[ϕ]. The network is trained via denoising score matching, minimizing the deviation between the predicted score and the analytically available conditional score derived from the VE kernel.
The score network architecture is a fully convolutional U-Net with circular (periodic) padding, maintaining translation invariance and enabling evaluation across arbitrary lattice sizes.
Validation and Numerical Results
Scalar and Two-Point Observables
Comprehensive validation was conducted across phases in both 2D (L=128, λ=0.022) and 3D (L=64, λ=0.9). Generated ensembles were compared against reference FA-HMC–Wolff ensembles (Fourier-accelerated HMC plus cluster updates) for scalar observables and the momentum-space propagator z>10.
The diffusion model accurately reproduces scalar quantities (order parameter, susceptibility, Binder cumulant, action density) in the symmetric, near-critical, and broken regimes Figure 1. Notably, discrepancies are concentrated in the zero-mode sector (magnetization) and, in 3D, in the action density, especially near criticality. The susceptibility excess in the broken phase is attributable to residual broadening of magnetization peaks, quantitatively measured and understood as intra-sector zero-mode variance.

Figure 1: Joint distribution of z>11 and action density z>12 for three z>13 values in 2D; diffusion-model contours closely track FA-HMC–Wolff references across phases.
Spatial translation invariance is maintained by the convolutional architecture, confirmed by uniformity in single-site cumulant heatmaps Figure 2.

Figure 2: Single-site cumulants z>14 and z>15 display no spatial artifacts, confirming translation invariance of generated ensembles.
Momentum-Space Propagator
Critical and noncritical scaling of the momentum-space propagator was faithfully reproduced by the diffusion models, including anomalous Ising critical exponents and infrared plateau behavior in symmetric phases (Figures 4, 5, 6, 7, 8, 9). Early training epochs showed over- or underestimation in infrared or ultraviolet modes, but convergence led to errors concentrated only in the zero-mode.

Figure 3: Momentum-space propagator z>16 at the near-critical point in 2D (z>17), showing convergence to FA-HMC–Wolff reference with correct critical scaling.

Figure 4: Propagator z>18 at near-criticality in 3D (z>19), with subpercent deviations from reference, tracking Ising scaling.
Diagnostics: Score Quality, MALA Acceptance, and Effective Sample Size
Score Network Quality
Direct comparison of the learned score with the negative action gradient on reference configurations at small diffusion time reveals rapid alignment in direction (cosine similarity ϕ40), with magnitude convergence occurring more slowly (Figure 5, 11). This indicates the network reconstructs the score field structure from data alone.

Figure 5: Score network cosine similarity, mean-squared error, and magnitude ratio diagnostics for 2D ϕ41 model, showing rapid directional learning.

Figure 6: Same diagnostics for 3D ϕ42 model, confirming efficacy across dimensions and phases.
MALA Acceptance Rate
Using the learned score in a Metropolis-adjusted Langevin proposal for exact Boltzmann sampling, acceptance rates above 65% (across all training sizes and phases) were achieved at late training, with robust performance in cross-volume scenarios (Figures 12, 13). Early training acceptance was near zero.

Figure 7: MALA acceptance rate evolution in 2D; rates plateau above 85% after training.

Figure 8: MALA acceptance rate for 3D models, replicating high post-training acceptance irrespective of phase.
HMC-Referenced Effective Sample Size
Observable-level HMC-referenced MSE-based effective sample size (ESS) diagnostics show that for scalar and two-point observables, the generated ensemble reaches ESS ratios near unity, indicating that ϕ43 diffusion samples are as efficient as ϕ44 independent target samples (Figures 14, 15).

Figure 9: Efficiency ratios for scalar and propagator observables in 2D critical regime; ratio approaches unity post-training.

Figure 10: Corresponding efficiency ratio in 3D near criticality, with action density remaining a limiting observable due to residual bias.
Cross-Volume Generalization
Fully convolutional score networks allow for cross-ϕ45 generalization. Training on multiple small lattice sizes enables sampling on much larger, unseen volumes without retraining. In both 2D and 3D, diffusion models trained on small sizes (ϕ46) successfully transfer to ϕ47 and ϕ48, reproducing propagators and scalar observables comparable or superior to in-distribution training, with exceptions only in the broken-phase susceptibility (Figures 19, 20, 21, 22, 23, 24, 25).

Figure 11: Single-size extrapolation baselines at ϕ49; multi-N(0,Σ2(t)1)0 trained model outperforms all single-size baselines.
Cross-volume models yield high MALA acceptance rates at unseen sizes, confirming that convolutional architecture, with circular padding, encodes scale generalization effectively.
Implications and Future Directions
This study establishes score-based diffusion models as an effective non-Markovian sampling framework for interacting lattice field theories, particularly near criticality. The demonstrated cross-volume generalization provides a viable strategy for large-volume sampling, significantly reducing training costs. Practical implications include initializing Markov-chain samplers with diffusion-generated configurations, potentially eliminating lengthy thermalization phases.
Theoretically, the results highlight that incorporating physical locality and convolutional symmetry into diffusion-model architectures is critical for scaling and phase transfer. Residual biases in zero-mode sectors identify avenues for further improvements via advanced noise schedules or multi-scale architectures (e.g., EDM2).
Prospective developments include:
- Application of EDM/EDM2 parameterizations for improved training stabilization and reduced reverse SDE steps.
- Use of diffusion models for initializing HMC or cluster algorithms in equilibrium and dynamical field theories.
- Extension to gauge theories (QCD) using gauge-equivariant convolutional architectures.
- Embedding physical time for dynamical sampling, opening non-equilibrium and real-time lattice field theory research.
Conclusion
Variance-exploding score-based diffusion models, when equipped with fully convolutional architectures and validated against rigorous diagnostics, provide an efficient, scalable generative sampling solution for lattice field theories near criticality. Cross-volume training allows knowledge transfer from ensembles generated on small lattices to large volumes, as confirmed by observable-level comparisons and diagnostics. The practical and theoretical advantages illustrated here lay the groundwork for further adoption in broader classes of quantum field theories and advanced lattice simulations.