- The paper introduces a hybrid diffusion-warm MCMC method that achieves an order-of-magnitude reduction in thermalization time for the XY model.
- It combines temperature-conditioned diffusion models on small lattices with a few Wolff cluster updates to accurately restore long-range correlations.
- Experiments confirm enhanced accuracy for key observables such as spin-spin correlations, energy, and helicity modulus across various system sizes.
Diffusion-Warm MCMC for Efficient Thermalization in the XY Model
Introduction
The XY model, as a paradigmatic continuous spin system in condensed matter physics, presents substantial computational challenges for sampling equilibrium configurations in large system sizes. Traditional Markov chain Monte Carlo (MCMC) methods, such as those based on Wolff cluster updates, exhibit exponential thermalization times at low temperatures and are fundamentally limited by poor scaling with system size, particularly across the Berezinskii-Kosterlitz-Thouless (BKT) phase transition. Recent attempts at leveraging machine learning, especially diffusion-based score models, have achieved accurate sampling in fixed or closely related system sizes but fail to generalize in size or capture long-range correlations at scale. This paper introduces a robust hybrid—diffusion-warm MCMC—that exploits the strengths of both approaches: diffusion models deliver strong initializations, which, when followed by only a small number of Wolff MCMC updates, yield equilibrium ensembles orders of magnitude faster.
Methodology: Temperature-Conditional Diffusion Models and Size Extrapolation
A rotationally invariant, temperature-conditioned score-based diffusion model is trained on moderate-sized (20×20) XY lattices and across a range of inverse temperatures. The model leverages classifier-free guidance (CFG) to increase sample diversity and consistency with the target distribution, with score networks architected via periodic convolutions and sine-cosine embeddings for angular data, ensuring full compatibility with the hypertorus geometry of the XY phase space.
Inference on larger lattices is performed by seeding random configurations and applying the learned reverse SDE, followed by k steps of Wolff MCMC. The necessity of the MCMC correction arises from the inability of small-lattice-trained diffusion models to accurately capture long-wavelength spin waves and global thermodynamic correlations in larger systems. The paper demonstrates that this combination rapidly achieves equilibrium across a broad range of observables.

Figure 1: End-to-end depiction of the diffusion-warmed MCMC pipeline, including reverse diffusion and cluster updates.
Results: Rapid Thermalization and Accurate Physical Observables
Thermalization performance is characterized using a suite of physical observables: spin-spin correlation functions C(r), energy, magnetization, and the helicity modulus. The critical empirical finding is that diffusion-warmed MCMC achieves an order of magnitude reduction in thermalization time relative to conventional MCMC starting from randomly initialized states. This improvement holds at all examined system sizes and both sides of the BKT transition, including the challenging low-temperature regime where equilibration times are otherwise prohibitive.


Figure 2: Diffusion-warm sampling drastically reduces the required steps for thermalization and total runtime versus canonical MCMC, especially for larger lattices near the critical point.
Analysis of the spin-spin correlation function shows that naively extrapolating the diffusion model to larger systems yields C(r) curves that deviate from correctly thermalized samples, particularly in the long-range regime and at low temperature (where correlations are typically underestimated). However, only a modest number (20-40) of Wolff cluster updates rectify the discrepancies and accurately restore equilibrium long-range order.

Figure 3: Spin-spin correlations for diffusion model and MCMC at three temperatures reveal under/overestimation outside the training regime.

Figure 4: Wolff updates on diffusion samples rapidly correct long-range correlations, matching those from full thermalization.
In addition, the study examines the helicity modulus, a global observable highly sensitive to long-range order and phase coherence. For lattices larger than the training size, direct diffusion samples underestimate the helicity modulus at low temperature, but as with C(r), convergence is restored via a short Wolff MCMC refinement.

Figure 5: Equilibrium observables, including energy and helicity modulus, are correctly recovered for all system sizes with only 30 cluster updates post-diffusion.
Theoretical Implications and Comparison to Baselines
The results demonstrate that generative models are now able to efficiently warm-start physically meaningful ensembles for continuous-variable lattice systems, moving beyond the domain of discrete-spin models. Compared to tiling-based initializations—a common baseline for scaling generative samples—diffusion-warmed approaches better capture spatial correlations and avoid spurious periodic artifacts. Unlike direct variational or importance-sampling approaches, the diffusion-warm method is robust to mode collapse and does not require explicit normalization of sample densities.
The findings also challenge reported limitations in size generalization for learned models [see "Efficient Identification of Critical Transitions via Flow Matching" (Lee et al., 21 Aug 2025)], demonstrating that size extrapolation is viable when combined with a modest number of physically informed MCMC corrections.
Practical Significance and Outlook
For Gibbs sampling in the XY model and, by extension, other continuous-state or O(n) spin systems, the diffusion-warm strategy offers a scalable path to equilibrium sampling in parameter regimes where traditional MCMC becomes intractable. This unlocks large-lattice studies of phase transitions and topological phenomena without resorting to specialized hardware or massive compute resources.
The architectural elements—periodic convolutions, angle embeddings, temperature conditioning, and classifier-free guidance—provide a blueprint for adapting diffusion-warm MCMC to broader classes of physical systems, including frustrated lattices, molecular conformer ensembles, and possibly quantum many-body problems. Furthermore, the results highlight the need for future work on diffusion processes directly respecting the manifold structure and for models trained to extrapolate simultaneously in temperature and system size.
Conclusion
Diffusion-warm MCMC, as formulated in this work, enables efficient and accurate sampling of large-scale equilibrium configurations in the XY model via generative diffusion initialization and minimal MCMC refinement. This hybrid paradigm substantially improves thermalization speed by nearly an order of magnitude and generalizes beyond the training regime where previous deep generative approaches falter. The implications are significant for computational approaches to condensed matter, statistical mechanics, and potentially neural complex-valued networks, with further opportunities in extending to other continuous-variable models and exploring scaling limits in higher dimensions.