---
title: Sample-Based Neural Diagonalization (SND)
url: https://www.emergentmind.com/topics/sample-based-neural-diagonalization-snd
type: topic
---

# Sample-Based Neural Diagonalization (SND)

Sample-based Neural Diagonalization (SND) denotes a class of neural-network-assisted strategies for reducing the cost of diagonalization in quantum many-body computation. In the literature represented by Stratis et al., SND is a feed-forward eigenvalue surrogate used inside Monte Carlo sample generation for the spin–fermion, or double-exchange, model, replacing repeated exact diagonalization of a fermionic matrix by direct prediction of its spectrum [2206.07753]. In later work on sample-based diagonalization (SBD), SND instead denotes an autoregressive neural sampler that selects basis configurations for projected diagonalization, while adaptive-basis SND (AB-SND) augments this with a learnable basis transformation [2508.12724]. This suggests that the unifying idea is not a single canonical architecture, but the use of learned samples or learned spectra to make spectrally informed computation tractable.

## 1. Scope of the term and problem classes

The term SND is used in two distinct settings. In the spin–fermion setting, the computational bottleneck is the repeated evaluation of all single-particle eigenvalues for a Hamiltonian that depends on a classical spin configuration. In the SBD setting, the bottleneck is the selection of a small but effective basis subspace whose projected Hamiltonian yields an accurate ground-state estimate. In both cases, the learned model is coupled directly to a diagonalization-related step rather than being used only as a generic regressor or classifier.

| Setting | Learned object | Computational role |
|---|---|---|
| Spin–fermion model | Full spectrum $\{\hat\lambda_i(S)\}$ or free energy $\hat F(S)$ | Replace exact diagonalization during Monte Carlo proposals |
| SBD for quantum Ising models | Autoregressive sampler $p_\theta(x)$ | Select basis states for projected diagonalization |
| AB-SND | Sampler plus basis transformation | Concentrate the ground state before projected diagonalization |

A common source of confusion is to assume that SND names a single fixed algorithm. In the cited literature, it instead names two related but non-identical constructions: one predicts eigenvalues of a configuration-dependent Hamiltonian, and the other learns how to sample basis states that make a projected Hamiltonian accurate. The overlap lies in the reliance on neural networks to extract spectrally relevant information from samples.

## 2. Spin–fermion SND as neural eigenvalue prediction

In the spin–fermion formulation, the system consists of $N$ classical unit spins $S=\{S_1,\dots,S_N\}$ with $S_i\in\mathbb{R}^3$ and $|S_i|=1$, coupled to spinful fermions on the same sites. The single-orbital spin–fermion, or double-exchange, Hamiltonian is
$$
H(S) = -t \sum_{\langle i,j\rangle,\sigma} (c_{i\sigma}^\dagger c_{j\sigma} + h.c.)
+ \frac{J}{2} \sum_{i,\alpha,\beta,\gamma} S_i^\gamma c_{i\alpha}^\dagger \sigma^\gamma_{\alpha\beta} c_{i\beta},
$$
with $t=1$, Hund coupling $J$, Pauli matrices $\sigma^\gamma$, and nearest-neighbor pairs $\langle i,j\rangle$. For each classical-spin update in a Monte Carlo sweep, one must recompute all single-particle eigenvalues $\epsilon_i(S)$ of the $2N\times 2N$ fermion matrix, and standard dense diagonalization scales as $\mathcal{O}((2N)^3)\sim\mathcal{O}(N^3)$ [2206.07753].

Three surrogate models are compared in this formulation: a linear Heisenberg model fitting $F(S)$ as $\sum_{|i-j|=r} J_r S_i\!\cdot\!S_j$, an $N1$ feed-forward network predicting free energy directly, and an $N1Eigenvalues$ model predicting the full spectrum. The SND construction is the eigenvalue model. Its input is the flattened real vector
$$
x\in\mathbb{R}^{3N}: x=[S_1^x,S_1^y,S_1^z,\dots,S_N^z].
$$
The architecture is fully connected with one hidden layer of size $h$, Softplus activation $f(y)=\ln(1+e^y)$, and a linear output layer of size $2N$ interpreted as predicted eigenvalues. After inference, the $2N$ outputs are sorted so that $\hat\lambda_1\le \dots \le \hat\lambda_{2N}$.

Training uses exact-diagonalization data generated at each temperature $T=1/\beta$ from three independent Metropolis–Hastings chains. After a warm-up of $1000\cdot N$ proposals, the procedure generates $N_s=10^4$ samples for training and $N_s=10^3$ for validation and test, with thinning by saving one in every $N/r$ proposals, where $r$ is the acceptance ratio. Each saved spin configuration is labeled by its full eigenvalue set and free energy. Data augmentation exploits translation invariance and global $SO(3)$ rotation invariance using a discrete Euler-angle grid,
$$
\alpha\in\{0,\pi\},\quad \beta\in\{0,\pi/2,\pi,3\pi/2\},\quad \gamma\in\{0,\pi\},
$$
excluding the identity, which augments each raw sample by up to 23 non-trivial variants.

The training loss for the eigenvalue model is the mean absolute error per eigenvalue,
$$
L = \frac{1}{2N}\sum_{i=1}^{2N} |\hat\lambda_i-\lambda_i^{exact}|,
$$
while mean squared error on a held-out test set is monitored as a secondary metric. Optimization uses Adam with initial learning rate $10^{-3}$ and exponential decay
$$
lr(t)=lr(0)\cdot 0.9995^t,
$$
for $2\cdot 10^4$ epochs. Recommended hyperparameters are specified for both one and two dimensions: for 1D with $N=20$, $N1$ uses $h=60, n_b=20$ and $N1Eigenvalues$ uses $h=80, n_b=50$; for 2D with $N=6\times 6=36$, $N1$ uses $h=108, n_b=20$ and $N1Eigenvalues$ uses $h=144, n_b=20$.

## 3. SND as learned basis selection in sample-based diagonalization

In the later SBD formulation, the Hilbert space is expressed in the computational $(\sigma^z)$ basis $\{|x\rangle\equiv|x_1x_2\dots x_N\rangle\}$ with $x_i\in\{0,1\}$. A standard testbed is the transverse-field Ising model
$$
H \;=\; -\sum_{\langle i,j\rangle}J_{ij}\,Z_iZ_j \;-\; h\sum_i X_i,
$$
whose full Hilbert dimension is $2^N$. One selects a subset of $S$ basis states
$$
\mathcal{S}=\{|x^{(1)}\rangle,\dots,|x^{(S)}\rangle\},
$$
defines the projector
$$
P_{\mathcal{S}}=\sum_{l=1}^S |x^{(l)}\rangle\langle x^{(l)}|,
$$
and forms the projected Hamiltonian
$$
H_{\mathcal{S}} = P_{\mathcal{S}} H P_{\mathcal{S}}, \qquad
[H_{\mathcal{S}}]_{lm}=\langle x^{(l)}|H|x^{(m)}\rangle.
$$
Diagonalizing the $S\times S$ matrix $H_{\mathcal{S}}$ yields eigenvalues $E_{\mathcal{S}}^{(0)}\le E_{\mathcal{S}}^{(1)}\le \dots$, with $E_{\mathcal{S}}^{(0)}$ a variational upper bound to the true ground-state energy $E_0$ and $E_{\mathcal{S}}^{(0)}\to E_0$ as $S\to 2^N$ [2508.12724].

The central issue is how to choose the $S$ basis states so that $E_{\mathcal{S}}^{(0)}$ is accurate while $S\ll 2^N$. Standard SBD samples configurations from a heuristic distribution, such as squared amplitudes of a variational neural-quantum-state, but this fails when the true ground state is delocalized in the computational basis. SND replaces the fixed heuristic sampler by an autoregressive neural network $p_\theta(x)$ trained to focus on bitstrings with large ground-state weight, with the explicit target
$$
p_\theta(x)\approx |\psi_0(x)|^2, \qquad \psi_0(x)=\langle x|\psi_0\rangle.
$$

Training proceeds in $K$ batches of $S$ sampled bitstrings each. For batch $k$, the sampled set $\mathcal{S}^{(k)}=\{x^{(k,1)},\dots,x^{(k,S)}\}$ defines a subspace Hamiltonian whose lowest eigenvalue is $E^{(k)}$. The loss is
$$
L(\theta)
=
\frac{1}{K}\sum_{k=1}^K
\Bigl(E^{(k)}-\overline{E}\Bigr)
\sum_{l=1}^S
\log p_\theta\bigl(x^{(k,l)}\bigr),
$$
with $\overline{E}=\tfrac{1}{K}\sum_k E^{(k)}$. The corresponding REINFORCE-style gradient is
$$
\nabla_\theta L
\approx
\frac{1}{K}\sum_{k=1}^K
\bigl(E^{(k)}-\overline{E}\bigr)
\sum_{l=1}^S
\nabla_\theta \log p_\theta\bigl(x^{(k,l)}\bigr).
$$
Minimizing this objective shifts probability mass toward sample sets that yield low projected energies.

AB-SND extends this by introducing a parameterized basis transformation optimized jointly with the sampler. The paper considers single-spin and non-overlapping two-spin rotations, which are tractable on classical computers, and also more expressive global unitaries that can be implemented using quantum circuits. The stated purpose is to make the ground-state wave function more concentrated, thereby enlarging the regime in which sample-based diagonalization remains effective.

## 4. Neural architectures, observables, and algorithmic use

The two SND formulations differ sharply in architecture. In the spin–fermion case, the network is a shallow fully connected regressor with a single hidden layer and linear spectral output. In the SBD case, the network is a Transformer encoder with causal masking, using 2 layers, 4 heads, and embedding size 64 to realize the autoregressive factorization
$$
p_\theta(x_1,\dots,x_N)=\prod_{i=1}^N p_\theta(x_i\mid x_{<i}).
$$
At step $i$, the model sees $(x_1,\dots,x_{i-1})$ plus optional conditioning for AB-SND, and output logits are converted to probabilities by a softmax, optionally temperature-scaled during inference [2508.12724].

In the spin–fermion setting, SND is embedded directly into thermodynamic estimation. Given exact eigenvalues $\{\epsilon_i(S)\}$, the fermionic free energy is
$$
F(\beta,S)= -\frac{1}{\beta}\sum_{i=1}^{2N}\ln[1+e^{-\beta(\epsilon_i(S)-\mu)}].
$$
SND replaces $\epsilon_i(S)$ by $\hat\lambda_i(S)$ and computes
$$
\hat F(S)= -\frac{1}{\beta}\sum_i \ln[1+e^{-\beta(\hat\lambda_i-\mu)}].
$$
The resulting surrogate free energy is used in the Metropolis acceptance step,
$$
\min[1,\exp(-\beta(\hat F'-\hat F_{old}))],
$$
and the same replacement is propagated into the formulas for average internal energy, specific heat, magnetization $M$, and staggered magnetization $M_s$ [2206.07753].

In the SBD setting, by contrast, the learned model is not predicting eigenvalues of the full Hamiltonian. It is selecting a basis in which a smaller projected Hamiltonian can be assembled and diagonalized. The optimization target is therefore energetic rather than reconstructive: the reported relative error is
$$
\epsilon = \frac{|E_{est}-E_{exact}|}{|E_{exact}|}.
$$
This distinction matters conceptually. In the spin–fermion formulation, the network acts as a surrogate for a spectral map $S\mapsto \{\epsilon_i(S)\}$; in the SBD formulation, it acts as a proposal mechanism for the subspace itself.

## 5. Benchmarks, accuracy, and scaling behavior

For the spin–fermion model, the reported results show a strong separation between one and two dimensions. In 1D with $N=20$, $N1Eigenvalues$ reaches test-set $\mathrm{MSE}\lesssim 10^{-6}$ and mean absolute error per eigenvalue of order $10^{-4}$–$10^{-5}$. Over $T\in[0.02,1.0]$, the reported errors are $\lesssim 1\%$. In 2D with $N=36$, similar MSE levels are obtained at all temperatures, with slightly higher errors at the lowest $T$; over $T\in[0.05,1.0]$, $N1Eigenvalues$ yields errors $\lesssim 2\%$ down to $T=0.05$. For thermodynamic observables in 1D, all models reproduce $\langle E\rangle$, $C_V$, $|M|$, and $|M_s|$ within statistical error using $10^5$ Monte Carlo samples. In 2D, only $N1Eigenvalues$, and to a lesser extent the RKKY-inspired Heisenberg model, remain accurate for $|M|$ and $|M_s|$ below $T\sim 0.1$, whereas the free-energy network $N1$ fails in the low-temperature regime despite low MSE on $F$. The computational comparison is $\mathcal{O}(N^3)$ per update step for exact diagonalization versus $\mathcal{O}(N\cdot h)$ for SND inference; for $N=36$ and $h=144$, the reported proposal-stage speed-up is $\gtrsim 10$–$100\times$, with end-to-end acceleration depending on how many proposals are retained versus discarded [2206.07753].

For SBD-based SND and AB-SND, the benchmarks cover 1D ferromagnetic TFIM with periodic boundaries up to $N=100$, 2D square-lattice TFIM with open boundaries at $N=100$, and a 2D Edwards–Anderson spin glass with $J_{ij}\sim \mathcal{N}(0,1)$ at $N=64$. The reported dependence on sample count $S$ shows that, for $h=0.5$ where the ground state is concentrated, all SBD variants improve as $S$ grows. Standard SBD from NQS sampling saturates around $\epsilon\sim 10^{-2}$ when $S\sim 10^4$ in the 2D-EAM. SND matches NQS-SBD for small $h$ but fails for the disordered spin-glass regime, with error $\gtrsim 10^{-1}$. AB-SND with single-spin rotations outperforms both, reaching $\epsilon<10^{-3}$ with $S\sim 10^4$. As a function of transverse field, standard methods fail as $h\to\infty$, while AB-SND maintains small $\epsilon$ from $h\ll 1$ to $h\gg 1$, with only a peak near the quantum phase transition $h\approx h_c$ and still reported error $\lesssim 10^{-2}$ for $S=10^4$. For scaling with system size, the paper reports that achieving $\epsilon<1\%$ in the 1D TFIM at $h=0.5$ requires exponentially growing $S$ for standard NQS-SBD beyond $N\approx 60$, whereas AB-SND with adaptive single-spin rotations needs essentially constant $S$ up to $N=100$. Additional results indicate that non-overlapping two-spin adaptive rotations further reduce $\epsilon$ near criticality, and that a proof-of-concept shallow VQE-style circuit on 6 qubits with 16 samples yields $\epsilon<10^{-3}$ except extremely close to $h_c$ [2508.12724].

## 6. Limitations, misconceptions, and extensions

A common misconception is that low scalar free-energy error is sufficient for reliable thermodynamics. The spin–fermion results explicitly counter this in two dimensions: the free-energy network $N1$ can exhibit low MSE on $F$ and still fail in the low-$T$ regime for $|M|$ and $|M_s|$. The corresponding interpretation given in the source is that SND works best when the target is the full spectrum, because a richer target retains more information than the scalar free energy. The same source also states that simple 1D chains can be captured even by low-dimensional Heisenberg fits, whereas in 2D the spectrum network is crucial. Reported limitations include overfitting of $N1$ in 2D, training cost scaling as $\mathcal{O}(N^2\cdot N_s)$ because each label contains $2N$ eigenvalues, and a practical limit around $N\sim 10^2$ [2206.07753].

In the SBD line of work, a different misconception would be to equate SND with exact reconstruction of wave-function amplitudes. The stated objective is energy-based: SND and AB-SND target low projected energies rather than direct amplitude reconstruction. Reported limitations include duplicate samples as $S$ grows, non-convex optimization landscapes, and degradation of plain SND near critical points or in spin-glass regimes. The paper notes that increasing the effective temperature in the final softmax dramatically raises the yield of unique samples while barely affecting $\epsilon$. It further suggests advanced RL or second-order optimizers for difficult landscapes, deeper entangling circuits on real quantum hardware for greater accuracy, and extensions to excited states via deflated diagonalization or to time evolution by projecting the propagator [2508.12724].

The two lines of work also point to different extension paths. In the spin–fermion setting, retraining on spectra is proposed for other quantum-classical hybrids such as Holstein and Kondo-lattice models, with the additional suggestion that one may pre-train at high temperature and reuse the spectrum network across temperatures, and that equivariant architectures such as graph-neural networks could incorporate lattice symmetries more directly. In the SBD setting, AB-SND indicates that basis adaptation can be as important as sampling itself when the target state is delocalized. Taken together, these results support a narrow but technically significant conclusion: within the current literature, SND is best understood as a neural mechanism for concentrating spectral relevance—either by predicting the spectrum attached to a sampled configuration, or by sampling the basis states on which a projected diagonalization should be performed.

Source: https://www.emergentmind.com/topics/sample-based-neural-diagonalization-snd