---
title: Transformer-Based Neural-Network Quantum States
url: https://www.emergentmind.com/topics/transformer-based-neural-network-quantum-states
type: topic
---

# Transformer-Based Neural-Network Quantum States

Searching arXiv for recent and foundational papers on transformer-based neural-network quantum states.
Transformer-based neural-network quantum states (NQS) are representations of many-body wave functions or density operators in which a Transformer maps discrete configurations, patch tokens, occupation strings, or informationally complete POVM outcomes to amplitudes, phases, determinants, or probability distributions. Within this umbrella, reported constructions include autoregressive wave functions with exact sampling, Vision-Transformer and patch-based log-amplitude ansätze optimized by variational Monte Carlo (VMC), probabilistic POVM formulations for unitary and Liouvillian dynamics, and fermionic architectures in which self-attention predicts backflow-inspired orbitals or corrections to a reference state [1912.11052], [2009.05580], [2208.01758], [2603.02316]. The resulting models have been used for quantum circuits, open systems, frustrated spin models, ab initio chemistry, impurity models, lattice gauge theory, anyonic chains, and correlated spin-fermion lattice models.

## 1. Representational paradigms

A central line of work represents the wave function autoregressively. In the gauge-invariant and anyonic constructions, the state is written as
\[
\Psi(\sigma_1,\dots,\sigma_N)=\prod_{i=1}^N \psi(\sigma_i\mid \sigma_{<i})
=\prod_{i=1}^N a_i(\sigma_i\mid \sigma_{<i})\,e^{i\theta_i(\sigma_i\mid \sigma_{<i})},
\]
which is exactly normalizable and exactly sampleable in one forward pass [2101.07243]. The multi-purpose Transformer Quantum State extends this factorization by conditioning on physical parameters \(J\), with amplitude obtained from conditional probabilities and phase from an additive phase head; this is the basis for pretraining on families of Hamiltonians rather than a single instance [2208.01758]. In electronic-structure settings, an analogous decomposition is used with \(q_\theta(s)=\prod_i p_\theta(s_i\mid s_{<i})\) and \(\psi_\theta(s)=\sqrt{q_\theta(s)}\,e^{i\phi_\theta(s)}\), optionally combined with a reference state of weight \(\alpha\) [2412.12248]. In ab initio quantum chemistry, the amplitude squared is modeled autoregressively by Transformer decoder layers, while the phase is provided by a separate MLP [2306.16705].

A second line of work models \(\log \Psi_\theta(\sigma)\) directly rather than through normalized conditionals. Vision-Transformer-based ansätze for long-range Ising chains output a real logarithmic amplitude and define \(\Psi_\theta(s)=\exp[\log \Psi_\theta(s)]\), so that Born-rule sampling proceeds by Metropolis-Hastings on ratios \(p_\theta(s')/p_\theta(s)\) [2407.04773]. Deep ViT encoder states for frustrated two-dimensional magnets and the convolutional transformer wave function similarly map a spin configuration to a complex or amplitude-plus-phase log-wave-function representation, again coupled to Monte Carlo sampling and local-energy evaluation rather than exact ancestral sampling [2310.05715], [2503.10462]. A related complex-valued ViT ansatz for impurity models computes \(\log \psi_{\boldsymbol\theta}(s)\) through Transformer encoder blocks followed by a \(\log\cosh\) aggregation [2408.13050].

A third line begins not from a wave function in the computational basis but from an informationally complete POVM representation. Carrasquilla et al. map an \(N\)-qubit state \(\rho\) to
\[
P(\vec a)=\mathrm{Tr}[M^{(\vec a)}\rho],
\]
where \(M^{(\vec a)}=M^{(a_1)}\otimes\cdots\otimes M^{(a_N)}\), and show that a local unitary induces a local quasi-stochastic update of \(P\) [1912.11052]. The open-system extension maps the Lindblad equation to a linear system of ODEs for \(P(x)\), enabling Transformer-based simulation of mixed-state dynamics and steady states in the POVM basis [2009.05580].

Fermionic models also motivate determinant-based formulations. In the Ancilla Layer Model study, the Transformer acts on composite local tokens and outputs spin-configuration-dependent single-particle orbitals \(\Phi_{r,\alpha}(\boldsymbol s)\), while the many-body amplitude is a Slater determinant \(\Psi_\theta(\boldsymbol s)=\det[n\star \Phi(\boldsymbol s)]\) [2603.02316]. This places transformer NQS at the intersection of autoregressive neural wave functions, direct log-amplitude models, and backflow-Slater constructions rather than in a single canonical form.

## 2. Architectural patterns

Reported architectures span encoder-only, decoder-only, autoregressive, and ViT-style designs. The multi-purpose TQS uses an encoder-only transformer with one-hot or scaled one-hot inputs for spins and physical parameters, sinusoidal positional encodings for spin positions, learned positional embeddings for parameter tokens, causal masking, and two output heads for conditional probabilities and phase increments [2208.01758]. The exact circuit-simulation approach uses an autoregressive Transformer encoder with just one layer, \(h=8\) heads, \(d_k=d_v=d_{\text{model}}/h\), residual connections, LayerNorm, and a final linear layer plus softmax producing \(P_\theta(a_k\mid a_{<k})\), with \(d_{\text{model}}\in\{16,32\}\) [1912.11052]. For open systems, the causal Transformer is stacked with \(n_{\text{layers}}\in\{1,2\}\), \(d_{\text{model}}\in\{32,64\}\), \(n_{\text{heads}}=4\), and positional vectors that may be sinusoidal or learned [2009.05580].

Patch-based and ViT-style tokenizations dominate in two-dimensional lattice problems. Deep ViT NQS for the \(J_1\)-\(J_2\) Heisenberg model partition an \(L\times L\) lattice into non-overlapping \(b\times b\) patches, project each patch into a \(d\)-dimensional token, process the resulting sequence by \(n_l\) encoder layers with multi-head factored attention and feed-forward networks, and collapse the final token sequence to a complex \(\log \Psi_\theta(\sigma)\) through a \(\log\cosh\) output head [2310.05715]. The thermodynamic-limit study likewise uses patch embeddings, \(n_\ell=8\) Transformer layers, \(h=12\) heads, \(d=72\), and a scalar complex readout \(A_\theta(\boldsymbol\sigma)\) with \(\Psi_\theta(\boldsymbol\sigma)=\exp[A_\theta(\boldsymbol\sigma)]\) [2602.02665]. In impurity models, each occupation-number configuration is reshaped into a binary image with \(N_p=4N_o\) horizontal patches, embedded into \(\mathbb C^d\), and processed by standard MHSA and MLP blocks with learned positional encodings [2408.13050].

Several studies modify attention to encode lattice structure directly. The long-range Ising-chain ViT imposes circulant attention matrices and cyclic symmetrization instead of explicit positional encodings, thereby enforcing periodic translational invariance [2407.04773]. The composite spin-fermion ansatz for the Ancilla Layer Model uses a factored site-only attention matrix with a learned or analytic distance-decay bias \(B_{ij}\), no query/key projections on content, \(n_\ell=4\) self-attention layers, \(h=12\), \(d=72\), and a feed-forward width \(4d=288\) with GELU activations [2603.02316]. The Spatial Attention mechanism inserts a learned inverse length scale \(\gamma\) into each attention head through the factor \(e^{-\gamma d(i,j)}\), where \(d(i,j)\) is Euclidean patch distance [2602.02665]. The scaling-law study uses the same distance-dependent kernel as the defining structural bias of the transformer Ansatz [2606.02794].

Hybrid convolution-attention variants have also appeared. The convolutional transformer wave function stacks Vis-Transformer-style blocks of the form Norm \(\rightarrow\) ConvUnit \(\rightarrow\) MHSA \(\rightarrow\) IRFFN, with \(3\times 3\) depthwise convolutions, \(1\times 1\) pointwise convolutions, relative positional encoding \(P_{ij}=p_{\Delta x,\Delta y}\), and a final pair-complex activation that yields \(\log A_\theta(\sigma)\) and \(\phi_\theta(\sigma)\) [2503.10462]. In electronic systems, a decoder-only Transformer with depth \(N_{\rm dec}\), embedding dimension \(d_{\rm emb}\), and \(N_h\) heads is used to parameterize corrections around a reference state, and in the largest reported runs the values were \(d_{\rm emb}=300\), \(N_{\rm dec}=4\), and \(N_h=10\) [2412.12248].

## 3. Optimization and training algorithms

The dominant optimization principle is VMC energy minimization. Across spin, fermionic, and chemistry settings, the loss is
\[
E(\theta)=\frac{\langle \Psi_\theta|H|\Psi_\theta\rangle}{\langle \Psi_\theta|\Psi_\theta\rangle},
\qquad
E_{\rm loc}(s)=\sum_{s'} H_{s,s'}\,\frac{\Psi_\theta(s')}{\Psi_\theta(s)},
\]
with gradients obtained from the log-derivative trick or from stochastic reconfiguration (SR) [2208.01758], [2407.04773], [2603.02316]. For exact-sampling autoregressive states, Monte Carlo does not require a Markov chain: samples are drawn token by token from the normalized conditional distributions [2101.07243]. In non-autoregressive log-amplitude models, sampling typically uses Metropolis-Hastings with local spin updates, global magnetization reversals, or problem-specific moves [2503.10462]. Ab initio chemistry combines autoregressive sampling with a data-centric parallelization scheme, heuristic parallel batch autoregressive sampling, CUDA local-energy kernels, and AdamW parameter updates [2306.16705].

SR and related natural-gradient methods have become central for large-scale transformer NQS. The large-scale ViT work derives the exact identity
\[
\delta\theta=\tau\,X\,(X^T X+\lambda I)^{-1}f,
\]
which replaces inversion of the \(P\times P\) covariance matrix by inversion of a \(2M\times 2M\) matrix and makes SR practical for \(P=267{,}720\) real parameters on the \(10\times10\) square-lattice \(J_1\)-\(J_2\) Heisenberg model [2310.05715]. The thermodynamic-limit and scaling-law studies use SR with MARCH or SPRING, \(M=2^{14}\) samples per step, and diagonal shift \(\lambda=10^{-4}\) [2602.02665], [2606.02794]. In the physics-informed electronic-state framework, \(\theta\) is optimized with SOAP rather than ADAM, while the single scalar \(\alpha_0=\tanh^{-1}(2\alpha-1)\) is updated by SGD [2412.12248].

Several transformer NQS depart from direct VMC. In the exact circuit-simulation approach, training is gate by gate: after applying the quasi-stochastic matrix \(O^{(i+1)}\) to the current modeled distribution, a fresh Transformer is fitted by minimizing the reverse KL divergence
\[
\mathrm{KL}(P^{(e)}_{i+1}\,\|\,P_{\theta_{i+1}}),
\]
with gradients estimated from exact samples and updates performed by Adam [1912.11052]. For open systems, the Liouvillian can be treated either dynamically, through a forward-backward trapezoid update and the loss \(\mathcal C_{\rm dyn}\), or variationally, through the steady-state objective \(\mathcal C_{\rm var}=\|L\,P_\theta\|_1\), optimized by ADAM [2009.05580]. Impurity models employ a subspace-expansion scheme rather than Metropolis: an active subspace \(\mathcal S\) of the \(N_s\) most probable bit strings is iteratively enlarged by all Hamiltonian-connected states, truncated, and then used for SR in the restricted space [2408.13050].

Pretraining has been introduced to address difficult optimization landscapes. In the hybrid numerical/experimental framework, a patched autoregressive transformer is first optimized by a data-driven loss
\[
\mathcal L_{\rm pre}
=\alpha\{\mathcal D_{\rm KL}\ \text{or}\ \mathcal W\}(q_Z\|\lvert\Psi\rvert^2)+\beta C_{\rm corr},
\]
which combines computational-basis snapshots with correlations from a second measurement basis, and only then refined by Hamiltonian-driven VMC [2406.00091]. This two-stage procedure was designed to improve access to the sign structure and to accelerate convergence in two-dimensional systems.

## 4. Symmetry, inductive bias, and interpretability

A persistent theme is that physical symmetry is usually enforced explicitly. In gauge theories and anyonic chains, local constraints are built into the autoregressive generation process by a binary mask \(M(\sigma_{<k},z)\) that zeros out gauge-violating or fusion-forbidden next tokens before local renormalization [2101.07243]. This construction guarantees that every sampled prefix remains in the physical Hilbert space, and it yields exact representations for ground and excited states of the 2D and 3D toric codes and the X-cube fracton model [2101.07243]. In open-system simulations on two-dimensional lattices, autoregressive ordering breaks geometric symmetries, so translational and rotational symmetry are partially restored by averaging over up to eight snake-like orderings, producing a String State ansatz [2009.05580]. In frustrated spin systems, exact lattice symmetries are often imposed by explicit projection,
\[
\Psi_\theta^{\rm sym}(\boldsymbol\sigma)=\frac{1}{|G|}\sum_{g\in G}\Psi_\theta(g\cdot \boldsymbol\sigma),
\]
with \(G\) taken as translations and point-group operations such as \(C_{6v}\) or \(C_{4v}\) [2602.02665]. Symmetry projection over translations, point-group operations, and spin parity is also used in deep ViT NQS for the square-lattice \(J_1\)-\(J_2\) model [2310.05715].

Inductive biases are likewise encoded directly into the architecture or basis. The long-range ViT imposes circulant attention and cyclic-shift symmetrization, rather than relying on learned positional structure, to respect periodic translational invariance [2407.04773]. The Spatial Attention mechanism introduces a single learned length scale \(\gamma\) per head per layer, so that long-range patch coupling is suppressed at initialization and correlations are built up from short to long range [2602.02665]. The convolutional transformer wave function places a depthwise convolution before attention and an inverted-residual feed-forward block after attention, explicitly combining short-range and global couplings [2503.10462]. In the Ancilla Layer Model, each local state \(s_i=(n_{i\uparrow},n_{i\downarrow},S^z_{1,i},S^z_{2,i})\) is compressed to an integer token \(t_i\in\{0,\dots,15\}\), allowing self-attention to operate naturally on a composite local Hilbert space of dimension \(\mathcal V=16\) [2603.02316].

Interpretability has been pursued most directly through basis design. Sobral, Perle, and Scheurer construct a physics-informed second-quantized basis around either a Hartree-Fock or strong-coupling reference state and write
\[
\ket{\Psi_{\{\theta,\alpha\}}=\alpha\ket{\rm RS}
+\sqrt{1-\alpha^2}\sum_{s\neq\rm RS}\psi_\theta(s)\ket{s}.
\]
In this setting, the single scalar \(\alpha\) is an interpretable measure of how product-like the ground state is, low-order excitations \(j\le 2\) carry \(>90\%\) of the probability in the HF basis, and PCA of the hidden representation orders configurations by excitation class [2412.12248]. Reported results indicate that symmetry and interpretability generally arise from masks, projections, distance biases, or basis engineering rather than from generic self-attention alone. A related issue concerns sign structure: while some architectures use explicit phase heads or complex \(\log\Psi\), the large-scale triangular-lattice study found that the overlap with the classical three-sublattice sign and the Huse-Elser sign rule decays exponentially in \(L^2\), which indicates an intrinsically non-local sign structure [2602.02665].

## 5. Applications and benchmarked performance

Transformer NQS have been benchmarked on both exact-probabilistic and variational tasks. In quantum-circuit simulation, Carrasquilla et al. simulated GHZ and linear graph-state circuits up to \(60\) qubits and a \(6\)-qubit VQE circuit for the transverse-field Ising model. Small \(2\)-qubit circuits converged to machine precision, with fidelity approaching \(1\) and KL divergence approaching \(10^{-8}\); for GHZ and graph circuits up to \(60\) qubits, the classical fidelity \(F_c\) dropped roughly linearly in \(N\), reaching \(\approx0.9\) at \(N=60\) with \(d_{\rm model}=16\) and improving if \(d_{\rm model}=32\) [1912.11052]. In open systems, the probabilistic Transformer closely tracked exact QuTiP results for 1D Heisenberg dynamics up to \(N=40\), matched \(O(10^{-2})\) accuracy in \(3\times3\) 2D Heisenberg steady states with symmetry-averaged String States, and outperformed RBM-MCMC on the \(16\)-site 1D TFIM fixed point in the regime \(g/\gamma\in[1,2.5]\) [2009.05580].

In spin models, the range of systems is broader. The long-range Ising-chain ViT computed the full phase diagram for chain lengths up to \(N=200\), obtained \(J_c\approx-2.09\), \(\nu\approx1.08\), \(\beta\approx0.19\) on the ferromagnetic side and \(J_c\approx+4.75\), \(\nu\approx1.17\), \(\beta\approx0.09\) on the antiferromagnetic side at \(\alpha=2.5\), and achieved \(V\)-scores typically \(O(10^{-4})\), while an RBM under a fixed \(3\) min training budget remained near \(V\approx10^{-2}\) close to criticality [2407.04773]. For the \(10\times10\) square-lattice \(J_1\)-\(J_2\) Heisenberg model at \(J_2/J_1=0.5\), the deep ViT optimized with large-scale SR reached \(E/N=-0.497634(1)\) after \(\sim14{,}000\) SR steps, a reported state-of-the-art variational energy for this benchmark [2310.05715]. The Spatial Attention study extended VMC simulations of the triangular-lattice Heisenberg antiferromagnet to clusters up to \(42\times42\), obtained \(e_0\approx-0.55168(2)\) and \(\mathcal M_0=0.148(1)\), and also reported state-of-the-art energies for a \(J_1\)-\(J_2\) Heisenberg model on a \(20\times20\) square lattice [2602.02665]. The convolutional transformer wave function matched or slightly improved upon CNN(GELU) on the \(6\times6\) \(J_1\)-\(J_2\) model at equal cost, reached \(\varepsilon_{\rm rel}\simeq5\times10^{-5}\) and \(\sigma^2/N\simeq3\times10^{-5}\) on the \(10\times10\) problem after \(\sim10^4\) MinSR steps, and remained accurate in real-time dynamics of the 2D TFIM up to \(tJ\simeq1.0\) on \(6\times6\) and \(\simeq0.8\) on \(8\times8\) [2503.10462].

Gauge, anyonic, and composite spin-fermion models form another major application area. The gauge-invariant and anyonic autoregressive framework reproduced \(E_0=-2L^2\) for the \(11\times11\) toric code up to sampling noise, reached \(E_{\rm var}=-2.41421356\pm10^{-6}\) for the \(1+1\)D \(\mathrm{U}(1)\) quantum link model on \(6\) unit cells, matched tensor-network results up to \(160\) cells, identified the 2D \(\mathbb Z_2\) gauge-theory transition near \(h\approx0.34\), and extracted \(c=0.703\pm0.005\) for the \(\mathrm{SU}(2)_3\) anyonic chain [2101.07243]. For the Ancilla Layer Model, the backflow-Slater Transformer kept the relative energy error below \(10^{-4}\) on a \(42\)-site chain at doping \(\delta\simeq0.2857\) for \(J_K\) up to \(5.0\), and it characterized LL, LL\(^*\), and LE phases through structure factors, spin gaps, and central charges [2603.02316].

Electronic-structure and impurity problems have provided a distinct benchmark suite. The NNQS-Transformer for ab initio quantum chemistry matched or slightly outperformed NAQS on seven STO-3G molecules between \(14\) and \(30\) qubits, achieved chemical accuracy for BeH\(_2\) and average error to the FCI complete-basis limit below \(2.4\times10^{-3}\) Ha for H\(_2\) in \(56\)- and \(92\)-qubit bases, and demonstrated strong scaling with efficiency \(\ge84\%\) from \(4\) to \(32\) GPUs for a \(120\)-qubit benzene problem [2306.16705]. The physics-informed Transformer for electronic quantum states recovered ED energies with relative error \(\delta E_{\rm TQS}/E_{\rm ED}\lesssim10^{-3}\) across \(N_e\le12\) and \(0.04<t/U<0.2\), required only \(n_U^f\approx1\,812\) distinct configurations after convergence for \(N_e=30\), and converged in \(\sim2{,}000\) epochs on a single NVIDIA H100 in a few hours [2412.12248]. In impurity models, the ViT-NQS reached subspace relative errors of \(2\times10^{-3}\), \(5\times10^{-4}\), and \(1\times10^{-4}\) for \(N_s=512,1024,2048\) in the single-orbital Anderson model, achieved MPS-level accuracy in the three-orbital Anderson model with \(\mathcal O(10^3)\) parameters instead of \(\mathcal O(10^8)\), and reproduced core-level X-ray absorption spectra through a restricted excitation space [2408.13050].

## 6. Scaling behavior, limitations, and open questions

A recent line of work formulates scaling laws for transformer NQS in terms of the \(V\)-score,
\[
V=N\,\frac{\langle H^2\rangle-\langle H\rangle^2}{\langle H\rangle^2},
\]
and the total training compute \(f\). For transformer wave functions on square and triangular lattices up to \(20\times20\), the reported scaling law is
\[
V(f,N)=A\,f^{-\alpha}N^\beta,
\]
with data collapse after rescaling the compute axis as \(fN^{-\beta/\alpha}\) [2606.02794].

| Hamiltonian | \(\alpha\) | \(\beta\) |
|---|---:|---:|
| Square Heisenberg (\(J_2/J_1=0\)) | \(1.29(2)\) | \(1.31(3)\) |
| Square \(J_1\)-\(J_2\) (\(J_2/J_1=0.5\)) | \(0.86(3)\) | \(0.93(4)\) |
| Triangular Heisenberg (\(J_2/J_1=0\)) | \(0.54(1)\) | \(0.46(2)\) |
| Triangular \(J_1\)-\(J_2\) (\(J_2/J_1=0.125\)) | \(0.40(1)\) | \(0.42(2)\) |

Because \(\alpha\approx\beta\) in all reported cases, the study interprets the transformer Ansatz as size-consistent for the systems considered, and the decrease of \(\alpha\) with frustration is presented as a quantitative measure of representational difficulty [2606.02794]. This establishes scaling laws as a benchmarking framework for transformer-based variational states rather than only a description of isolated numerical experiments.

At the same time, recurrent limitations remain visible across the literature. Self-attention in its naive form scales as \(O(N^2)\) in the number of tokens, which is explicitly identified as a bottleneck in convolutional transformer wave functions and motivates approximate attention or hybrid conv-transformer designs for larger lattices [2503.10462]. In chemistry, Hamiltonian evaluation remains \(O(N^4)\) in the worst case and heuristic batch-autoregressive sampling can suffer from load imbalance at very large scale [2306.16705]. Near criticality, the physics-informed electronic-state approach requires larger \(n_U\) and larger networks to resolve many excitation classes, making sampling the bottleneck [2412.12248]. In open systems, one-dimensional autoregressive orderings break two-dimensional lattice symmetries unless additional string averaging is introduced [2009.05580]. Hybrid pretraining was proposed precisely because direct Hamiltonian optimization can face a rugged and complicated loss landscape in expressive transformer NQS [2406.00091].

Open directions named in the cited work include real-space topological models, excited states, spin and gauge symmetries, approximate or sparse attention for larger lattices, mixed-state and dynamics pretraining, and extensions to higher-dimensional fermionic problems [2412.12248], [2208.01758]. Taken together, the existing literature presents transformer-based NQS not as a single ansatz but as a methodological family whose practical success depends on the interaction between representation, symmetry handling, optimization algorithm, and the physical basis in which the model is defined.

Source: https://www.emergentmind.com/topics/transformer-based-neural-network-quantum-states