---
title: Projector Variational Ansatz (PVA) in Many-Body Physics
url: https://www.emergentmind.com/topics/projector-variational-ansatz-pva
type: topic
---

# Projector Variational Ansatz (PVA) in Many-Body Physics

Projector Variational Ansatz (PVA) denotes a family of variational constructions in which a reference state is acted on by a parametrized projector or projector-like operator whose parameters are optimized to approximate a ground state. In the literature represented here, the term covers matrix-product projected states (MPPS), tensor-product projected states (TPPS), projector–variational formulations for non-linear wave functions, auxiliary-field projector slices, and a one-ancilla quantum-circuit ansatz. Across these variants, the projector is not introduced in isolation: it is paired with an input state that already contains selected physical structure, such as long-range entanglement, fermionic sign information, or a mean-field broken-symmetry pattern, while the variational projector supplies controlled corrections [1201.6121, 1407.4107, 1610.09326, 2308.08594, 2606.07084].

## 1. Formal definitions and historical lineages

The earliest formulation in this set is the one-dimensional MPO-based projected wave function of 2012. There the ansatz is
\[
|\Psi_{\rm PVA}\rangle=\Bigl(\prod_{i=1}^N P_i\Bigr)|\Psi_0\rangle
=\Bigl(\prod_{i=1}^N \sum_{s_i,s'_i}W_i^{s_i,s'_i}\,|s_i\rangle\langle s'_i|\Bigr)|\Psi_0\rangle,
\]
with a reference state $|\Psi_0\rangle$ and a chain of local projectors in matrix-product-operator form. This formulation is called Matrix-Product Projected States (MPPS) in the paper [1201.6121].

Sikora et al. generalized the construction to higher-dimensional tensor networks in 2014. For a lattice basis $|\sigma\rangle\equiv|\sigma_1,\dots,\sigma_N\rangle$, the TPPS form is
\[
\Psi_{\rm PVA}(\sigma)\equiv \langle \sigma|P_T|\phi_0\rangle
=\mathrm{tTr}\prod_{i=1}^N P[i]^{\sigma_i}\times\langle \sigma|\phi_0\rangle
\equiv W(\sigma),
\]
where $P_T\equiv\otimes_i P[i]$ is written as a tensor-product projector and $\mathrm{tTr}$ contracts the virtual indices. In that paper the same construction is called TPPS, or tensor-product projected states [1407.4107].

Later work broadened the meaning of PVA beyond diagonal tensor-network reweighting. Schwarz, Alavi and Booth recast projector Monte Carlo as minimization of a positive-definite Lagrangian,
\[
\mathcal{L}[\Psi(\{Z_\sigma\})]
=
\langle \Psi|\hat H|\Psi\rangle
-
E_0\bigl(\langle \Psi|\Psi\rangle-A\bigr),
\]
so that gradient descent on a non-linear ansatz generalizes imaginary-time projection [1610.09326]. The 2023 auxiliary-field formulation built the wave function from a fixed number $N_l$ of optimized projector slices,
\[
|\psi(\alpha)\rangle
=
\sum_{\{x^1,\dots,x^{N_l}\}}
\Bigl[\prod_{l=1}^{N_l} B^l(x^l;\alpha)\Bigr]\,|\phi_0\rangle,
\]
and also wrote it as
\[
|\psi(\alpha)\rangle = \prod_{l=1}^{N_l} e^{-\tau^l\mathcal{H}^l(\alpha)}|\phi_0\rangle.
\]
This form was presented as a variational replacement for a long stochastic auxiliary-field projection [2308.08594].

A distinct quantum-computing usage appeared in 2026. Dumontier et al. defined a depth-$M$ unitary on $n+1$ qubits,
\[
A(\vec\theta,\vec\delta,\vec\phi)\equiv A_M\cdots A_1,
\]
with one ancilla qubit and layer generators
\[
G_k(\theta,\delta,\phi)=e^{i\phi(I\otimes X)}\,e^{i\delta(I\otimes Z)}\,e^{i\theta(H_k\otimes Z)}.
\]
The effective projector on the data register is
\[
P(\vec\theta,\vec\delta,\vec\phi)
=
(I\otimes\langle 0|)\,A(\vec\theta,\vec\delta,\vec\phi)\,(I\otimes|0\rangle),
\]
followed by post-selection or amplitude amplification [2606.07084].

| Formulation | Defining expression | Representative paper |
|---|---|---|
| MPPS | $|\Psi_{\rm PVA}\rangle=(\prod_i P_i)|\Psi_0\rangle$ | [1201.6121] |
| TPPS | $\Psi_{\rm PVA}(\sigma)=\langle \sigma|P_T|\phi_0\rangle$ | [1407.4107] |
| Non-linear projector–variational form | minimize $\mathcal L[\Psi(\{Z_\sigma\})]$ | [1610.09326] |
| Auxiliary-field slice PVA | $|\psi(\alpha)\rangle=\sum_{\{x^l\}}[\prod_l B^l(x^l;\alpha)]|\phi_0\rangle$ | [2308.08594] |
| Quantum PVA | $|\psi\rangle=P(\vec\theta,\vec\delta,\vec\phi)|\psi_{\rm init}\rangle$ | [2606.07084] |

Taken together, these formulations suggest a unifying interpretation: PVA is less a single ansatz than a design principle in which projector structure and prior physical information are distributed between a reference state and a variationally optimized filtering layer.

## 2. Reference states, local projectors, and circuit blocks

The reference state is central in every PVA variant. The 2012 and 2014 papers explicitly allow $|\Psi_0\rangle$ or $|\phi_0\rangle$ to be a Slater determinant, a Jastrow–Slater state, a mean-field BCS state, a Jastrow wave function for bosons, a Slater determinant or BCS/Pfaffian for fermions, or a Hartree–Fock or mean-field spin-density-wave state. The 2016 formulation uses a reference Slater determinant $|\Phi\rangle$ inside a Correlator Product State (CPS). The 2023 auxiliary-field construction starts from a simple reference Slater determinant such as RHF or UHF, while the 2026 quantum version acts on an easily prepared initial state $|\psi_{\rm init}\rangle$ [1201.6121, 1407.4107, 1610.09326, 2308.08594, 2606.07084].

In the MPO and tensor-product formulations, the local projectors carry auxiliary indices that build correlations across sites. In one dimension, each site $i$ has matrices
\[
W_i^{s_i,s'_i}= \bigl[W_i^{s_i,s'_i}[\alpha_{i-1},\alpha_i]\bigr],
\quad
\alpha_{i-1},\alpha_i=1,\dots,\chi.
\]
For periodic chains one contracts all $\{\alpha_i\}$, and when $\chi=1$ the ansatz reduces to a single-site product-state reweighting of $|\Psi_0\rangle$. In two dimensions, each square-lattice tensor $P[i]^{\sigma_i}$ is rank-4 in its virtual legs, with components $P[i]^{\sigma_i}_{\alpha_i\beta_i\gamma_i\delta_i}$ and bond dimension $D$; graphically each site carries four virtual legs and one physical leg [1201.6121, 1407.4107].

The CPS version retains the projector logic but replaces tensor legs by overlapping local correlators. Each correlator is diagonal in the computational basis,
\[
\hat C_\lambda
=
\sum_{\mathbf n_\lambda}
C_{\mathbf n_\lambda}\,\hat P_{\mathbf n_\lambda},
\qquad
\hat P_{\mathbf n_\lambda}=|\mathbf n_\lambda\rangle\langle \mathbf n_\lambda|,
\]
and the full wave function is
\[
|\Psi_{\mathrm{CPS}}\rangle
=
\Bigl(\prod_\lambda \hat C_\lambda\Bigr)|\Phi\rangle.
\]
Variational parameters may include both correlator amplitudes and orbital coefficients defining the RHF, UHF, or GHF reference [1610.09326].

The 2023 auxiliary-field formulation parameterizes each slice by three sets of free variables: $\lambda_i^l$, the strength of the on-site HS field at site $i$; $\tau_\sigma^l$, an overall time-step scale for spin channel $\sigma$; and $t_{ij}^{l\sigma}$, the matrix elements of an effective one-body kinetic Hamiltonian. The paper states that none of these parameters is constrained to be uniform or Hermitian, and that $\tau_\sigma^l$ can be positive or negative [2308.08594].

In the 2026 quantum-circuit version, one layer consists of three operations on the data-plus-ancilla register: the signal operator
\[
W_{k_j}(\theta_j)=e^{i\theta_j(H_{k_j}\otimes Z_a)},
\]
a phase shift on the ancilla
\[
S_{Z,a}(\delta_j)=e^{i\delta_j(I\otimes Z_a)},
\]
and a signal-processing rotation
\[
S_{X,a}(\phi_j)=e^{i\phi_j(I\otimes X_a)}.
\]
The ancilla is measured in the $Z$ basis, and only shots with ancilla outcome $0$ are retained [2606.07084].

## 3. Variational objectives and optimization procedures

For MPO- and tensor-based PVA, the optimization is cast in standard variational Monte Carlo language. In the TPPS formulation, the variational energy is
\[
E[\Psi]
=
\frac{\langle \Psi|H|\Psi\rangle}{\langle \Psi|\Psi\rangle}
=
\frac{\sum_\sigma |W(\sigma)|^2 E_{\rm loc}(\sigma)}
{\sum_\sigma |W(\sigma)|^2},
\]
with local energy
\[
E_{\rm loc}(\sigma)
=
\sum_{\sigma'}
\frac{W(\sigma')}{W(\sigma)}
\langle \sigma'|H|\sigma\rangle.
\]
Configurations are sampled with probability $p(\sigma)\propto |W(\sigma)|^2$, using Metropolis acceptance $\min\{1,|W(\sigma_{\rm new})|^2/|W(\sigma_{\rm old})|^2\}$. Optimization proceeds through logarithmic derivatives
\[
O_k(\sigma)=\partial \ln W(\sigma)/\partial \theta_k
\]
and the stochastic reconfiguration update
\[
\sum_l S_{kl}\,\delta\theta_l=-\gamma f_k,
\]
with covariance matrix $S_{kl}=\langle O_k^*O_l\rangle-\langle O_k^*\rangle\langle O_l\rangle$ and force $f_k=\langle O_k^*E_{\rm loc}\rangle-\langle O_k^*\rangle\langle E_{\rm loc}\rangle$ [1407.4107]. The 2012 MPO formulation uses the same VMC structure, writes the local energy as a stochastic average over connected configurations, and presents the SR update as a natural-gradient method [1201.6121].

The 2016 projector–variational formulation changes the perspective from direct minimization of $E$ to descent on the Lagrangian $\mathcal L$. For a linear expansion it recovers the standard first-order imaginary-time propagator,
\[
|\Psi^{(k+1)}\rangle
=
\bigl[\hat I-\tau_k(\hat H-E^{(k)}\hat I)\bigr]|\Psi^{(k)}\rangle,
\]
which the paper identifies with the power-method or projector-QMC master equation. For non-linear ansätze, parameter updates are implemented by stochastic gradients. To accelerate convergence, the method employs Nesterov’s momentum and RMSprop, with $\rho\approx 0.9$, $\epsilon\sim 10^{-6}$, and a parameter-wise learning rate $\tau_\sigma^{(k)}=\eta/\sqrt{E^{(k)}[g_\sigma^2]+\epsilon}$ [1610.09326].

The 2023 auxiliary-field PVA evaluates
\[
E(\alpha)=\frac{\langle \psi(\alpha)|H|\psi(\alpha)\rangle}{\langle \psi(\alpha)|\psi(\alpha)\rangle}
=
\frac{\langle E_L\cdot S\rangle_\rho}{\langle S\rangle_\rho},
\]
where $\rho(x,x')=|\langle x'|x\rangle|$, $S(x,x')=\mathrm{sign}\,\langle x'|x\rangle$, and $E_L(x,x')=\langle x'|H|x\rangle/\langle x'|x\rangle$. Reverse-mode automatic differentiation is used to obtain
\[
\partial_\alpha E
=
\frac{\langle S[\partial_\alpha E_L+(E_L-E)O]\rangle_\rho}{\langle S\rangle_\rho},
\qquad
O=\partial_\alpha \ln[\rho\cdot S].
\]
The paper states that typical runs use $\sim 2\times 10^5$ samples per gradient step, $\sim 180$ total steps, and $120$–$150$ walkers, with a small-batch stochastic gradient descent schedule such as $0.01$ for the first $80$ steps and $0.001$ for the next $100$ steps, together with gradient clipping for stability [2308.08594].

In the 2026 quantum setting, the cost function is the post-selected energy
\[
C(\vec\theta,\vec\delta,\vec\phi)
=
\langle \psi_{\rm init}|A^\dagger \Pi_0 (H_p\otimes I)\Pi_0 A|\psi_{\rm init}\rangle,
\qquad
\Pi_0=I\otimes |0\rangle\langle 0|.
\]
Adaptive growth of the circuit uses a commutator-gradient criterion: the gradient with respect to a new layer’s $\theta$ at $\theta=0$ is proportional to
\[
\partial C/\partial \theta_j|_{0}\propto \langle \phi_j|[H_p,H_{k_j}]|\phi_j\rangle.
\]
The algorithm then appends the operator with the largest gradient magnitude and re-optimizes all parameters [2606.07084].

## 4. Entanglement structure, sign encoding, and systematic improvability

A defining claim of the MPO and TPPS literature is that the projector need not generate all correlations from a product state. In the 2012 paper, the MPO projectors are said to improve the short-range entanglement of a given trial wave function while the long-range entanglement is contained in the initial guess. The 2014 TPPS paper sharpens this point: a pure tensor-product state with small bond dimension $D$ can only faithfully represent area-law entangled states, whereas an input state $|\phi_0\rangle$ that already contains long-range entanglement can considerably reduce the bond dimensions needed for comparable accuracy. The paper states in particular that area-law states in gapped two-dimensional systems can be captured by TPS with $D\sim O(10)$, while critical or Fermi-sea states that violate a strict area law can be handled by encoding their algebraic entanglement in $|\phi_0\rangle$ and keeping the projector network light, with $D\lesssim 3$–$4$ [1201.6121, 1407.4107].

For fermions, the same logic is applied to the sign structure. In the TPPS construction, $|\phi_0\rangle$ is chosen as a Slater determinant or BCS/Pfaffian whose amplitude already carries the correct sign structure under particle exchanges. Because $P_T$ is diagonal in the $\sigma$ basis, it multiplies each configuration amplitude by a positive weight $\mathrm{tTr}\prod P[i]^{\sigma_i}$ and therefore preserves all fermionic signs coming from $\langle \sigma|\phi_0\rangle$ [1407.4107].

Systematic improvability enters explicitly in the auxiliary-field formulation. There the number of slices $N_l$ is user-controlled, and the paper states that in the limit $N_l\to\infty$ the product $\prod_l e^{-\tau^l\mathcal H^l}$ can approach the true imaginary-time projector $e^{-\beta H}$. The energy variance
\[
\sigma^2\equiv \langle H^2\rangle-\langle H\rangle^2
\]
is monitored as a diagnostic, with $\sigma^2\to 0$ for an exact eigenstate; the reported trend is that $\sigma^2$ falls as $N_l$ grows. Even with $N_l=4$ slices, the paper reports variational energies within $10^{-3}t$ of exact or best-known values and $\sigma^2\approx O(10^{-1})$ [2308.08594].

The same 2023 work attributes to the optimized projector slices a capacity for automatic symmetry restoration and local-order detection. Starting from a symmetry-broken UHF reference with large staggered magnetization on a $10\times 10$ lattice, optimization of four slices makes the local $\langle S_i^z\rangle$ collapse to $O(10^{-2})$ on every site, recovering full SU(2) and translation symmetry. On a $12\times 4$ cylinder at doping $n=5/6$ with edge antiferro pinning fields, the ansatz reproduces the staggered magnetization profile and hole-density modulation nearly pixel-for-pixel in agreement with DMRG or AFQMC, without prespecifying a stripe wavelength [2308.08594].

A common misconception is that PVA is merely another name for a conventional MPS or TPS. The MPPS and TPPS papers make the distinction explicit: conventional MPS/TPS start from product states, whereas projector variational forms use a nontrivial parent state to carry long-range structure that the projector then refines [1201.6121, 1407.4107].

## 5. Reported benchmarks across lattice, molecular, and extended systems

The one-dimensional MPPS benchmarks were carried out for the spinless-fermion $t$–$V$ model on chains up to $L=202$. At $V/t=1$, which is the critical Luttinger-liquid regime, an MPPS with bond dimension $\chi=5$ reaches an energy error of order $10^{-4}$, whereas a conventional MPS needs $\chi\sim 20$–$30$ for comparable accuracy. At the critical point $V/t=2$, MPPS still significantly outperforms MPS at equal $\chi$. For density–density correlations, MPPS$(\chi=10)$ reproduces the algebraic decay $R^{-2\kappa}$ up to approximately $50$ sites, while MPS$(\chi=10)$ correlations saturate after approximately $5$ sites. In the gapped charge-density-wave phase $V/t>2$, MPPS captures the order parameter more accurately than MPS with the same $\chi$ [1201.6121].

The TPPS benchmarks of Sikora et al. cover two-dimensional bosons and fermions. For hardcore bosons in the half-filled $t$–$V$ model on a $4\times 4$ lattice, a simple two-parameter Jastrow state has approximately $10$–$20\%$ error near $V\approx 2$, TPS with $D=2$ yields approximately $2$–$3\%$ error, and TPPS with $D=2$ reduces the error to approximately $1\%$, matching TPS with $D=3$. On an $8\times 8$ lattice, TPPS with $D=2$ gives relative energy errors that agree with SSE QMC within $\lesssim 1\%$, and the structure factor $S(\pi,\pi)$ improves markedly over plain TPS with $D=2$. For spinless fermions in the same model, the SL-only TPPS is exact at $V=0$ and retains less than $5\%$ error up to $V\approx 2$; adding a Jastrow factor reduces the error further. For the half-filled two-dimensional Hubbard model on $4\times 4$, the paper compares SL, $d$-BCS, and SDW input states; without projection, SDW is best for large $U$, while with increasing $D$ all projected states converge toward the exact energy, and for $D=2$ the TPPS-SL and TPPS-$d$BCS energies are within $\lesssim 1\%$ of exact diagonalization at $U/t=4,10$ [1407.4107].

The non-linear projector–variational framework of Schwarz, Alavi and Booth was demonstrated on lattice and ab initio systems. For the half-filled two-dimensional Hubbard model on $98$ sites with $U=8t$, using overlapping $5$-site correlators and stochastically optimized GHF orbitals, the method converged to $97.9\%$ of the correlation energy reported by large-scale GFMC. For the one-dimensional Hubbard model on $22$ sites with $U=4t$, enlarging CPS plaquettes and allowing RHF, UHF, or GHF references pushed the energies toward exact DMRG values, and with $3$-site correlators plus an RHF reference the method yielded lower energy than previous VMC Linear-Method results. For the symmetric dissociation of an $\mathrm{H}_{50}$ chain in STO-6G, the method recovered more than $99\%$ of the DMRG correlation at stretched bonds and achieved absolute error not exceeding $1.1\,\mathrm{kcal/mol}$ per atom near equilibrium. For a periodic graphene sheet with $4\times 4$ $k$-point sampling, a double-$\zeta$ basis, an active space of $32$ localized C $2p_z$ orbitals, and about $67\,584$ correlator parameters, the sampled $2$-RDM showed only nearest-neighbour antiferromagnetic correlations surviving [1610.09326].

The 2023 auxiliary-field PVA was benchmarked on the two-dimensional Hubbard model for cylindrical and fully periodic supercells. The paper states that for $U=4$ the PVA energies are within $10^{-3}t$ of essentially exact AFQMC and slightly lower than current variational-Monte-Carlo and fixed-node DMC results, while at $U=8$ the energies remain competitive, within $\le 10^{-2}t$ of the best published variational states. The reported variances drop rapidly with $N_l$, which the authors use as evidence for systematic improvability [2308.08594].

The 2026 Projector Quantum Variational Ansatz was benchmarked against standard ADAPT-VQE on stretched-geometry molecules. On $\mathrm{H}_4$, PVA reached $\Delta E<1.6\times 10^{-3}\,\mathrm{Ha}$ in $8$ layers versus $15$ layers for ADAPT. On LiH, both methods reached chemical accuracy in $4$ layers, but PVA converged faster thereafter. On $\mathrm{BeH}_2$, PVA required $52$ layers while standard ADAPT-VQE stalled at $\gtrsim 200$ layers. On $\mathrm{H}_6$, PVA was approximately $2\times$ faster early, though both plateaued at a similar chemical-accuracy layer. Finite-shot SPSA runs with $5\times 10^5$ shots on $\mathrm{H}_4$ and LiH showed monotonic convergence to chemical accuracy with stable post-selection probability [2606.07084].

## 6. Relations to adjacent methods, limitations, and open directions

PVA repeatedly appears in direct comparison with alternative ansätze. The 2012 paper contrasts MPPS with Jastrow–Slater, conventional MPS/TPS, and correlator-product or entangled-plaquette states, arguing that the projected reference-state strategy can reach the same accuracy with much smaller bond or correlator size. The 2014 TPPS paper similarly emphasizes reduced bond dimension, a variational upper-bound guarantee through Monte Carlo sampling and SR, natural handling of fermion signs via the input wave function, and the ability to capture non–area-law and topological states provided the input state encodes them. In the 2023 work, the method is positioned against long AFQMC projections, current VMC states, and fixed-node DMC. In the 2026 work, the circuit ansatz is explicitly related to both ISQ-QSP and ADAPT-VQE: with a fixed choice of $H_{k_j}=H_p$ and QSP phase sequence it reproduces the standard ISQ-QSP filter, whereas setting $\delta_j=\phi_j=0$ for all $j$ makes each block reduce to $e^{i\theta_j H_{k_j}}$, exactly the ADAPT-VQE form [1201.6121, 1407.4107, 2308.08594, 2606.07084].

The limitations are equally explicit. In higher-dimensional tensor-network PVA, exact contraction is exponentially expensive in general; the 2012 paper points to TERG or corner-transfer-matrix schemes with cost scaling typically as $\chi^{D+1}$ or higher depending on lattice geometry, while the 2014 TPPS paper gives a contraction cost of approximately $O(ND^3D_f^3)$ for TRG with cutoff $D_f$. The 2014 paper also notes that optimization may get stuck in local minima for many parameters and that accuracy is limited by both the orbital entanglement in $|\phi_0\rangle$ and the expressivity of the projector network. The 2023 auxiliary-field variant notes matrix-product stability issues as $N_l$ grows, including re-orthonormalization and low-rank factorization concerns, and identifies a possible infinite-variance issue in the denominator $\langle S\rangle$; for $N_l\le 4$ it is reported as negligible, but longer slices may require “bridge-link” fixes. The same paper also remarks that sign-problem constraints could appear for truly large $\beta$-equivalent projections, even though no exponential sign problem arises for modest $N_l$ in the reported benchmarks. In the quantum-circuit setting, each PVA layer adds two extra CNOTs for the ancilla-controlled block relative to qubit-pool ADAPT, and post-selection or amplitude amplification introduces an additional resource trade-off [1407.4107, 2308.08594, 2606.07084].

Several extensions are stated directly in the tensor-network literature. The 2014 TPPS paper lists the use of symmetric tensors or global quantum-number projections to reduce parameters, hybridization with MERA layers in $P_T$, application to frustrated magnets, spin liquids, and topological Hubbard models, and time-dependent extensions for dynamic correlation functions. The 2023 auxiliary-field paper suggests a route to quantum chemistry by treating the ansatz as a systematically improvable non-orthogonal expansion of Slater determinants at polynomial cost. The 2026 quantum version suggests a corresponding route on near-term hardware: rather than constructing a state transition directly, the ansatz uses an ancilla to flag the good subspace and then extracts the data-register state by post-selection or amplitude amplification [1407.4107, 2308.08594, 2606.07084].

A plausible implication of this lineage is that PVA serves as a bridge concept between several traditions that are often treated separately—tensor-network VMC, projector Monte Carlo, auxiliary-field methods, and variational quantum circuits. What remains invariant across these settings is the central architectural choice: the variational object is a projector or projector-like filter, and its practical success depends on how effectively the reference state absorbs the long-range or sign-structured part of the many-body problem.

Source: https://www.emergentmind.com/topics/projector-variational-ansatz-pva