Papers
Topics
Authors
Recent
Search
2000 character limit reached

Ensemble Variational Quantum Eigensolver

Updated 12 July 2026
  • Ensemble VQE is a framework that optimizes a collection of quantum states collectively to target multiple eigenstates or related problem instances.
  • It encompasses various formulations—state-averaged, ancilla-purified, KD-VQE, and cross-instance methods—each tailored for specific excitation or ground state challenges.
  • Practical insights include trade-offs in weighting strategies, measurement overhead, and parameter transfer, affecting overall convergence and optimization landscape.

Ensemble variational quantum eigensolver (ensemble VQE) denotes a family of VQE formulations in which the optimization target is not a single variational state but an ensemble of states, an ensemble-defined low-energy subspace, or an ensemble of related problem instances. In the recent literature, the term is used for at least four distinct constructions: state-averaged excited-state VQE based on a shared unitary acting on orthonormal inputs, ancilla-purified single-circuit methods that determine many eigenpairs simultaneously, Boltzmann-weighted multi-candidate ground-state solvers such as Knowledge Distillation Inspired VQE (KD-VQE), and simultaneous-solving schemes that transfer parameters and measurements across multiple Hamiltonians (Rajamani et al., 22 Sep 2025, Xu et al., 2022, Li, 6 May 2025, Hutchings et al., 2024). This suggests that “ensemble VQE” is best understood as a unifying perspective on collective variational optimization rather than as a single canonical algorithm.

1. Conceptual scope

The literature uses “ensemble VQE” for related but non-identical objectives. In one line of work, the ensemble is a set of orthonormal trial states generated by applying one shared parametrized unitary to multiple reference states; the optimization target is an ensemble energy or Hamiltonian subtrace over that subspace. In another, the ensemble is a population of trial wavefunctions whose measurement resources are redistributed during optimization according to a Boltzmann law. In a third, the ensemble is the collection of states encountered along a seed VQE trajectory and reused to warm start multiple target instances (Rajamani et al., 22 Sep 2025, Li, 6 May 2025, Hutchings et al., 2024).

Variant Ensemble object Core mechanism
State-averaged / weighted ensemble VQE {U(θ)Φj}\{U(\theta)\lvert \Phi_j\rangle\} Minimize ensemble energy or subtrace
Ancilla-purified one-circuit method Orthogonal states labeled by ancillas Optimize one purified circuit and diagonalize subspace
KD-VQE Multiple trial wavefunctions Boltzmann-weighted shot allocation with virtual annealing
Simultaneous-solving ensemble Seed trajectory across related Hamiltonians Warm starts, pruning, and measurement reuse

This plurality matters because the strengths and failure modes depend on what is being ensembled. Subspace-based formulations are primarily designed for multiple low-lying eigenstates and excited-state structure. KD-VQE addresses exploration, variance control, and robustness in ground-state search. Cross-instance ensemble methods focus on transfer across related Hamiltonians rather than simultaneous optimization of a single spectrum.

2. Variational formulations

A central formulation is the weighted-ensemble objective

EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,

where wjwj+1w_j \geq w_{j+1}, H^\hat{H} is the target Hamiltonian, and EjEj+1E_j \leq E_{j+1} are exact eigenvalues. Writing Ψj(θ)=U(θ)Φj\lvert \Psi_j(\theta)\rangle = U(\theta)\lvert \Phi_j\rangle, the same objective can be expressed through the ensemble density operator

ρ(θ)=j=0K1wjΨj(θ)Ψj(θ),EC(θ)=Tr ⁣[ρ(θ)H^].\rho(\theta)=\sum_{j=0}^{K-1} w_j \ket{\Psi_j(\theta)}\bra{\Psi_j(\theta)}, \qquad E_C(\theta)=\operatorname{Tr}\!\left[\rho(\theta)\hat{H}\right].

For equal weights, the cost becomes the normalized Hamiltonian subtrace

ET(θ)=1Kj=0K1Ψj(θ)H^Ψj(θ)=1Kj=0K1Ej(θ),E_T(\theta)=\frac{1}{K}\sum_{j=0}^{K-1}\bra{\Psi_j(\theta)}\hat{H}\ket{\Psi_j(\theta)} =\frac{1}{K}\sum_{j=0}^{K-1}E_j(\theta),

which is invariant under rotations within the optimized KK-dimensional subspace (Rajamani et al., 22 Sep 2025).

A closely related subspace formulation appears in the ancilla-purified one-circuit method. There the objective is

C(θ)=k=0K1wkψk(θ)Hψk(θ),C(\theta)=\sum_{k=0}^{K-1} w_k \,\langle \psi_k(\theta) \lvert H \rvert \psi_k(\theta)\rangle,

with projector

EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,0

This is a generalized Rayleigh–Ritz bound over an optimized orthogonal subspace rather than a single vector (Xu et al., 2022).

KD-VQE adopts a different ensemble view. For EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,1 parameterized trial states EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,2 with EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,3, it defines a temperature-dependent mixed-state ansatz

EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,4

with EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,5, EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,6, Boltzmann weights

EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,7

and ensemble energy

EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,8

which satisfies EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,9 (Li, 6 May 2025).

3. State-averaged and subspace ensemble VQE

In state-averaged ensemble VQE, the same unitary wjwj+1w_j \geq w_{j+1}0 is applied to orthonormal initial states wjwj+1w_j \geq w_{j+1}1, so the prepared states remain orthonormal because unitaries preserve inner products. This eliminates the need for explicit overlap penalties when the shared-unitary construction is used. The equi-ensemble case wjwj+1w_j \geq w_{j+1}2 is distinguished because the objective becomes the subtrace wjwj+1w_j \geq w_{j+1}3, so optimization targets the optimal wjwj+1w_j \geq w_{j+1}4-dimensional eigensubspace rather than a prescribed ordering of circuit outputs (Rajamani et al., 22 Sep 2025).

The weighted-ensemble variant uses unequal nonnegative weights. The cited comparison argues that weighted optimization requires the ordering of weights wjwj+1w_j \geq w_{j+1}5 to align a priori with the eventual energy ordering of the prepared states, whereas equal weights remove that requirement. The same study further reports that unequal weights can lead to mode collapse toward strongly weighted states, sharp local minima or plateaus, and sensitivity to ordering guesses near avoided crossings or conical intersections; by contrast, the equi-ensemble landscape is described as smoother and focused on the subspace itself (Rajamani et al., 22 Sep 2025).

These differences were quantified in two benchmarks. For formaldimine with wjwj+1w_j \geq w_{j+1}6, equi-ensemble optimization reached the exact eigensubspace within wjwj+1w_j \geq w_{j+1}7 Hartree in approximately wjwj+1w_j \geq w_{j+1}8 iterations with 1-GUCCSD and approximately wjwj+1w_j \geq w_{j+1}9 iterations with 2-GUCCSD. Under ideal ordering, weighted 1-GUCCSD reached H^\hat{H}0 error H^\hat{H}1 Hartree and H^\hat{H}2 error H^\hat{H}3 Hartree in approximately H^\hat{H}4 iterations, while weighted 2-GUCCSD reduced both to H^\hat{H}5 Hartree but needed approximately H^\hat{H}6 iterations. Under incorrect ordering, weighted 1-GUCCSD produced H^\hat{H}7 errors up to H^\hat{H}8 Hartree even when H^\hat{H}9 remained EjEj+1E_j \leq E_{j+1}0 Hartree, and weighted 2-GUCCSD required approximately EjEj+1E_j \leq E_{j+1}1 iterations to recover EjEj+1E_j \leq E_{j+1}2 Hartree accuracy (Rajamani et al., 22 Sep 2025).

For the linear EjEj+1E_j \leq E_{j+1}3 chain in the Q-DFT setting with EjEj+1E_j \leq E_{j+1}4, equi-ensemble plus classical diagonalization reduced errors by more than two orders of magnitude for EjEj+1E_j \leq E_{j+1}5 Å and by approximately one order of magnitude for EjEj+1E_j \leq E_{j+1}6 Å, with approximately EjEj+1E_j \leq E_{j+1}7 fewer iterations in most regimes. The same study reported that weighted ensembles exhibited large, non-democratic errors dominated by the highest-energy state and exceeded EjEj+1E_j \leq E_{j+1}8 Hartree at EjEj+1E_j \leq E_{j+1}9 Å for the highest-energy state under the tested weighting scheme (Rajamani et al., 22 Sep 2025).

4. Ancilla-purified one-circuit realization

An important single-circuit realization of ensemble VQE introduces ancilla qubits to purify the trial states and enforce orthogonality by construction. The setup uses Ψj(θ)=U(θ)Φj\lvert \Psi_j(\theta)\rangle = U(\theta)\lvert \Phi_j\rangle0 physical qubits for the system and Ψj(θ)=U(θ)Φj\lvert \Psi_j(\theta)\rangle = U(\theta)\lvert \Phi_j\rangle1 ancillas, with ancilla Hilbert-space dimension Ψj(θ)=U(θ)Φj\lvert \Psi_j(\theta)\rangle = U(\theta)\lvert \Phi_j\rangle2. Choosing Ψj(θ)=U(θ)Φj\lvert \Psi_j(\theta)\rangle = U(\theta)\lvert \Phi_j\rangle3 suffices to label Ψj(θ)=U(θ)Φj\lvert \Psi_j(\theta)\rangle = U(\theta)\lvert \Phi_j\rangle4 target states. The initial state is prepared as Bell pairs between ancillas and selected physical qubits,

Ψj(θ)=U(θ)Φj\lvert \Psi_j(\theta)\rangle = U(\theta)\lvert \Phi_j\rangle5

equivalently

Ψj(θ)=U(θ)Φj\lvert \Psi_j(\theta)\rangle = U(\theta)\lvert \Phi_j\rangle6

A shared Ψj(θ)=U(θ)Φj\lvert \Psi_j(\theta)\rangle = U(\theta)\lvert \Phi_j\rangle7 acts only on the physical qubits, producing

Ψj(θ)=U(θ)Φj\lvert \Psi_j(\theta)\rangle = U(\theta)\lvert \Phi_j\rangle8

and orthogonality follows from

Ψj(θ)=U(θ)Φj\lvert \Psi_j(\theta)\rangle = U(\theta)\lvert \Phi_j\rangle9

No penalty terms or recursive deflation stages are required (Xu et al., 2022).

Diagonal energies are obtained by measuring the physical Hamiltonian and classically sorting shots by ancilla label. Off-diagonal subspace matrix elements are extracted through ancilla-only operators,

ρ(θ)=j=0K1wjΨj(θ)Ψj(θ),EC(θ)=Tr ⁣[ρ(θ)H^].\rho(\theta)=\sum_{j=0}^{K-1} w_j \ket{\Psi_j(\theta)}\bra{\Psi_j(\theta)}, \qquad E_C(\theta)=\operatorname{Tr}\!\left[\rho(\theta)\hat{H}\right].0

with

ρ(θ)=j=0K1wjΨj(θ)Ψj(θ),EC(θ)=Tr ⁣[ρ(θ)H^].\rho(\theta)=\sum_{j=0}^{K-1} w_j \ket{\Psi_j(\theta)}\bra{\Psi_j(\theta)}, \qquad E_C(\theta)=\operatorname{Tr}\!\left[\rho(\theta)\hat{H}\right].1

Similarly,

ρ(θ)=j=0K1wjΨj(θ)Ψj(θ),EC(θ)=Tr ⁣[ρ(θ)H^].\rho(\theta)=\sum_{j=0}^{K-1} w_j \ket{\Psi_j(\theta)}\bra{\Psi_j(\theta)}, \qquad E_C(\theta)=\operatorname{Tr}\!\left[\rho(\theta)\hat{H}\right].2

The final eigenpairs follow from the generalized eigenproblem

ρ(θ)=j=0K1wjΨj(θ)Ψj(θ),EC(θ)=Tr ⁣[ρ(θ)H^].\rho(\theta)=\sum_{j=0}^{K-1} w_j \ket{\Psi_j(\theta)}\bra{\Psi_j(\theta)}, \qquad E_C(\theta)=\operatorname{Tr}\!\left[\rho(\theta)\hat{H}\right].3

Because ρ(θ)=j=0K1wjΨj(θ)Ψj(θ),EC(θ)=Tr ⁣[ρ(θ)H^].\rho(\theta)=\sum_{j=0}^{K-1} w_j \ket{\Psi_j(\theta)}\bra{\Psi_j(\theta)}, \qquad E_C(\theta)=\operatorname{Tr}\!\left[\rho(\theta)\hat{H}\right].4, the off-diagonal reconstruction uses ρ(θ)=j=0K1wjΨj(θ)Ψj(θ),EC(θ)=Tr ⁣[ρ(θ)H^].\rho(\theta)=\sum_{j=0}^{K-1} w_j \ket{\Psi_j(\theta)}\bra{\Psi_j(\theta)}, \qquad E_C(\theta)=\operatorname{Tr}\!\left[\rho(\theta)\hat{H}\right].5 ancilla observables, which scales as ρ(θ)=j=0K1wjΨj(θ)Ψj(θ),EC(θ)=Tr ⁣[ρ(θ)H^].\rho(\theta)=\sum_{j=0}^{K-1} w_j \ket{\Psi_j(\theta)}\bra{\Psi_j(\theta)}, \qquad E_C(\theta)=\operatorname{Tr}\!\left[\rho(\theta)\hat{H}\right].6 in the worst case (Xu et al., 2022).

The method was demonstrated for the transverse Ising model

ρ(θ)=j=0K1wjΨj(θ)Ψj(θ),EC(θ)=Tr ⁣[ρ(θ)H^].\rho(\theta)=\sum_{j=0}^{K-1} w_j \ket{\Psi_j(\theta)}\bra{\Psi_j(\theta)}, \qquad E_C(\theta)=\operatorname{Tr}\!\left[\rho(\theta)\hat{H}\right].7

with ρ(θ)=j=0K1wjΨj(θ)Ψj(θ),EC(θ)=Tr ⁣[ρ(θ)H^].\rho(\theta)=\sum_{j=0}^{K-1} w_j \ket{\Psi_j(\theta)}\bra{\Psi_j(\theta)}, \qquad E_C(\theta)=\operatorname{Tr}\!\left[\rho(\theta)\hat{H}\right].8, ρ(θ)=j=0K1wjΨj(θ)Ψj(θ),EC(θ)=Tr ⁣[ρ(θ)H^].\rho(\theta)=\sum_{j=0}^{K-1} w_j \ket{\Psi_j(\theta)}\bra{\Psi_j(\theta)}, \qquad E_C(\theta)=\operatorname{Tr}\!\left[\rho(\theta)\hat{H}\right].9, and periodic chains up to ET(θ)=1Kj=0K1Ψj(θ)H^Ψj(θ)=1Kj=0K1Ej(θ),E_T(\theta)=\frac{1}{K}\sum_{j=0}^{K-1}\bra{\Psi_j(\theta)}\hat{H}\ket{\Psi_j(\theta)} =\frac{1}{K}\sum_{j=0}^{K-1}E_j(\theta),0. For ET(θ)=1Kj=0K1Ψj(θ)H^Ψj(θ)=1Kj=0K1Ej(θ),E_T(\theta)=\frac{1}{K}\sum_{j=0}^{K-1}\bra{\Psi_j(\theta)}\hat{H}\ket{\Psi_j(\theta)} =\frac{1}{K}\sum_{j=0}^{K-1}E_j(\theta),1 with ET(θ)=1Kj=0K1Ψj(θ)H^Ψj(θ)=1Kj=0K1Ej(θ),E_T(\theta)=\frac{1}{K}\sum_{j=0}^{K-1}\bra{\Psi_j(\theta)}\hat{H}\ket{\Psi_j(\theta)} =\frac{1}{K}\sum_{j=0}^{K-1}E_j(\theta),2, the four lowest eigenvalues converged rapidly and the errors dropped exponentially at nearly the same rate for all states. For ET(θ)=1Kj=0K1Ψj(θ)H^Ψj(θ)=1Kj=0K1Ej(θ),E_T(\theta)=\frac{1}{K}\sum_{j=0}^{K-1}\bra{\Psi_j(\theta)}\hat{H}\ket{\Psi_j(\theta)} =\frac{1}{K}\sum_{j=0}^{K-1}E_j(\theta),3 with ET(θ)=1Kj=0K1Ψj(θ)H^Ψj(θ)=1Kj=0K1Ej(θ),E_T(\theta)=\frac{1}{K}\sum_{j=0}^{K-1}\bra{\Psi_j(\theta)}\hat{H}\ket{\Psi_j(\theta)} =\frac{1}{K}\sum_{j=0}^{K-1}E_j(\theta),4, only six of the eight intended lowest states were captured because the chosen initial physical basis had insufficient overlap with all eight lowest eigenstates; increasing to ET(θ)=1Kj=0K1Ψj(θ)H^Ψj(θ)=1Kj=0K1Ej(θ),E_T(\theta)=\frac{1}{K}\sum_{j=0}^{K-1}\bra{\Psi_j(\theta)}\hat{H}\ket{\Psi_j(\theta)} =\frac{1}{K}\sum_{j=0}^{K-1}E_j(\theta),5 resolved the rank deficiency and recovered all eight lowest eigenvalues with errors below ET(θ)=1Kj=0K1Ψj(θ)H^Ψj(θ)=1Kj=0K1Ej(θ),E_T(\theta)=\frac{1}{K}\sum_{j=0}^{K-1}\bra{\Psi_j(\theta)}\hat{H}\ket{\Psi_j(\theta)} =\frac{1}{K}\sum_{j=0}^{K-1}E_j(\theta),6 at ET(θ)=1Kj=0K1Ψj(θ)H^Ψj(θ)=1Kj=0K1Ej(θ),E_T(\theta)=\frac{1}{K}\sum_{j=0}^{K-1}\bra{\Psi_j(\theta)}\hat{H}\ket{\Psi_j(\theta)} =\frac{1}{K}\sum_{j=0}^{K-1}E_j(\theta),7 and ET(θ)=1Kj=0K1Ψj(θ)H^Ψj(θ)=1Kj=0K1Ej(θ),E_T(\theta)=\frac{1}{K}\sum_{j=0}^{K-1}\bra{\Psi_j(\theta)}\hat{H}\ket{\Psi_j(\theta)} =\frac{1}{K}\sum_{j=0}^{K-1}E_j(\theta),8 (Xu et al., 2022).

5. KD-VQE and virtual annealing

KD-VQE is an ensemble VQE strategy that maintains and optimizes a population of trial wavefunctions in parallel while reallocating shots dynamically according to a Boltzmann distribution defined by a virtual temperature. The procedure begins at high temperature, where the distribution is nearly uniform, and gradually lowers the temperature so that resources condense onto low-energy candidates. The per-iteration shot allocation under total budget ET(θ)=1Kj=0K1Ψj(θ)H^Ψj(θ)=1Kj=0K1Ej(θ),E_T(\theta)=\frac{1}{K}\sum_{j=0}^{K-1}\bra{\Psi_j(\theta)}\hat{H}\ket{\Psi_j(\theta)} =\frac{1}{K}\sum_{j=0}^{K-1}E_j(\theta),9 is

KK0

and candidates receiving negligible shots can be pruned (Li, 6 May 2025).

The workflow has three stages: reduced subspace identification using conservation laws or symmetries of KK1; mixed-state ansatz preparation with multiple candidates in the reduced subspace; and variational optimization with virtual annealing. The paper specifies initialization by choosing ensemble size KK2, initial parameters KK3, shot budget KK4, and a large initial temperature KK5. Each iteration allocates shots according to KK6, measures KK7, estimates gradients, and updates parameters with simple gradient descent,

KK8

while the temperature is annealed multiplicatively,

KK9

Parameter-shift is compatible, with

C(θ)=k=0K1wkψk(θ)Hψk(θ),C(\theta)=\sum_{k=0}^{K-1} w_k \,\langle \psi_k(\theta) \lvert H \rvert \psi_k(\theta)\rangle,0

The discussion also notes that C(θ)=k=0K1wkψk(θ)Hψk(θ),C(\theta)=\sum_{k=0}^{K-1} w_k \,\langle \psi_k(\theta) \lvert H \rvert \psi_k(\theta)\rangle,1 and that reaching additive error C(θ)=k=0K1wkψk(θ)Hψk(θ),C(\theta)=\sum_{k=0}^{K-1} w_k \,\langle \psi_k(\theta) \lvert H \rvert \psi_k(\theta)\rangle,2 requires C(θ)=k=0K1wkψk(θ)Hψk(θ),C(\theta)=\sum_{k=0}^{K-1} w_k \,\langle \psi_k(\theta) \lvert H \rvert \psi_k(\theta)\rangle,3 shots, so concentrating shots on promising candidates reduces variance where it matters most (Li, 6 May 2025).

The demonstration used the two-site Fermi–Hubbard Hamiltonian at half filling,

C(θ)=k=0K1wkψk(θ)Hψk(θ),C(\theta)=\sum_{k=0}^{K-1} w_k \,\langle \psi_k(\theta) \lvert H \rvert \psi_k(\theta)\rangle,4

with C(θ)=k=0K1wkψk(θ)Hψk(θ),C(\theta)=\sum_{k=0}^{K-1} w_k \,\langle \psi_k(\theta) \lvert H \rvert \psi_k(\theta)\rangle,5, C(θ)=k=0K1wkψk(θ)Hψk(θ),C(\theta)=\sum_{k=0}^{K-1} w_k \,\langle \psi_k(\theta) \lvert H \rvert \psi_k(\theta)\rangle,6, and four qubits after fermion-to-qubit mapping. The reduced subspace enforced particle-number conservation, six half-filling trial states C(θ)=k=0K1wkψk(θ)Hψk(θ),C(\theta)=\sum_{k=0}^{K-1} w_k \,\langle \psi_k(\theta) \lvert H \rvert \psi_k(\theta)\rangle,7 were used as seeds, and a hardware-efficient ansatz with parameterized single-qubit gates and entangling CNOTs was combined with a Fourier transform that diagonalized the hopping term. The ensemble size was C(θ)=k=0K1wkψk(θ)Hψk(θ),C(\theta)=\sum_{k=0}^{K-1} w_k \,\langle \psi_k(\theta) \lvert H \rvert \psi_k(\theta)\rangle,8, the per-iteration budget was C(θ)=k=0K1wkψk(θ)Hψk(θ),C(\theta)=\sum_{k=0}^{K-1} w_k \,\langle \psi_k(\theta) \lvert H \rvert \psi_k(\theta)\rangle,9 shots, the starting temperature was EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,00 with EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,01, the temperature decayed as EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,02 each step, and candidates with EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,03 shots were pruned. Particle-number conservation at half filling was enforced with the penalty

EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,04

Empirically, EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,05 was filtered out after approximately EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,06 steps; EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,07 and EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,08 were removed around approximately EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,09 steps; EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,10 and EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,11 were trapped near local minima with energies approximately EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,12 and were pruned around approximately EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,13 steps; beyond approximately EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,14 steps, shots concentrated entirely on EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,15, whose energy converged toward the exact ground-state energy of approximately EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,16 (Li, 6 May 2025).

6. Cross-instance ensembles, warm starts, and measurement reuse

A separate usage of ensemble VQE concerns solving many related VQE problems in concert and sharing information between them so that each instance benefits from work done on others. The concrete procedure runs a seed VQE on one Hamiltonian, stores the parameter trajectory EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,17, and evaluates those parameter settings on target Hamiltonians to select warm starts. For a Hamiltonian decomposition EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,18, the standard objective is

EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,19

For a family EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,20, the same measured EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,21 can be linearly recombined whenever the Pauli basis overlaps, since

EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,22

The operational method in the cited work does not jointly optimize a multi-instance loss; instead, it uses the seed trajectory to initialize each target independently (Hutchings et al., 2024).

The algorithm stores the seed trajectory, subsamples “every 10th parameter set,” evaluates target energies EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,23, and selects the best EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,24 from the first half of the trajectory as the warm start. The stated reason for pruning is that late-stage iterates become tightly clustered, with tiny parameter displacements and negligible energy changes, and therefore carry little new information for transfer. The paper formalizes the intuition through possible gradient-norm, energy-change, and clustering criteria, but the implemented heuristic is simply to ignore the second half of the trajectory (Hutchings et al., 2024).

The experimental setting used MaxCut Ising Hamiltonians

EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,25

with random graphs of edge probability EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,26, Qiskit TwoLocal ansätze with linear entanglement and EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,27 to EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,28 layers, and SLSQP as the optimizer; SPSA performed markedly worse in the reported experiments. The study considered EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,29 and EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,30 qubits, with EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,31 target graphs for EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,32–EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,33 nodes and EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,34 for EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,35 nodes, on Qiskit Aer simulator, repeated EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,36 times. The main empirical result was negative but informative: random initialization performed about as well on average as both warm-start variants, and restricting selection to the first half of the seed trajectory did not degrade performance relative to using the full trajectory. The paper attributes this outcome to the use of independent random graphs, which likely reduced the value of transfer (Hutchings et al., 2024).

7. Advantages, limitations, and points of dispute

A recurrent misconception is that ensemble VQE is synonymous with excited-state VQE. The surveyed literature does not support that restriction. In one class of methods, the ensemble is an excited-state subspace optimized with a shared unitary; in KD-VQE, it is a population of trial ground-state candidates with adaptive shot allocation; and in simultaneous-solving approaches, it is a reusable trajectory across related instances. The common structure is collective information management, but the target object differs across formulations (Rajamani et al., 22 Sep 2025, Li, 6 May 2025, Hutchings et al., 2024).

A second misconception is that orthogonality enforcement necessarily requires explicit overlap penalties and Hadamard- or SWAP-type measurements. That is not true for shared-unitary formulations on orthonormal inputs or for ancilla purification. In those settings, orthogonality is preserved implicitly by unitarity, and overlap measurements are optional or only needed for noise-robust post-processing through a generalized eigenproblem (Xu et al., 2022, Rajamani et al., 22 Sep 2025).

The principal controversy in the recent literature concerns equi-ensemble versus weighted-ensemble optimization. The cited comparison argues that equal weights are preferable because they convert the objective into the Hamiltonian subtrace, eliminate ordering assumptions, and focus the optimization on the target subspace. Weighted ensembles are reported to be sensitive to state-ordering guesses, especially near strong vibronic couplings, avoided crossings, and conical intersections; the same study therefore concludes that “the equi-ensemble is the way to go” for the tested chemistry benchmarks (Rajamani et al., 22 Sep 2025). A plausible implication is that the debate is not about whether ensembles are useful, but about which ensemble objective yields the most stable optimization landscape.

Limitations are equally method-specific. Ancilla-purified subspace methods require ancilla capacity EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,37 and off-diagonal measurement overhead that scales as EC(θ)=j=0K1wjΦjU^(θ)H^U^(θ)Φjj=0K1wjEj,E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j} \geq \sum_{j=0}^{K-1} w_j E_j,38 in the worst case. KD-VQE depends on ensemble seeding, temperature schedule, and pruning thresholds; if no candidate spans the basin of the true ground state, the algorithm condenses to the best available local minimum. Cross-instance warm-start ensembles depend strongly on instance similarity; in the reported MaxCut study, transfer from a seed trajectory did not outperform random initialization on average. Taken together, these results indicate that ensemble VQE is not uniformly superior to single-state VQE. Its effectiveness depends on whether the ensemble structure matches the physics of the problem, the measurement model, and the optimizer landscape (Xu et al., 2022, Li, 6 May 2025, Hutchings et al., 2024).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Ensemble Variational Quantum Eigensolver (VQE).