Ensemble VQE is a framework that optimizes a collection of quantum states collectively to target multiple eigenstates or related problem instances.
It encompasses various formulations—state-averaged, ancilla-purified, KD-VQE, and cross-instance methods—each tailored for specific excitation or ground state challenges.
Practical insights include trade-offs in weighting strategies, measurement overhead, and parameter transfer, affecting overall convergence and optimization landscape.
Ensemblevariational quantum eigensolver (ensemble VQE) denotes a family of VQE formulations in which the optimization target is not a single variational state but an ensemble of states, an ensemble-defined low-energy subspace, or an ensemble of related problem instances. In the recent literature, the term is used for at least four distinct constructions: state-averaged excited-state VQE based on a shared unitary acting on orthonormal inputs, ancilla-purified single-circuit methods that determine many eigenpairs simultaneously, Boltzmann-weighted multi-candidate ground-state solvers such as Knowledge Distillation Inspired VQE (KD-VQE), and simultaneous-solving schemes that transfer parameters and measurements across multiple Hamiltonians (Rajamani et al., 22 Sep 2025, Xu et al., 2022, Li, 6 May 2025, Hutchings et al., 2024). This suggests that “ensemble VQE” is best understood as a unifying perspective on collective variational optimization rather than as a single canonical algorithm.
1. Conceptual scope
The literature uses “ensemble VQE” for related but non-identical objectives. In one line of work, the ensemble is a set of orthonormal trial states generated by applying one shared parametrized unitary to multiple reference states; the optimization target is an ensemble energy or Hamiltonian subtrace over that subspace. In another, the ensemble is a population of trial wavefunctions whose measurement resources are redistributed during optimization according to a Boltzmann law. In a third, the ensemble is the collection of states encountered along a seed VQE trajectory and reused to warm start multiple target instances (Rajamani et al., 22 Sep 2025, Li, 6 May 2025, Hutchings et al., 2024).
Variant
Ensemble object
Core mechanism
State-averaged / weighted ensemble VQE
{U(θ)∣Φj⟩}
Minimize ensemble energy or subtrace
Ancilla-purified one-circuit method
Orthogonal states labeled by ancillas
Optimize one purified circuit and diagonalize subspace
KD-VQE
Multiple trial wavefunctions
Boltzmann-weighted shot allocation with virtual annealing
Simultaneous-solving ensemble
Seed trajectory across related Hamiltonians
Warm starts, pruning, and measurement reuse
This plurality matters because the strengths and failure modes depend on what is being ensembled. Subspace-based formulations are primarily designed for multiple low-lying eigenstates and excited-state structure. KD-VQE addresses exploration, variance control, and robustness in ground-state search. Cross-instance ensemble methods focus on transfer across related Hamiltonians rather than simultaneous optimization of a single spectrum.
2. Variational formulations
A central formulation is the weighted-ensemble objective
where wj≥wj+1, H^ is the target Hamiltonian, and Ej≤Ej+1 are exact eigenvalues. Writing ∣Ψj(θ)⟩=U(θ)∣Φj⟩, the same objective can be expressed through the ensemble density operator
This is a generalized Rayleigh–Ritz bound over an optimized orthogonal subspace rather than a single vector (Xu et al., 2022).
KD-VQE adopts a different ensemble view. For EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,1 parameterized trial states EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,2 with EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,3, it defines a temperature-dependent mixed-state ansatz
which satisfies EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,9 (Li, 6 May 2025).
3. State-averaged and subspace ensemble VQE
In state-averaged ensemble VQE, the same unitary wj≥wj+10 is applied to orthonormal initial states wj≥wj+11, so the prepared states remain orthonormal because unitaries preserve inner products. This eliminates the need for explicit overlap penalties when the shared-unitary construction is used. The equi-ensemble case wj≥wj+12 is distinguished because the objective becomes the subtrace wj≥wj+13, so optimization targets the optimal wj≥wj+14-dimensional eigensubspace rather than a prescribed ordering of circuit outputs (Rajamani et al., 22 Sep 2025).
The weighted-ensemble variant uses unequal nonnegative weights. The cited comparison argues that weighted optimization requires the ordering of weights wj≥wj+15 to align a priori with the eventual energy ordering of the prepared states, whereas equal weights remove that requirement. The same study further reports that unequal weights can lead to mode collapse toward strongly weighted states, sharp local minima or plateaus, and sensitivity to ordering guesses near avoided crossings or conical intersections; by contrast, the equi-ensemble landscape is described as smoother and focused on the subspace itself (Rajamani et al., 22 Sep 2025).
These differences were quantified in two benchmarks. For formaldimine with wj≥wj+16, equi-ensemble optimization reached the exact eigensubspace within wj≥wj+17 Hartree in approximately wj≥wj+18 iterations with 1-GUCCSD and approximately wj≥wj+19 iterations with 2-GUCCSD. Under ideal ordering, weighted 1-GUCCSD reached H^0 error H^1 Hartree and H^2 error H^3 Hartree in approximately H^4 iterations, while weighted 2-GUCCSD reduced both to H^5 Hartree but needed approximately H^6 iterations. Under incorrect ordering, weighted 1-GUCCSD produced H^7 errors up to H^8 Hartree even when H^9 remained Ej≤Ej+10 Hartree, and weighted 2-GUCCSD required approximately Ej≤Ej+11 iterations to recover Ej≤Ej+12 Hartree accuracy (Rajamani et al., 22 Sep 2025).
For the linear Ej≤Ej+13 chain in the Q-DFT setting with Ej≤Ej+14, equi-ensemble plus classical diagonalization reduced errors by more than two orders of magnitude for Ej≤Ej+15 Å and by approximately one order of magnitude for Ej≤Ej+16 Å, with approximately Ej≤Ej+17 fewer iterations in most regimes. The same study reported that weighted ensembles exhibited large, non-democratic errors dominated by the highest-energy state and exceeded Ej≤Ej+18 Hartree at Ej≤Ej+19 Å for the highest-energy state under the tested weighting scheme (Rajamani et al., 22 Sep 2025).
4. Ancilla-purified one-circuit realization
An important single-circuit realization of ensemble VQE introduces ancilla qubits to purify the trial states and enforce orthogonality by construction. The setup uses ∣Ψj(θ)⟩=U(θ)∣Φj⟩0 physical qubits for the system and ∣Ψj(θ)⟩=U(θ)∣Φj⟩1 ancillas, with ancilla Hilbert-space dimension ∣Ψj(θ)⟩=U(θ)∣Φj⟩2. Choosing ∣Ψj(θ)⟩=U(θ)∣Φj⟩3 suffices to label ∣Ψj(θ)⟩=U(θ)∣Φj⟩4 target states. The initial state is prepared as Bell pairs between ancillas and selected physical qubits,
∣Ψj(θ)⟩=U(θ)∣Φj⟩5
equivalently
∣Ψj(θ)⟩=U(θ)∣Φj⟩6
A shared ∣Ψj(θ)⟩=U(θ)∣Φj⟩7 acts only on the physical qubits, producing
∣Ψj(θ)⟩=U(θ)∣Φj⟩8
and orthogonality follows from
∣Ψj(θ)⟩=U(θ)∣Φj⟩9
No penalty terms or recursive deflation stages are required (Xu et al., 2022).
Diagonal energies are obtained by measuring the physical Hamiltonian and classically sorting shots by ancilla label. Off-diagonal subspace matrix elements are extracted through ancilla-only operators,
Because ρ(θ)=j=0∑K−1wj∣Ψj(θ)⟩⟨Ψj(θ)∣,EC(θ)=Tr[ρ(θ)H^].4, the off-diagonal reconstruction uses ρ(θ)=j=0∑K−1wj∣Ψj(θ)⟩⟨Ψj(θ)∣,EC(θ)=Tr[ρ(θ)H^].5 ancilla observables, which scales as ρ(θ)=j=0∑K−1wj∣Ψj(θ)⟩⟨Ψj(θ)∣,EC(θ)=Tr[ρ(θ)H^].6 in the worst case (Xu et al., 2022).
The method was demonstrated for the transverse Ising model
with ρ(θ)=j=0∑K−1wj∣Ψj(θ)⟩⟨Ψj(θ)∣,EC(θ)=Tr[ρ(θ)H^].8, ρ(θ)=j=0∑K−1wj∣Ψj(θ)⟩⟨Ψj(θ)∣,EC(θ)=Tr[ρ(θ)H^].9, and periodic chains up to ET(θ)=K1j=0∑K−1⟨Ψj(θ)∣H^∣Ψj(θ)⟩=K1j=0∑K−1Ej(θ),0. For ET(θ)=K1j=0∑K−1⟨Ψj(θ)∣H^∣Ψj(θ)⟩=K1j=0∑K−1Ej(θ),1 with ET(θ)=K1j=0∑K−1⟨Ψj(θ)∣H^∣Ψj(θ)⟩=K1j=0∑K−1Ej(θ),2, the four lowest eigenvalues converged rapidly and the errors dropped exponentially at nearly the same rate for all states. For ET(θ)=K1j=0∑K−1⟨Ψj(θ)∣H^∣Ψj(θ)⟩=K1j=0∑K−1Ej(θ),3 with ET(θ)=K1j=0∑K−1⟨Ψj(θ)∣H^∣Ψj(θ)⟩=K1j=0∑K−1Ej(θ),4, only six of the eight intended lowest states were captured because the chosen initial physical basis had insufficient overlap with all eight lowest eigenstates; increasing to ET(θ)=K1j=0∑K−1⟨Ψj(θ)∣H^∣Ψj(θ)⟩=K1j=0∑K−1Ej(θ),5 resolved the rank deficiency and recovered all eight lowest eigenvalues with errors below ET(θ)=K1j=0∑K−1⟨Ψj(θ)∣H^∣Ψj(θ)⟩=K1j=0∑K−1Ej(θ),6 at ET(θ)=K1j=0∑K−1⟨Ψj(θ)∣H^∣Ψj(θ)⟩=K1j=0∑K−1Ej(θ),7 and ET(θ)=K1j=0∑K−1⟨Ψj(θ)∣H^∣Ψj(θ)⟩=K1j=0∑K−1Ej(θ),8 (Xu et al., 2022).
5. KD-VQE and virtual annealing
KD-VQE is an ensemble VQE strategy that maintains and optimizes a population of trial wavefunctions in parallel while reallocating shots dynamically according to a Boltzmann distribution defined by a virtual temperature. The procedure begins at high temperature, where the distribution is nearly uniform, and gradually lowers the temperature so that resources condense onto low-energy candidates. The per-iteration shot allocation under total budget ET(θ)=K1j=0∑K−1⟨Ψj(θ)∣H^∣Ψj(θ)⟩=K1j=0∑K−1Ej(θ),9 is
K0
and candidates receiving negligible shots can be pruned (Li, 6 May 2025).
The workflow has three stages: reduced subspace identification using conservation laws or symmetries of K1; mixed-state ansatz preparation with multiple candidates in the reduced subspace; and variational optimization with virtual annealing. The paper specifies initialization by choosing ensemble size K2, initial parameters K3, shot budget K4, and a large initial temperature K5. Each iteration allocates shots according to K6, measures K7, estimates gradients, and updates parameters with simple gradient descent,
K8
while the temperature is annealed multiplicatively,
K9
Parameter-shift is compatible, with
C(θ)=k=0∑K−1wk⟨ψk(θ)∣H∣ψk(θ)⟩,0
The discussion also notes that C(θ)=k=0∑K−1wk⟨ψk(θ)∣H∣ψk(θ)⟩,1 and that reaching additive error C(θ)=k=0∑K−1wk⟨ψk(θ)∣H∣ψk(θ)⟩,2 requires C(θ)=k=0∑K−1wk⟨ψk(θ)∣H∣ψk(θ)⟩,3 shots, so concentrating shots on promising candidates reduces variance where it matters most (Li, 6 May 2025).
with C(θ)=k=0∑K−1wk⟨ψk(θ)∣H∣ψk(θ)⟩,5, C(θ)=k=0∑K−1wk⟨ψk(θ)∣H∣ψk(θ)⟩,6, and four qubits after fermion-to-qubit mapping. The reduced subspace enforced particle-number conservation, six half-filling trial states C(θ)=k=0∑K−1wk⟨ψk(θ)∣H∣ψk(θ)⟩,7 were used as seeds, and a hardware-efficient ansatz with parameterized single-qubit gates and entangling CNOTs was combined with a Fourier transform that diagonalized the hopping term. The ensemble size was C(θ)=k=0∑K−1wk⟨ψk(θ)∣H∣ψk(θ)⟩,8, the per-iteration budget was C(θ)=k=0∑K−1wk⟨ψk(θ)∣H∣ψk(θ)⟩,9 shots, the starting temperature was EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,00 with EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,01, the temperature decayed as EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,02 each step, and candidates with EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,03 shots were pruned. Particle-number conservation at half filling was enforced with the penalty
Empirically, EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,05 was filtered out after approximately EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,06 steps; EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,07 and EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,08 were removed around approximately EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,09 steps; EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,10 and EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,11 were trapped near local minima with energies approximately EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,12 and were pruned around approximately EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,13 steps; beyond approximately EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,14 steps, shots concentrated entirely on EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,15, whose energy converged toward the exact ground-state energy of approximately EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,16 (Li, 6 May 2025).
6. Cross-instance ensembles, warm starts, and measurement reuse
A separate usage of ensemble VQE concerns solving many related VQE problems in concert and sharing information between them so that each instance benefits from work done on others. The concrete procedure runs a seed VQE on one Hamiltonian, stores the parameter trajectory EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,17, and evaluates those parameter settings on target Hamiltonians to select warm starts. For a Hamiltonian decomposition EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,18, the standard objective is
For a family EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,20, the same measured EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,21 can be linearly recombined whenever the Pauli basis overlaps, since
The operational method in the cited work does not jointly optimize a multi-instance loss; instead, it uses the seed trajectory to initialize each target independently (Hutchings et al., 2024).
The algorithm stores the seed trajectory, subsamples “every 10th parameter set,” evaluates target energies EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,23, and selects the best EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,24 from the first half of the trajectory as the warm start. The stated reason for pruning is that late-stage iterates become tightly clustered, with tiny parameter displacements and negligible energy changes, and therefore carry little new information for transfer. The paper formalizes the intuition through possible gradient-norm, energy-change, and clustering criteria, but the implemented heuristic is simply to ignore the second half of the trajectory (Hutchings et al., 2024).
The experimental setting used MaxCut Ising Hamiltonians
with random graphs of edge probability EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,26, Qiskit TwoLocal ansätze with linear entanglement and EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,27 to EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,28 layers, and SLSQP as the optimizer; SPSA performed markedly worse in the reported experiments. The study considered EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,29 and EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,30 qubits, with EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,31 target graphs for EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,32–EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,33 nodes and EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,34 for EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,35 nodes, on Qiskit Aer simulator, repeated EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,36 times. The main empirical result was negative but informative: random initialization performed about as well on average as both warm-start variants, and restricting selection to the first half of the seed trajectory did not degrade performance relative to using the full trajectory. The paper attributes this outcome to the use of independent random graphs, which likely reduced the value of transfer (Hutchings et al., 2024).
7. Advantages, limitations, and points of dispute
A recurrent misconception is that ensemble VQE is synonymous with excited-state VQE. The surveyed literature does not support that restriction. In one class of methods, the ensemble is an excited-state subspace optimized with a shared unitary; in KD-VQE, it is a population of trial ground-state candidates with adaptive shot allocation; and in simultaneous-solving approaches, it is a reusable trajectory across related instances. The common structure is collective information management, but the target object differs across formulations (Rajamani et al., 22 Sep 2025, Li, 6 May 2025, Hutchings et al., 2024).
A second misconception is that orthogonality enforcement necessarily requires explicit overlap penalties and Hadamard- or SWAP-type measurements. That is not true for shared-unitary formulations on orthonormal inputs or for ancilla purification. In those settings, orthogonality is preserved implicitly by unitarity, and overlap measurements are optional or only needed for noise-robust post-processing through a generalized eigenproblem (Xu et al., 2022, Rajamani et al., 22 Sep 2025).
The principal controversy in the recent literature concerns equi-ensemble versus weighted-ensemble optimization. The cited comparison argues that equal weights are preferable because they convert the objective into the Hamiltonian subtrace, eliminate ordering assumptions, and focus the optimization on the target subspace. Weighted ensembles are reported to be sensitive to state-ordering guesses, especially near strong vibronic couplings, avoided crossings, and conical intersections; the same study therefore concludes that “the equi-ensemble is the way to go” for the tested chemistry benchmarks (Rajamani et al., 22 Sep 2025). A plausible implication is that the debate is not about whether ensembles are useful, but about which ensemble objective yields the most stable optimization landscape.
Limitations are equally method-specific. Ancilla-purified subspace methods require ancilla capacity EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,37 and off-diagonal measurement overhead that scales as EC(θ)=j=0∑K−1wj⟨Φj∣U^†(θ)H^U^(θ)∣Φj⟩≥j=0∑K−1wjEj,38 in the worst case. KD-VQE depends on ensemble seeding, temperature schedule, and pruning thresholds; if no candidate spans the basin of the true ground state, the algorithm condenses to the best available local minimum. Cross-instance warm-start ensembles depend strongly on instance similarity; in the reported MaxCut study, transfer from a seed trajectory did not outperform random initialization on average. Taken together, these results indicate that ensemble VQE is not uniformly superior to single-state VQE. Its effectiveness depends on whether the ensemble structure matches the physics of the problem, the measurement model, and the optimizer landscape (Xu et al., 2022, Li, 6 May 2025, Hutchings et al., 2024).
“Emergent Mind helps me see which AI papers have caught fire online.”
Philip
Creator, AI Explained on YouTube
Sign up for free to explore the frontiers of research
Discover trending papers, chat with arXiv, and track the latest research shaping the future of science and technology.Discover trending papers, chat with arXiv, and more.