---
title: Ensemble Variational Quantum Eigensolver
url: https://www.emergentmind.com/topics/ensemble-variational-quantum-eigensolver-vqe
type: topic
---

# Ensemble Variational Quantum Eigensolver

Ensemble variational quantum eigensolver (ensemble VQE) denotes a family of VQE formulations in which the optimization target is not a single variational state but an ensemble of states, an ensemble-defined low-energy subspace, or an ensemble of related problem instances. In the recent literature, the term is used for at least four distinct constructions: state-averaged excited-state VQE based on a shared unitary acting on orthonormal inputs, ancilla-purified single-circuit methods that determine many eigenpairs simultaneously, Boltzmann-weighted multi-candidate ground-state solvers such as Knowledge Distillation Inspired VQE (KD-VQE), and simultaneous-solving schemes that transfer parameters and measurements across multiple Hamiltonians [2509.17982][2206.11036][2505.03998][2410.21413]. This suggests that “ensemble VQE” is best understood as a unifying perspective on collective variational optimization rather than as a single canonical algorithm.

## 1. Conceptual scope

The literature uses “ensemble VQE” for related but non-identical objectives. In one line of work, the ensemble is a set of orthonormal trial states generated by applying one shared parametrized unitary to multiple reference states; the optimization target is an ensemble energy or Hamiltonian subtrace over that subspace. In another, the ensemble is a population of trial wavefunctions whose measurement resources are redistributed during optimization according to a Boltzmann law. In a third, the ensemble is the collection of states encountered along a seed VQE trajectory and reused to warm start multiple target instances [2509.17982][2505.03998][2410.21413].

| Variant | Ensemble object | Core mechanism |
|---|---|---|
| State-averaged / weighted ensemble VQE | \( \{U(\theta)\lvert \Phi_j\rangle\} \) | Minimize ensemble energy or subtrace |
| Ancilla-purified one-circuit method | Orthogonal states labeled by ancillas | Optimize one purified circuit and diagonalize subspace |
| KD-VQE | Multiple trial wavefunctions | Boltzmann-weighted shot allocation with virtual annealing |
| Simultaneous-solving ensemble | Seed trajectory across related Hamiltonians | Warm starts, pruning, and measurement reuse |

This plurality matters because the strengths and failure modes depend on what is being ensembled. Subspace-based formulations are primarily designed for multiple low-lying eigenstates and excited-state structure. KD-VQE addresses exploration, variance control, and robustness in ground-state search. Cross-instance ensemble methods focus on transfer across related Hamiltonians rather than simultaneous optimization of a single spectrum.

## 2. Variational formulations

A central formulation is the weighted-ensemble objective
\[
E_C(\theta)=\sum_{j=0}^{K-1} w_j \bra{\Phi_j}\hat{U}^\dagger(\theta)\hat{H}\hat{U}(\theta)\ket{\Phi_j}
\geq \sum_{j=0}^{K-1} w_j E_j,
\]
where \(w_j \geq w_{j+1}\), \(\hat{H}\) is the target Hamiltonian, and \(E_j \leq E_{j+1}\) are exact eigenvalues. Writing \(\lvert \Psi_j(\theta)\rangle = U(\theta)\lvert \Phi_j\rangle\), the same objective can be expressed through the ensemble density operator
\[
\rho(\theta)=\sum_{j=0}^{K-1} w_j \ket{\Psi_j(\theta)}\bra{\Psi_j(\theta)},
\qquad
E_C(\theta)=\operatorname{Tr}\!\left[\rho(\theta)\hat{H}\right].
\]
For equal weights, the cost becomes the normalized Hamiltonian subtrace
\[
E_T(\theta)=\frac{1}{K}\sum_{j=0}^{K-1}\bra{\Psi_j(\theta)}\hat{H}\ket{\Psi_j(\theta)}
=\frac{1}{K}\sum_{j=0}^{K-1}E_j(\theta),
\]
which is invariant under rotations within the optimized \(K\)-dimensional subspace [2509.17982].

A closely related subspace formulation appears in the ancilla-purified one-circuit method. There the objective is
\[
C(\theta)=\sum_{k=0}^{K-1} w_k \,\langle \psi_k(\theta) \lvert H \rvert \psi_k(\theta)\rangle,
\]
with projector
\[
P(\theta)=\sum_{k=0}^{K-1} \lvert \psi_k(\theta)\rangle\langle \psi_k(\theta)\rvert,
\qquad
C(\theta)=\mathrm{Tr}\big(P(\theta)H\big)\geq \sum_{i=0}^{K-1}E_i.
\]
This is a generalized Rayleigh–Ritz bound over an optimized orthogonal subspace rather than a single vector [2206.11036].

KD-VQE adopts a different ensemble view. For \(K\) parameterized trial states \(\{\lvert\psi(\theta_i)\rangle\}_{i=1}^K\) with \(E_i(\theta_i)=\langle\psi(\theta_i)\lvert H\rvert\psi(\theta_i)\rangle\), it defines a temperature-dependent mixed-state ansatz
\[
\rho(T)=\frac{1}{Z}\sum_{i=1}^K e^{-\beta E_i}\lvert\psi(\theta_i)\rangle\langle\psi(\theta_i)\rvert,
\]
with \(\beta=1/T\), \(Z=\sum_{j=1}^K e^{-\beta E_j}\), Boltzmann weights
\[
P_i=\frac{e^{-\beta E_i}}{Z},
\]
and ensemble energy
\[
E_{\mathrm{mix}}(T)=\operatorname{tr}(\rho(T)H)=\sum_{i=1}^K P_i E_i,
\]
which satisfies \(E_{gs}\leq E_{\mathrm{mix}}(T)\) [2505.03998].

## 3. State-averaged and subspace ensemble VQE

In state-averaged ensemble VQE, the same unitary \(U(\theta)\) is applied to orthonormal initial states \(\{\lvert \Phi_j\rangle\}\), so the prepared states remain orthonormal because unitaries preserve inner products. This eliminates the need for explicit overlap penalties when the shared-unitary construction is used. The equi-ensemble case \(w_j=1/K\) is distinguished because the objective becomes the subtrace \(E_T(\theta)\), so optimization targets the optimal \(K\)-dimensional eigensubspace rather than a prescribed ordering of circuit outputs [2509.17982].

The weighted-ensemble variant uses unequal nonnegative weights. The cited comparison argues that weighted optimization requires the ordering of weights \(w_j\) to align a priori with the eventual energy ordering of the prepared states, whereas equal weights remove that requirement. The same study further reports that unequal weights can lead to mode collapse toward strongly weighted states, sharp local minima or plateaus, and sensitivity to ordering guesses near avoided crossings or conical intersections; by contrast, the equi-ensemble landscape is described as smoother and focused on the subspace itself [2509.17982].

These differences were quantified in two benchmarks. For formaldimine with \(K=2\), equi-ensemble optimization reached the exact eigensubspace within \(10^{-9}\) Hartree in approximately \(300\) iterations with 1-GUCCSD and approximately \(500\) iterations with 2-GUCCSD. Under ideal ordering, weighted 1-GUCCSD reached \(E_T\) error \(\approx 10^{-6}\) Hartree and \(E_C\) error \(\approx 10^{-5}\) Hartree in approximately \(300\) iterations, while weighted 2-GUCCSD reduced both to \(\approx 10^{-9}\) Hartree but needed approximately \(900\) iterations. Under incorrect ordering, weighted 1-GUCCSD produced \(E_C\) errors up to \(10^{-2}\) Hartree even when \(E_T\) remained \(\approx 10^{-6}\) Hartree, and weighted 2-GUCCSD required approximately \(1400\) iterations to recover \(\approx 10^{-9}\) Hartree accuracy [2509.17982].

For the linear \( \mathrm{H}_{16} \) chain in the Q-DFT setting with \(K=8\), equi-ensemble plus classical diagonalization reduced errors by more than two orders of magnitude for \(R<2\) Å and by approximately one order of magnitude for \(R>2\) Å, with approximately \(2\times\) fewer iterations in most regimes. The same study reported that weighted ensembles exhibited large, non-democratic errors dominated by the highest-energy state and exceeded \(10^{-1}\) Hartree at \(R=0.5\) Å for the highest-energy state under the tested weighting scheme [2509.17982].

## 4. Ancilla-purified one-circuit realization

An important single-circuit realization of ensemble VQE introduces ancilla qubits to purify the trial states and enforce orthogonality by construction. The setup uses \(N_p\) physical qubits for the system and \(N_a\) ancillas, with ancilla Hilbert-space dimension \(M=2^{N_a}\). Choosing \(N_a=\lceil \log_2 K\rceil\) suffices to label \(K\) target states. The initial state is prepared as Bell pairs between ancillas and selected physical qubits,
\[
\lvert \psi_{\rm ini}\rangle
=
\bigotimes_{i=1}^{N_a}\lvert B\rangle_{(p_i,a_i)}
\bigotimes_{j=N_a+1}^{N_p}\lvert 0\rangle_{p_j},
\qquad
\lvert B\rangle_{(p_i,a_i)}
=
\frac{\lvert 0\rangle_{p_i}\lvert 0\rangle_{a_i}+\lvert 1\rangle_{p_i}\lvert 1\rangle_{a_i}}{\sqrt{2}},
\]
equivalently
\[
\lvert \psi_{\rm ini}\rangle
=
\frac{1}{\sqrt{M}}\sum_{\alpha=0}^{M-1}\lvert \alpha\rangle_p\otimes \lvert \alpha\rangle_a.
\]
A shared \(U(\theta)\) acts only on the physical qubits, producing
\[
\lvert \Psi(\theta)\rangle
=
\frac{1}{\sqrt{M}}\sum_{\alpha=0}^{M-1}\lvert \psi_\alpha(\theta)\rangle\otimes \lvert \alpha\rangle_a,
\qquad
\lvert \psi_\alpha(\theta)\rangle = U(\theta)\lvert \alpha\rangle_p,
\]
and orthogonality follows from
\[
\langle \psi_i(\theta)\vert \psi_j(\theta)\rangle=\delta_{ij}.
\]
No penalty terms or recursive deflation stages are required [2206.11036].

Diagonal energies are obtained by measuring the physical Hamiltonian and classically sorting shots by ancilla label. Off-diagonal subspace matrix elements are extracted through ancilla-only operators,
\[
H_{ij}=M\langle \Psi(\theta)\vert H\otimes \lvert j\rangle_a\langle i\rvert_a \vert \Psi(\theta)\rangle
= M\sum_{\mu} V_{ji,\mu}\, h_\mu,
\]
with
\[
h_\mu=\langle \Psi(\theta)\vert H\otimes \hat{A}_\mu\vert \Psi(\theta)\rangle.
\]
Similarly,
\[
S_{ij}=M\sum_{\mu} V_{ji,\mu}\, s_\mu,
\qquad
s_\mu=\langle \Psi(\theta)\vert I\otimes \hat{A}_\mu\vert \Psi(\theta)\rangle.
\]
The final eigenpairs follow from the generalized eigenproblem
\[
H\,\mathbf{c}=\lambda\,S\,\mathbf{c},
\qquad
\lvert E_i\rangle_p=\sum_{k=0}^{K-1} c_k^{(i)}\lvert \psi_k(\theta)\rangle.
\]
Because \(N_a=\lceil \log_2 K\rceil\), the off-diagonal reconstruction uses \(4^{N_a}\) ancilla observables, which scales as \(K^2\) in the worst case [2206.11036].

The method was demonstrated for the transverse Ising model
\[
H=-J\sum_i \sigma_i^z \sigma_{i+1}^z + h\sum_i \sigma_i^x,
\]
with \(J=1\), \(h=0.5\), and periodic chains up to \(N=8\). For \(K=4\) with \(N_a=2\), the four lowest eigenvalues converged rapidly and the errors dropped exponentially at nearly the same rate for all states. For \(K=8\) with \(N_a=3\), only six of the eight intended lowest states were captured because the chosen initial physical basis had insufficient overlap with all eight lowest eigenstates; increasing to \(N_a=4\) resolved the rank deficiency and recovered all eight lowest eigenvalues with errors below \(10^{-6}\) at \(L=6\) and \(m \gtrsim 300\) [2206.11036].

## 5. KD-VQE and virtual annealing

KD-VQE is an ensemble VQE strategy that maintains and optimizes a population of trial wavefunctions in parallel while reallocating shots dynamically according to a Boltzmann distribution defined by a virtual temperature. The procedure begins at high temperature, where the distribution is nearly uniform, and gradually lowers the temperature so that resources condense onto low-energy candidates. The per-iteration shot allocation under total budget \(S\) is
\[
s_i = S\cdot P_i = S\cdot \frac{e^{-\beta E_i}}{Z},
\]
and candidates receiving negligible shots can be pruned [2505.03998].

The workflow has three stages: reduced subspace identification using conservation laws or symmetries of \(H\); mixed-state ansatz preparation with multiple candidates in the reduced subspace; and variational optimization with virtual annealing. The paper specifies initialization by choosing ensemble size \(K\), initial parameters \(\theta_i\), shot budget \(S\), and a large initial temperature \(T_0\). Each iteration allocates shots according to \(P_i\), measures \(E_i\), estimates gradients, and updates parameters with simple gradient descent,
\[
\theta_i \leftarrow \theta_i - \eta \nabla_{\theta_i} E_i,
\]
while the temperature is annealed multiplicatively,
\[
T_{k+1}=\alpha T_k,\qquad 0<\alpha<1.
\]
Parameter-shift is compatible, with
\[
\partial_\theta E(\theta)=\frac{1}{2}\big[E(\theta+\pi/2)-E(\theta-\pi/2)\big].
\]
The discussion also notes that \(\mathrm{Var}[\hat{E}_i]\propto 1/s_i\) and that reaching additive error \(\epsilon\) requires \(O(1/\epsilon^2)\) shots, so concentrating shots on promising candidates reduces variance where it matters most [2505.03998].

The demonstration used the two-site Fermi–Hubbard Hamiltonian at half filling,
\[
H = -t \sum_{\sigma\in\{\uparrow,\downarrow\}}
(c_{1\sigma}^\dagger c_{2\sigma}+c_{2\sigma}^\dagger c_{1\sigma})
+ U \sum_{i=1}^{2} n_{i\uparrow}n_{i\downarrow},
\]
with \(t=1\), \(U=1\), and four qubits after fermion-to-qubit mapping. The reduced subspace enforced particle-number conservation, six half-filling trial states \(\psi_I,\ldots,\psi_{VI}\) were used as seeds, and a hardware-efficient ansatz with parameterized single-qubit gates and entangling CNOTs was combined with a Fourier transform that diagonalized the hopping term. The ensemble size was \(K=6\), the per-iteration budget was \(S=10^4\) shots, the starting temperature was \(k_B T=25\) with \(k_B=1\), the temperature decayed as \(T\to 0.95T\) each step, and candidates with \(s_i<100\) shots were pruned. Particle-number conservation at half filling was enforced with the penalty
\[
\epsilon_i(\theta_i)=
\langle \psi_i(\theta_i)\lvert H\rvert \psi_i(\theta_i)\rangle
-\lambda \big|\langle \psi_i(\theta_i)\lvert \hat{n}\rvert \psi_i(\theta_i)\rangle -2\big|,
\qquad \lambda>0.
\]
Empirically, \(\psi_I\) was filtered out after approximately \(60\) steps; \(\psi_{II}\) and \(\psi_{III}\) were removed around approximately \(80\) steps; \(\psi_{IV}\) and \(\psi_V\) were trapped near local minima with energies approximately \(0\) and were pruned around approximately \(85\) steps; beyond approximately \(85\) steps, shots concentrated entirely on \(\psi_{VI}\), whose energy converged toward the exact ground-state energy of approximately \(-1.56\) [2505.03998].

## 6. Cross-instance ensembles, warm starts, and measurement reuse

A separate usage of ensemble VQE concerns solving many related VQE problems in concert and sharing information between them so that each instance benefits from work done on others. The concrete procedure runs a seed VQE on one Hamiltonian, stores the parameter trajectory \(\{\theta^{(t)}\}\), and evaluates those parameter settings on target Hamiltonians to select warm starts. For a Hamiltonian decomposition \(H=\sum_k c_k P_k\), the standard objective is
\[
E(\theta)=\langle \psi(\theta)\lvert H\rvert \psi(\theta)\rangle
=\sum_k c_k \langle P_k\rangle_\theta.
\]
For a family \(\{H_i\}_{i=1}^M\), the same measured \(\langle P_k\rangle_\theta\) can be linearly recombined whenever the Pauli basis overlaps, since
\[
E_i(\theta)=\sum_k c_{ik}\langle P_k\rangle_\theta.
\]
The operational method in the cited work does not jointly optimize a multi-instance loss; instead, it uses the seed trajectory to initialize each target independently [2410.21413].

The algorithm stores the seed trajectory, subsamples “every 10th parameter set,” evaluates target energies \(E_i(\theta^{(t)})\), and selects the best \( \theta^{(t)} \) from the first half of the trajectory as the warm start. The stated reason for pruning is that late-stage iterates become tightly clustered, with tiny parameter displacements and negligible energy changes, and therefore carry little new information for transfer. The paper formalizes the intuition through possible gradient-norm, energy-change, and clustering criteria, but the implemented heuristic is simply to ignore the second half of the trajectory [2410.21413].

The experimental setting used MaxCut Ising Hamiltonians
\[
H=\sum_{i<j} w_{ij} Z_i Z_j,
\]
with random graphs of edge probability \(\tilde{p}=1/2\), Qiskit TwoLocal ansätze with linear entanglement and \(2\) to \(5\) layers, and SLSQP as the optimizer; SPSA performed markedly worse in the reported experiments. The study considered \(n=5,6,7,\) and \(10\) qubits, with \(k=9\) target graphs for \(5\)–\(7\) nodes and \(k=2\) for \(10\) nodes, on Qiskit Aer simulator, repeated \(10\) times. The main empirical result was negative but informative: random initialization performed about as well on average as both warm-start variants, and restricting selection to the first half of the seed trajectory did not degrade performance relative to using the full trajectory. The paper attributes this outcome to the use of independent random graphs, which likely reduced the value of transfer [2410.21413].

## 7. Advantages, limitations, and points of dispute

A recurrent misconception is that ensemble VQE is synonymous with excited-state VQE. The surveyed literature does not support that restriction. In one class of methods, the ensemble is an excited-state subspace optimized with a shared unitary; in KD-VQE, it is a population of trial ground-state candidates with adaptive shot allocation; and in simultaneous-solving approaches, it is a reusable trajectory across related instances. The common structure is collective information management, but the target object differs across formulations [2509.17982][2505.03998][2410.21413].

A second misconception is that orthogonality enforcement necessarily requires explicit overlap penalties and Hadamard- or SWAP-type measurements. That is not true for shared-unitary formulations on orthonormal inputs or for ancilla purification. In those settings, orthogonality is preserved implicitly by unitarity, and overlap measurements are optional or only needed for noise-robust post-processing through a generalized eigenproblem [2206.11036][2509.17982].

The principal controversy in the recent literature concerns equi-ensemble versus weighted-ensemble optimization. The cited comparison argues that equal weights are preferable because they convert the objective into the Hamiltonian subtrace, eliminate ordering assumptions, and focus the optimization on the target subspace. Weighted ensembles are reported to be sensitive to state-ordering guesses, especially near strong vibronic couplings, avoided crossings, and conical intersections; the same study therefore concludes that “the equi-ensemble is the way to go” for the tested chemistry benchmarks [2509.17982]. A plausible implication is that the debate is not about whether ensembles are useful, but about which ensemble objective yields the most stable optimization landscape.

Limitations are equally method-specific. Ancilla-purified subspace methods require ancilla capacity \(N_a=\lceil \log_2 K\rceil\) and off-diagonal measurement overhead that scales as \(K^2\) in the worst case. KD-VQE depends on ensemble seeding, temperature schedule, and pruning thresholds; if no candidate spans the basin of the true ground state, the algorithm condenses to the best available local minimum. Cross-instance warm-start ensembles depend strongly on instance similarity; in the reported MaxCut study, transfer from a seed trajectory did not outperform random initialization on average. Taken together, these results indicate that ensemble VQE is not uniformly superior to single-state VQE. Its effectiveness depends on whether the ensemble structure matches the physics of the problem, the measurement model, and the optimizer landscape [2206.11036][2505.03998][2410.21413].

Source: https://www.emergentmind.com/topics/ensemble-variational-quantum-eigensolver-vqe