- The paper shows that finite-shot statistical noise, rather than overlap-matrix ill-conditioning, is the dominant source of failure in quantum Krylov ground-state energy estimates.
- The authors introduce imaginary and unitary filtering, which flag unreliable generalized eigenvalue solutions without reference energies, using thresholds tied to chemical accuracy.
- Adaptive sampling-aware regularization, especially the Lee et al. threshold, outperforms other methods, while QBKS-U generally delivers better accuracy and lower variance than QBKS-H.
Overview
This communication by Oliveira, Ziems, and Glaser examines the numerical stability of quantum Krylov subspace (QKS) algorithms for ground-state energy estimation, with a specific focus on the widely cited concern that the generalized eigenvalue problem (GEVP) arising in these methods becomes severely ill-conditioned as the Krylov subspace grows. The authors' central finding is a revision of this narrative: while ill-conditioning is real and observable in ideal, noiseless simulations, it is not the dominant failure mode in realistic finite-shot settings. There, statistical fluctuations in the measured matrix elements dominate, and reliable solutions cannot be extracted without regularization or filtering. The paper also introduces two diagnostic metrics — imaginary filtering and unitary filtering — that flag unreliable GEVP solutions without any knowledge of the true eigenspectrum.
Methodological setting
The authors consider Krylov subspaces built from B initial references {∣ψ0(b)​⟩} and a generator V^, projecting H^Ψ=EΨ onto the subspace to obtain the GEVP TΦ=ΛSΦ, where T contains expectation values of either H^ or the time-evolution operator e−iH^τ, and S is the overlap matrix. Two algorithm variants are studied: QBKS-H (with f(H^)=H^) and QBKS-U (with {∣ψ0(b)​⟩}0), the latter being preferred for early fault-tolerant hardware since its circuits are more amenable to near-term implementation.
The test systems are two strongly correlated molecules — triangular BeH{∣ψ0(b)​⟩}1 and rectangular H{∣ψ0(b)​⟩}2 — represented in an STO-3G basis with Jordan–Wigner mapping, chosen precisely because they are multi-configurational and not amenable to single-reference classical treatments. All simulations assume error-free circuit implementations; only numerical (round-off/linear-dependence) and statistical (finite-shot) errors are considered.
The proposed reliability metrics
Two observations motivate the new diagnostics:
- In QBKS-H, the Ritz variational principle guarantees a real lower bound on the ground-state energy only if {∣ψ0(b)​⟩}3 is Hermitian and {∣ψ0(b)​⟩}4 is Hermitian positive definite. When noise or ill-conditioning destroys positive definiteness, eigenvalues acquire imaginary parts. Rather than discarding these, the authors propose using {∣ψ0(b)​⟩}5 as an error indicator, with a chemical-accuracy threshold of {∣ψ0(b)​⟩}6 Hartree.
- In QBKS-U, exact eigenvalues of the unitary propagator must have unit modulus. Deviations {∣ψ0(b)​⟩}7 from unity therefore directly signal error accumulation, with a threshold scaled as {∣ψ0(b)​⟩}8 to match the target phase precision.
Both metrics are purely a posteriori: they require no reference energies and no modification of the GEVP itself.
Noiseless analysis: condition number is not predictive
In ideal simulations, increasing the number of initial references {∣ψ0(b)​⟩}9 does reduce V^0 at fixed subspace size, as expected from the greater linear independence of the block-Krylov basis. However, the ground-state energy error shows no corresponding improvement across widely varying condition numbers at small subspace sizes. This is a notable negative result: V^1 alone is insufficient to assess solution accuracy. Moreover, following equicost lines (equal numbers of distinct circuits to measure), a single reference with more iterations reaches chemical accuracy at lower cost than multiple references with fewer iterations — despite the latter's smaller condition numbers. This contradicts the intuition that minimizing V^2 via multireference constructions is the efficient strategy, and the authors accordingly restrict subsequent analysis to single-reference QKS-H and QKS-U.
Regularization via singular-value thresholding of V^3 behaves as expected but is threshold-sensitive: a moderate threshold (V^4) mitigates ill-conditioning and smooths convergence, whereas an aggressive threshold (V^5) over-truncates the Krylov basis and degrades accuracy. Unitary filtering, applied post hoc, removes most spurious solutions without altering the subspace or the condition number.
Regarding the evolution time step V^6, larger steps yield more accurate energies per subspace size (consistent with prior work), and unregularized runs become unstable after a few iterations, with anomalous energy spikes correlating strongly with unitarity violations. Filtering eliminates all such spikes, producing smooth convergence; for long time steps and small subspaces, filtered results match or exceed the best regularization threshold tested. A caveat: for large Krylov dimensions, regularization outperforms filtering alone, so the two techniques are complementary rather than interchangeable.
Finite-shot regime: statistics, not conditioning, dominate
With V^7 shots per matrix element (V^8 each for real and imaginary parts), the picture changes qualitatively. Sampling noise lowers V^9 substantially relative to the noiseless case — noise effectively breaks the near-linear dependencies — yet accurate solutions become impossible without regularization. The unregularized ground-state energy error exceeds chemical accuracy for both molecules and both algorithms, and is reliably flagged as erroneous by the proposed metrics. This is the paper's strongest claim against the conventional framing: in realistic settings the GEVP is not ill-conditioned, yet still fails without mitigation, because statistical fluctuations in H^Ψ=EΨ0 and H^Ψ=EΨ1 corrupt the algebraic structure.
Four regularization strategies are compared: a fixed singular-value threshold; an adaptive "elbow" detection on the singular-value spectrum (via the Kneedle method); a noise- and iteration-dependent threshold adapted from Kirby's analysis; and the sampling-error-aware threshold of Lee et al. The elbow method performs worst, yielding the least accurate mean energies and largest standard deviations. The Lee et al. adaptive threshold achieves the highest accuracy for both algorithms, reaching chemical accuracy at smaller subspace dimensions. At equal shot budgets, QKS-U shows moderately better mean accuracy and lower variance than QKS-H, though with less smooth step-wise convergence.
Critically, across all regularization schemes the resulting condition numbers are of similar magnitude even when accuracies differ markedly — reinforcing that H^Ψ=EΨ2 is not a suitable reliability metric, whereas H^Ψ=EΨ3 and H^Ψ=EΨ4 track solution quality closely.
Limitations and open questions
Several caveats bound the scope of these conclusions. First, all simulations assume error-free quantum circuit execution; coherent gate errors, which do not average out like shot noise, are not modeled and could alter the relative importance of numerical versus statistical instabilities. Second, the reliability metrics assess whether a solution is plausible, not whether it is accurate: a value passing the filter on an unconverged calculation is not guaranteed to lie within chemical accuracy, and the authors note that filtering alone is insufficient to achieve chemical accuracy under shot noise without accompanying regularization. Third, the study is restricted to two small strongly correlated molecules in a minimal basis; scaling behavior with system size and basis complexity remains untested. Fourth, the choice of initial reference relies on knowledge of the true ground state overlap ordering, which would not be available in practice. Finally, the filtering thresholds are calibrated to a chemical-accuracy target; generalization to other precision regimes is asserted rather than demonstrated.
Conclusion
This work reframes the stability problem of quantum Krylov subspace methods. Ill-conditioning of the overlap-matrix GEVP, often cited as a fundamental obstacle, is shown to be manageable through standard singular-value truncation and, importantly, is superseded by statistical noise as the dominant error source once finite sampling is considered. The introduced imaginary and unitary filters provide practical, assumption-free diagnostics that reliably identify spurious solutions in both regimes. The results suggest that efforts to improve quantum Krylov performance should prioritize sampling-error mitigation and regularization design over condition-number minimization, and they leave open the question of how these findings extend to hardware-level coherent noise and larger correlated systems.