---
title: Active Sampling SQD for Energy Estimation
url: https://www.emergentmind.com/topics/active-sampling-sample-based-quantum-diagonalization-as-sqd
type: topic
---

# Active Sampling SQD for Energy Estimation

Active Sampling Sample-based Quantum Diagonalization (AS-SQD) is a sample-based, perturbation-guided subspace method for estimating ground-state energies of quantum many-body Hamiltonians from finite samples of computational-basis measurements produced by an imperfect quantum device. It was introduced as a formulation of Sample-based Quantum Diagonalization (SQD) as an active learning problem: given an initial set of measured bitstrings, the method asks which additional basis states should be added to the effective subspace in order to most efficiently improve the ground-state energy estimate. In AS-SQD, the Hamiltonian is restricted to a selected set of computational-basis states and classically diagonalized, while basis growth is guided by an acquisition function derived from Epstein–Nesbet second-order perturbation theory rather than by passive reuse of only sampled configurations or by random expansion [2603.13536].

## 1. Problem setting and motivating constraints

AS-SQD is motivated by the operating conditions of near-term quantum devices, where access to a quantum state is typically limited to a finite multiset of computational-basis measurement outcomes rather than to amplitudes or full tomography. If a device prepares a state \(|\psi\rangle\), the observed bitstrings \(b \in \{0,1\}^n\) are distributed according to
\[
p(b) = |\langle b|\psi\rangle|^2.
\]
The setting of interest further includes imperfect state preparation with excited-state contamination. The paper studies a mixture model
\[
p(b) = (1-\eta)|\langle b|\psi_0\rangle|^2 + \eta |\langle b|\psi_1\rangle|^2,
\]
where \(|\psi_0\rangle\) is the ground state, \(|\psi_1\rangle\) the first excited state, and \(\eta\) the contamination rate; the main experiments use \(80\%\) ground state and \(20\%\) first excited state [2603.13536]. State preparation and measurement (SPAM) errors and gate noise further distort the empirical distribution.

The target quantity is the ground-state energy \(E_0\) of an \(n\)-qubit Pauli Hamiltonian
\[
H = \sum_{\ell=1}^L c_\ell P_\ell,\qquad P_\ell \in \{I,X,Y,Z\}^{\otimes n},\qquad c_\ell\in\mathbb{R},
\]
under the restriction that one does not perform full state tomography, does not measure each Hamiltonian term exhaustively as in naive VQE, and does not assume access to an ideal pure ground state [2603.13536].

This problem formulation places AS-SQD within a broader family of quantum-centric subspace methods in which the quantum processor is used primarily as a configuration sampler and the decisive spectral computation is transferred to a classical restricted-space diagonalization [2603.13536]. A plausible implication is that AS-SQD is best viewed not as a replacement for all ground-state algorithms, but as a post-sampling inference procedure tailored to finite-shot, noisy, and contaminated data.

## 2. SQD background and the failure modes addressed by AS-SQD

SQD begins from a set of computational-basis states
\[
S = \{ |s_1\rangle, |s_2\rangle, \dots, |s_m\rangle \}
\]
and defines the restricted Hamiltonian
\[
H_S = [\langle s_i|H|s_j\rangle]_{i,j=1}^m.
\]
Diagonalizing \(H_S\) yields the lowest restricted eigenvalue \(E_S\) and an approximate ground state
\[
|\psi_S\rangle = \sum_{s\in S} c_s |s\rangle,
\]
with variational energy
\[
E(S) = \min_{|\psi\rangle\in \mathrm{span}(S)} \frac{\langle\psi|H|\psi\rangle}{\langle\psi|\psi\rangle}.
\]
Because \(m\) is much smaller than \(2^n\), this restricted diagonalization can be inexpensive even when the full Hilbert space is not [2603.13536].

In naive SQD, the initial subspace is built by taking the top-\(K\) most frequent measured bitstrings and diagonalizing once. The paper identifies three failure modes of this strategy. First, finite-shot bias causes important low-probability basis states to be missed. Second, excited-state contamination pushes the most frequent measured configurations away from the true ground-state support. Third, once the initial subspace is fixed, there is no systematic mechanism to add energetically relevant states that were not directly observed [2603.13536].

AS-SQD addresses these pathologies by replacing passive subspace construction with adaptive basis-state acquisition. This intervention is motivated by the observation that random expansion becomes inefficient as system size grows, while simple reuse of measured configurations inherits the statistical and physical biases of the raw sample [2603.13536].

The broader SQD literature sharpens the significance of this design choice. A critical assessment of SQD for Heisenberg and Hubbard models found that even probability-ordered inclusion of computational-basis configurations exhibits exponential growth in the number of configurations required to reach fixed fidelity thresholds, and argued that this reflects intrinsic delocalization of the wavefunction in the computational basis rather than mere sampling inefficiency [2605.02494]. That result concerns probability-ordered inclusion rather than the Hamiltonian-coupling criterion used in AS-SQD, but it establishes an important backdrop: active basis acquisition must be judged not only by finite-shot improvements but also by how effectively it exploits Hamiltonian structure.

## 3. Active learning formulation and perturbation-theoretic acquisition

AS-SQD casts SQD as an active learning problem in Hilbert space. Given a current subspace \(S\) and its lowest eigenpair \((E_S,|\psi_S\rangle)\), the method generates a candidate set of external basis states connected to \(S\) under the Hamiltonian,
\[
\mathcal{N}(s) = \{ |k\rangle : \langle k|H|s\rangle \neq 0\},\qquad
C(S) = \Big( \bigcup_{s\in S} \mathcal{N}(s) \Big)\setminus S,
\]
and asks which \(k\in C(S)\) should be added next [2603.13536].

The ranking criterion is derived from Epstein–Nesbet second-order perturbation theory. For an external basis state \(|k\rangle\), the paper introduces the EN-type energy contribution
\[
\Delta E_k^{(2)} \approx \frac{|\langle k|H|\psi_S\rangle|^2}{E_S - H_{kk}},
\qquad
H_{kk}=\langle k|H|k\rangle,
\]
and then defines the acquisition score by its magnitude,
\[
a(k) = \frac{|\langle k|H|\psi_S\rangle|^2}{|E_S - H_{kk}|}.
\]
In implementation, a regularized form is used,
\[
a(k) = \frac{|\nu_k|^2}{\max(|E_S - H_{kk}|,\epsilon)},
\qquad
\nu_k = \sum_{s\in D} c_s \langle k|H|s\rangle,
\]
where the sum is restricted to the dominant support
\[
D = \{ s\in S : |c_s|^2 > \tau \}.
\]
The numerator measures Hamiltonian coupling to the current approximate ground state, while the denominator penalizes candidates whose diagonal energies are far from the current energy estimate [2603.13536].

The paper emphasizes that this is “physics-guided” basis acquisition because strong matrix elements \(|\langle k|H|\psi_S\rangle|\) directly indicate participation in the low-energy manifold, while energetic proximity encoded by \(|E_S-H_{kk}|\) suppresses very high-energy directions. An ablation study further shows that the coupling term is the dominant signal: the full EN-like score is slightly better, but ranking only by \(|\langle k|H|\psi_S\rangle|^2\) already captures most of the advantage, whereas denominator-only and diagonal-only criteria perform poorly [2603.13536].

This structure makes AS-SQD closely analogous to selected CI methods such as CIPSI and ASCI, except that the basis states are computational-basis bitstrings and the importance estimate is built from measured data and the Pauli Hamiltonian rather than from a conventional determinant expansion [2603.13536].

## 4. Algorithmic workflow and computational structure

The initialization in the reported implementation uses the top \(K=50\) most frequent bitstrings as the initial basis \(S_0\). At each iteration, AS-SQD constructs the restricted Hamiltonian
\[
H_S = [\langle s_i| H | s_j\rangle]_{i,j=1}^{|S|},
\]
computes the lowest eigenpair, identifies the dominant support \(D\), generates 1-hop candidates connected to \(D\), scores them with the acquisition function, and adds the top \(B=20\) candidates. The experiments use \(T=10\) iterations, so the final subspace size is at most \(K + BT = 250\) basis states [2603.13536].

Hamiltonian matrix elements are computed directly from the Pauli decomposition: each Pauli string maps a computational-basis state to another basis state up to a phase, so \(\langle s_i|H|s_j\rangle\) can be accumulated efficiently. Classical diagonalization scales as \(O(|S|^3)\); with \(|S|\) kept around \(10^2\), the cost is reported as inexpensive even for 16 qubits [2603.13536].

Candidate generation is deliberately local. In the main experiments, only 1-hop neighbors are considered:
\[
C = \left( \bigcup_{s\in D} \mathcal{N}(s) \right) \setminus S.
\]
The paper also tested 2-hop candidates and found that this hurt convergence under a fixed budget because the candidate pool became too large and included many weakly relevant states [2603.13536]. This suggests that a strict 1-hop neighborhood functions as a useful inductive bias rather than a mere implementation convenience.

The method requires no further quantum queries after the initial sample collection. Finite-shot noise and contamination only affect the initial basis construction; all subsequent expansion, scoring, and diagonalization steps are classical and Hamiltonian-driven [2603.13536]. That property distinguishes AS-SQD from sampling schemes that repeatedly query the device during the adaptive loop.

A related but distinct direction is SQD with amplitude amplification, where the quantum sampling distribution itself is actively reshaped by suppressing already-seen bitstrings. That approach, SQD-AA, is framed as an active sampling variant of SQD and achieves reduced query complexity relative to direct sampling for model distributions and molecular examples [2605.02565]. AS-SQD instead keeps the measured data fixed and makes the active choice at the level of basis-state inclusion. The two strategies therefore act at different layers of the SQD pipeline.

## 5. Benchmarks, empirical performance, and ablation results

The principal benchmarks in the AS-SQD paper are disordered one-dimensional Heisenberg and transverse-field Ising (TFIM) chains. The Heisenberg Hamiltonian is
\[
H = J\sum_{i=1}^{n} \left(X_i X_{i+1} + Y_i Y_{i+1} + Z_i Z_{i+1}\right) + \sum_{i=1}^{n} h_i Z_i,
\]
with \(J=1\), periodic boundary conditions, and random fields \(h_i \sim \mathcal{N}(0,h^2)\) with \(h=0.5\). The TFIM benchmark is
\[
H = -J\sum_{i=1}^{n} Z_i Z_{i+1} - h_x\sum_{i=1}^{n} X_i + \sum_{i=1}^{n} g_i Z_i,
\]
with \(J=1\), \(h_x=1\), and \(g_i \sim \mathcal{N}(0,0.5^2)\) [2603.13536].

The simulations use system sizes \(n\in\{8,10,12,16\}\), five disorder realizations per size, and \(N_{\text{shots}}=2000\)–3000 bitstrings drawn from the contaminated distribution with \(\eta=0.2\). Performance is reported in terms of the median absolute energy error
\[
\mathrm{Err} = |E_{\text{est}} - E_0|.
\]
The comparisons include standard SQD with no expansion, random SQD with random additions from the connectivity graph, and AS-SQD with EN-inspired acquisition [2603.13536].

At \(n=8\), both random SQD and AS-SQD achieve near machine precision, reflecting the small Hilbert space. For \(n=10,12,16\), standard SQD errors grow rapidly, random SQD offers limited improvement and saturates, and AS-SQD achieves substantially lower median errors across all sizes. At \(n=16\), the paper reports that AS-SQD can approximate the ground energy very accurately using only \(\sim 250\) basis states out of the \(2^{16}=65{,}536\)-dimensional Hilbert space [2603.13536].

Hardware validation is performed on IBM Quantum with noisy Trotterized circuits and SPAM errors. For Heisenberg chains up to \(n\le 12\), AS-SQD consistently outperforms standard SQD and random SQD. At \(n=8\), it recovers the exact ground energy within numerical precision, approximately \(10^{-14}\), from hardware samples despite substantial SPAM and gate noise [2603.13536]. The paper attributes this robustness to the fact that noise-induced bitstrings often have high diagonal energies \(H_{kk}\) or weak couplings to the emergent low-energy manifold, so their acquisition scores remain small.

The ablation results further delimit what matters algorithmically. The full EN-inspired score and a coupling-only score perform similarly well and clearly outperform denominator-only, diagonal-only, standard SQD, and random SQD. The candidate-horizon study shows that 2-hop expansion slows convergence under fixed budget, reinforcing the claim that local, Hamiltonian-coupling-guided exploration is the effective bias [2603.13536].

## 6. Relation to adjacent SQD developments, limitations, and outlook

AS-SQD sits within an expanding family of SQD refinements. Extended SQD for molecular excited states augments the original sampled subspace with single and double, or higher, excitations of the dominant SQD configurations and diagonalizes in the enlarged determinant space, improving excited-state accuracy over both plain SQD and QSE(SD) while keeping the additional work classical [2411.00468]. Cluster-adaptive SQD replaces a single global configuration-recovery reference by cluster-specific references to better handle multimodal determinant distributions in strongly correlated systems, yielding lower variational energies than standard SQD in stretched \(\mathrm{N}_2\) and [2Fe-2S] benchmarks [2603.09346]. PIGen-SQD adds generative machine learning and perturbative screening to configuration recovery, using a physics-informed RBM-driven search to reduce diagonalization cost while maintaining chemical accuracy under strong correlation [2512.06858]. These developments suggest that the SQD framework admits orthogonal improvements in basis acquisition, recovery, generative exploration, and post-sampling enrichment.

The principal limitations stated for AS-SQD are also explicit. Benchmarks are limited to 16 qubits, though the conceptual framework is not tied to that scale. The method assumes that the Hamiltonian is known in Pauli-string form and benefits from locality and sparsity. It uses only computational-basis samples and a fixed measurement dataset; “active sampling” refers to active basis-state selection, not to adaptive remeasurement. The denominator of the EN score requires a regularization parameter \(\epsilon\) when \(E_S \approx H_{kk}\), though this is described as a minor numerical issue [2603.13536].

Broader SQD analyses introduce a more structural caveat. For Heisenberg and Hubbard lattices, probability-ordered inclusion of configurations sampled from the exact ground state still requires an exponentially growing number of configurations to reach fixed energy fidelities, tracking the exponential growth of the effective support \(N_{\mathrm{eff}} = e^S\) derived from the Shannon entropy of the computational-basis distribution [2605.02494]. This does not directly invalidate AS-SQD, because AS-SQD ranks Hamiltonian-connected candidates rather than merely ordering configurations by probability. It does, however, indicate that computational-basis sparsity cannot be assumed generically.

The paper itself suggests several natural extensions: adaptive measurement allocation, more advanced perturbative or uncertainty-aware acquisition functions, application to molecular electronic structure and periodic solids, integration with VQE or QSE, and larger hardware experiments possibly combined with explicit error mitigation [2603.13536]. A plausible implication is that the most powerful future variants may combine several existing ideas: Hamiltonian-guided basis acquisition as in AS-SQD, quantum-side distribution shaping as in SQD-AA, mode-aware recovery as in cluster-adaptive SQD, and structure-aware postselection or generative modeling as in later SQD variants [2605.02565] [2603.09346] [2512.06858].

In its original formulation, however, AS-SQD is specifically the demonstration that finite-shot, contaminated bitstring data can be converted into accurate low-energy estimates by treating basis expansion as an active learning problem and ranking new basis states with an Epstein–Nesbet-inspired perturbative acquisition function [2603.13536].

Source: https://www.emergentmind.com/topics/active-sampling-sample-based-quantum-diagonalization-as-sqd