- The paper introduces stabilizer scars as verifiable witnesses of quantum advantage, combining a polynomial-dimensional simulable subspace with fidelity transfer to hard-to-simulate states.
- The paper proves direct fidelity estimation requires polynomial resources, with copy costs scaling as O(d_s²), and demonstrates size-independent failure rates in simulations up to 142 qubits.
- The paper shows scar and typical-state fidelities closely match under weak local depolarizing noise, while noting open challenges involving correlated noise, logical errors, and inefficient sampling.
Motivation and problem statement
Establishing quantum advantage in quantum simulation remains difficult for a structural reason: the same property that makes a computation classically intractable—highly entangled, non-integrable dynamics—also makes its output exponentially costly to verify. Early supremacy demonstrations based on sampling from hard distributions suffer from exactly this tension (2607.14212). The paper under discussion proposes an escape route tailored to non-equilibrium quantum simulation: use models hosting stabilizer scars, a class of quantum many-body scars (QMBS) whose subspace is spanned by stabilizer states. Within this polynomially large subspace, dynamics are classically simulable and direct fidelity estimation (DFE) is efficient; crucially, the authors then argue that the measured scar-subspace fidelity bounds the fidelity of classically intractable simulations run on the same device.
The protocol rests on three ingredients: (i) an exact QMBS subspace S of dimension ds polynomial in qubit number n, embedded in an exponentially large Hilbert space of thermalizing states; (ii) efficient DFE for any state in S; and (iii) equivalence, under local depolarizing noise, between the average fidelity of scar states and that of typical (non-scar) states evolved by the same circuit.
Efficient direct fidelity estimation from stabilizer structure
DFE estimates F(ρ,σ)=∑kχρ(k)χσ(k) by sampling Paulis Wk with probability Pρ(k)=χρ2(k) and measuring χσ experimentally. With exact expectation values, s=⌈1/(ϵ2δ)⌉ samples suffice independent of system size. The bottleneck is the shot cost per sample, which scales inversely with ∣Tr[ρWki]∣2—exponentially small for typical states, rendering DFE hopeless there.
The paper's central technical result is that this quantity is only inverse-polynomial small throughout the scar subspace. Using the packing bound that any state in a ds0-dimensional space has stabilizer fidelity at least ds1, combined with the Haug–Piroli relation between stabilizer fidelity and stabilizer Rényi entropy (SRE), they obtain
ds2
which implies ds3. Since ds4 is bounded and hence sub-Gaussian (Hoeffding), the distribution concentrates around this lower bound, so MCMC-based DFE rarely proposes observables with exponentially small expectation values. Via the Leone et al. bound involving the 0-SRE—which scales only as ds5 because only ds6 Paulis have nonzero support on ds7—the total copy cost obeys
ds8
This is a strong claim: fidelity of arbitrary superpositions within the scar subspace is verifiable with resources scaling polynomially (ds9) in system size, even though the ambient Hilbert space is exponentially large. Numerically, Metropolis-Hastings DFE at n0 achieves size-independent failure probability up to n1 qubits, and both worst-case and empirically optimal shot costs scale as n2. One caveat the authors concede: the n3 versus n4 discrepancy between their two bounds may reflect looseness in the Haug–Piroli prefactor, and their uniform-proposal MCMC sampler is itself not efficient—a point deferred to forthcoming work.
Fidelity transfer to classically non-simulable states
The pivotal assumption enabling benchmarking beyond simulability is that the dominant noise is local depolarization, n5. For Haar-typical inputs the average channel fidelity is exactly n6. For inputs drawn uniformly from the scar subspace, the authors compute n7 via projected Kraus operators n8 and prove two-sided bounds bracketing n9:
S0
tightened at larger S1 using exact combinatorial evaluation of S2 over products of the stabilizer generators S3, S4 (nonzero traces occur only when the number of S5 factors is zero or even, with closed-form values S6 in the former case).
Implication: in the weak-noise regime (S7 or asymptotically large S8), the fidelity of classically non-simulable thermalizing states equals that of the efficiently measurable scar states, so a scar-benchmark measurement certifies simulation quality where classical verification is impossible. Levy's lemma extends this from the average to all but an exponentially small fraction of Haar-random states, via an elementary proof that the channel fidelity is Lipschitz continuous with constant 4. The honest limitation, stated plainly by the authors: away from weak noise the upper bound exceeds S9 by subleading corrections, so the scar benchmark may overestimate fidelity of typical states at constant F(ρ,σ)=∑kχρ(k)χσ(k)0; and result (iii) is proven only as an ensemble average, with circuit-level evidence at F(ρ,σ)=∑kχρ(k)χσ(k)1 qubits serving as numerical support rather than proof.
Prototype model and circuit-level evidence
The concrete testbed is a non-integrable Ising model on an F(ρ,σ)=∑kχρ(k)χσ(k)2 lattice with periodic boundaries and parity constraint, dual to F(ρ,σ)=∑kχρ(k)χσ(k)3 lattice gauge theory. For odd F(ρ,σ)=∑kχρ(k)χσ(k)4 it hosts a QMBS subspace of dimension F(ρ,σ)=∑kχρ(k)χσ(k)5, spanned by explicit stabilizer basis states built from F(ρ,σ)=∑kχρ(k)χσ(k)6 and F(ρ,σ)=∑kχρ(k)χσ(k)7 generators; the underlying commutant algebra arises from a conserved single-particle/single-hole fermionic sector under Jordan-Wigner transformation. Notably, scar and non-scar initial states are prepared by the same Bell-pair circuit acting on different input bitstrings, so preparation complexity is identical.
Classical emulation of noisy circuits (Qiskit, single-qubit depolarizing noise after each gate, F(ρ,σ)=∑kχρ(k)χσ(k)8) shows that fidelities of scar-evolved and non-scar-evolved outputs are nearly indistinguishable already at F(ρ,σ)=∑kχρ(k)χσ(k)9, converging to the analytical average-fidelity curves, with the difference Wk0 decreasing with circuit depth Wk1. Shot-sampled MCMC chains (Wk2 shots per Pauli, Wk3 samples, Wk4) achieve precision substantially better than the theoretical bound, consistent across varied shot budgets.
Limitations and open questions
Three caveats are acknowledged explicitly. First, the fidelity-equivalence claim (iii) is established on average over input ensembles; whether it holds for the structured subset of states produced by actual circuit evolution is supported numerically but not proven. Second, the analysis assumes local depolarizing noise; extension to other error models is open, although the authors argue generic noise is unlikely to preserve the commutant of a nontrivial simulated Hamiltonian. Third—and most pointedly—it is unclear whether local stochastic noise is a valid effective description for logical qubits, given recent indications of non-Markovian logical noise and the poor understanding of logical errors under approximate decoding. Additionally, the demonstrated MCMC proposal scheme is not efficient, and the tightness of the SRE-based copy bound depends on an unproven refinement of the Haug–Piroli inequality.
Conclusion
This work converts the classical simulability of stabilizer scars from a curiosity into a verification resource: polynomial-cost DFE inside a Wk5 stabilizer subspace, plus a proven average-fidelity equivalence under local depolarization, yields a scalable benchmark whose measured value constrains the fidelity of classically intractable simulations of the same circuits. The approach is demonstrated analytically and numerically on a Wk6 LGT-dual Ising chain up to 142-qubit DFE emulation. Its ultimate value hinges on hardware implementation and on whether the noise assumptions survive fault-tolerant encoding—questions the paper leaves open rather than resolves.