Empirical Quantum Advantage (EQA)
- EQA is defined as the experimentally validated superiority of quantum computations over classical methods using rigorous task-specific metrics and resource comparisons.
- It operationalizes quantum advantage by benchmarking full workflows in areas like quantum chemistry, learning, and optimization rather than isolated circuit performance.
- EQA frameworks demand strict fairness protocols, error-mitigation strategies, and statistical evidence to distinguish genuine quantum benefits from measurement artifacts.
Searching arXiv for papers on Empirical Quantum Advantage and closely related formulations to ground the article in the cited literature. Searching specifically for the MARL paper and broader EQA framework papers referenced in the source material. Empirical Quantum Advantage (EQA) denotes a class of claims in which a quantum computation is not merely theoretically superior in asymptotic complexity, but is shown experimentally or in controlled empirical studies to outperform classical computation on a specified task under explicit correctness criteria, resource accounting, and comparison protocols. Recent work operationalizes EQA in several compatible ways: as a platform-agnostic and empirically verifiable superiority claim with validated correctness and a task-specific metric (Lanes et al., 25 Jun 2025), as “measurable gains on actual hardware over state-of-the-art classical methods” (Shah et al., 30 Sep 2025), and, in especially stringent settings, as exceeding a mathematically proven classical ceiling under identical constraints (Dahia et al., 14 May 2026). The term therefore refers less to a single universal benchmark than to a family of experimentally grounded standards for establishing quantum advantage.
1. Definitions and conceptual scope
The modern literature treats EQA as an operational notion rather than a purely asymptotic one. In the broadest framework, a task exhibits empirical quantum advantage when correctness can be rigorously validated and a quantum workflow demonstrates superior efficiency, cost-effectiveness, or accuracy relative to classical computation alone, with the superiority condition expressed through an advantage factor
together with an acceptance threshold on output quality (Lanes et al., 25 Jun 2025). A closely related formulation in “quantum utility” defines practical quantum advantage relative to the best classical device of similar size, weight, and cost, and extends the comparison to time, energy, and accuracy under explicit SWaP-C constraints (Herrmann et al., 2023).
Application-specific literatures refine this template. In quantum chemistry, EQA is framed as solving a chemically relevant task faster or at lower resource than the best classical method at matched accuracy and robustness, with explicit attention to wall-clock time, measurement cost, and reproducibility (Chan, 2024). In quantum learning and sensing, the perspective literature organizes ideal quantum advantages around five properties—predictability, typicality, robustness, verifiability, and usefulness—and emphasizes that empirical adoption will increasingly depend on time-to-solution, energy, and reliability rather than asymptotic exponents alone (Huang et al., 7 Aug 2025). In biomarker discovery, the Q4Bio program stresses that EQA is a moving target because classical baselines improve, so any claimed crossover must be revisited as both classical and quantum methods advance (Shah et al., 30 Sep 2025).
A particularly strict version appears in decentralized quantum multi-agent reinforcement learning. There, EQA is defined as performance that cannot be matched by any classical decentralized strategy, demonstrated by exceeding a known classical performance bound under identical no-communication constraints (Dahia et al., 14 May 2026). This formulation makes explicit that not every empirical quantum performance gain counts as EQA in the strongest sense: some studies report better outcomes than classical baselines without a proven classical ceiling, and therefore describe empirical advantage but not formal EQA.
2. Operational criteria, fairness, and statistical evidence
Across domains, EQA depends on the quality of the comparison protocol as much as on the quantum result itself. The general framework papers require transparent task selection, explicit acceptance criteria, matched end-to-end cost models, and disclosure of compilation, calibration, mitigation, and post-processing overheads (Lanes et al., 25 Jun 2025). The strongest empirical studies therefore do not compare isolated kernel evaluations or circuit depths, but full workflows.
Fairness protocols vary with the task. In biomarker discovery, the Q4Bio pipeline uses strict train/test separation, site-stratified splits for pathomics, the same feature pool for classical and hybrid selectors, simple downstream learners such as logistic regression and SVC, and shared computational budgets for classical solvers (Shah et al., 30 Sep 2025). In healthcare quantum-kernel studies on electronic health records, the quantum and classical models are both SVMs, differing only in the kernel, with matched regularization sweeps and fixed configuration grids over feature counts and training sizes (Krunic et al., 2021). In decentralized MARL, fairness is made unusually explicit: execution remains strictly decentralized with no runtime communication; classical and quantum CHSH policies are matched at four trainable parameters each; Bell-state preparation adds zero trainable parameters; and multiple seeds, standardized training steps, and variance bands are reported (Dahia et al., 14 May 2026).
Optimization-oriented EQA introduces a further statistical layer when the true optimum is not classically available. In error-mitigation-aware benchmarking, the central quantity is the “confidence of advantage,”
where is a certified classical interval containing the unknown optimum and is the finite-shot quantum energy estimator (Demarty et al., 26 Jan 2026). This approach makes empirical advantage probabilistic and resource-aware: error mitigation may reduce bias while increasing variance, so potential advantage is meaningful only relative to shot budget, circuit depth, and mitigation overhead.
This emphasis on protocol design reflects a common concern in the literature: algorithmic coincidence, metric cherry-picking, and underpowered baselines can create the appearance of advantage without isolating a genuinely quantum mechanism. Much of the EQA literature can be read as an attempt to eliminate those ambiguities.
3. Mechanisms by which EQA is established
Empirical studies associate EQA with several distinct quantum resources. One mechanism is nonclassical correlation from entanglement under a proven classical ceiling. In the CHSH nonlocal game, Alice and Bob receive and win when . The classical win rate is bounded by $0.75$, while the Tsirelson limit is
Entangled QMARL agents consistently exceed $0.75$ and approach 0, whereas unentangled quantum circuits match the classical baseline, isolating entanglement as the active coordination mechanism (Dahia et al., 14 May 2026). The advantage concentrates on the hard input pair 1, precisely where classical decentralized strategies cannot satisfy the parity constraint.
A second mechanism is coherent multi-copy processing of quantum data. In “learning from experiments,” the quantum-memory model stores multiple copies coherently and applies collective measurements before a final classical readout, whereas the classical-memory model destructively measures each copy and retains only classical outcomes. The paper proves exponential sample-complexity separations for predicting properties of quantum states, for QPCA-type tasks, and for learning approximate dynamical models, and reports demonstrations on up to 40 superconducting qubits and 1300 quantum gates (Huang et al., 2021). The claimed advantage is not primarily in faster classical post-processing, but in extracting more information per physical experiment.
A third mechanism is constrained quantum generation in subspaces that classical samplers handle poorly. In CE-QAOA for constrained optimization, the algorithm operates inside the one-hot product space 2, initializes each block in a 3 state, preserves feasibility with a two-local XY mixer, and wraps constant-depth sampling with a deterministic classical checker. The resulting PHQC solver returns the best observed feasible solution in 4 time, achieves a 5 shot-complexity reduction when 6 locations are fixed, and yields an 7 minimax separation against a classical baseline restricted to raw bitstring sampling (Onah et al., 18 Nov 2025).
A fourth mechanism is efficient learning combined with classically hard inference. In generative quantum models, instantaneously-deep QNNs and tomographically complete shallow QNNs are designed so that training is efficient, free of barren plateaus or proliferating local minima, yet inference can reproduce deep-circuit sampling beyond the reach of classical simulation. On a 68-qubit superconducting processor, the reported experiments show efficient learning of classically intractable probability distributions and learning of compressed circuits for accelerated physical simulation, with mapped beyond-classical regimes extending to 34,304 shallow qubits (Huang et al., 10 Sep 2025). A related but more targeted result concerns learning from noisy quantum data: coherent quantum processing already separates from fixed-measurement schemes at 30–40 noisy qubits, and matching the noisy coherent protocol would require months or years of measurements under the optimized measure-first baselines studied (Danaci et al., 20 May 2026).
4. Empirical domains and representative case studies
EQA has been pursued across a heterogeneous set of application areas, and the evidentiary standard varies substantially by domain. In precision oncology, the Q4Bio program embeds a quantum-in-the-loop HRQAOA feature-selection subroutine into a classical biomarker-discovery workflow. EQA is defined there as measurable hardware gains over state-of-the-art classical methods, and the resource analysis identifies a plausible EQA window when dense third-order PCBO instances exceed 8, given observed exponential scaling of exact classical solvers and polynomial HRQAOA overhead for reducing 9 (Shah et al., 30 Sep 2025). The paper does not claim a realized crossover at that scale, but it formalizes how such a crossover would be assessed.
In constrained optimization, CE-QAOA provides a more directly algorithmic notion of empirical advantage. On QOPTLib traveling-salesman instances with 0–1, noiseless simulations at depth 2 recover global optima with polynomial shot budgets and coarse parameter grids tied to problem size (Onah et al., 18 Nov 2025). The reported advantages, however, are benchmark-specific and simulation-based rather than hardware-demonstrated.
Healthcare quantum-kernel studies present a mixed record. In electronic-health-record prediction of six-month persistence on biologic therapy in rheumatoid arthritis, quantum-kernel SVMs on IBM hardware showed EQA in 92% of configuration-grid points for non-probability-based F1, in 33% for probability-based balanced accuracy, in 8% for probability-based F1, and in 0% for non-probability-based balanced accuracy (Krunic et al., 2021). The result is therefore metric-dependent and regime-dependent rather than universal. A gene-expression study on AML versus ALL classification is more skeptical: it adopts F1, balanced accuracy, PTRI, and geometric difference as EQA indicators, but the reported predictive metrics do not show consistent or statistically supported superiority of the quantum kernel over the classical baseline (Ghosh et al., 2024).
Game-theoretic finance supplies a distinct empirical template. In quantum trading games implemented on an ion-trap quantum computer, the two-player Prisoner’s Dilemma yields a quantum correlated equilibrium payoff of 3 per player versus a classical Nash equilibrium payoff of 1, while Chicken yields 2 versus the classical mixed-strategy value of 1 and 2 versus the classical correlated-equilibrium value of 3 (Khan et al., 27 Jan 2025). These experiments establish higher-paying equilibria on hardware, but the paper does not report shot counts, confidence intervals, or formal statistical tests, so the empirical strength of the claim is narrower than in studies with explicit uncertainty quantification.
Not all promising hardware demonstrations reach EQA. In quantum policy evaluation on hardware, the theoretical quadratic sample-complexity improvement over classical Monte Carlo is clear, but the study does not implement a classical baseline, does not report time-to-solution comparisons, and therefore does not demonstrate strict empirical quantum advantage (Hein et al., 9 Sep 2025).
5. Verification, controversy, and common failure modes
A recurrent theme in the EQA literature is that advantage claims are unusually fragile to baseline choice, metric selection, and verification strategy. The review literature on quantum versus classical computational advantage documents multiple historical claims that were later weakened or refuted by classical simulation advances, especially in random circuit sampling and photonic sampling (LaRose, 2024). A more argumentative assessment concludes that finite-fidelity random circuit sampling has achieved quantum advantage under a physics-level standard of evidence, especially in weak-noise regimes where measured XEB tracks circuit-averaged fidelity, but it also acknowledges that the task is “nearly useless,” that verification relies on proxies and extrapolations, and that earlier claims were later simulated classically (Hangleiter, 10 Mar 2026). The difference between these positions is not primarily factual; it reflects different thresholds for what counts as validated correctness.
The perspective literature generalizes this concern by warning against “pseudo-advantages.” Hidden data-access assumptions, omitted preprocessing cost, unfair baselines, postselection, and metric shopping can all create apparent superiority that disappears under normalized resource budgets (Huang et al., 7 Aug 2025). In application studies, this problem often appears as metric sensitivity. The EHR quantum-kernel paper reports strong EQA by non-probability-based F1 and none by non-probability-based balanced accuracy (Krunic et al., 2021); the gene-expression study finds heuristic indications of favorable quantum regions through PTRI and geometric difference, yet does not establish consistent predictive superiority (Ghosh et al., 2024). Such results do not negate EQA, but they show that the phenomenon is often local in task space.
Another recurring issue is the gap between asymptotic and empirical claims. The hardware quantum-policy-evaluation paper is explicit that the theoretical 4 versus 5 sample-complexity separation does not by itself establish EQA; on current devices, circuit depth, oracle-construction overhead, and noise prevent an observed hardware advantage over classical Monte Carlo (Hein et al., 9 Sep 2025). The same distinction motivates recent formal work on instance-level quantum advantage, which defines “queasy” instances as those whose quantum instance complexity is significantly smaller than their classical instance complexity, thereby shifting attention from worst-case asymptotics to single-instance empirical regimes (Buhrman et al., 2 Oct 2025).
6. Limitations and future directions
Most existing EQA evidence is constrained by noise, finite shots, and limited scale. The QMARL CHSH result is compelling precisely because it uses noiseless simulators, provable classical ceilings, matched parameter counts, and controlled ablations; transferring that standard to general Dec-POMDPs, noisy hardware, and larger agent counts remains open (Dahia et al., 14 May 2026). The Q4Bio biomarker-discovery program likewise identifies hardware limitations, sparsification compromises, and the need for faster correlation-dictionary construction as central determinants of whether the predicted crossover region can actually be reached (Shah et al., 30 Sep 2025).
A major methodological direction is finite-shot, mitigation-aware benchmarking. The optimization benchmarking framework centered on
6
shows how EQA claims can be recast as confidence statements that explicitly include noise, mitigation bias, and variance inflation (Demarty et al., 26 Jan 2026). This is particularly important for near-term optimization, where the optimum is not classically accessible and advantage must be established relative to certified classical intervals rather than exact answers.
Another direction is the tightening of reporting standards. The platform-agnostic framework papers recommend publishing instance sets, generation procedures, baseline algorithms and software versions, compilation details, mitigation overheads, confidence intervals, and reproducibility artifacts, together with end-to-end quantum and classical pipeline times rather than isolated kernel or circuit timings (Lanes et al., 25 Jun 2025). This suggests that future EQA work will increasingly resemble benchmark science rather than proof-of-principle demonstration.
Finally, several literatures converge on the need for stronger verification and more useful target tasks. The historical reviews point toward classically verifiable advantage, fault-tolerant demonstrations, and the “100 logical qubits” regime as the next major milestone beyond noisy random-circuit sampling (Hangleiter, 10 Mar 2026). In application domains, a plausible implication is that the most durable EQA claims will come either from tasks with provable classical bounds, as in CHSH, or from workflows where correctness can be certified and quantum benefit survives updated classical baselines, finite-shot statistics, and realistic overhead accounting.