- The paper introduces the shots-to-approximate-solution metric STS(r) and a shell model that estimate the measurement shots needed to reach a target approximation ratio with 99% confidence.
- The paper finds a positive excitation-matched concentration advantage in 86 of 96 instances, with annealing benefits plateauing at approximately 2,500–5,000 ns and model fits achieving KL divergence below 0.1 for over 97% of tests.
- The paper shows that exact MIS solutions require shots that grow exponentially with system size at rates of 0.035–0.040, while 90%-quality solutions require only order-unity shots, highlighting where quantum concentration helps and where classical postprocessing is sufficient.
Overview
This paper introduces an operational benchmarking framework for neutral-atom quantum optimization, centered on a shots-to-approximate-solution metric STS(r) that quantifies how many physical measurement shots an annealing protocol requires to obtain at least one postprocessed output at a target approximation ratio r (with confidence 0.99). The framework is applied to maximum independent set (MIS) instances encoded on programmable Rydberg-atom arrays on QuEra Aquila, with system sizes up to 125 sites. Two methodological contributions anchor the analysis: a one-parameter "shell model" of the postprocessed output distribution, and an excitation-matched random baseline that separates genuine structural concentration from trivial excitation-density effects (2608.12858).
The STS metric and shell model
After deterministic postprocessing—feasibility projection, greedy maximization, and ℓ-swap local improvement—each shot yields a valid independent set whose deficit from the MIS size α defines a Hamming-shell index j=α−∣S∣. The empirical shell distribution is modeled as
πj(β)∝dα−je−βj,
where dα−j is the exact count of independent sets of size α−j, computed by a geometry-aware transfer-matrix dynamic program. This form is motivated both by a Jaynes-style maximum-entropy argument and by an algorithmic locality derivation: for product-measure inputs and a finite-radius postprocessing pipeline, defect factorization yields the degeneracy-weighted exponential form with corrections of order O(j2/N), so the model is most accurate precisely in the low-j region controlling near-exact targets. The success probability at ratio r0 is the cumulative shell weight up to cutoff r1, giving directly r2.
The authors are careful about interpretation: r3 measures the density of postprocessing-irreparable defects in the raw output and would be only a monotone proxy for any effective annealing temperature, not a temperature itself. Because the same Boltzmann form arises for any product-measure input—including the non-thermal Bernoulli baseline—the shell distribution alone cannot distinguish thermal from purely algorithmic origins; resolving this requires raw-bitstring statistics.
Excitation-matched baseline and measured advantage
A central concern is that apparent improvements in postprocessed quality can arise merely from variations in per-shot Rydberg excitation density r4. The paper therefore constructs a null model by applying identical postprocessing to Bernoulli bitstrings matched to the measured density, defining the residual r5. The restriction to one-point matching is deliberate: matching blockade-violation counts would build part of the structure under test into the null model, rendering a vanishing residual uninterpretable.
Experiments across four size groups (r6–r7, 30 instances each, binned by the hardness proxy r8) show that r9 rises from negative or near-zero values at short annealing times to a positive plateau around ℓ0–ℓ1 ns. At the plateau, group means of ℓ2 range from ℓ3 (small) to ℓ4 (xlarge), and the residual is positive in 86 of 96 individual instances in bins 1–4. Notably, ℓ5 itself is essentially size-independent (regression slope consistent with zero), consistent with its interpretation as an intensive log-probability per defect; the extensive growth of shot cost enters through the degeneracies ℓ6 rather than through ℓ7. The apparent decrease of ℓ8 with size is carried entirely by the small group and is not read as evidence of systematic decay.
Model validation proceeds along two axes: monotonic sensitivity of fitted ℓ9 to controlled bit-flip degradation of mock MIS solutions, and goodness-of-fit via KL divergence, with over 97% of tested instances below the empirical threshold α0. A model-free cross-check by direct counting confirms the two-regime structure independently of the shell model: at α1 the direct-count cost is 2 shots in every size group.
Two-regime scaling and large-deviation interpretation
The key result is a sharp target dependence of the operational benefit. For the exact target (α2), α3 grows approximately exponentially with α4, with fitted growth rates α5–α6 nearly indistinguishable across groups; within the shell-model description, increasing α7 reduces the shot count chiefly as a prefactor effect rather than a resolvable change of slope. For the relaxed target α8, the required shot count remains order-unity across the entire size range for both annealing data and the matched random baseline.
This two-regime structure is rationalized through a large-deviation analysis: writing the intensive shell depth α9 and assuming an entropy-density form j=α−∣S∣0, the success probability obeys j=α−∣S∣1, where the rate function j=α−∣S∣2 vanishes once the accepted cutoff j=α−∣S∣3 exceeds the dominant-shell location j=α−∣S∣4 satisfying j=α−∣S∣5. A positive j=α−∣S∣6 shifts j=α−∣S∣7 downward and reduces j=α−∣S∣8 exponentially in j=α−∣S∣9 for near-exact targets, but has weak effect once πj(β)∝dα−je−βj,0.
Directly measured rate functions support this picture partially: πj(β)∝dα−je−βj,1 at the exact target is intensive to within 8% across all size groups (πj(β)∝dα−je−βj,2), so πj(β)∝dα−je−βj,3 grows linearly with πj(β)∝dα−je−βj,4—a measured trend, not an extrapolation. However, the predicted kink at πj(β)∝dα−je−βj,5 is absent at these sizes because πj(β)∝dα−je−βj,6–πj(β)∝dα−je−βj,7 remains of order unity; the Laplace regime requiring πj(β)∝dα−je−βj,8 sets in near πj(β)∝dα−je−βj,9, roughly twice the present array sizes.
Classical reference costs
The paper is explicit that no end-to-end quantum speedup is claimed. The King's-lattice instances admit PTASs and are exactly solvable here by transfer-matrix dynamic programming with subexponential dα−j0 cost—optimal under the exponential-time hypothesis. All 120 experimental instances are solved exactly in under four seconds single-core (median 97 ms at dα−j1–dα−j2). At the relaxed target, one to two greedy passes of the classical pipeline suffice throughout, costing milliseconds. The measured shot counts therefore characterize how the protocol concentrates probability near the MIS manifold, not competitive runtime cost.
Limitations and open questions
Several caveats bear directly on the results. First, the shell model systematically overestimates the exact-hit probability dα−j3 by factors of roughly 1.2–3 in larger groups, because dα−j4 counts all maximum independent sets whereas the pipeline reaches only 1-swap-stable ones; shell-model dα−j5 should be read as a lower bound, and replacing degeneracies with stable-set counts is identified as a natural refinement. Second, target-ratio quantization (dα−j6) prevents resolution of the continuous crossover at dα−j7, and fixed-ratio targets correspond to physically different relaxations across size groups. Third, whether dα−j8 saturates or decays at large dα−j9 remains open beyond α−j0. Fourth, the physical meaning of α−j1—its connection to diabatic excitations or Kibble–Zurek-type scaling filtered by the pipeline's repair radius—is unresolved and would require raw-bitstring correlation-length measurements. Finally, the entropy-density description underlying the large-deviation argument is supported only over a narrow accessible window.
Conclusion
The paper establishes a density-controlled, approximation-aware diagnostic for neutral-atom quantum optimization. Its main quantitative findings—a positive excitation-matched concentration advantage in ~90% of instances, exponential shot-cost growth at exact targets with intensive rate function, and order-unity cost at relaxed targets—are internally consistent between shell-model fits and model-free direct counting. The framework clarifies precisely where quantum dynamics contributes operationally (near-exact targets) and where classical postprocessing alone suffices, while honestly bounding the scope of the claims against inexpensive classical solvers on the same graph family.