- The paper introduces joint fiducial grouping, which estimates multiple Pauli-transfer coefficients from one compatible preparation–measurement setting while preserving unbiasedness and finite-sample guarantees.
- The method guarantees setting compression and can reduce channel uses when target weight is concentrated within commuting groups, with fSim simulations reaching up to 2.70× setting compression and 1.43× shot improvement.
- The grouped estimator lowers estimation error in simulated noise and provides an inverse-free, context-sensitive reward for reinforcement-learning gate calibration, although hardware validation and practical runtime savings remain open.
Direct fidelity estimation (DFE) provides a context-preserving route to estimating the entanglement fidelity of a quantum channel, but its standard formulation assigns one preparation–measurement configuration per sampled input–output Pauli pair, an overhead that grows rapidly for non-Clifford targets. The paper under review introduces joint fiducial grouping, which partitions the support of the target Pauli transfer matrix into blocks whose input and output Paulis are simultaneously (qubit-wise) commuting, so that a single fiducial pair yields unbiased estimates of all Pauli-transfer coefficients in the block. The authors derive finite-sample guarantees showing that grouping never increases the number of distinct input–output settings and can reduce the required channel uses when target weight is concentrated within compatible groups. They validate the framework on the continuously parameterized two-qubit gate fSim(θ,φ) and deploy the grouped estimator as a reward signal in reinforcement-learning-based gate calibration (2608.18548).
Background: direct channel fidelity estimation
For an n-qubit implemented channel E and unitary target U with d=2n, the entanglement fidelity is the normalized Hilbert–Schmidt overlap of their superoperators,
Fe(U,E)=d21α,β∑χU(α,β)χE(α,β),
where χA(α,β)=d1Tr[PαA(Pβ)] are Pauli transfer coefficients in the phase-free Pauli basis. Standard DFE samples pairs from Ω={(α,β):χU=0} with probability pα,β=χU(α,β)2/d2; each sampled pair requires preparing eigenstates of Pβ, applying n0, and measuring n1. Because only local Pauli fiducials bracket the circuit window, the protocol preserves the execution context — unlike randomized benchmarking, whose twirling destroys information about how a specific circuit shapes noise, and unlike context-aware fidelity estimation (CAFE), which requires physically implementing the reference inverse n2 and hence suffers a bootstrapping problem when the characterized gates are themselves imperfect.
The cost structure of DFE depends on the target's nonstabilizerness: Clifford targets have n3 and admit n4 sampling, whereas generic full-support non-Clifford targets approach n5. This makes fractional gates such as n6 — used on superconducting processors for Hamiltonian simulation — a natural but demanding testbed.
Joint fiducial grouping
A valid group n7 satisfies n8 and n9 for all members; in practice the stronger qubit-wise commuting (QWC) condition is imposed so that fiducials reduce to single-qubit rotations. Groups are constructed via a greedy clique-cover heuristic on a compatibility graph weighted by E0, adapting measurement-grouping techniques from variational-algorithm observable estimation to channel certification.
The grouped estimator samples whole groups with probability E1, where E2 collects the target coefficients of the group. A single shot prepares a uniformly random common eigenbasis state of the commuting inputs, applies E3 once, and measures all commuting output observables simultaneously; the resulting parities E4 are unbiased estimators of each coefficient, so the aggregate E5 is an unbiased single-shot estimator of E6. Notably, the authors retain the empirical mean rather than median-of-means aggregation, since median-of-means introduces bias while Hoeffding's inequality suffices given bounded single-shot contributions.
The central result is a theorem establishing additive error E7 with failure probability at most E8 using E9 sampled groups and per-group shot allocations derived from the target coefficients. The expected number of channel uses obeys
U0
where the effective group size U1 satisfies U2 and U3. Two consequences follow directly:
- Guaranteed setting compression: the number of distinct input–output settings drops from U4 to U5, giving compression factor U6 independent of the implemented channel.
- Conditional shot compression: the improvement factor U7 exceeds unity only when target weight within compatible groups is non-uniform.
The quantity U8 admits an entropic interpretation as U9, the Rényi-d=2n0 effective support of the target weight conditioned on the group. This refines the known connection between DFE hardness and stabilizer Rényi entropies: what matters here is not the global nonstabilizerness of d=2n1 but the concentration of weight inside each commuting block. A corollary sharpens the trade-off: for Clifford targets every conditional distribution is uniform, so d=2n2 and no worst-case shot saving exists (d=2n3), even though full configuration compression d=2n4 remains available. Grouping therefore benefits structured non-Clifford dynamics most, and the paper is explicit that non-Cliffordness alone does not imply a shot advantage.
Numerical validation on fractional gates
Mapping the predicted advantages across the full d=2n5 landscape of d=2n6 (with a d=2n7 coefficient cutoff), the leading shot improvement d=2n8 ranges from 1 to 1.43, concentrated on ridges where target weight clusters within compatible groups, while broad regions remain near unity. The setting compression d=2n9 ranges from 1.67 at Clifford corners up to 2.70, with a plateau near 2.5; along lines of constant support the counts are fixed (e.g., Fe(U,E)=d21α,β∑χU(α,β)χE(α,β),0 along Fe(U,E)=d21α,β∑χU(α,β)χE(α,β),1), even as coefficients vary continuously.
Fixed-point simulations at Fe(U,E)=d21α,β∑χU(α,β)χE(α,β),2 with Fe(U,E)=d21α,β∑χU(α,β)χE(α,β),3 compare grouped and standard DFE under ideal readout and symmetric 5% readout error with linear-inversion mitigation. Under both conditions the grouped estimator exhibits lower RMSE across the precision sweep, driven primarily by variance reduction rather than bias shift. The authors caution that this is an estimator-level comparison at common requested Fe(U,E)=d21α,β∑χU(α,β)χE(α,β),4, not an equal-total-budget comparison, and that the advantage is conditional on the grouping structure and allocation rule — linear-inversion REM amplifies finite-shot fluctuations by factors Fe(U,E)=d21α,β∑χU(α,β)χE(α,β),5, and grouping merely amortizes these fluctuations over shared configurations without making the assignment inversion intrinsically more accurate.
Inverse-free repeated-cycle estimation versus CAFE
Substituting DFE into CAFE's repeated-window construction yields an inverse-free alternative: local Pauli fiducials estimate Fe(U,E)=d21α,β∑χU(α,β)χE(α,β),6 at each depth without any two-design preparation or physical reference inverse, confining auxiliary errors to high-fidelity single-qubit primitives. Under a controlled Qiskit Aer noise model including coherent residuals and readout assignment (mitigation disabled), CAFE shows lower trial-to-trial variability in the fitted one-cycle fidelity (standard deviation Fe(U,E)=d21α,β∑χU(α,β)χE(α,β),7 vs. Fe(U,E)=d21α,β∑χU(α,β)χE(α,β),8 for grouped DFE at matched budgets), but larger bias and RMSE (Fe(U,E)=d21α,β∑χU(α,β)χE(α,β),9 vs. χA(α,β)=d1Tr[PαA(Pβ)]0). The fitted depth-independent SPAM term is also larger for CAFE (χA(α,β)=d1Tr[PαA(Pβ)]1 vs. χA(α,β)=d1Tr[PαA(Pβ)]2), consistent with its entangling reference operations contributing error even at zero forward repetitions. Grouped DFE thus delivers the most accurate fitted cycle fidelity at matched budget, while CAFE concentrates estimates more tightly — a distinction between variance and accuracy the paper states plainly. A further practical advantage for parametric gates is that varying target parameters updates only the classical sampling distribution, not the physical fiducial template.
Reinforcement-learning-based calibration
Finally, the grouped estimator serves as the reward in a PPO-driven, model-free calibration of the five-parameter ansatz χA(α,β)=d1Tr[PαA(Pβ)]3 against three representative fSim targets. Readout noise affects the estimated reward trajectory more than the final mean-policy infidelity, which remains comparable between ideal and mitigated conditions. The clearest benefit appears in cumulative executed configurations: grouping lowers the count by reusing compatible bases within each action evaluation, reducing compilation and controller-loading pressure — relevant because process-fidelity optimization can itself become circuit-prohibitive. The paper notes this count is a proxy for instruction overhead, not wall-clock runtime, which depends on the control stack.
Limitations and open questions
Several caveats are stated explicitly. All evidence is simulation-based; quantifying the resource accounting on real control hardware and compilation policies remains open, particularly whether FPGA-based runtime parameter selection converts setting compression into actual upload savings. Truncation of small Pauli-transfer coefficients introduces a bias bounded by χA(α,β)=d1Tr[PαA(Pβ)]4, computable from the target alone but nonzero. Shot compression is absent for flat within-group weight distributions, and the REM-related RMSE advantage is conditional on grouping structure and allocation rule. The comparison against CAFE rests on a specific simulator noise model rather than calibrated device data. Open directions include a derandomized, target-biased grouped analogue of classical-shadow process tomography that could reuse grouped configurations across nearby calibration iterations, and extension to overlapping groups, which may further reduce sampling cost.
Conclusion
This work adapts measurement-grouping methodology from variational observable estimation to channel certification, yielding a grouped DFE estimator that preserves the finite-sample guarantees of the original protocol while provably compressing input–output settings and conditionally reducing channel uses according to a Rényi-χA(α,β)=d1Tr[PαA(Pβ)]5 effective-support criterion. The fSim case studies delineate precisely where the gains arise — concentrated within-group weight, not mere non-Cliffordness — and demonstrate the estimator's utility as a context-sensitive, inverse-free reward for machine-learning-driven gate calibration. The framework offers an auditable resource trade-off for closed-loop characterization pipelines, pending experimental validation on hardware control stacks.