Leakage-Aware Randomized Benchmarking
- Leakage-aware randomized benchmarking is a protocol that explicitly models population leakage from the computational subspace to enable refined error assessments.
- It extends standard RB by incorporating modified decay models—including single-, bi-, multi-, or matrix-exponential forms—to accurately capture survival, leakage, and seepage rates.
- The approach employs Clifford-based protocols and spectral techniques to distinguish coherent from incoherent errors, leading to more precise gate fidelity estimations.
Leakage-aware randomized benchmarking (RB) comprises randomized benchmarking protocols that explicitly treat population transfer outside a designated computational subspace rather than assuming that all errors remain internal to that subspace. In the cited literature, leakage is modeled by embedding the computational space in a larger Hilbert space, modifying the RB experiment and its fitting model so that decay parameters can estimate survival, leakage, seepage, and subspace fidelity even when standard RB would misinterpret the data. Across the main protocol families, the central technical theme is that leakage changes the sequence-length dependence from the standard single-exponential form to single-, bi-, multi-, or matrix-exponential decays, with the precise form determined by the subspace structure, the gate set, and the noise model (Wallman et al., 2014, Wood et al., 2017, Helsen et al., 2020, Chen et al., 31 Jan 2025).
1. Subspace structure and leakage metrics
Leakage-aware RB starts from an explicit decomposition of the physical Hilbert space into a computational part and a leakage part. The notation varies across papers, but the structure is the same: , , or . Leakage is the unwanted transfer of population from the computational subspace into the leakage subspace, while seepage denotes population flowing back into the computational subspace (Wallman et al., 2014, Wood et al., 2017, Chen et al., 31 Jan 2025).
A basic quantity is the survival rate of a state under a channel with respect to ,
Averaging over states in gives
and for the whole space this average survival is denoted . In the same framework, incoherent leakage is
0
while coherent leakage is
1
with
2
These definitions separate irreversible population loss from reversible population exchange (Wallman et al., 2014).
A complementary formulation introduces state leakage
3
together with channel leakage and seepage rates
4
In this notation, 5 is the average population lost from 6 and 7 is the average population returning from 8 (Wood et al., 2017).
The same 2017 framework also defines coherent leakage through off-diagonal blocks between computational and leakage sectors: 9 with the bound
0
where 1. Channel-level coherent leakage and seepage obey analogous bounds,
2
This makes explicit that population transfer and coherence between subspaces are distinct aspects of leakage (Wood et al., 2017).
A more recent parameterization emphasizes the computational sector directly. For a channel 3, the computational population is
4
the leakage rate is 5, the depolarizing parameter is
6
and the computational error is 7. The average gate fidelity on computational inputs is then
8
This formulation is designed for fidelity estimation in the presence of leakage rather than leakage-rate estimation alone (Chen et al., 31 Jan 2025).
2. Modified decay laws and leakage signatures
A defining result of leakage-aware RB is that leakage alters the functional form of the RB decay. In an extended qutrit model, leakage errors lead to a sum of exponentials rather than a single exponential. With initial state and measurement restricted to the qubit subspace, and with an expanded single-qubit Clifford group that includes both 9 and 0 rotations, the average fidelity decay takes the form
1
For unital leakage errors, 2, so the asymptotic fidelity is 3, whereas in the absence of leakage it saturates at 4. The change in asymptote is therefore a diagnostic of leakage in that model (Epstein et al., 2013).
A scalable protocol for explicitly estimating leakage rates was developed by modifying the RB randomization so that gates are sampled from a group
5
where 6 and 7 are unitary 1-designs on 8 and 9, respectively. The experiment prepares 0 in 1, applies a random sequence, and measures the probability of being in 2 after 3 steps. The resulting averaged decay depends on the type of leakage (Wallman et al., 2014).
For incoherent leakage, with 4, the sequence average follows
5
where 6 is SPAM-dependent and 7 quantifies gate dependence. For coherent leakage, with 8, the decay is bi-exponential,
9
and
0
The distinction between a single exponential and a two-exponential decay is the operational signature that separates incoherent leakage from coherent population transfer (Wallman et al., 2014).
This directly contradicts a common simplification in which leakage is treated as a small correction to ordinary RB. The cited works instead show that the structure of the decay itself can be altered, including non-scalar matrix-exponential behavior in more general settings (Epstein et al., 2013, Helsen et al., 2020).
3. Clifford-based leakage benchmarking protocols
One major line of work modifies standard Clifford RB so that leakage and seepage can be estimated together with fidelity. In leakage randomized benchmarking (LRB) for a Clifford gate set, the protocol prepares 1 in 2, applies a random sequence of 3 Clifford gates together with a final recovery Clifford, and measures computational-subspace populations. The leakage-population decay is fit to
4
with
5
When leakage is weak, the approximations
6
connect the fit directly to leakage and seepage. A second fit,
7
yields the average gate fidelity
8
This protocol is intended to estimate 9, 0, and 1 simultaneously for a Clifford gate set (Wood et al., 2017).
A different but related construction extends the Clifford group itself to randomize phases on leakage levels: 2 The purpose of this extension is to destroy residual coherence in the non-computational subspace so that the effective dynamics reduce to populations on a 3-dimensional diagonal subspace. After twirling, the error channel becomes a real, non-negative 4 Markov matrix, and the sequence fidelity has the generic form
5
This formulation admits arbitrary gate-dependence in the error channels and does not require perturbative gate-independence assumptions (Chasseur et al., 2015).
The same protocol analyzes the variance of sequence fidelities and states that, to leading order,
6
It is also compatible with interleaved randomized benchmarking and can be expanded to benchmarking of arbitrary gates. For an interleaved gate 7, the extracted gate error is
8
The setting is described as relevant for superconducting transmon qubits, among other systems (Chasseur et al., 2015).
Taken together, these Clifford-based approaches establish two complementary viewpoints: one emphasizes direct estimation of leakage and seepage rates together with subspace fidelity, while the other emphasizes a complete RB protocol whose effective twirl converts leakage dynamics into a classical Markov problem (Wood et al., 2017, Chasseur et al., 2015).
4. General RB theory and representation-theoretic extensions
A general framework for randomized benchmarking later placed leakage-aware protocols inside a broader theory of RB. In this framework, the implementation map 9 that associates a gate to a quantum channel is not restricted to the computational subspace and may be trace-non-increasing or non-trace-preserving, thereby encompassing leakage and loss. The core structural statement is that, under realistic conditions, the RB signal is well described by
0
a linear combination of matrix exponential decays. In leakage-aware scenarios, additional subspaces appear, so there can be multiple decay channels associated with leakage-space population and return processes (Helsen et al., 2020).
The same framework states that accurate interpretation of these decay rates requires Markovianity and time-independence, closeness on average in diamond norm to a reference representation,
1
with 2, and uniform or well-controlled sampling over the group. When these conditions hold, dominant decay rates can be used as figures of merit even in the presence of leakage (Helsen et al., 2020).
A notable technical contribution of this framework is post-processing that isolates decay components without requiring physical inversion gates. For a subrepresentation 3, the filtered data are
4
where 5 is a filter function. The same paper recommends MUSIC/ESPRIT and related spectral-estimation methods for extracting multiple decay rates, which is particularly relevant when leakage produces matrix-exponential or heavily mixed exponential behavior (Helsen et al., 2020).
A parallel development extended character randomized benchmarking to non-multiplicity-free groups and derived a new leakage RB protocol for more general groups of gates. The key relaxation is that the group need only preserve the computational-leakage subspace splitting 6, rather than factor into independently sampled 2-designs on the computational and leakage sectors. In Liouville form, the leakage and seepage rates are
7
and the subspace average fidelity is
8
For the trivial irrep, the decay is fit in a form that yields
9
and more explicitly
0
Under additional assumptions, the computational-subspace fidelity can be extracted as
1
This representation-theoretic extension broadens leakage RB to groups that do not satisfy the multiplicity-free restriction of earlier character RB (Claes et al., 2020).
5. Regime-specific methods and universal-gate-set variants
Recent work has emphasized that there is no single leakage-aware fit model that is optimal across all experimental regimes. One 2025 treatment develops several methods for measuring fidelity with randomized benchmarking in the presence of leakage errors, explicitly organized by error regime and by whether leakage detection or post-selection is available. It also states that, when leakage is substantial, standard randomized benchmarking can cause a large overestimation of fidelity (Chen et al., 31 Jan 2025).
Two leakage-detection strategies are identified there: explicitly addressing leakage states and ancilla- or gadget-based leakage detection. The same work gives regime-dependent survival-probability formulas, including short-sequence linearization, an EXP-LIN model when computational error dominates, double-exponential fitting in the no-seepage regime, and basis-randomized fitting for pure population transfer (Chen et al., 31 Jan 2025).
| Regime | Method | Representative fit |
|---|---|---|
| Short sequences | Comp. survival | 2 |
| Comp. error dominant | EXP-LIN | 3 |
| Comp. error dominant | LPS | 4 |
| No seepage | 2EXP | 5 |
| No seepage | LPS | 6 |
| Population transfer | Basis randomization | 7 |
The same source associates retention models with the post-selected cases, namely 8 in the computational-error-dominant regime and 9 in the no-seepage regime (Chen et al., 31 Jan 2025).
A different 2023 development addresses universal gate sets and multi-qubit systems by proposing leakage randomized benchmarking (LRB) and interleaved LRB (iLRB) that require only access to the computational Pauli group and do not require operations that twirl the leakage subspace. The formalism maps leakage dynamics to a classical Markov chain over invariant subspaces. For the noise channel 0, the average leakage and seepage are
1
and the averaged sequence outcome is
2
where 3 is the transition matrix of the induced Markov process (Wu et al., 2023).
In the single-site leakage model, 4 reduces to an 5 matrix,
6
while in the cross-talk-free independent leakage model the computational-subspace survival factorizes over qubits. For a target gate benchmarked via iLRB, random Paulis are interleaved with the gate of interest, and for two-qubit examples such as iSWAP and CZ the paper gives analytic relations between fitted decay constants and leakage/seepage rates. This extends leakage-aware benchmarking beyond Clifford-only settings and into generic 7-site gates under stated noise assumptions (Wu et al., 2023).
6. Numerical validation, practical significance, and limitations
The principal motivation for leakage-aware RB is that leakage can invalidate the interpretation of ordinary RB data. Multiple papers therefore evaluate robustness numerically. In the 2014 survival-rate protocol, simulations with weakly gate-dependent Pauli noise gave 8 for incoherent leakage, and simulations of a three-level shelving system gave 9 for coherent leakage, both in good agreement with the predicted exponential forms (Wallman et al., 2014).
Earlier analysis of realistic RB error models concluded that, in almost all cases considered, benchmarking provides better than a factor-of-two estimate of average error rate, including leakage to higher levels, provided the leakage-aware sum-of-exponentials model is used when appropriate. The same study stated that only a small number of trials are required for high-confidence estimation if the standard error of the fidelity measurements is small, and reported reliable extraction at modest sampling such as 00 in the leakage examples (Epstein et al., 2013).
The 2023 LRB/iLRB framework reported simulations showing robustness to strong depolarizing preparation noise and measurement bit-flip errors, as well as accurate recovery of theoretical leakage and seepage rates for two-qubit gates under realistic device-inspired models. The 2025 regime-based study numerically demonstrated the methods for two-qubit randomized benchmarking and then applied them to previously shared data from Quantinuum systems. In that case, leakage-aware methods generally estimated higher (worse) 2Q infidelity than simple or legacy RB analysis by more than one standard deviation, although the absolute difference remained small; the same analysis noted that the short-sequence method underestimates infidelity when longer sequences are used because the linear approximation breaks down (Wu et al., 2023, Chen et al., 31 Jan 2025).
Several limitations recur across the literature. Many derivations assume Markovian and time-independent noise, and some variants additionally assume gate-independence or specific commutation properties of the target-gate noise (Helsen et al., 2020, Wu et al., 2023). More general leakage patterns can produce multi-exponential or matrix-exponential signals that are difficult to fit directly, which is why filtering and spectral-estimation methods such as MUSIC/ESPRIT were proposed in the general framework (Helsen et al., 2020). Another recurrent point is that 01 alone is insufficient for full characterization: thermal relaxation, erasure, and unitary leakage can share the same sum while having different physical behavior and different implications for error correction (Wood et al., 2017).
Within these constraints, leakage-aware RB has developed from modified Clifford twirls and survival-rate protocols into a family of subspace-sensitive benchmarking methods. The common conclusion across the cited work is that leakage must be treated as a first-class error mechanism rather than as a perturbation of ordinary RB fits, because it changes both the operational figures of merit and the mathematical form of the benchmarking signal (Wallman et al., 2014, Chasseur et al., 2015, Claes et al., 2020, Chen et al., 31 Jan 2025).