---
title: 'Quantum Volume (QV): Benchmarking NISQ Devices'
url: https://www.emergentmind.com/topics/quantum-volume-qv
type: topic
---

# Quantum Volume (QV): Benchmarking NISQ Devices

Quantum Volume (QV) is a system-level, single-number metric quantifying the balanced interplay of width, depth, fidelity, connectivity, and compilation quality for a noisy intermediate-scale quantum (NISQ) device. QV captures the largest size of a random, square quantum circuit (equal width and depth) that a device can execute such that the measured heavy-output probability (HOP) exceeds a fixed threshold, typically 2/3. Critically, QV is designed to be architecture-agnostic and holistic, folding hardware, compiler, and device variability into a single operational benchmark [2203.03816].

## 1. Formal Definition and Benchmarking Protocol

Formally, a device’s quantum volume is
\[
\mathrm{QV}=2^d
\]
where $d$ is the largest integer for which the device can reliably run random $m$-qubit, depth-$m$ circuits (with $m=d$) and achieve mean heavy-output probability (HOP) above threshold, with high statistical confidence. The “heavy set” for each random circuit $U$ is defined as
\[
H_U=\{x: p_U(x)>p_\text{median}\}
\]
where $p_U(x)=|\langle x|U|0\rangle|^2$ and $p_\text{median}$ is the median of all outcome probabilities. The heavy-output probability for each random instance is the fraction of measurement outcomes in the heavy set. The QV pass criterion requires that over $k$ circuits (commonly $k\geq100$), the mean HOP $\geq 2/3$, with the lower $2\sigma$ bound and the $99\%$ confidence lower bound also exceeding $2/3$ [2203.03816, 2110.14108].

Circuit construction proceeds layer-wise: every layer consists of a random permutation of all $m$ qubits and a set of disjoint two-qubit SU(4) random unitaries. If $m$ is odd, one qubit is idle each layer. The device "passes" for width $m$ if the statistical pass criteria above are satisfied [1811.12926, 2203.02108].

## 2. Holistic System Characterization

Quantum Volume is expressly designed to be a holistic system benchmark [2303.02108, 1811.12926, 2008.08571]. It probes the interplay and limitations arising from:

- **Gate fidelities:** Imperfect single- and two-qubit gate errors accumulate with increasing depth, reducing HOP at larger circuit sizes.
- **Qubit connectivity:** Non-all-to-all connectivity increases SWAP overhead during random permutations, compounding error and reducing effective maximum $m$.
- **Coherence time:** Maximum executable depth at a given width is ultimately coherence limited.
- **Compiler optimization:** Improved gate decompositions, noise-adaptive routing, and advanced transpiler routines can compress circuit depth and decrease cumulative error.
- **Variability:** QV is sensitive to gate calibration drifts, cross-talk, and device-to-device as well as subset-to-subset performance variations [2203.03816].

This comprehensive scope is achieved by essentially requiring the device to successfully execute deep, wide, entangling, and randomly structured circuits—which collectively simulate realistic complex workloads (in a worst-case sense)—at scale.

## 3. Statistical and Practical Considerations

The QV test is operationalized as follows:

- For each width $m$, generate $k\geq100$ random depth-$m$ circuits.
- For each circuit, obtain HOP by comparing measured bitstrings to the classically precomputed heavy set.
- Aggregate results over all circuits and test if mean HOP and confidence intervals clear the $2/3$ threshold [2110.14808].

The ideal mean HOP is $(1+\ln2)/2\approx0.846$ for large $m$ (Porter-Thomas limit). In the fully depolarized limit, HOP falls to 1/2. The threshold $2/3$ is chosen to robustly exclude trivial classical sampling [2110.14808]. For physically feasible QV tests (up to $m\sim 8$ classically), the test is computationally bounded by classical simulation for heavy set identification; beyond this, scalable mirror-circuit or parity-constrained variants are required [2502.02575, 2303.02108].

Typical QV experiments on major platforms include IBM Q, IonQ, Rigetti, OQC, Quantinuum, with measured values (2022) ranging from $QV=2$ (OQC Lucy) to $QV\geq512$ (Quantinuum H1-2). Device-to-device and subset-to-subset variability is significant—only specific qubit subsets satisfy the QV pass for a given $m$, and the same subset may drift in/out of pass region over time [2203.03816].

## 4. Compilation, Transpilation, and Error Mitigation Impact

Compilation and classical pre-processing are major QV determinants. Enhanced transpilation (qubit subset enumeration, noise-adaptive layout, advanced routing, and pulse-level schedule optimization) has been shown to substantially boost achievable QV compared to black-box, default transpilation [2008.08571, 2203.03816]. Custom passmanagers, such as IBM’s QV passmanager (incorporating CPLEX routing and pulse optimization), yield fewer compilation failures and higher maximal $m$, at the cost of significant classical resources (upwards of 100,000 CPU-hours for $m\leq7$).

Error mitigation techniques (notably zero-noise extrapolation and dynamical decoupling) have demonstrably increased effective QV by one or more increments (e.g., $QV$ from $2^5$ to $2^6$ on IBM devices), as measured either with the canonical or mirror quantum volume protocol [2203.05489, 2306.15863]. The “effective quantum volume” must be reported together with applied mitigation and shot overhead to maintain a fair metric [2303.02108, 2306.15863].

## 5. Variants, Extensions, and Limitations

Original QV specifies “square circuits” (width = depth). Miller et al. extend the concept with Quantum Volumetric Classes, QV-$k$, targeting random circuits of width $n$ and depth $n^k$. These variants (QV-2, QV-3, etc.) better capture the scaling of circuit depths for workloads with $d \propto n^\alpha$, aligning volumetric benchmarking with algorithmic structure [2207.02315, 2502.00113]. In practical benchmarking (91% of surveyed quantum algorithms), square (QV-1), quadratic (QV-2), and cubic (QV-3) circuit families suffice.

QV is resource intensive: typical experiments require $>1000$ random circuits and $>100$ shots/circuit. Scalability to $m\gg8$ is classically blocked by exponential cost of the heavy set identification. Parity-constrained QV variants eliminate the O($2^m$) simulation overhead by structuring the random circuits to have known heavy output subspaces, enabling benchmarking of devices with size far beyond classical simulation [2502.02575]. Mirror Quantum Volume (MQV) and related techniques use inversion-based circuits to avoid this bottleneck, but may mask some error types [2303.02108].

QV has several caveats: it does not directly reflect algorithmic performance for non-scrambling, highly structured workloads; it ignores circuit speed (addressed by metrics such as CLOPS); its single-number nature gives only coarse granularity; and its accuracy depends on transparent reporting of compiler/mapping details [2203.03816, 2110.14108].

## 6. Theoretical Foundations and Modeling

Quantum Volume reflects a rigorous theoretical interplay between native error rates, connectivity, and gate durations. The effective error per circuit layer governs the achievable $m$; for simple depolarizing models, the cross-over point at which mean HOP drops below $2/3$ is set by cumulative error per layer exceeding approximately $1/(3m)$ [2412.12959]. For NISQ architectures, QV can be estimated analytically from gate error rates, connectivity exponent $m$, and circuit depth scaling [2502.00113]. Extensions to fault-tolerant architectures incorporate error-corrected logical gates and magic-state distillation overheads to predict QV-k in the logical regime.

Platform-specific QV definitions exist for photonic and measurement-based architectures. For photonic MBQC with Gottesman-Kitaev-Preskill encoding, QV is analytically derived as a function of GKP squeezing and photon transmission efficiency via error-mapping to effective Pauli error probability, then standard QV pass criteria [2208.11724].

Topologically, Quantum Volume appears in the context of two-dimensional insulators as the Brillouin zone integral of the square root of the quantum-metric determinant; in this setting, QV variation tracks topological phase transitions and bounds the count of symmetry-protected boundary states [2501.17671].

## 7. Current Best Practices and Outlook

Best practice prescriptions for maximizing QV include:

- Enumerate connected qubit subsets and select for highest HOP.
- Employ aggressive compiler optimizations: noise-adaptive layouts, advanced routing, pulse-level schedule.
- Exploit dynamic decoupling and readout error mitigation wherever permissible.
- Document and report error mitigation overheads and compilation-level optimizations for transparency [2203.03816, 2303.02108].

Principal open challenges include: standardizing compiler stacks for reproducibility, reducing classical overhead for heavy set computations (especially for $m>8$), integrating QV full-stack analysis with error-corrected logical layers, benchmarking non-square (rectangular) algorithmic workloads, and systematically correlating QV to application performance metrics [2203.03816].

Quantum Volume remains the prevailing system-level metric for quantifying NISQ device capabilities. Its robustness, platform-independence, and sensitivity to both hardware and software-stack optimizations have made it foundational for benchmarking and progress tracking in quantum computing [2203.03816, 1811.12926, 2110.14108].

Source: https://www.emergentmind.com/topics/quantum-volume-qv