Quantum Circuit Unoptimization
- Quantum circuit unoptimization is a method that transforms a circuit into an equivalent one with increased cost metrics through either accidental mapping errors or deliberate redundancy insertion.
- It supports practical applications such as compiler benchmarking, digital zero-noise extrapolation, and the evaluation of hardware-induced mapping overhead on devices like IBM QX2 and QX4.
- Recent advances formalize unoptimization as an algorithmic primitive that preserves functional equivalence while providing a controlled noise-scaling parameter and exposing nonlocal Clifford+T identities.
Searching arXiv for the cited papers and closely related work on quantum circuit unoptimization, circuit optimization on IBM devices, ZNE via unoptimization, and routing/mapping. Quantum circuit unoptimization denotes the use or deliberate construction of quantum circuits whose implementation cost is higher than necessary while the intended computation is preserved. In the NISQ compilation literature, the term first arose in connection with circuits mapped to hardware without exploiting device-specific coupling constraints, cancellation identities, or optimal logical-to-physical qubit assignments, thereby introducing unnecessary SWAPs, CNOT ladders, redundant single-qubit rotations, and excess depth. A later line of work elevated unoptimization into an explicit algorithmic primitive: given a circuit , construct an equivalent circuit such that but the circuit is syntactically more complex, enabling compiler benchmarking, circuit-equivalence formulations, and digital zero-noise extrapolation (ZNE). The topic therefore spans both harmful implementation overhead and controlled redundancy insertion, with the common invariant that functional equivalence is preserved while resource usage changes (Sisodia et al., 2018, Mori et al., 2023, Pelofske et al., 8 Mar 2025).
1. Conceptual definition and cost model
In its most general form, quantum circuit unoptimization is a transformation from a circuit to a circuit such that the implemented unitary is unchanged up to global phase while one or more cost metrics increase. A formal statement used in the literature is an unoptimization primitive satisfying equivalence, , together with redundancy, , for a chosen cost function and . The corresponding cost functions include total gate count , two-qubit gate count 0, and circuit depth 1. In IBM-device studies, “circuit cost” is often the ordered pair consisting of gate count and number of levels, where “levels” is used synonymously with depth (Mori et al., 2023, Sisodia et al., 2018).
The literature uses the term in two related senses. In the first, unoptimization is accidental: compilation and mapping fail to respect hardware structure, so the realized circuit is more expensive than necessary. In the second, unoptimization is deliberate: identities, decompositions, and structurally nontrivial rewrites are inserted on purpose, keeping the same unitary while increasing syntactic complexity. The latter is explicitly described as “the inverse operation of circuit optimization,” but it is not arbitrary corruption. Under ideal execution, observables are unchanged; only the internal gate-level implementation differs (Mori et al., 2023, Pelofske et al., 8 Mar 2025).
This distinction is central because the empirical consequences differ. Accidental unoptimization typically depresses fidelity and weakens nonclassicality witnesses on noisy hardware. Deliberate unoptimization, by contrast, can be useful when the increased gate count is exploited as a controlled noise-scaling parameter, as an adversarial input for compiler evaluation, or as a structured means of exposing optimization identities that current compilers fail to use (Sisodia et al., 2018, Pelofske et al., 8 Mar 2025, Mori et al., 24 Sep 2025).
2. Hardware-induced unoptimization on IBM superconducting devices
A canonical NISQ manifestation of unoptimization appears on IBM’s 5-qubit devices QX2 and QX4, whose native gate model allows all single-qubit Clifford+T gates directly but restricts two-qubit operations to directed CNOT edges. Unsupported CNOTs must therefore be realized through additional sequences, often involving SWAP-mediated movement and direction-correction patterns. The principal sources of unoptimization identified in this setting are unnecessary SWAPs caused by poor initial mapping, non-minimal Clifford+T decompositions, excessive CNOTs induced by directionality constraints, and redundant single-qubit rotations that could be commuted, merged, or canceled (Sisodia et al., 2018).
The optimization strategy developed for these devices is architecture-aware and exhaustive at the 5-qubit scale. It precomputes minimal transformations for every unsupported CNOT, stores them in lookup tables specialized to QX2 or QX4, evaluates all 2 logical-to-physical qubit permutations, replaces every non-native CNOT with its best template, and then applies local simplifications such as 3, 4, commuting-gate reorderings, and template reductions. For reversible benchmarks, the same philosophy is applied after synthesis into Clifford+T form. This compilation discipline directly targets the pair 5 (Sisodia et al., 2018).
The quantitative effect is substantial. In a three-qubit Mermin-inequality circuit implemented on QX2, the original circuit from Alsina and Latorre had 6 gates and 7 levels, whereas the optimized form had 8 gates and 9 levels, corresponding to a 0 gate-count reduction and a 1 level reduction. The fidelity with respect to the ideal target state increased from 2 to 3. For a second Mermin circuit, the cost changed from 4 gates and 5 levels to 6 gates and 7 levels, with fidelity increasing from 8 to 9. Across other Clifford+T tasks, reported gate-count reductions include 0–1 for non-destructive discrimination circuits and 2 for teleportation-related tasks; reversible benchmarks on QX4 show reductions such as 3 gates and 4 levels for alu-v1_28 and 5 gates and 6 levels for alu-v4_37 (Sisodia et al., 2018).
The Mermin-inequality case makes the operational meaning of unoptimization especially clear. For the tripartite witness
7
the classical bound is 8 and the quantum mechanical maximum is 9. The original circuit yielded 0 with 1 shots in the earlier experiment, and 2 when re-run unoptimized with 3 shots; the optimized circuit produced 4 with 5 shots. The violation above the classical limit therefore increased from 6 to 7, illustrating how accidental unoptimization weakens observable nonclassicality on noisy devices (Sisodia et al., 2018).
More general routing work places this IBM-specific phenomenon into a broader transformation problem. On constrained QPUs, additional SWAPs and, on directed architectures, sometimes Hadamards for CNOT direction reversal must be inserted to preserve functionality, which increases circuit size and depth. A simulated-annealing plus heuristic-search method for this problem reported an average reduction of 8 in transformed circuit size on IBM Q20 relative to the state-of-the-art baseline in that study, and an average reduction of 9 on IBM QX5 versus an A*-based method. In this sense, hardware-induced unoptimization is the practical consequence of suboptimal mapping and routing rather than an isolated pathology of a particular benchmark (1908.08853).
3. Unoptimization as an algorithmic primitive and equivalence framework
The 2023 formalization of quantum circuit unoptimization turns the phenomenon into a constructive primitive. An unoptimizer 0 acts on a circuit 1 and a classical recipe 2, encoded as a finite string of elementary operations from a finite set 3, and outputs a redundant circuit 4 such that 5. The recipe is reversible in the sense that, given 6 and the recipe, the original circuit can be recovered. The elementary operations used in this framework are gate insertion, gate swapping, gate decomposition, and gate synthesis (Mori et al., 2023).
The central local gadget is an elementary recipe cycle. One selects a pair of two-qubit gates 7 and 8 sharing exactly one qubit, inserts an arbitrary two-qubit unitary 9 and its inverse 0 between them, and then replaces 1 by a conjugated three-qubit gadget
2
so that it can be swapped left through 3 while preserving the overall unitary. The resulting composite is then decomposed to two-qubit-or-smaller structure using KAK, CSD, or QSD, and locally synthesized again to discourage trivial cancellation. Repetition of this cycle systematically increases circuit depth and gate count while keeping unitary equivalence (Mori et al., 2023).
Two pair-selection strategies were studied. In the random strategy 4, candidate gate pairs are sampled afresh; in the concatenated strategy 5, subsequent elementary recipes preferentially reuse structure created by the previous cycle. Numerical simulations on random 6-qubit circuits with 7 to 8, 9 random instances per 0, and iteration budget 1 showed that the unoptimization ratio 2 grows with 3 for both strategies. After compiler optimization, Pytket compresses strongly under 4, with 5 regardless of 6, but becomes much less effective under 7, whereas Qiskit exhibits only modest compression under both strategies. The three-qubit merge ratio 8 remains near 9 under 0 but grows from approximately 1 to 2 under 3, correlating with the residual difficulty of Pytket optimization (Mori et al., 2023).
The same work proposes the Quantum Circuit Equivalence Test (QCE). In its broad form, the decision problem asks whether two circuits 4 and 5 satisfy 6. The paper states that this problem is contained both in the NP and BQP classes but is not trivially included in the P class. The NP-style perspective uses the unoptimization recipe itself as a witness: if a finite sequence of allowed equivalence-preserving operations transforms one circuit into the other, the transformation can be checked efficiently. The BQP-style perspective uses fidelity estimation, for example by constructing 7 and estimating the overlap on 8, or by performing a SWAP test on the prepared output states. A promised version of the problem is then placed in NP 9 BQP (Mori et al., 2023).
This formalization reorients unoptimization from an implementation defect into a controlled algebraic operation. A plausible implication is that it supplies a common language for adversarial benchmark generation, equivalence certification under restricted rewrite systems, and compiler-stress constructions that remain exact under ideal semantics (Mori et al., 2023).
4. Digital zero-noise extrapolation by deliberate unoptimization
A distinct application emerged when unoptimization was used as a digital noise-scaling mechanism for ZNE. In this setting, the input is a circuit 0 compiled to a target basis such as U3 and CX, and the output is a circuit 1 satisfying 2 under ideal execution but having greater gate count and depth so that noise is amplified in a controlled way on imperfect hardware. The constraints are explicit: measurements, qubit allocations, and algorithmic inputs such as angles or graph structure are unchanged; only the internal gate-level implementation is modified (Pelofske et al., 8 Mar 2025).
The algorithm is an iterative four-step recipe. First, two existing two-qubit gates 3 and 4 are selected, either by a concatenated strategy that deterministically picks the first pair sharing a common qubit or by a random strategy that samples pairs uniformly at random. A random two-qubit unitary 5 is then used to insert the identity-like sequence 6 between them. Second, 7 is swapped with 8, producing a gate 9 that captures the algebraic effect of the swap. Third, composite gates are decomposed into the target elementary basis. Fourth, a synthesis pass, such as Qiskit transpilation at optimization_level=3, recompiles the new gate fabric. Repeating the recipe 00 times yields 01, a family of unitary-equivalent circuits with monotonically increasing cost in typical cases (Pelofske et al., 8 Mar 2025).
The noise-scaling variable is defined from post-transpilation gate counts. If 02 is the total number of elementary gates in the original circuit and 03 is the count after 04 iterations, the paper defines
05
Alternative scale proxies based on two-qubit gate count, depth, or scheduled time are also discussed. The point is not merely to increase noise, but to create many distinct circuit variants at the same or similar noise scales. If 06 variants are produced for scale 07, the averaged observable
08
is used to mitigate bias from localized hardware noise and structured error propagation (Pelofske et al., 8 Mar 2025).
The method was evaluated in depolarizing-noise simulations for two benchmarks. The first used 09-qubit quantum volume circuits with heavy output probability (HOP) as the observable. The second used 10-qubit 11 QAOA circuits for Max-Cut on random 12-regular graphs, with 13 or cut value as the observable. The simulations used depolarizing rates 14, 15 shots per circuit per noise scale, 16 unoptimization iterations per recursive sequence, and both concatenated and random selection strategies. ZNE employed polynomial fitting in the scale parameter, with both linear and quadratic fits considered (Pelofske et al., 8 Mar 2025).
The reported outcome is that unoptimization-based ZNE can approximately recover signal from noisy quantum simulations. For quantum volume, the noisy HOP drifts toward 17 as the scale increases, while polynomial extrapolation recovers the noiseless HOP more accurately; quadratic fits generally yield lower error than linear fits. For QAOA, extrapolated cost values move closer to the ideal than baseline noisy execution. Averaging over 18 independent recursive sequences improves RMSE relative to single-sequence ZNE, and the best-performing configuration is quadratic extrapolation with the random selection strategy and averaging over 19 distinct unoptimization recursions (Pelofske et al., 8 Mar 2025).
This use of unoptimization reverses the usual NISQ intuition. Extra gates remain harmful for direct execution, but they become useful when interpreted as a digitally tunable parameter for extrapolation. The method also resists simplification by contemporary server-side transpilers, because the insert-swap-decompose-synthesize pattern does not collapse back to the original circuit under common peephole optimization passes (Pelofske et al., 8 Mar 2025).
5. Structured redundancy, multi-product commutation, and compiler benchmarking
A further development uses unoptimization to expose nonlocal Clifford+T identities not exploited by current compilers. The relevant setting is sequential Pauli-based computation (PBC), obtained by pushing Clifford gates to the front of a Clifford+T circuit so that the non-Clifford portion becomes a sequence of 20 Pauli rotations,
21
Within this form, T-count is the central fault-tolerant resource, and local layer-based optimizers can merge equal axes only when commutation structure permits (Mori et al., 24 Sep 2025).
The key new identity is the multi-product commutation relation (MCR). For mutually distinct Pauli axes 22, the pairs 23 and 24 satisfy MCR when 25, all cross-pairs anticommute, and 26. Under these conditions, the composite blocks commute even though the individual cross terms do not: 27 A constructive theorem shows that if 28 and 29, then the unique fourth axis completing the relation is 30. Appendix B of the same work derives the count of valid MCR axis combinations,
31
which grows exponentially with 32 (Mori et al., 24 Sep 2025).
Unoptimization enters as a controlled method for building benchmark circuits whose optimal T-count is known in advance. Starting from 33, which has 34, the unoptimizer repeatedly inserts MCR-based eight-gate identities, augments them with auxiliary 35 pairs when necessary, and performs MCR-justified swaps to create redundancy while preserving equivalence up to global phase, 36. This procedure is repeated 37 times for 38 to 39 qubits, with 40 random instances per 41, and the resulting circuits are exactly decomposed into Clifford+T (Mori et al., 24 Sep 2025).
The evaluation metric is the compiler’s ability to recover the known original T-count. The paper defines the absolute T-count gap 42 and the normalized recovery rate
43
where 44 means perfect recovery. Four toolchains were tested: Pytket (RemoveRedundancies), PyZX (full_reduce), TMerge, and FastTODD. With full MCR unoptimization, modest reductions appear only at small sizes: for 45, FastTODD achieves approximately 46 reduction, PyZX and TMerge approximately 47, and Pytket approximately 48. As 49 increases, all reduction rates decay; by 50, all achieve 51 reduction. FastTODD also introduces more than 52 ancillas at 53, with compilation times exceeding 54 hours and no T-count improvement. Without MCR-based swaps, all compilers produce 55 reduction across all 56 (Mori et al., 24 Sep 2025).
The immediate conclusion drawn in that work is that MCR-aware transformations are not incorporated into the tested compilers. More broadly, the result shows that unoptimization can be used not merely to make circuits longer, but to inject algebraically meaningful redundancy whose removal requires nonlocal identities outside present optimization repertoires (Mori et al., 24 Sep 2025).
6. Routing overhead, mapping theory, and uncomplexity bounds
Although deliberate unoptimization emphasizes equivalence-preserving redundancy insertion, a parallel literature studies the inverse problem: how to quantify and minimize the hardware-induced overhead that appears when ideal circuits are mapped onto constrained processors. In one formulation, a QPU is represented by a coupling graph 57, a logical-to-physical mapping 58 determines executability of front-layer CNOTs, and the transformation cost is dominated by inserted SWAPs and, for directed architectures, Hadamards used to reverse CNOT direction. For IBM Q20, the CNOT distance is
59
where each hop corresponds to one SWAP implemented with 60 CNOTs. For IBM QX5, the analogous cost depends on directionality and can be 61 or 62 depending on whether a final directed segment supports the desired CNOT orientation (1908.08853).
The simulated-annealing framework for initial placement minimizes an energy
63
where 64 is a near-term lookahead subset of the circuit. Routing then proceeds with a double look-ahead heuristic over SWAP and direction-flip candidates, combining exact short-horizon gate costs with a weighted estimate of the remaining workload. The paper gives polynomial-time and 65-space guarantees for the resulting procedure and explicitly interprets the reduction of inserted gates as mitigation of unoptimization, because fewer additional two-qubit operations imply lower error accumulation and depth growth (1908.08853).
A more abstract formulation uses quantum circuit uncomplexity to derive a lower bound on minimal SWAP overhead. Here the interaction graph (IG) of the circuit and the coupling graph (CG) of the processor are encoded as Gibbs-state density matrices,
66
and routing is modeled by doubly-stochastic channels over subsystem permutations together with projective edge erasures. The relevant distance is the quantum Jensen–Shannon divergence,
67
with metric 68. Minimizing this distance over admissible permutation channels yields the SWAP uncomplexity
69
interpreted as a “lightcone bound” for feasible routing paths (Steinberg et al., 2024).
This bound was compared with Qiskit and with brute-force methods on over 70 realistic benchmark experiments, and the paper reports that neither the brute-force method nor the Qiskit compiler surpasses it. The same work introduces an initial placement method based on VF2 subgraph isomorphism and graph-edit distance, selecting the processor partition nearest to a graph isomorphism between interaction and coupling graphs. In this perspective, unoptimization is not inserted by hand; it is quantified as the structural mismatch between a circuit’s interaction requirements and the processor topology. Larger graph-edit distance and more missing edges imply greater routing uncomplexity and therefore larger unavoidable overhead (Steinberg et al., 2024).
7. Applications, limitations, and interpretive issues
The application space of quantum circuit unoptimization is unusually broad. On the experimental NISQ side, reducing accidental unoptimization strengthens fidelity and improves nonclassicality witnesses such as the Mermin inequality. In compiler engineering, deliberate unoptimization produces adversarial yet exact benchmarks that reveal the limits of local peephole optimization, three-qubit squashing, and T-count reduction passes. In error mitigation, deliberate unoptimization provides a practical digital route to ZNE with many distinct circuit variants per noise scale. Additional applications proposed in the literature include quantum advantageous machine-learning datasets and fidelity benchmarks based on identity-equivalent deep circuits (Sisodia et al., 2018, Mori et al., 2023, Pelofske et al., 8 Mar 2025).
Several misconceptions are corrected by this body of work. Quantum circuit unoptimization is not synonymous with arbitrary noise injection or semantic corruption: the deliberate variants preserve the ideal unitary by construction, often up to global phase, and in the ZNE setting they also preserve measurements, qubit allocations, and algorithmic inputs. Conversely, not all unoptimization is useful. On real hardware, accidental unoptimization is empirically detrimental because each added gate and each extra level increase exposure to stochastic error and decoherence. The same increase in cost becomes useful only when it is structured, controlled, and paired with a downstream task such as extrapolation or benchmarking (Sisodia et al., 2018, Pelofske et al., 8 Mar 2025).
The limitations are equally explicit. Unoptimization-based ZNE uses gate-count scaling as an approximation to physical noise scaling; real devices exhibit drift, crosstalk, leakage, and coherent calibration errors, so averaging over variants mitigates but does not eliminate bias. Very large scale factors can hit decoherence floors, making moderate scaling with more variants preferable to extreme depth. In the compiler-benchmarking setting, the effectiveness of adversarial circuits is basis- and toolchain-dependent, and future compiler passes may learn to recognize currently opaque patterns. In the MCR setting, the number of candidate structures grows exponentially, so brute-force global search is infeasible. In the equivalence-test setting, finite precision and gate-set restrictions complicate exact certification for continuous-parameter circuits (Pelofske et al., 8 Mar 2025, Mori et al., 24 Sep 2025, Mori et al., 2023).
Current research directions follow directly from these constraints. Proposed improvements include combining unoptimization-based ZNE with randomized compiling or Pauli twirling, exploring layerwise Richardson extrapolation on top of unoptimization, developing adaptive schedules for scale factors and variant generation, and integrating MCR-aware transformations into T-optimizing compilers. This suggests that quantum circuit unoptimization is evolving from a descriptive label for poor mapping into a systematic methodology for studying the boundary between semantic equivalence and implementation complexity across compilation, verification, mitigation, and fault-tolerant optimization (Pelofske et al., 8 Mar 2025, Mori et al., 24 Sep 2025).