Papers
Topics
Authors
Recent
Search
2000 character limit reached

Quantum Circuit for General Unitary: Improved T-count via Block Flattening and Dilation

Published 18 Aug 2026 in quant-ph | (2608.17846v1)

Abstract: Synthesizing arbitrary nn-qubit unitaries using as few non-Clifford gates as possible is a central problem in fault-tolerant quantum compilation. We present a Clifford+TT quantum circuit construction that approximately implements any classically specified unitary to within error εε and achieves a worst-case TT-count with leading exponential scaling of 2<sup>5n/42<sup>{5n/4} whenever log(1/ε)=poly(n)\log(1/ε)=\operatorname{poly}(n). This improves upon the best previous 2<sup>4n/32<sup>{4n/3} scaling. The key innovation lies in treating the target unitary as a single block-encoded object rather than a long product of simpler operations. A technique of block flattening controls the normalization while preserving an efficient implementation of the block encoding; subsequently, quantum singular value transformation maps its common singular value to one, thereby recovering the target unitary.

Authors (3)

Summary

  • The paper improves the worst-case Clifford+$T$ synthesis bound from $\widetilde{O}(2^{4n/3})$ to $O(d^{5/4}L^{5/8}\log d)$, where $d=2^n$ and $L=n+\log(1/\epsilon)$, when $L\le d$.
  • The method combines simultaneous block flattening with Boolean sign diagonals, Halmos dilation, a jointly compiled SELECT operation, and one-point QSVT amplification to reduce block-encoding normalization overhead.
  • The construction uses $O(d\sqrt{L})$ clean ancillas, while leaving a $\widetilde{O}(d^{1/4})$ gap above the known lower bound and relying on excluded randomized classical preprocessing to find flattening signs.

Overview

This paper addresses the problem of synthesizing an arbitrary nn-qubit unitary UU(2n)U \in U(2^n) as a Clifford+TT circuit with minimal TT-count, up to error ϵ\epsilon, in the clean-ancilla model of Tan. The authors—Yuan, Zhang, and Zi—establish a worst-case upper bound with leading scaling O~(25n/4)\widetilde{O}(2^{5n/4}) whenever log(1/ϵ)=poly(n)\log(1/\epsilon) = \mathrm{poly}(n), improving the best previous bound of O~(24n/3)\widetilde{O}(2^{4n/3}) due to Tan. The central conceptual departure is that all prior constructions decomposed the target into a long product of simpler unitaries; here, the target is treated as a single block-encoded object, and quantum singular value transformation (QSVT) is used to strip away the block-encoding normalization.

The main theorem states: for every effectively specified UU (entries approximable to any precision with certified error bounds), if L:=n+log(1/ϵ)d=2nL := n + \log(1/\epsilon) \le d = 2^n, there exists a Clifford+UU(2n)U \in U(2^n)0 circuit implementing UU(2n)U \in U(2^n)1 to error UU(2n)U \in U(2^n)2 with UU(2n)U \in U(2^n)3-count

UU(2n)U \in U(2^n)4

using UU(2n)U \in U(2^n)5 clean ancillas. In the high-precision regime UU(2n)U \in U(2^n)6, Tan's construction gives UU(2n)U \in U(2^n)7 and remains preferable. Combined with the known UU(2n)U \in U(2^n)8 lower bound at constant accuracy from Gosset–Kothari–Wu, this narrows the worst-case gap to a factor of UU(2n)U \in U(2^n)9. Classical preprocessing time for finding the required random signs is excluded from the resource count—an assumption worth noting when interpreting the result.

Simultaneous block flattening

The first ingredient transforms TT0 into a unitary TT1, where TT2 are diagonal Boolean phase oracles (TT3 diagonals) and TT4 is the Walsh transform, such that every TT5-block of TT6 has operator norm at most

TT7

simultaneously over all TT8 blocks. The proof is probabilistic: a random Rademacher diagonal TT9 makes every entry of TT0 small via Hoeffding tails plus a union bound; conditioned on this, each column slab of TT1 is an isometry with rows of squared norm at most TT2, so a second random Rademacher diagonal TT3 applied through the Walsh transform yields a matrix Rademacher series whose variance parameter is bounded by TT4. Tropp's rectangular matrix concentration inequality, followed by a union bound over all block pairs, gives total failure probability below TT5.

The point of flattening is quantitative: without it, direct block encoding would incur normalization TT6 rather than TT7, inflating both the QSVT degree and hence the overall TT8-count. The two oracles cost only TT9 ϵ\epsilon0 gates in total and are undone at the end.

Block encoding via dilation and a joint SELECT

Each normalized block ϵ\epsilon1 is a contraction, completed to a unitary by its Julia–Halmos dilation on a doubled space of dimension ϵ\epsilon2. All ϵ\epsilon3 dilations are packed into a single uniformly controlled unitary,

ϵ\epsilon4

acting on registers ϵ\epsilon5. A circuit ϵ\epsilon6 consisting of Hadamards on ϵ\epsilon7 and ϵ\epsilon8, the SELECT, and a swap of ϵ\epsilon9 and O~(25n/4)\widetilde{O}(2^{5n/4})0 then satisfies

O~(25n/4)\widetilde{O}(2^{5n/4})1

Because O~(25n/4)\widetilde{O}(2^{5n/4})2 is a scalar multiple of a unitary, all its singular values equal O~(25n/4)\widetilde{O}(2^{5n/4})3—a property essential to what follows. Compiling O~(25n/4)\widetilde{O}(2^{5n/4})4 via Tan's UCU lemma costs O~(25n/4)\widetilde{O}(2^{5n/4})5 O~(25n/4)\widetilde{O}(2^{5n/4})6 gates and ancillas per query, where O~(25n/4)\widetilde{O}(2^{5n/4})7.

Robust one-point amplification and the exponent 5/4

Since the encoded object has a single common singular value O~(25n/4)\widetilde{O}(2^{5n/4})8, a degree-O~(25n/4)\widetilde{O}(2^{5n/4})9 odd polynomial can be constructed (via a scaled Chebyshev polynomial log(1/ϵ)=poly(n)\log(1/\epsilon) = \mathrm{poly}(n)0 plus the QSP completion lemma of Gilyén et al.) satisfying log(1/ϵ)=poly(n)\log(1/\epsilon) = \mathrm{poly}(n)1 exactly while remaining bounded on log(1/ϵ)=poly(n)\log(1/\epsilon) = \mathrm{poly}(n)2. QSVT then maps log(1/ϵ)=poly(n)\log(1/\epsilon) = \mathrm{poly}(n)3 back to log(1/ϵ)=poly(n)\log(1/\epsilon) = \mathrm{poly}(n)4 using log(1/ϵ)=poly(n)\log(1/\epsilon) = \mathrm{poly}(n)5 alternating queries to log(1/ϵ)=poly(n)\log(1/\epsilon) = \mathrm{poly}(n)6 and log(1/ϵ)=poly(n)\log(1/\epsilon) = \mathrm{poly}(n)7 interleaved with signal rotations.

A notable technical contribution is the robustness analysis: if the compiled query satisfies log(1/ϵ)=poly(n)\log(1/\epsilon) = \mathrm{poly}(n)8 with log(1/ϵ)=poly(n)\log(1/\epsilon) = \mathrm{poly}(n)9, the amplified clean block deviates from O~(24n/3)\widetilde{O}(2^{4n/3})0 by O~(24n/3)\widetilde{O}(2^{4n/3})1, and ancilla leakage is controlled via a polar-decomposition argument combined with Markov's inequality on O~(24n/3)\widetilde{O}(2^{4n/3})2, yielding full clean-isometry error O~(24n/3)\widetilde{O}(2^{4n/3})3. Choosing O~(24n/3)\widetilde{O}(2^{4n/3})4 balances these errors.

Balancing the two competing terms in the total cost O~(24n/3)\widetilde{O}(2^{4n/3})5—the amplification overhead decreasing in O~(24n/3)\widetilde{O}(2^{4n/3})6, the SELECT compilation cost increasing in O~(24n/3)\widetilde{O}(2^{4n/3})7—at O~(24n/3)\widetilde{O}(2^{4n/3})8 yields the final O~(24n/3)\widetilde{O}(2^{4n/3})9-count UU0 with UU1 ancillas. The improvement over Tan's exponent UU2 comes precisely from controlling the block-encoding normalization and the selected-dilation dimension simultaneously.

Extension to multiplexed unitaries

An appendix extends the framework to multiplexed unitaries UU3 over UU4 branches of dimension UU5. A strengthened "common-sign" flattening lemma shows one pair of sign diagonals works for all branches simultaneously, with failure probability bounded by union over UU6 entries and UU7 triples. The resulting optimized UU8-count is piecewise:

Regime UU9-count
L:=n+log(1/ϵ)d=2nL := n + \log(1/\epsilon) \le d = 2^n0 L:=n+log(1/ϵ)d=2nL := n + \log(1/\epsilon) \le d = 2^n1
L:=n+log(1/ϵ)d=2nL := n + \log(1/\epsilon) \le d = 2^n2 L:=n+log(1/ϵ)d=2nL := n + \log(1/\epsilon) \le d = 2^n3

For L:=n+log(1/ϵ)d=2nL := n + \log(1/\epsilon) \le d = 2^n4 this improves the L:=n+log(1/ϵ)d=2nL := n + \log(1/\epsilon) \le d = 2^n5 target-side term of the direct compiler, with the two bounds agreeing at L:=n+log(1/ϵ)d=2nL := n + \log(1/\epsilon) \le d = 2^n6.

Limitations and open questions

Several caveats qualify the results. First, the construction assumes the unitary is effectively specified with certified error bounds, and the randomized classical preprocessing needed to find the flattening signs is excluded from the resource count; the paper does not bound this preprocessing cost explicitly beyond asserting constant success probability per random pair. Second, the improvement holds only in the regime L:=n+log(1/ϵ)d=2nL := n + \log(1/\epsilon) \le d = 2^n7; for exponentially small error (L:=n+log(1/ϵ)d=2nL := n + \log(1/\epsilon) \le d = 2^n8), Tan's L:=n+log(1/ϵ)d=2nL := n + \log(1/\epsilon) \le d = 2^n9 construction remains the state of the art, so the new technique does not dominate uniformly across all accuracy scalings. Third, the ancilla count UU(2n)U \in U(2^n)00 is substantial, and no tradeoff analysis against reduced workspace is provided. Finally, the gap between the UU(2n)U \in U(2^n)01 upper bound and the UU(2n)U \in U(2^n)02 lower bound leaves open whether arbitrary unitaries admit UU(2n)U \in U(2^n)03 UU(2n)U \in U(2^n)04-count synthesis, or whether a stronger lower bound exists—the question the authors identify as central.

Conclusion

The paper improves the worst-case Clifford+UU(2n)U \in U(2^n)05 UU(2n)U \in U(2^n)06-count for general unitary synthesis from UU(2n)U \in U(2^n)07 to UU(2n)U \in U(2^n)08 in the broad-accuracy regime UU(2n)U \in U(2^n)09, by combining simultaneous block flattening via Boolean sign diagonals and Walsh transforms, Halmos dilation of normalized blocks assembled into a jointly compiled SELECT, and exact one-point QSVT amplification with a robustness analysis against compiled-query errors. The same machinery yields improved bounds for multiplexed unitaries. The remaining multiplicative gap of UU(2n)U \in U(2^n)10 above the known lower bound defines the outstanding problem in this line of work.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.