Quantum Circuit for General Unitary: Improved T-count via Block Flattening and Dilation
Published 18 Aug 2026 in quant-ph | (2608.17846v1)
Abstract: Synthesizing arbitrary n-qubit unitaries using as few non-Clifford gates as possible is a central problem in fault-tolerant quantum compilation. We present a Clifford+T quantum circuit construction that approximately implements any classically specified unitary to within error ε and achieves a worst-case T-count with leading exponential scaling of 2<sup>5n/4 whenever log(1/ε)=poly(n). This improves upon the best previous 2<sup>4n/3 scaling. The key innovation lies in treating the target unitary as a single block-encoded object rather than a long product of simpler operations. A technique of block flattening controls the normalization while preserving an efficient implementation of the block encoding; subsequently, quantum singular value transformation maps its common singular value to one, thereby recovering the target unitary.
The paper improves the worst-case Clifford+$T$ synthesis bound from $\widetilde{O}(2^{4n/3})$ to $O(d^{5/4}L^{5/8}\log d)$, where $d=2^n$ and $L=n+\log(1/\epsilon)$, when $L\le d$.
The method combines simultaneous block flattening with Boolean sign diagonals, Halmos dilation, a jointly compiled SELECT operation, and one-point QSVT amplification to reduce block-encoding normalization overhead.
The construction uses $O(d\sqrt{L})$ clean ancillas, while leaving a $\widetilde{O}(d^{1/4})$ gap above the known lower bound and relying on excluded randomized classical preprocessing to find flattening signs.
Overview
This paper addresses the problem of synthesizing an arbitrary n-qubit unitary U∈U(2n) as a Clifford+T circuit with minimal T-count, up to error ϵ, in the clean-ancilla model of Tan. The authors—Yuan, Zhang, and Zi—establish a worst-case upper bound with leading scaling O(25n/4) whenever log(1/ϵ)=poly(n), improving the best previous bound of O(24n/3) due to Tan. The central conceptual departure is that all prior constructions decomposed the target into a long product of simpler unitaries; here, the target is treated as a single block-encoded object, and quantum singular value transformation (QSVT) is used to strip away the block-encoding normalization.
The main theorem states: for every effectively specified U (entries approximable to any precision with certified error bounds), if L:=n+log(1/ϵ)≤d=2n, there exists a Clifford+U∈U(2n)0 circuit implementing U∈U(2n)1 to error U∈U(2n)2 with U∈U(2n)3-count
U∈U(2n)4
using U∈U(2n)5 clean ancillas. In the high-precision regime U∈U(2n)6, Tan's construction gives U∈U(2n)7 and remains preferable. Combined with the known U∈U(2n)8 lower bound at constant accuracy from Gosset–Kothari–Wu, this narrows the worst-case gap to a factor of U∈U(2n)9. Classical preprocessing time for finding the required random signs is excluded from the resource count—an assumption worth noting when interpreting the result.
Simultaneous block flattening
The first ingredient transforms T0 into a unitary T1, where T2 are diagonal Boolean phase oracles (T3 diagonals) and T4 is the Walsh transform, such that every T5-block of T6 has operator norm at most
T7
simultaneously over all T8 blocks. The proof is probabilistic: a random Rademacher diagonal T9 makes every entry of T0 small via Hoeffding tails plus a union bound; conditioned on this, each column slab of T1 is an isometry with rows of squared norm at most T2, so a second random Rademacher diagonal T3 applied through the Walsh transform yields a matrix Rademacher series whose variance parameter is bounded by T4. Tropp's rectangular matrix concentration inequality, followed by a union bound over all block pairs, gives total failure probability below T5.
The point of flattening is quantitative: without it, direct block encoding would incur normalization T6 rather than T7, inflating both the QSVT degree and hence the overall T8-count. The two oracles cost only T9 ϵ0 gates in total and are undone at the end.
Block encoding via dilation and a joint SELECT
Each normalized block ϵ1 is a contraction, completed to a unitary by its Julia–Halmos dilation on a doubled space of dimension ϵ2. All ϵ3 dilations are packed into a single uniformly controlled unitary,
ϵ4
acting on registers ϵ5. A circuit ϵ6 consisting of Hadamards on ϵ7 and ϵ8, the SELECT, and a swap of ϵ9 and O(25n/4)0 then satisfies
O(25n/4)1
Because O(25n/4)2 is a scalar multiple of a unitary, all its singular values equal O(25n/4)3—a property essential to what follows. Compiling O(25n/4)4 via Tan's UCU lemma costs O(25n/4)5 O(25n/4)6 gates and ancillas per query, where O(25n/4)7.
Robust one-point amplification and the exponent 5/4
Since the encoded object has a single common singular value O(25n/4)8, a degree-O(25n/4)9 odd polynomial can be constructed (via a scaled Chebyshev polynomial log(1/ϵ)=poly(n)0 plus the QSP completion lemma of Gilyén et al.) satisfying log(1/ϵ)=poly(n)1 exactly while remaining bounded on log(1/ϵ)=poly(n)2. QSVT then maps log(1/ϵ)=poly(n)3 back to log(1/ϵ)=poly(n)4 using log(1/ϵ)=poly(n)5 alternating queries to log(1/ϵ)=poly(n)6 and log(1/ϵ)=poly(n)7 interleaved with signal rotations.
A notable technical contribution is the robustness analysis: if the compiled query satisfies log(1/ϵ)=poly(n)8 with log(1/ϵ)=poly(n)9, the amplified clean block deviates from O(24n/3)0 by O(24n/3)1, and ancilla leakage is controlled via a polar-decomposition argument combined with Markov's inequality on O(24n/3)2, yielding full clean-isometry error O(24n/3)3. Choosing O(24n/3)4 balances these errors.
Balancing the two competing terms in the total cost O(24n/3)5—the amplification overhead decreasing in O(24n/3)6, the SELECT compilation cost increasing in O(24n/3)7—at O(24n/3)8 yields the final O(24n/3)9-count U0 with U1 ancillas. The improvement over Tan's exponent U2 comes precisely from controlling the block-encoding normalization and the selected-dilation dimension simultaneously.
Extension to multiplexed unitaries
An appendix extends the framework to multiplexed unitaries U3 over U4 branches of dimension U5. A strengthened "common-sign" flattening lemma shows one pair of sign diagonals works for all branches simultaneously, with failure probability bounded by union over U6 entries and U7 triples. The resulting optimized U8-count is piecewise:
Regime
U9-count
L:=n+log(1/ϵ)≤d=2n0
L:=n+log(1/ϵ)≤d=2n1
L:=n+log(1/ϵ)≤d=2n2
L:=n+log(1/ϵ)≤d=2n3
For L:=n+log(1/ϵ)≤d=2n4 this improves the L:=n+log(1/ϵ)≤d=2n5 target-side term of the direct compiler, with the two bounds agreeing at L:=n+log(1/ϵ)≤d=2n6.
Limitations and open questions
Several caveats qualify the results. First, the construction assumes the unitary is effectively specified with certified error bounds, and the randomized classical preprocessing needed to find the flattening signs is excluded from the resource count; the paper does not bound this preprocessing cost explicitly beyond asserting constant success probability per random pair. Second, the improvement holds only in the regime L:=n+log(1/ϵ)≤d=2n7; for exponentially small error (L:=n+log(1/ϵ)≤d=2n8), Tan's L:=n+log(1/ϵ)≤d=2n9 construction remains the state of the art, so the new technique does not dominate uniformly across all accuracy scalings. Third, the ancilla count U∈U(2n)00 is substantial, and no tradeoff analysis against reduced workspace is provided. Finally, the gap between the U∈U(2n)01 upper bound and the U∈U(2n)02 lower bound leaves open whether arbitrary unitaries admit U∈U(2n)03 U∈U(2n)04-count synthesis, or whether a stronger lower bound exists—the question the authors identify as central.
Conclusion
The paper improves the worst-case Clifford+U∈U(2n)05 U∈U(2n)06-count for general unitary synthesis from U∈U(2n)07 to U∈U(2n)08 in the broad-accuracy regime U∈U(2n)09, by combining simultaneous block flattening via Boolean sign diagonals and Walsh transforms, Halmos dilation of normalized blocks assembled into a jointly compiled SELECT, and exact one-point QSVT amplification with a robustness analysis against compiled-query errors. The same machinery yields improved bounds for multiplexed unitaries. The remaining multiplicative gap of U∈U(2n)10 above the known lower bound defines the outstanding problem in this line of work.