Catalyst Towers in Surface Codes
- Catalyst Towers are catalytic circuit gadgets in surface-code architectures that employ phase catalysis to implement arbitrary single-qubit Z-rotations with lower T-count and runtime.
- They exist in two variants—in-circuit towers integrated within data circuits and independent towers functioning as dedicated resource factories—each offering unique trade-offs in measurement depth and ancilla overhead.
- Explicit surface-code layouts and resource scaling analyses show that while catalyst towers achieve significant runtime reductions at moderate code distances, increased ancilla requirements can limit benefits at larger distances.
Catalyst towers are a family of catalytic circuit gadgets for implementing arbitrary single-qubit -rotations in a surface-code architecture by exchanging additional ancilla qubits and routing area for reduced -count, reduced -depth, and shorter runtime. In the formulation analyzed in "Space and Time Cost of Continuous Rotations in Surface Codes" (Sun et al., 8 Aug 2025), catalyst towers appear in two principal variants—in-circuit towers and independent towers—and are evaluated not only at the level of asymptotic gate counts but also through explicit surface-code layouts, physical-qubit footprints, spacetime volume, and runtime in practical application settings.
1. Concept and theoretical basis
All catalyst-tower constructions in the cited treatment are based on the elementary phase catalysis gadget of Gidney and Fowler, which, given a single resource state and a catalyst state , applies two rotations to arbitrary data qubits while perfectly recovering the catalyst. By chaining these gadgets in layers, one obtains a tower that amortizes the cost of producing or applying related rotations across a family of angles (Sun et al., 8 Aug 2025).
Two constructions are distinguished. In-circuit towers directly apply to data qubits at amortized cost
rather than as in conventional gate synthesis. Here
is the 0-count of synthesizing one 1 to error 2. The stated trade-off is an increased measurement depth by 3.
Independent towers act as higher-level factories that consume 4 states and produce a spectrum of resource states 5. Each rotation is then applied by RUS teleportation, with success probability 6 per layer and angle doubling on failure. The expected teleportation depth scales only logarithmically in 7, and both 8-count and 9-depth are asymptotically smaller than for serial gate synthesis (Sun et al., 8 Aug 2025).
The basic gate-complexity formulas are presented as:
0
1
and, for independent towers with 2 layers and 3 repetitions,
4
5
These expressions formalize the central design principle: catalyst towers replace repeated high-cost synthesis of related rotations with a structured catalytic process whose marginal cost grows more slowly, while shifting burden into ancilla provisioning, layout complexity, and routing.
2. Variants and operational mechanisms
The distinction between in-circuit and independent towers is architectural rather than merely algebraic. In-circuit towers are embedded directly in the computational circuit and apply the required rotation family in situ. Their main advantage is low amortized 6-count relative to naively synthesizing each rotation independently. Their main limitation, as presented, is that the measurement depth increases linearly with the number of layers, so runtime improvements are not unconditional (Sun et al., 8 Aug 2025).
Independent towers externalize rotation preparation into separate factory-like subsystems. They generate a buffer of resource states that can subsequently be consumed via RUS teleportation. In the reported analysis, this shifts the design toward large-scale parallelism: data qubits can be arranged so that patch edges are exposed to either distillation factories or tower outputs, enabling massive parallel teleportation. The expected teleportation depth remains 7, which is the principal reason the independent construction can outperform standard Clifford+8 synthesis in runtime-sensitive regimes.
The paper also notes that, in both tower types, the corner notation of Gidney is used to represent Logical-AND gates. This places catalyst towers in a broader lineage of fault-tolerant constructions that use specialized logical gadgets to reduce non-Clifford overhead.
A common misconception is that lower 9-count alone determines the superior implementation strategy. The analysis rejects this simplification explicitly: because towers require additional ancilla qubits and routing area, the cost function to optimize should be total runtime or total space rather than isolated 0 metrics. This suggests that catalyst towers are best understood as space–time trade-off constructions, not universally dominant replacements for Clifford+1 synthesis.
3. Surface-code realization
The surface-code layouts are developed under a concrete implementation model: all logical qubits live on 2 patches, with Litinski’s lattice surgery and auto-corrected magic factories. This matters because the resource trade-offs are not abstract; they are tied to patch geometry, routing, and distillation throughput (Sun et al., 8 Aug 2025).
For in-circuit towers in the phase-oracle circuit, data qubits—described as copies of the input register—sit in two horizontal stripes. Each tower occupies 7 rows of patches per layer: four for catalyst and seed, two for Logical-AND ancillae, plus routing space. Magic factories, specifically AutoCCZ states, feed 3 states into the tower via blue ancilla rectangles.
For independent towers, the layout is spatially decoupled from the data region. Independent towers are placed in a separate region around the data. Each 4-layer tower requires approximately 5 logical patches, including routing, to produce the desired buffer of resource states. The data arrangement is chosen so that each patch edge is exposed to either a distillation factory or a tower output, enabling the cited large-scale parallel teleportation.
The implementation picture can be summarized as follows.
| Variant | Functional role | Layout characteristic |
|---|---|---|
| In-circuit tower | Applies 6 directly to data qubits | 7 rows of patches per layer |
| Independent tower | Produces resource states for later teleportation | Separate region around the data |
| Conventional synthesis | Synthesizes rotations serially in Clifford+7 form | No tower-specific ancilla structure |
The significance of these layouts is methodological. They translate catalytic rotation schemes into explicit surface-code resource accounting, making comparison with conventional synthesis depend on physical architecture rather than only logical circuit identities.
4. Resource model and scaling laws
The cost model introduces physical-qubit and spacetime metrics that combine distillation demand with logical-patch occupancy. Let 8 be the total 9 states consumed by a subroutine of depth 0, and let 1 be the number of logical qubits including ancillae. If a 2-factory has footprint 3 qubits and latency 4 cycles, then the number of physical qubits for distillation is
5
the data qubit footprint is
6
and total physical qubits are
7
The spacetime volume is
8
and, more generally,
9
The running time of the subroutine is approximated by
0
where 1 is the measurement depth in code-time steps (Sun et al., 8 Aug 2025).
Within this model, the relevant depth scalings are:
- Clifford+2 synthesis: 3, approximately 32 steps in the reported example.
- In-circuit tower: 4.
- Independent tower: 5, approximately 17 steps in the phase-oracle example.
- Hamming-weight phasing: 6 and 7, with applicability restricted to parallel, identical angles.
This formulation shows why catalyst towers do not have a single monotone advantage. Lower 8 can reduce factory demand, but increased 9 raises the 0 and 1 terms. A plausible implication is that the dominant implementation changes with code distance because distillation savings and ancilla overhead scale differently.
5. Code-distance regimes and comparative performance
The reported resource comparisons identify three code-distance regimes in a phase-oracle circuit with 2 pieces, 3, and 4, using a 5 factory with 6 qubits, 7 cycles, and 8 (Sun et al., 8 Aug 2025).
For this case, the logical-qubit counts are given as
9
for in-circuit towers, and
0
for independent towers.
The regime structure is explicit:
| Code-distance regime | Main conclusion |
|---|---|
| Small 1 | Independent towers achieve the lowest 2 |
| Medium 3 | Independent towers still halve runtime with modest 4-overhead |
| Large 5 | Conventional synthesis yields the smallest 6 and lowest 7 |
The same section states that in-circuit towers always use the least 8 but run slower, with 9, and offer no volume reduction at large 0.
These comparisons directly constrain generalization. Catalyst towers are not presented as asymptotically optimal in all regimes; rather, they are advantageous in specific fault-tolerant operating windows. The paper further emphasizes that larger 1 suppresses the logical error rate,
2
but also raises 3 and 4. As a result, the 5 extra logical qubits per layer required by towers become increasingly expensive at large code distance.
6. Applications, trade-offs, and interpretation
Two practical application examples are used to ground the analysis: a parallel piecewise phase oracle for option pricing and a variational Gaussian-state preparation circuit (Sun et al., 8 Aug 2025).
For the phase-oracle case, the application summary reports:
- Conventional: 6, 7 qubits, 8 qubit·cycles.
- Independent towers: 9, 0 qubits, 1 qubit·cycles.
- Speedup of 2 in runtime and a 30% reduction in volume at 3.
For the Gaussian variational circuit, consisting of 60 copies of 35 4 rotations with 5, the reported figures are:
- Canonical: 6 per repetition, 7.
- Independent control scheme: 8, 9.
- Independent excess scheme: 00, 01.
Assuming 300 logical data qubits plus routing in a 1:3 ratio and 4200 qubits to store RUS states, the paper states:
- For 02, the excess scheme uses the fewest physical qubits and is approximately 03 faster than synthesis.
- For 04, synthesis begins to win on space, though towers still cut runtime by a factor of approximately 05.
- At 06, the excess scheme uses approximately 07 qubits versus 08 for synthesis.
The main trade-off parameters are then listed explicitly: code distance 09, ancilla overhead, 10-factory choice, routing factor, and repetition count 11. The dependence on repetition count is especially important: independent towers amortize startup cost only if the same 12 is used many times, with amplitude estimation given as an example of a regime with 13.
A recurrent misunderstanding is that catalyst towers should replace conventional synthesis whenever continuous rotations are numerous. The reported evidence is narrower. The conclusions are sensitive to application scenarios and parameter choices, and conventional Clifford+14 synthesis may prove more efficient at large code distances. Conversely, catalyst towers may be particularly advantageous for early fault-tolerant quantum applications, where low and medium code distances are assumed and a spacetime trade-off is needed to reduce the runtime of individual circuit runs, especially in scenarios involving high circuit repetition counts (Sun et al., 8 Aug 2025).