- The paper develops an exhaustive, translation-invariant SWAP-routing search that produces nearest-neighbor syndrome-extraction schedules of 15 or 18 ticks per round, optimal within the tested routing class.
- Strictly local routing reduces circuit-level thresholds by roughly 2.1–2.5 times under SI1000 noise and about three times under standard depolarizing noise, reaching 0.110%–0.134% and 0.168% for representative cases.
- Routed tile codes can offset their lower thresholds through higher encoding efficiency, using 52% of the surface-code qubits at physical error rate 2×10⁻⁴ under SI1000 and 30%–47% fewer qubits under standard noise at 10⁻³.
Tile codes are a family of planar quantum low-density parity-check (qLDPC) codes with weight-6 stabilizers and open boundary conditions, introduced as a hardware-oriented relaxation of bivariate bicycle (BB) codes. The paper under review addresses a central obstacle to their practical deployment: although tile-code stabilizers are local within small windows (3×3 or 4×3) on a two-dimensional lattice, they are not strictly nearest-neighbor, so a naive syndrome-extraction circuit requires mid-range ancilla–data connectivity. The authors develop an exhaustive search algorithm for SWAP-based routing that implements syndrome extraction using only nearest-neighbor gates on a square lattice, then quantify the circuit-level threshold penalty of this constraint and show that, despite it, routed tile codes can outperform the rotated surface code in physical qubits per logical qubit at sufficiently low error rates.
Code structure and noise models
The four tile-code families studied here are specified by polynomial pairs (f,g) over the Laurent polynomial ring R=F2[x±1,y±1], following Haah's module formalism specialized to qubits. Qubits sit on edges of an Lx×Ly open-boundary lattice (L qubits on horizontal edges, R qubits on vertical edges); each bulk stabilizer has weight 6, while boundary stabilizers arise from truncation. A grafting procedure removes redundant boundary-edge qubits, reducing n below 2LxLy. The four families span encoding efficiencies kd2/n from $1.78$ to $4.00$, with representative instances including 4×30, 4×31, 4×32, and 4×33. Notably, the logical dimension 4×34 is fixed by the bulk polynomials alone and does not vary with lattice size — unlike BB codes, where 4×35 depends on code dimensions.
Two gate-dependent Pauli noise models are simulated: a standard depolarizing model following Bravyi et al., where all operations carry error rate 4×36, and the SI1000 model adapted to CNOT-native circuits, with measurement errors at 4×37, reset errors at 4×38, single-qubit and gate-idle errors at 4×39, resonator idle at (f,g)0, and native SWAP assigned (f,g)1 based on tunable-coupler measurements. The SI1000 treatment extends prior work by assigning idle noise to ancillae of reduced-weight boundary checks, which are idle during some CNOT ticks — a situation absent from the periodic-boundary BB circuits.
Routing via exhaustive search
Without locality constraints, the depth-8 schedule of Bravyi et al. transfers directly to tile codes and is provably minimal for weight-6 checks. Under strict nearest-neighbor constraints, the authors parameterize routing strategies as uniform "routing words" (f,g)2: all ancillae execute the same SWAP sequence, with independently chosen starting offsets for (f,g)3- and (f,g)4-ancillae, interleaved with CNOTs applied to adjacent support qubits. Even rounds run the time-reverse of odd rounds, so all SWAPs cancel after two rounds and no frame tracking is required.
Valid schedules are found by a three-phase exhaustive search: (1) coverage verification that each ancilla visits every support qubit; (2) enumeration of CNOT orderings subject to the CSS specialization of condition (b') of Gehér et al., requiring an even number of shared qubits on which the X-CNOT precedes the Z-CNOT for every overlapping check pair; and (3) sub-tick packing of disjoint CNOTs into parallel layers. The resulting optimal schedules require 15 ticks per round for the (f,g)5 and (f,g)6 families and 18 for the other two; because the search is exhaustive over this class, these tick counts are optimal among uniform translation-invariant routing words. The algorithm relies only on translation invariance and applies to other transversal-invariant code families.
Circuit-level thresholds
Thresholds were estimated with Stim simulations decoded by BP+OSD (min-sum scaling factor 0.5, OSD-CS order 7, up to 2000 iterations), fitting logical error rates to the quadratic finite-size scaling ansatz (f,g)7 with (f,g)8, treating (f,g)9 as a family- and model-specific fit parameter. Only R=F2[x±1,y±1]0-stabilizer syndromes were decoded (correcting R=F2[x±1,y±1]1-errors); the authors argue the R=F2[x±1,y±1]2-error thresholds coincide, since both noise channels are invariant under a transversal-Hadamard frame change mapping the circuit onto the boundary-dual code, and they verified numerically that the periodic-boundary versions have isomorphic R=F2[x±1,y±1]3 and R=F2[x±1,y±1]4 Tanner graphs. A decoder comparison appendix shows BP+OSD outperforming BP+LSD by up to R=F2[x±1,y±1]5 and BP+UFD by up to R=F2[x±1,y±1]6 in sub-threshold logical error rate at R=F2[x±1,y±1]7, justifying the choice.
The headline numbers are stark. Under SI1000 without routing, thresholds range from R=F2[x±1,y±1]8 to R=F2[x±1,y±1]9 across the four families; with SWAP routing they fall to Lx×Ly0–Lx×Ly1, a reduction factor of roughly Lx×Ly2–Lx×Ly3. Under the standard model the penalty is larger, about a factor of three (e.g., Lx×Ly4 down to Lx×Ly5 for the Lx×Ly6 family), which the authors attribute to the larger idle-error contribution in that model. These results quantify directly, rather than assuming, the price of enforcing strictly local connectivity on otherwise high-efficiency planar qLDPC codes.
A lower threshold does not by itself determine hardware cost, so the authors compare physical qubits per logical qubit needed to reach target per-logical, per-round error rates Lx×Ly7. Block error probabilities are converted via Lx×Ly8 with Lx×Ly9 and n0 rounds, and distance suppression factors n1 are extracted from fits of n2 against n3. The baseline is the rotated surface code with MWPM decoding, simulated under identical circuit-level noise conventions (n4 qubits per logical qubit). Two larger instances, n5 and n6, supplement the threshold-study instances; deep sub-threshold points use shot budgets up to n7.
Under SI1000, the routed n8 family is more expensive than the surface code at n9 (a factor of 2LxLy0 at 2LxLy1, about 2LxLy2 of its routed threshold), but the tile-to-surface qubit ratio falls monotonically to 2LxLy3 at 2LxLy4, crossing unity near 2LxLy5. Below this crossover the advantage grows continuously as 2LxLy6 decreases. Under the standard depolarizing model the picture is more favorable still: the routed tile code leads at 2LxLy7 for every target considered, requiring 2LxLy8 fewer qubits than the surface code at 2LxLy9, kd2/n0 fewer at kd2/n1, and kd2/n2 fewer at kd2/n3. This demonstrates concretely that the encoding-rate gain of tile codes can compensate for both the reduced threshold and the full SWAP-routing overhead once the physical error rate is sufficiently far below threshold.
Limitations and open questions
The paper is explicit about several caveats. First, optimality of the routing schedules holds only within the class of uniform, translation-invariant routing words; non-uniform protocols or co-optimized qubit placement could yield shorter schedules. Second, all quantitative thresholds depend on BP+OSD with fixed hyperparameters, and alternative decoders or tuning could shift estimates, particularly under biased or operation-dependent noise. Third, the finite-size scaling treats kd2/n4 as an effective fit parameter because code distance does not scale linearly with lattice dimensions, which weakens the universality interpretation of the extracted thresholds. Fourth, the footprint comparison mixes decoders (BP+OSD for tile codes, MWPM for the surface code), the standard choice for each but not a controlled variable. Finally, only kd2/n5-basis decoding was performed; the claimed equality of kd2/n6-error thresholds rests on a symmetry argument verified for the periodic-boundary Tanner graphs, with boundary effects acknowledged to affect sub-threshold rates. Whether non-uniform routing or decoder improvements can raise the routed thresholds toward the unrouted values remains an open question raised by these results.
Conclusion
This work closes a concrete gap between planar qLDPC code constructions and strictly two-dimensional superconducting hardware. By exhaustively searching SWAP-based routing words, the authors produce explicit, provably optimal-within-class nearest-neighbor syndrome-extraction circuits for four tile-code families, and measure the cost of locality: a two-to-three-fold threshold reduction under realistic noise models. The resource-footprint analysis converts this penalty into an actionable operating regime — routed tile codes beat the rotated surface code in qubit efficiency below kd2/n7 under SI1000, and already at kd2/n8 under standard depolarizing noise. The routing-search procedure itself, requiring only translation invariance, constitutes a reusable compilation tool for other translation-invariant code families on planar grids.