Papers
Topics
Authors
Recent
Search
2000 character limit reached

Walking Cat Architecture

Updated 3 July 2026
  • Walking Cat Architecture is a fault-tolerant quantum computing blueprint that employs cat-state ancillas and LDPC codes to facilitate scalable quantum operations on trapped-ion platforms.
  • It leverages innovative constructs such as cat factories, a walking cat protocol, and a 2D QCCD micro-architecture to enable fast, parallelizable logical measurements and T-gate executions.
  • The design demonstrates practical applicability with benchmarks in Hamiltonian simulation and Shor’s algorithm, highlighting trade-offs in code overhead, throughput, and micro-architectural complexity.

The Walking Cat architecture defines a comprehensive and practical fault-tolerant quantum computing blueprint for trapped-ion platforms, integrating cat-state-based logical measurements, LDPC code-based quantum error correction, a 2D QCCD micro-architecture, fast streaming decoders, and hierarchical compilation. The architecture is characterized by its use of “cat factories”—dedicated regions producing large entangled cat states as key ancillas—and its compatibility with experimentally demonstrated hardware elements. The approach emphasizes high-throughput, parallelizable operations, and scalability, with configurations enabling hundreds of logical qubits and execution of millions of TT gates per day using only thousands of physical qubits (Tripier et al., 21 Apr 2026).

1. Cat States and the Walking Cat Protocol

At its core, the Walking Cat architecture employs cat states—ww-qubit GHZ-type entangled states CATw=(0w+1w)/2|\mathrm{CAT}_w\rangle = (|0\rangle^{\otimes w} + |1\rangle^{\otimes w})/\sqrt{2} —as ancillas for fault-tolerantly measuring multi-qubit Pauli operators P=i=1wPi\overline{P} = \prod_{i=1}^w P_i. Logical measurements proceed via (a) entangling each PiP_i-rotated data qubit with a cat ancilla qubit (using CNOT/CZ), (b) measuring all cat qubits in the XX basis, and (c) evaluating the parity of measurement outcomes.

Cat factories produce cat states via a doubling-tree CNOT circuit in log2w+1\lceil \log_2 w \rceil + 1 rounds, followed by mm rounds of ZZZZ-check verification with ancillas. Only cat states passing all checks (rejection probability (2m+1)wp\sim (2m+1)\,w\,p for ww0 the gate error rate) are distributed for consumption.

The "walking cat" protocol shuttles verified cat qubits under target memory/data blocks along QCCD junctions, merges them for logical measurements, and returns them for possible reuse. For inter-block measurements, two cat states are stitched using Bell pairs generated in Bell factories.

2. Quantum LDPC Codes and Error Model

Logical qubits are encoded in quantum LDPC codes with parameters:

  • [[70, 6, 9]]: 6 logical qubits, distance 9, code rate ww1
  • [[102, 22, 9]]: 22 logical qubits, distance 9, code rate ww2

A standard circuit-level noise scaling for code distance ww3 models the logical error per operation as

ww4

with ww5 the two-qubit physical error rate, ww6 the code threshold, and ww7. At ww8, one finds ww9–CATw=(0w+1w)/2|\mathrm{CAT}_w\rangle = (|0\rangle^{\otimes w} + |1\rangle^{\otimes w})/\sqrt{2}0.

3. Fault-Tolerance Protocols

Syndrome Extraction

The three-ring framework arranges data and ancilla qubits in concentric rings on the 2D QCCD array. Ancillas undergo cyclic shifts along three nested “loops” (short, medium, long) to enact stabilizer CNOTs without all-to-all ion connectivity. Each syndrome-extraction cycle (SEC) combines parallel CNOTs and ancilla shuttling, requiring CATw=(0w+1w)/2|\mathrm{CAT}_w\rangle = (|0\rangle^{\otimes w} + |1\rangle^{\otimes w})/\sqrt{2}1 physical-operation cycles (POCs), CATw=(0w+1w)/2|\mathrm{CAT}_w\rangle = (|0\rangle^{\otimes w} + |1\rangle^{\otimes w})/\sqrt{2}2s each, yielding a SEC time of approximately CATw=(0w+1w)/2|\mathrm{CAT}_w\rangle = (|0\rangle^{\otimes w} + |1\rangle^{\otimes w})/\sqrt{2}3ms.

Decoding

A streaming beam search decoder converts syndrome bits CATw=(0w+1w)/2|\mathrm{CAT}_w\rangle = (|0\rangle^{\otimes w} + |1\rangle^{\otimes w})/\sqrt{2}4 to detector outcomes CATw=(0w+1w)/2|\mathrm{CAT}_w\rangle = (|0\rangle^{\otimes w} + |1\rangle^{\otimes w})/\sqrt{2}5, rendering the parity-check matrix into a banded staircase form. Using a sliding-window (window size CATw=(0w+1w)/2|\mathrm{CAT}_w\rangle = (|0\rangle^{\otimes w} + |1\rangle^{\otimes w})/\sqrt{2}6, commit size CATw=(0w+1w)/2|\mathrm{CAT}_w\rangle = (|0\rangle^{\otimes w} + |1\rangle^{\otimes w})/\sqrt{2}7), the inner decoder tackles CATw=(0w+1w)/2|\mathrm{CAT}_w\rangle = (|0\rangle^{\otimes w} + |1\rangle^{\otimes w})/\sqrt{2}8 block subproblems and commits corrections for three SECs at a time. The achieved logical error rate is within CATw=(0w+1w)/2|\mathrm{CAT}_w\rangle = (|0\rangle^{\otimes w} + |1\rangle^{\otimes w})/\sqrt{2}9 of an optimal global decoder, down to P=i=1wPi\overline{P} = \prod_{i=1}^w P_i0. Latency per SEC is approximately P=i=1wPi\overline{P} = \prod_{i=1}^w P_i1ms (Q70), P=i=1wPi\overline{P} = \prod_{i=1}^w P_i2ms (Q102); the 99\%-tile reaction time is P=i=1wPi\overline{P} = \prod_{i=1}^w P_i3ms (Q70), P=i=1wPi\overline{P} = \prod_{i=1}^w P_i4ms (Q102).

Clifford-Frame Tracking

Logical Pauli measurements and Clifford gates are implemented by updating a Clifford-frame matrix P=i=1wPi\overline{P} = \prod_{i=1}^w P_i5 for each block and by measuring appropriate conjugated operators. This Pauli-by-measurement/Clifford-by-frame scheme requires no physical Pauli/Clifford gates.

4. Micro-Architecture and Parallelism

The underlying hardware comprises a 2D QCCD chip, partitioned into:

  • “Gate” rows for one- and two-qubit RF gates
  • “Optical” rows for measurement/reset
  • Ion transport junctions for shuttling

Memory blocks organize data, ancilla, and beacon rows in a folded three-ring structure, with each block containing P=i=1wPi\overline{P} = \prod_{i=1}^w P_i6 qubits plus ten-qubit reservoirs. Cat factories possess four rows of P=i=1wPi\overline{P} = \prod_{i=1}^w P_i7 wells for in-place cat state preparation and verification.

Movement strategies include block-swapping (long-ring), cyclic translation of ancillas (medium-ring), and intra-row permutations (short-ring). RF antennas beneath all gate rows enable billion-scale single/two-qubit gate parallelism. Measurement/reset operations in optical rows are pipelined with shuttling.

Cat-state routing involves shuttling batches of P=i=1wPi\overline{P} = \prod_{i=1}^w P_i8 qubits under target blocks via vertical injection, accordion-like horizontal transport, and merging. The routing cost for P=i=1wPi\overline{P} = \prod_{i=1}^w P_i9, PiP_i0 is PiP_i1 POC.

5. Compilation, Decoding, and Resource Estimates

A hierarchical compiler (Qualtran/QREF style) expands algorithmic specifications to logical instructions—Pauli preparations, logical measurements (LM1, LM2), frame-tracked Cliffords, and magic-state injections. Magic-state injection/verification employs IS+H primitives (logical PiP_i2 via cat-based PiP_i3 measurement).

The decoder operates continuously during SECs, delivering logical outcomes with PiP_i4 ms latency.

Representative resource estimates include:

Instance Memory Code Magic Factory Code Logical Qubits Physical Qubits
Dense PiP_i5 [[102,22,9]] PiP_i6CH₂ [[54,2,10]] 110 2,514

Throughput for PiP_i7-gates (double PiP_i8 via paired PiP_i9 states) is XX0 ms latency, yielding XX1 XX2-gates/s (XX3/day). One million XX4-gates/day is supported with XX5–XX6 physical qubits.

6. Application Benchmarks and Performance

The architecture supports quantum simulations and factoring tasks with practical runtimes:

  • Hamiltonian simulation (100-site Heisenberg model, degree 7): Sixth-order Trotter with 10 steps (XX7 2nd-order steps, XX8 sequential XX9QZ layers plus controlled-phase estimation). Using log2w+1\lceil \log_2 w \rceil + 10 physical qubits:
    • 80 sites: log2w+1\lceil \log_2 w \rceil + 1110 hours
    • 100 sites: log2w+1\lceil \log_2 w \rceil + 1218 hours
    • 120 sites: log2w+1\lceil \log_2 w \rceil + 1353 hours
    • Aggregate time for iterative QPE (50 shots): log2w+1\lceil \log_2 w \rceil + 14 month
  • Shor’s algorithm: Factoring a 30-bit integer (1,071,514,531) requires 34 Q70 and 9 CH₂ codes, runtime log2w+1\lceil \log_2 w \rceil + 15 day. 10/20/40-bit instances require log2w+1\lceil \log_2 w \rceil + 161 hour to 2 days, respectively.

Simulation results at log2w+1\lceil \log_2 w \rceil + 17, leakage log2w+1\lceil \log_2 w \rceil + 18, and loss log2w+1\lceil \log_2 w \rceil + 19 yield logical errors per SEC of mm0 (Q70) and mm1 (Q102).

Magic-state factories:

7. Architectural Variants and Trade-Offs

Three principal architectural realizations are distinguished:

  • Simple: single [[70,6,9]] code for both memory and magic state distillation; minimal code diversity but higher magic factory overhead and reduced throughput
  • Fast: magic-state distillation using CH₂ in [[54,2,10]], memory in [[70,6,9]], and ZZZZ0–frame tracking for optimal ZZZZ1 throughput
  • Dense: memory using [[102,22,9]] for maximal logical qubits per physical qubit

Trade-offs relate primarily to code overhead, throughput, and micro-architectural complexity. The architecture demonstrates that hundreds of logical qubits and millions of logical gates per day are attainable with near-term trapped-ion hardware, enabling classically intractable simulations and practical Shor’s algorithm execution (Tripier et al., 21 Apr 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Walking Cat Architecture.