Papers
Topics
Authors
Recent
Search
2000 character limit reached

Efficient CBP Block for Quantum Decoding

Updated 12 July 2026
  • Efficient CBP Block is a decoding method that replaces global tensor-network contraction with blockwise belief propagation using local MPS messages.
  • The approach leverages Bethe free energy approximations to accurately estimate coset probabilities while preserving degenerate maximum-likelihood decoding benefits.
  • Its performance, enhanced by parallel execution and structured block decomposition, depends on the balance between block size and code distance for optimal efficiency.

“Efficient CBP Block” denotes the use of block belief propagation as an approximate tensor-network contraction engine inside a degenerate quantum maximum-likelihood decoder for the surface code. In this construction, the decoding problem is kept in the exact logical-coset framework of tensor-network decoding, but the expensive global contraction step is replaced by a blockwise message-passing procedure that uses matrix-product-state representations only locally inside each block. The resulting decoder is intended to preserve the accuracy advantages of degenerate maximum-likelihood decoding while gaining efficiency, parallelism, and potential suitability for real-time decoding (Kaufmann et al., 2024).

1. Decoding setting and logical-coset formulation

The method is formulated for stabilizer-code decoding under a Pauli noise channel

N(ρ)=fPnπ(f)fρf,N(\rho)=\sum_{f\in P_n}\pi(f)\, f\rho f^\dagger,

where degeneracy is handled by cosets of the stabilizer group GG: two errors have the same action on the code space iff they differ by an element of GG. After syndrome measurement, the post-measurement state is

ρsβπ(fsLβG)fsLβρLβfs,\rho_s \propto \sum_\beta \pi(f_s L_\beta G)\, f_s L_\beta \rho L_\beta^\dagger f_s^\dagger,

so degenerate quantum maximum-likelihood decoding (DQMLD) chooses the coset fsLβGf_sL_\beta G with maximal probability (Kaufmann et al., 2024).

For the surface code on a d×dd\times d open square lattice, the checks are

Au=ueZe,Bp=epXe,A_u=\prod_{u\in e} Z_e,\qquad B_p=\prod_{e\in p} X_e,

with

n=d2+(d1)2,m=2d(d1),n=d^2+(d-1)^2,\qquad m=2d(d-1),

so

k=nm=1k=n-m=1

logical qubit. The logical operators are {Iˉ,Xˉ,Yˉ,Zˉ}\{\bar I,\bar X,\bar Y,\bar Z\}. Decoding is reduced to evaluating the four logical cosets associated with these operators, which the tensor-network formulation represents by four planar tensor networks GG0. Their contractions are the coset probabilities used for decoding. The central task is therefore not syndrome processing in isolation, but repeated and accurate estimation of GG1 for GG2.

2. Replacement of global tensor-network contraction by blockwise belief propagation

The original tensor-network decoder approximates each GG3 by the boundary-MPS method: the network is contracted column-by-column, the partial contraction is represented as an MPS of bond dimension GG4, each new MPO column doubles the bond dimension, and the MPS is then truncated back to GG5. This is accurate but expensive, with cost

GG6

for fixed GG7, and GG8 must grow with distance near threshold (Kaufmann et al., 2024).

The Efficient CBP Block construction replaces exactly this global bMPS contraction step with blockBP. The tensor network GG9 is partitioned into disjoint blocks, the blocks become nodes of a coarse-grained graph, and ordinary BP is run on that block graph. Because direct block-to-block messages would be large, each message GG0 is represented as an MPS with truncation bond dimension GG1. The conceptual update rule remains the BP tensor-network update

GG2

but in blockBP the local object being contracted is the entire subnetwork of a block together with the incoming MPS messages from neighboring blocks. The outgoing message is then obtained approximately by local bMPS contraction on that block.

This replacement changes the computational organization of the decoder. Instead of one large two-dimensional contraction, the work is decomposed into many smaller contractions that can be executed in parallel. The method therefore remains a belief-propagation decoder in the DQMLD framework, but one whose propagated objects are block-level tensor-network messages rather than scalar or edge-local BP messages. A plausible implication is that its efficiency derives from exploiting the planar tensor-network structure at block granularity rather than treating contraction as a monolithic global operation.

3. Bethe free energy and approximation of coset probabilities

After message convergence, the approximation to the total contraction is obtained from the Bethe free energy of the BP fixed point. The key identity used by the decoder is

GG3

with incoming messages normalized so that

GG4

This product formula is the quantity used to approximate GG5, and hence the logical-coset probabilities (Kaufmann et al., 2024).

The fixed-point interpretation is made explicit through the Bethe free-energy expression

GG6

with edge and vertex marginals

GG7

In the decoder, this decomposition justifies replacing direct global evaluation of GG8 by a product of local block contractions assembled from normalized blockwise messages.

The significance of this step is methodological. The decoder is not merely using BP as a heuristic support estimator; it is using BP fixed points to produce an approximate contraction value for each logical tensor network, thereby retaining the coset-probability semantics required by DQMLD. This is the mechanism by which the method preserves the degeneracy-aware objective of tensor-network decoding while reducing runtime bottlenecks.

4. Algorithmic procedure and convergence control

The input is a syndrome string GG9 of size ρsβπ(fsLβG)fsLβρLβfs,\rho_s \propto \sum_\beta \pi(f_s L_\beta G)\, f_s L_\beta \rho L_\beta^\dagger f_s^\dagger,0. One first chooses any Pauli string ρsβπ(fsLβG)fsLβρLβfs,\rho_s \propto \sum_\beta \pi(f_s L_\beta G)\, f_s L_\beta \rho L_\beta^\dagger f_s^\dagger,1 that produces the syndrome. For each logical class ρsβπ(fsLβG)fsLβρLβfs,\rho_s \propto \sum_\beta \pi(f_s L_\beta G)\, f_s L_\beta \rho L_\beta^\dagger f_s^\dagger,2, one constructs the corresponding tensor network ρsβπ(fsLβG)fsLβρLβfs,\rho_s \propto \sum_\beta \pi(f_s L_\beta G)\, f_s L_\beta \rho L_\beta^\dagger f_s^\dagger,3, partitions it into ρsβπ(fsLβG)fsLβρLβfs,\rho_s \propto \sum_\beta \pi(f_s L_\beta G)\, f_s L_\beta \rho L_\beta^\dagger f_s^\dagger,4 blocks, initializes the MPS messages to represent a uniform distribution, and iterates BP up to a maximum number of rounds (Kaufmann et al., 2024).

At each round, each block computes its outgoing MPS message using the BP update rule together with local bMPS contraction and truncation bond ρsβπ(fsLβG)fsLβρLβfs,\rho_s \propto \sum_\beta \pi(f_s L_\beta G)\, f_s L_\beta \rho L_\beta^\dagger f_s^\dagger,5. The outgoing message is normalized to left-canonical form with unit ρsβπ(fsLβG)fsLβρLβfs,\rho_s \propto \sum_\beta \pi(f_s L_\beta G)\, f_s L_\beta \rho L_\beta^\dagger f_s^\dagger,6 norm. Convergence is monitored by the average message distance

ρsβπ(fsLβG)fsLβρLβfs,\rho_s \propto \sum_\beta \pi(f_s L_\beta G)\, f_s L_\beta \rho L_\beta^\dagger f_s^\dagger,7

The implementation uses damping,

ρsβπ(fsLβG)fsLβρLβfs,\rho_s \propto \sum_\beta \pi(f_s L_\beta G)\, f_s L_\beta \rho L_\beta^\dagger f_s^\dagger,8

with ρsβπ(fsLβG)fsLβρLβfs,\rho_s \propto \sum_\beta \pi(f_s L_\beta G)\, f_s L_\beta \rho L_\beta^\dagger f_s^\dagger,9, and a two-layer checkerboard schedule: black blocks send to white blocks, then white blocks send to black blocks. All blocks in one color class can therefore update in parallel.

Two convergence thresholds, fsLβGf_sL_\beta G0, govern termination and trust in the approximation. If, for a given coset, fsLβGf_sL_\beta G1, the BP loop can stop early. If fsLβGf_sL_\beta G2, the Bethe approximation is considered trustworthy enough and the decoder computes fsLβGf_sL_\beta G3 using the product formula. If no coset reaches fsLβGf_sL_\beta G4, the decoder returns the coset with the smallest final fsLβGf_sL_\beta G5, based on the empirical observation that the correct coset often converges fastest. This convergence-based filtering is an important practical part of the method: it links numerical stability of the BP fixed point to decoding confidence.

5. Complexity, locality, and parallel scalability

The main efficiency claim of the construction is parallel rather than purely serial. If fsLβGf_sL_\beta G6 is the number of blocks, fsLβGf_sL_\beta G7 the number of message-passing rounds, and fsLβGf_sL_\beta G8 the cost of one block update, then a serial implementation costs

fsLβGf_sL_\beta G9

whereas a fully parallel implementation costs roughly

d×dd\times d0

For code distance d×dd\times d1, block size d×dd\times d2, and MPS bond dimension d×dd\times d3,

d×dd\times d4

Hence the serial runtime is

d×dd\times d5

comparable to or worse than the regular tensor-network decoder, but the parallel runtime can be as low as

d×dd\times d6

where the additive d×dd\times d7 accounts for pre- and post-processing (Kaufmann et al., 2024).

The authors also note an additional factor-of-4 reduction if all blocks within a color class are updated in parallel. The architecture is therefore naturally aligned with distributed or highly parallel hardware, because the dominant work is local block contraction and local message normalization rather than a global contraction sweep. The simulations used d×dd\times d8, and the method is presented as potentially suitable for real-time decoding precisely because BP-style algorithms are naturally distributed and require only a modest number of iterations.

This computational profile clarifies the meaning of “efficient” in the phrase Efficient CBP Block. The method does not remove the tensor-network formulation; it reorganizes it into block-local kernels with explicit parallel communication structure. Its efficiency is thus a property of decomposition and scheduling rather than of a weaker decoding objective.

6. Numerical behavior, operating regime, and limitations

The decoder was benchmarked in the code-capacity model with perfect stabilizer measurements under single-qubit depolarizing noise

d×dd\times d9

for Au=ueZe,Bp=epXe,A_u=\prod_{u\in e} Z_e,\qquad B_p=\prod_{e\in p} X_e,0, and lattice sizes Au=ueZe,Bp=epXe,A_u=\prod_{u\in e} Z_e,\qquad B_p=\prod_{e\in p} X_e,1. Comparisons were made against MWPM and the bMPS tensor-network decoder with Au=ueZe,Bp=epXe,A_u=\prod_{u\in e} Z_e,\qquad B_p=\prod_{e\in p} X_e,2 (Kaufmann et al., 2024).

The reported trends are regime dependent. For a large range of lattice sizes and noise levels, blockBP delivers a logical error probability that outperforms MWPM, sometimes by more than an order of magnitude. For small distances it is very competitive with the near-optimal tensor-network decoder; for Au=ueZe,Bp=epXe,A_u=\prod_{u\in e} Z_e,\qquad B_p=\prod_{e\in p} X_e,3, blockBP results are virtually indistinguishable from the bMPS decoder. As distance increases at fixed block size Au=ueZe,Bp=epXe,A_u=\prod_{u\in e} Z_e,\qquad B_p=\prod_{e\in p} X_e,4, however, accuracy degrades because the local block approximation becomes insufficient. Approximate crossover distances are reported as about Au=ueZe,Bp=epXe,A_u=\prod_{u\in e} Z_e,\qquad B_p=\prod_{e\in p} X_e,5 for Au=ueZe,Bp=epXe,A_u=\prod_{u\in e} Z_e,\qquad B_p=\prod_{e\in p} X_e,6, Au=ueZe,Bp=epXe,A_u=\prod_{u\in e} Z_e,\qquad B_p=\prod_{e\in p} X_e,7 for Au=ueZe,Bp=epXe,A_u=\prod_{u\in e} Z_e,\qquad B_p=\prod_{e\in p} X_e,8, and Au=ueZe,Bp=epXe,A_u=\prod_{u\in e} Z_e,\qquad B_p=\prod_{e\in p} X_e,9 for n=d2+(d1)2,m=2d(d1),n=d^2+(d-1)^2,\qquad m=2d(d-1),0. For n=d2+(d1)2,m=2d(d1),n=d^2+(d-1)^2,\qquad m=2d(d-1),1, blockBP often yields logical error rates an order of magnitude lower than MWPM.

These results place the method in a specific operating window. It is strongest when the chosen block size is large enough relative to the code distance that local block contractions still capture the dominant correlations. It becomes less effective when fixed-size blocks are asked to represent increasingly nonlocal structure. This suggests that the method is best understood as a structured approximation hierarchy between fully local BP and global tensor-network contraction: it keeps the degeneracy-aware coset formalism intact, but its numerical quality depends on how well a n=d2+(d1)2,m=2d(d1),n=d^2+(d-1)^2,\qquad m=2d(d-1),2 block decomposition resolves the relevant correlations.

In that sense, Efficient CBP Block is a decoder architecture rather than a single approximation formula. Its defining feature is the substitution of global contraction by blockwise BP with local MPS messages, and its main contribution is to make DQMLD substantially more parallelizable while retaining the core logical-coset semantics of tensor-network decoding.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Efficient CBP Block.