---
title: Tensor-Network Readout Error Mitigation
url: https://www.emergentmind.com/topics/tensor-network-readout-error-mitigation
type: topic
---

# Tensor-Network Readout Error Mitigation

Tensor-network readout error mitigation denotes a class of scalable methods that replace an exponentially large multiqubit readout model by a structured tensor network, typically a matrix product operator (MPO) in one dimension and a projected entangled-pair operator (PEPO) in two dimensions. In the narrow readout setting, the detector is modeled as a classical conditional probability channel \(\mathbf{\Lambda}\) such that \(\mathbf{P}_{\mathrm{noisy}}=\mathbf{\Lambda}\mathbf{P}_{\mathrm{ideal}}\), and mitigation amounts to characterizing \(\mathbf{\Lambda}\) and using it for inverse or likelihood-based correction. In a broader measurement-stage sense, related tensor-network schemes mitigate errors entirely through the final measurement layer, either by post-processing informationally complete measurement data with an inverse noise tensor network or by replacing the target observable with a surrogate observable computed in the Heisenberg picture [2606.25974] [2307.11740] [2602.05916].

## 1. Scope and formal setting

In the strict readout formulation, the matrix element \(\mathbf{\Lambda}_{\mathbf{x},\mathbf{y}}\) is the probability of observing output bit string \(\mathbf{x}=\{x_1,\dots,x_N\}\) when the actual ideal outcome is \(\mathbf{y}=\{y_1,\dots,y_N\}\). The noisy and ideal distributions then satisfy
\[
\mathbf{P}_{\mathrm{noisy}}=\mathbf{\Lambda}\,\mathbf{P}_{\mathrm{ideal}}.
\]
The standard scalable approximation is the uncorrelated model
\[
\mathbf{\Lambda}=\bigotimes_{k=1}^N \mathbf{\Lambda}^{[k]},
\]
but this neglects spatially correlated readout errors caused by crosstalk, shared control and measurement circuitry, and geometry-dependent detector effects. The central tensor-network premise is that such correlations are often local or short-ranged, so the full \(2^N\times 2^N\) readout matrix can be replaced by a low-bond-dimension tensor network without storing or inverting it densely [2606.25974].

A broader literature uses the same structural idea for output-noise mitigation rather than classical readout noise alone. Tensor-network error mitigation (TEM) constructs a tensor-network representation of the inverse of the global noise channel affecting the output state and applies that inverse to informationally complete measurement data in post-processing. In that framework, preparation and measurement errors can be relocated to the dynamics part, and the formalism assumes perfect initialization and measurements unless those errors are absorbed into an effective channel [2307.11740]. A still different measurement-stage approach computes in advance a surrogate observable \(\hat Y\) satisfying \(\mathcal{E}^\dagger(\hat Y)=X\), then measures \(\hat Y\) on the noisy state; that work explicitly assumes perfect initialization and perfect measurement, so it is adjacent to readout mitigation rather than a direct detector-noise inversion method [2602.05916].

This distinction matters because “tensor-network readout error mitigation” can refer either to correction of a terminal detector channel or, more broadly, to any mitigation mechanism implemented through the final measurement operator or final measurement data. The literature contains both usages.

## 2. Tensor-network representations

The most explicit tensor-network readout model treats \(\mathbf{\Lambda}\) itself as a double-layer MPO. The readout channel is written as
\[
\mathbf{\Lambda}_{\mathbf{x},\mathbf{y}}
=
|\mathbf{M}_{\mathbf{x},\mathbf{y}}|^2
=
\left|
\left(
M^{[1]}_{x_1,y_1} M^{[2]}_{x_2,y_2}\cdots M^{[N]}_{x_N,y_N}
\right)
\right|^2.
\]
Each local tensor carries physical indices \(x_k,y_k\in\{0,1\}\) and virtual indices of bond dimension \(\chi\). The squared-modulus construction guarantees \(\mathbf{\Lambda}_{\mathbf{x},\mathbf{y}}\ge 0\), while stochasticity is enforced through
\[
\sum_{\mathbf{x}} \mathbf{\Lambda}_{\mathbf{x},\mathbf{y}} = 1,\qquad \forall \mathbf{y}.
\]
The uncorrelated model is recovered at bond dimension \(\chi=1\), so tensor-network REM interpolates continuously between product readout models and correlated ones [2606.25974].

For two-dimensional layouts, the same construction becomes a PEPO,
\[
\mathbf{\Lambda}_{\mathbf{x},\mathbf{y}}
=
\left|
\mathrm{Tr}_{\{l_s,r_s,u_s,d_s\}}
\prod_s M^{[s]}_{x_s,y_s,l_s,r_s,u_s,d_s}
\right|^2,
\]
which is used in decoder-aware readout modeling on syndrome lattices. This extends the same locality assumption from chains to planar geometries [2606.25974].

Related tensor-network mitigation frameworks use different operator objects. In TEM, the relevant tensor network is an MPO in the Pauli transfer matrix (PTM) representation for the inverse global noise channel or its adjoint action on observables. In layerwise noise characterization, the unknown noise channel \(\mathcal N\) is represented as a locally purified density operator (LPDO), which guarantees complete positivity by construction and can be reshuffled into an MPO superoperator form [2307.11740] [2402.08556].

| Target object | Tensor-network form | Primary use |
|---|---|---|
| Correlated readout channel \(\mathbf{\Lambda}\) | Double-layer MPO or PEPO | Readout characterization and REM |
| Inverse global noise channel | MPO in PTM form | TEM post-processing |
| Surrogate measurement operator | MPO-generated observable | Heisenberg-picture pre-processing |

These constructions share a common structural assumption: the relevant channel, inverse channel, or observable remains compressible at moderate bond dimension because correlations are local enough, weak enough, or sufficiently structured. This suggests that tensor-network readout mitigation is best understood as one member of a larger family of locality-exploiting inverse-channel methods.

## 3. Characterization and calibration

In explicit readout-channel modeling, calibration data consist of pairs \(\{(\mathbf{x}^{(m)},\mathbf{y}^{(m)})\}_{m=1}^M\), where \(\mathbf{y}^{(m)}\) is a prepared computational-basis product state and \(\mathbf{x}^{(m)}\) is the observed noisy output. The MPO is trained by minimizing a penalized negative log-likelihood,
\[
\mathbb{L} =
-\frac{1}{M}\sum_{m=1}^M
\log
\frac{\mathbf{\Lambda}_{\mathbf{x}^{(m)},\mathbf{y}^{(m)}}}
{\sum_{\mathbf{x}}\mathbf{\Lambda}_{\mathbf{x},\mathbf{y}^{(m)}}}
+
\Delta\sum_{\mathbf{y}}
\left(
\sum_{\mathbf{x}}\mathbf{\Lambda}_{\mathbf{x},\mathbf{y}}-1
\right)^2,
\]
with \(\Delta=1\) in the reported experiments. The loss and gradients are evaluated by tensor-network contraction, and the paper uses Adam with \(\beta_1=0.9\), \(\beta_2=0.999\), and \(\eta=0.002\) [2606.25974].

On hardware, the framework was tested on the Baihua superconducting chip using five disjoint 6-qubit chains. Full \(N=6\) detector characterization served as reference. The MPO model converged to approximately
\[
d(\mathbf{\Lambda}_{\rm MPO},\mathbf{\Lambda}_{\rm Full})\approx 0.043,
\]
while the product model reached
\[
d(\mathbf{\Lambda}_{\rm Product},\mathbf{\Lambda}_{\rm Full})\approx 0.074.
\]
The total error strength was
\[
d(\mathbf{I},\mathbf{\Lambda}_{\rm Full})\approx 0.26.
\]
In synthetic correlated readout models up to \(N=20\), the reported scaling was
\[
d \sim M^{-0.35},
\]
and the number of samples needed to reach a fixed threshold \(d(\mathbf{\Lambda}_{\rm MPO},\mathbf{\Lambda}_{\rm Full})<\varepsilon\) scaled nearly linearly with \(N\) [2606.25974].

A broader tensor-network characterization program learns correlated quantum noise channels rather than classical readout matrices. There the experimentally implemented process is \(\mathcal E=\mathcal N\circ\mathcal U\), and \(\mathcal N\) is reconstructed as an LPDO from randomized informationally complete input states and local Pauli-basis measurements. The training objective is a Monte Carlo KL divergence plus a trace-preservation penalty,
\[
L(\boldsymbol\theta;S)=D_{\mathrm{KL}}(\boldsymbol\theta;S)+\eta\,\delta_{\mathrm{TP}}(\boldsymbol\theta),
\]
with \(\eta=1.2\) in the numerics. That work reports that linearly many random settings in \(n\) suffice in its regime, that \(10^3\) settings with \(10^3\) shots each characterize a 20-qubit brickwork depolarizing channel with error about \(10^{-4}\) in normalized Frobenius distance, and that the largest \(n=20\) training run with \(\chi_b=\chi_\kappa=4\) and \(N=10^6\) samples takes about one hour on a laptop [2402.08556].

The readout connection appears again in the treatment of SPAM. The main LPDO reconstruction assumes known preparation and measurement operators, but it can incorporate detector calibration by replacing ideal POVM effects \(\Pi\) with reconstructed noisy effects \(\tilde\Pi\). This makes tensor-network channel learning compatible with measurement-noise-aware calibration, although the primary object remains a quantum noise channel rather than a classical confusion matrix [2402.08556].

## 4. Mitigation workflows and application domains

Once a correlated readout MPO has been learned, the most direct mitigation route is inverse-based correction of observables. For a diagonal observable \(O\),
\[
\langle O \rangle_{\mathrm{ideal}}
=
\frac{1}{M}\sum_{m=1}^M
\left[
\sum_{\mathbf{y}} O(\mathbf{y})
\left(\mathbf{\Lambda}^{-1}\right)_{\mathbf{y},\mathbf{x}^{(m)}_{\rm noisy}}
\right].
\]
Rather than inverting a dense \(2^N\times 2^N\) matrix, the inverse is approximated by a variational MPO \(\mathbf{\Omega}\approx \mathbf{\Lambda}^{-1}\) obtained from
\[
\min_{\mathbf{\Omega}}\|\mathbf{\Lambda}\mathbf{\Omega}-\mathbf{I}\|_F^2.
\]
This is the basic readout-mitigation mechanism for nonlocal observables [2606.25974].

The same learned readout model can be used without explicit inversion. For global sampling tasks, an MPS model \(\mathbf{P}\) of the ideal distribution is trained under the noisy forward model by minimizing
\[
\mathbb{L}
=
-\frac{1}{M}\sum_{m=1}^M
\log
\frac{(\mathbf{\Lambda}\mathbf{P})(\mathbf{x}^{(m)}_{\rm noisy})}
{\sum_{\mathbf{y}}\mathbf{P}(\mathbf{y})}.
\]
After training, corrected bitstrings are generated by sequential sampling from the MPS. For cross-entropy benchmarking, the paper also gives a direct inverse-MPO estimator,
\[
F_{\rm XEB,REM}
=
\frac{2^N}{M}\sum_{m=1}^M
\sum_{\mathbf{y}}
\mathbf{P}_{\rm ideal}(\mathbf{y})
\left(\mathbf{\Lambda}^{-1}_{\rm MPO}\right)_{\mathbf{y},\mathbf{x}^{(m)}_{\rm noisy}}
-1.
\]
These two variants reflect a general split between direct inverse-channel estimators and latent clean-distribution inference under a noisy likelihood [2606.25974].

Random-measurement protocols can also absorb the tensor-network readout model. In classical shadows, the inverse readout channel is applied before the usual shadow inversion map \(\mathcal M^{-1}\). In learning-based tomography, the likelihood is written directly in terms of the noisy forward model,
\[
\mathbb{L}
=
-\frac{1}{M}\sum_{m=1}^M
\log
\left[
\sum_{\mathbf{y}}
\mathbf{\Lambda}_{\mathbf{x}^{(m)}_{\rm noisy},\mathbf{y}}
\frac{
\mathrm{Tr}\!\left(
\rho\,U^{(m)\dagger}\ket{\mathbf{y}}\!\bra{\mathbf{y}}U^{(m)}
\right)
}{
\mathrm{Tr}(\rho)
}
\right].
\]
This avoids a separate pre-correction step and turns readout mitigation into part of the state-learning objective [2606.25974].

The reported application range is broad. On a 7-qubit cluster-state experiment on Baihua, MPO REM combined with ZNE improved string-order estimates, especially for larger nonlocal strings. On a 7-qubit GHZ experiment, corrected samples moved the probabilities of \(0000000\) and \(1111111\) much closer to the ideal \(0.5\). In a 20-qubit classical-shadows benchmark for a finite-temperature Gibbs state of the 1D XY model, REM significantly improved estimates of \(\langle Z_{N/2}\rangle\). In 2D, the readout PEPO was fused directly with tensor-network decoder tensors, so decoding marginalized jointly over data errors and readout errors rather than correcting the syndrome first [2606.25974].

A recurrent trade-off is variance amplification. The inverse readout MPO is a quasi-probability object, so negative or large entries in \(\mathbf{\Lambda}^{-1}\) can increase estimator fluctuations. The 7-qubit string-order experiment explicitly reports larger variance after mitigation for this reason [2606.25974].

## 5. Relation to post-processing TEM and observable pre-processing

Tensor-network readout mitigation in the strict detector-channel sense sits next to two measurement-stage tensor-network paradigms. The first is TEM, a post-processing protocol that constructs a tensor-network representation of the inverse global noise channel affecting the output state and applies that inverse to informationally complete measurement outcomes. The central estimator is
\[
\bar O_{\rm n.m.}
=
\frac{1}{S}\sum_{\mathbf{k}}
\mathrm{tr}[{\cal M}(D_{\mathbf{k}})\, O]
=
\frac{1}{S}\sum_{\mathbf{k}}
\mathrm{tr}[D_{\mathbf{k}}\,{\cal M}^\dagger(O)],
\]
where \(D_{\mathbf{k}}\) are dual operators obtained from local informationally complete measurements. TEM therefore performs correction entirely in classical post-processing, but it is not a narrow readout-error method based on a classical assignment matrix [2307.11740].

TEM’s statistical advantage is one of the main reasons it remains relevant in discussions of tensor-network readout mitigation. For Pauli noise, the paper reports that the measurement overhead is quadratically smaller than in probabilistic error cancellation, with
\[
\gamma_{\rm TEM}\approx \sqrt{\gamma_{\rm PEC}}.
\]
In simulations up to 100 qubits and depth 100, both PEC and ZNE fail to produce accurate results by using \(\sim 10^5\) shots, while TEM succeeds. A later analysis formalizes TEM as Heisenberg-picture inversion of learned noise, proves that it asymptotically saturates the universal lower cost bound for unbiased estimation of Pauli observables under weak Pauli noise, and argues that sufficient bond dimension makes TEM behave similarly to an error correcting code of distance 3 [2307.11740] [2403.13542].

The second paradigm is observable pre-processing. There one computes a surrogate observable \(\hat Y\) such that
\[
\mathcal{E}^\dagger(\hat Y)=X,
\]
so that
\[
\mathrm{Tr}[X\rho]=\mathrm{Tr}[\hat Y\,\mathcal{E}(\rho)].
\]
In exact Pauli coordinates this is the linear solve \(\boldsymbol y=A^{-T}\boldsymbol x\) with \(y_0=-\boldsymbol c^T\boldsymbol y\), but the practical construction uses an MPO for \([\mathcal E^\dagger]^{-1}\) and compresses it after each layer. The most important simplification is the Dominant Component Approximation (DCA): for a Pauli target \(X=P_i\),
\[
\hat Y\approx [\mathcal{M}_L^\dagger]^{-1}_{i,i} P_i,
\]
so the mitigation reduces to measuring the same Pauli string and applying a scalar correction factor. The paper states that exact \(\hat Y\) saturates the quantum Cramér–Rao bound for unbiased estimators, that the method approaches the theoretical lower bound in measurement overhead, and that its classical cost can be roughly \(10^6\) times smaller than TEM in practical scenarios because only a diagonal element is needed instead of many dual-operator contractions [2602.05916].

These adjacent frameworks broaden the meaning of “readout-stage” tensor-network mitigation. They do not, however, remove the conceptual boundary between detector-noise correction and final-measurement engineering. TEM can absorb measurement noise into an effective global channel, and observable pre-processing acts entirely through the final measured operator, but neither is presented as a dedicated classical confusion-matrix remedy.

## 6. Assumptions, robustness, and open boundaries

All tensor-network readout mitigation schemes depend on structured noise. The explicit MPO/PEPO readout framework assumes local or short-range crosstalk, availability of calibration data, and a bond dimension \(\chi\) large enough to capture the relevant correlations. The TEM and observable-preprocessing variants add further assumptions: invertible noise maps, low enough noise that compression remains accurate, and geometries compatible with MPO compression, especially 1D nearest-neighbor circuits [2606.25974] [2602.05916].

Calibration quality is a separate constraint. A robustness analysis of inverse-channel mitigation under imperfect noise characterization predicts a threshold structure for PEC and TNEM/TEM in random local circuits: with space-time random disorder, there is a threshold in \(D\ge 2\), while in \(D=1\) mitigation fails at \(\mathcal O(1)\) time for any imperfection in the characterization of disorder. For readout-only mitigation, the paper explicitly cautions that these depth-dependent threshold statements do not transfer literally, because a single terminal readout layer does not accumulate disorder in time the same way. The safe conclusion is that inversion-based tensor-network mitigation remains practical only when the relevant local noise or readout model is sufficiently well characterized [2302.04278].

Device-dependent correlation structure also matters. A study of multiqubit readout correlations on IBM hardware found that correlations on IBMQ Melbourne are small compared to single-qubit readout errors but long-ranged and do not decay with inter-qubit distance, whereas on IBMQ Manhattan they are short-ranged and confined to neighboring qubits. This directly affects tensor-network design: small correlations help compression, but lack of spatial decay weakens the case for a strictly local tensor-network geometry aligned to hardware distance [2104.04607].

A further boundary concerns the detector model itself. Not all correlated readout errors are well described by a classical assignment matrix. A complementary non-tensor-network protocol models the readout device as a correlated POVM, includes coherent and classical detector errors, infers local correlation structure via overlapping detector tomography, and mitigates observables by reconstructing readout-mitigated states on connected noise clusters. That method avoids randomized measurements entirely, but it is not formulated as an MPO/PEPO algorithm. Its relevance is mainly conceptual: tensor-network readout mitigation and local-cluster POVM mitigation exploit the same core assumptions of locality, bounded correlation order, and observable-centered inference, but they operate on different mathematical objects [2503.24276].

The resulting picture is technically specific rather than universal. Tensor-network readout mitigation is strongest when readout correlations are local enough for low-bond representations, calibration is accurate, and the mitigation target can be written as efficient tensor contractions. It becomes harder when correlations are long-ranged, readout drift is significant, inverse channels are ill-conditioned, or coherent detector effects require a full operator-valued measurement model rather than a classical stochastic channel.

Source: https://www.emergentmind.com/topics/tensor-network-readout-error-mitigation