---
title: Involutory Fan-out Coupling in Quantum Circuits
url: https://www.emergentmind.com/topics/involutory-fan-out-coupling
type: topic
---

# Involutory Fan-out Coupling in Quantum Circuits

Involutory fan-out coupling denotes a class of quantum operations that distribute the logical value of one qubit across multiple qubits while remaining self-inverse, i.e., satisfying $U_{\rm fan}^2=I$. In the recent literature, this structure appears in several closely related forms: the canonical fan-out gate $|0\rangle\langle0|\otimes I^{\otimes n}+|1\rangle\langle1|\otimes X^{\otimes n}$, the subset-selective controlled-$U_k$ coupling used in direct quantum state tomography, phase-flip realizations generated by resonance engineering, and parity-equivalent constructions obtained from pairwise interactions. The shared involutory property is not merely algebraic. It underlies constant-depth implementations, enables repeated gate folding for zero-noise extrapolation, and supports scalable verification tasks such as single-circuit GHZ-state fidelity estimation [2604.04454], [2409.06989], [2203.01141].

## 1. Algebraic definition and involutory structure

The standard fan-out unitary acting on one control qubit and $n$ target qubits is
$$
U_{\rm fan}
=
|0\rangle\langle0|\otimes I^{\otimes n}
+
|1\rangle\langle1|\otimes X^{\otimes n},
$$
so that
$$
U_{\rm fan}\,\bigl|c\bigr\rangle\otimes|t_1,\dots,t_n\rangle
=
|c\rangle\otimes|t_1\oplus c,\dots,t_n\oplus c\rangle.
$$
Because $\bigl(X^{\otimes n}\bigr)^2=I$ and the projectors $|0\rangle\langle0|$ and $|1\rangle\langle1|$ are orthogonal, one readily checks that $U_{\rm fan}^2=I^{\otimes(n+1)}$. In this sense, fan-out is an involution, or self-inverse [2409.06989], [2007.04246].

A more selective version appears in direct quantum state tomography. Let $k\in\{0,1\}^n$ specify an arbitrary subset of system qubits and define
$$
U_k \equiv \bigotimes_{i=1}^n X_i^{k_i}.
$$
The fan-out coupling is then
$$
U_{\rm fan}\equiv |0\rangle\langle0|_m\otimes I_s + |1\rangle\langle1|_m\otimes U_k,
$$
with meter qubit $m$ and system register $s$. In the computational basis,
$$
U_{\rm fan}|0\rangle_m|x\rangle_s=|0\rangle_m|x\rangle_s,\qquad
U_{\rm fan}|1\rangle_m|x\rangle_s=|1\rangle_m|x\oplus k\rangle_s,
$$
and the matrix representation is block-diagonal,
$$
U_{\rm fan}=\operatorname{diag}(I_{2^n},U_k).
$$
Since $U_k^2=I_s$, the same calculation yields $U_{\rm fan}^2=I_m\otimes I_s$ [2604.04454].

The literature also contains a phase-flip form. In the resonance-engineered construction, the net action is
$$
U_{\rm fanout}:\;
\ket{x_1}\bigotimes_{i=2}^n\ket{b_i}\mapsto
\begin{cases}
\ket{0_1}\ket{b_2\cdots b_n},&x_1=0,\\
\ket{1_1}\,Z_2\otimes\cdots\otimes Z_n\,\ket{b_2\cdots b_n},&x_1=1,
\end{cases}
$$
which is described as a phase-flip fan-out. On the computational subspace this is equivalently written as
$$
\ket{x_1}\ket{b_2\cdots b_n}\mapsto
\ket{x_1}\ket{b_2\oplus x_1,\dots,b_n\oplus x_1},
$$
and involutivity again follows from $Z^2=I$ [2605.11073].

## 2. Constant-depth realizations

One constant-depth realization is purely unitary. Since $U_k=\prod_{i\,:\,k_i=1}X_i$, the controlled-$U_k$ fan-out coupling is a product of CNOT gates with a common control qubit and multiple targets. The critical observation is that all CNOTs sharing the same control commute, so they can be executed in one synchronous layer. In hardware with full connectivity, such as ion traps and Rydberg arrays, these CNOTs can literally fire in parallel. On nearest-neighbor superconducting layouts, the same logical constant-depth property can be preserved with mid-circuit measurement and feed-forward or with carefully placed SWAPs and routing [2604.04454].

A distinct superconducting implementation realizes constant depth through a teleportation-style GHZ protocol with four parallel time-steps. For a linear array of length $3n-2$, the circuit consists of: a parallel preparation layer, an entangling layer, a single measurement time-slice on $2(n-1)$ qubits, and a feedforward layer of conditional single-qubit rotations. The correction on the $q$th output qubit is
$$
R_q(z,x)=
\begin{cases}
Z^{z_q}\prod_{m=1}^q X^{x_m},&1\le q\le n-1,\\[4pt]
\prod_{m=1}^{n-1} X^{x_m},&q=n.
\end{cases}
$$
All four layers are fixed in number independent of $n$, so the overall depth is constant [2409.06989].

Global-interaction models provide another route. A minimal Hamiltonian for one-shot fan-out is
$$
H_{\mathit{fanout}}
=
\frac{\hbar}{2}\,\omega\,
\bigl(I-Z_c\bigr)\otimes\sum_{j=1}^N X_{t_j},
$$
with evolution for time $t=\pi/\omega$ yielding the fan-out unitary. In the $\ket0_c$ sector, nothing happens; in the $\ket1_c$ sector, each target undergoes an $R_x(\pi)=X$ [2007.04246].

Resonance engineering replaces direct multi-CNOT synthesis with Jaynes–Cummings interactions between multiple qubits and a common harmonic oscillator. In that construction,
$$
U_{\rm fanout}=e^{-iH_3t_1}\,e^{-iH_2t_2}\,e^{-iH_1t_1},
\qquad
t_1=\frac{\pi}{\Omega_c},\quad t_2=\frac{2\pi}{\Omega_t},
$$
and a strong anti-Jaynes–Cummings coupling creates an excitation-dependent blockade that produces a constant-depth fan-out [2605.11073].

## 3. Role in direct quantum state tomography

The most explicit use of involutory fan-out coupling in state characterization appears in the direct quantum state tomography scheme of Chang et al. The scheme combines strong-measurement estimation with a fan-out coupling architecture and enables mutually commuting interactions between system qubits and a single meter qubit. Because the interactions commute, the circuit depth is constant and independent of system size. The direct-tomography setting is especially suited to selective access to individual complex density-matrix elements, which is advantageous for sparse target states and some verification tasks [2604.04454].

In that framework, the fan-out coupling coherently routes information from an arbitrary subset of $n$ system qubits onto one meter qubit. The subset is specified by the bit string $k$, so the same algebraic primitive interpolates between single-target and many-target couplings without changing the circuit depth. This feature is central to the scheme’s claim of constant-depth tomography and verification.

The experimental validation reported four-qubit state reconstruction for $\mathrm{GHZ}_4$, $\lvert0000\rangle$, and $\lvert++++\rangle$. For the four-qubit tomography task, direct quantum state tomography used $2^4+1-1=31$ circuits, whereas standard Pauli tomography used $3^n=81$ circuits. The same work also demonstrated that for GHZ$_n$ fidelity estimation up to $n=20$, a single direct-tomography circuit with $k=(1,1,\dots,1)$ plus meter-$X$ measurement suffices for all $n$ [2604.04454].

These results place involutory fan-out coupling at the intersection of tomography and verification rather than only algorithmic depth reduction. A plausible implication is that the primitive is particularly attractive when only selected matrix elements or specific witnesses are required, rather than full informational completeness through a large measurement basis ensemble.

## 4. Noise scaling and error mitigation

The self-inverse property $U_{\rm fan}^2=I$ has direct operational consequences for error mitigation. In the tomography setting, inserting $U_{\rm fan}\cdot U_{\rm fan}=I$ between state preparation and measurement does not change the ideal protocol but doubles the noise on those gates. More generally, one can execute the sequence an odd number of times, such as $r=1,3,5$, and measure the same observable. If the noisy implementation has effective error rate $\epsilon$, then the expectation value at fold number $r$ obeys approximately
$$
E_r \approx E_{\rm ideal} - r\cdot \alpha\cdot \epsilon.
$$
Linear extrapolation of $E_r$ versus $r$ back to $r=0$ then yields an estimate of $E_{\rm ideal}$. In the reported implementation, measurements at $r=\{1,3,5\}$ were used, and Pauli twirling was applied around each two-qubit gate to depolarize coherent errors and make the extrapolation well controlled [2604.04454].

Readout mitigation was treated separately. On the superconducting processor used for direct tomography, readout error was mitigated by measuring $2\times2$ confusion matrices on each qubit and applying the inverse tensor product of these matrices to the raw counts. For the four-qubit state run, the typical calibration figures were: CZ error $\approx1.5\times10^{-3}$, single-qubit RX error $\approx2.3\times10^{-4}$, and readout error $\approx4.5\times10^{-3}$, with readout identified as dominant [2604.04454].

A complementary error picture is given by the constant-depth feedforward implementation. There, the total infidelity is modeled as arising from two-qubit and single-qubit control errors, mid-circuit readout and assignment errors, and decoherence during idling and feedforward latency, under an uncorrelated multiplicative success-probability model. The counts entering that model are $N_{\rm CZ}=3n-3$, $N_{\rm RO}=2n-2$, and an $N_{\rm idle}$ determined by the architecture variant under study [2409.06989].

## 5. Experimental realizations and quantitative performance

On IBM’s `ibm_aachen` superconducting processor, the direct-tomography implementation chose one qubit as the meter, typically a qubit with three neighbors, and used native CZ gates plus single-qubit RZ/SX rotations to decompose each CNOT. For $\mathrm{GHZ}_4$, the fidelity to the ideal GHZ state was $96.3(1)\%$ without readout mitigation and $98.99(6)\%$ with readout mitigation. Standard tomography gave $96.8(1)\%\rightarrow98.0(4)\%$ under the same conditions, and the cross-fidelity between direct quantum state tomography and standard quantum state tomography was $98$–$99\%$ across all states. For GHZ$_n$ fidelity estimation up to $n=20$, fidelity without mitigation dropped below the $0.5$ entanglement threshold at $n=20$, whereas combined readout error mitigation and zero-noise extrapolation pushed the 20-qubit fidelity to $53.2(3)\%$, certifying genuine multipartite entanglement [2604.04454].

The real-time-feedforward superconducting fan-out experiment demonstrated a quantum fan-out gate on up to four output qubits. All $Z$- and $X$-basis measurements were performed simultaneously through frequency-multiplexed readout. Each homodyne voltage was integrated for $\tau_{\rm meas}=400\,\mathrm{ns}$, thresholded in FPGAs, and routed to a central quantum system controller that produced the conditional recovery operations. The added runtime overhead from classical processing and signal routing was reported as a fixed $\tau_{\rm ff}\approx800\,\mathrm{ns}$, independent of $n$, and the full four-stage sequence fit in
$$
T_{\rm total}\approx352+208+400+928\ \mathrm{ns}=1888\ \mathrm{ns}.
$$
The same analysis extrapolated a scaling advantage over a unitary fan-out sequence beyond $25$ output qubits with feedforward control, or beyond $17$ output qubits if the classical feedforward latency is negligible [2409.06989].

Earlier superconducting proof-of-concept work implemented simultaneous cross-resonance fan-out on IBM Q Paris. Preparing $\ket1$ on qubit 3 and driving simultaneous cross-resonance to qubits 2 and 5 produced a 3-qubit GHZ-generation experiment in which serial CNOTs yielded $42\%\,\ket{000}+36\%\,\ket{111}$, while simultaneous fan-out yielded $31\%\,\ket{000}+29\%\,\ket{111}$ over $2\times8000$ shots. The simultaneous version ran in nearly half the time. In trapped-ion simulations using global Mølmer–Sørensen gates, $2$–$8$ qubits and $100\,000$ trajectories showed a $0.5$–$1.5\%$ fidelity gain over serial CNOT stacks [2007.04246].

The resonance-engineered proposal adds a large-system theoretical perspective. By exploiting permutation symmetry, the simulation complexity was reduced from exponential to polynomial, with total cost $\sum_{m=0}^{n-1}O(m)\sim O(n^2)$ instead of $O(4^n)$. Using this reduction, the authors simulated up to $n=100$ and reported agreement with the analytic infidelity bound $1-F\le(n-1)(\Omega_t/\Omega_s)^2$ [2605.11073].

## 6. Parity equivalence, compilation significance, and structural limits

Involutory fan-out coupling is closely related to parity. One exact construction uses an Ising-type Hamiltonian
$$
H_n=\sum_{1\le i<j\le n}J_{ij}Z_iZ_j
$$
on $n$ ancilla qubits, evolved for time $t=\pi/(4J)$. Under exact conditions on the couplings, this implements a diagonal unitary $U_n$ that can be converted in constant depth into the $(n+1)$-qubit parity gate
$$
P_n|x_1\cdots x_n,t\rangle
=
|x_1\cdots x_n,\;t\oplus(x_1\oplus\cdots\oplus x_n)\rangle,
$$
and hence into fan-out by surrounding parity on the target with Hadamards. The wrapper circuit $C_n$ uses one $U_n$, one $U_n^\dagger$, one CNOT onto the target, and $O(1)$ one-qubit Cliffords [2203.01141].

The exact coupling criterion is stringent. There must exist some $J>0$ such that every $J_{ij}$ is an odd integer multiple of $J$, and the graph whose edges satisfy $J_{ij}/J\equiv3\pmod 4$ must be Eulerian. Under inverse-square-law couplings, the resulting geometric constraints become restrictive: in $d=1$ no three distinct points work; in $d=2$ or $3$ one must avoid any triple forming a right angle; and no configuration of $n\ge5$ identical points in $\mathbb{R}^3$ can be even weakly isq-adequate. The maximal exact inverse-square realization in three dimensions is therefore the $4$-ancilla $+\,1$-target case [2203.01141].

From the compilation perspective, fan-out changes the scheduling model itself. In the conventional gate-by-gate, “exclusive-activation” model, one cannot apply more than one CNOT to the same control qubit in the same timestep. By contrast, the fan-out coupling primitive permits an arbitrary number of CNOTs sharing the same control in a single global operation. The cited synthesis results are concrete: shared-control single-qubit gates collapse to constant depth of $5$ layers, an arbitrary number of shared-control Toffolis can be scheduled in $12$ layers, and the $2k$-qubit SWAP test can be reduced to a total constant depth of $14$ layers regardless of $k$. More generally, a Controlled–$U$ of width $N$ and depth $D$ becomes a depth-$O(D)$ circuit with zero ancillas instead of $O(ND)$ under serialized CNOT layers [2007.04246].

A common misconception is that “constant depth” means a literal single physical gate layer on every architecture. The published implementations are more specific. In some settings, constant depth is obtained because all shared-control CNOTs commute and can occupy one synchronous layer; in others, it is achieved through mid-circuit measurement and feed-forward; and in still others, it is realized by global Hamiltonian evolution or resonance engineering. The invariant feature is not a unique hardware mechanism but the involutory fan-out action itself.

Source: https://www.emergentmind.com/topics/involutory-fan-out-coupling