---
title: Entanglement Forging in Quantum Simulations
url: https://www.emergentmind.com/topics/entanglement-forging-ef
type: topic
---

# Entanglement Forging in Quantum Simulations

Searching arXiv for recent and foundational papers on Entanglement Forging to ground the encyclopedia entry.
arxiv_search(query="entanglement forging", max_results=10, sort_by="relevance")
Entanglement Forging (EF) is a hybrid quantum–classical strategy for simulating a target many-body state or quantum process by replacing explicit entanglement across a bipartition with classical post-processing over smaller quantum computations. In its canonical form, EF expresses a state on a doubled Hilbert space as a short or structured sum of tensor-product states across a cut, measures local quantities on the fragments, and reconstructs global expectation values from those data. In electronic-structure settings, this often means mapping one qubit to one spatial orbital rather than one spin-orbital, thereby halving the qubit requirement by separating the $\alpha$ and $\beta$ spin sectors into distinct circuits; in broader settings, EF is a circuit-knitting or Schmidt-truncation technique that trades qubit count and circuit depth for additional measurements and classical recombination [2104.10220].

## 1. Conceptual and historical framework

The defining idea of EF is to approximate or represent a bipartite state by a sum of product states across a partition $A|B$. A generic form used across the literature is
$$
|\psi\rangle = \sum_k s_k\, |a_k\rangle \otimes |b_k\rangle,
$$
with $s_k \ge 0$ and $\sum_k s_k^2 = 1$, or, in a variational forged form,
$$
|\Psi\rangle = \sum_i c_i\, |\phi_i^A\rangle \otimes |\phi_i^B\rangle.
$$
The operational consequence is that one prepares only the fragment states and reconstructs the expectation values of operators that factor across the cut by classical combination of fragment correlators. In the chemistry-oriented formulation introduced for near-term hardware, this allowed ten spin-orbitals to be represented on five qubits by exploiting a spin-up versus spin-down partition and weak entanglement across that cut [2104.10220].

The 2021 formulation established EF as a method for “doubling the size” of quantum simulations on hardware by representing a $2N$-qubit state through $N$-qubit computations and classical post-processing. Later work generalized the paradigm in several directions. “Entropy-driven entanglement forging” used subsystem entropy, entanglement structure, and symmetries to choose the bipartition and forging rank [2409.04510]. “Entanglement Forging with generative neural network models” incorporated autoregressive neural networks to learn Schmidt-like coefficients and enable scalable sampling over many components [2205.00933]. “Hybrid Ground-State Quantum Algorithms based on Neural Schrödinger Forging” targeted the exponential bitstring summation bottleneck of Schrödinger-picture EF by adaptively selecting the most relevant bitstrings with a generative model [2307.02633]. Other works extended EF beyond ground-state variational simulation, including thermofield-double preparation [2311.10566], distributed quantum computation without quantum links [2409.02509], and quantum-centric electronic-structure workflows on IBM Heron hardware [2508.08229].

A recurring misconception is that EF is merely qubit tapering or a symmetry reduction. The literature instead treats EF as a distinct bipartition-based reconstruction strategy: tapering removes qubits through exact conserved quantities, whereas EF retains the physical degrees of freedom but shifts the burden of inter-partition entanglement from hardware to classical post-processing [2409.04510].

## 2. Mathematical structure and expectation-value reconstruction

The mathematical core of EF is a bipartite decomposition together with an estimator for expectation values of factorized operators. In one standard formulation, the state is written as
$$
|\psi\rangle = (U\otimes V)\sum_n \lambda_n\, |b_n\rangle \otimes |b_n\rangle,
$$
where $\{|b_n\rangle\}$ are computational-basis states, $U$ and $V$ are local unitaries, and the Schmidt coefficients satisfy $\sum_n \lambda_n^2 = 1$ [2104.10220]. The associated Schrödinger-picture forging identity introduces superposition states
$$
|\phi^p_{xy}\rangle = \frac{|x\rangle + i^p |y\rangle}{\sqrt{2}}, \qquad p\in\mathbb{Z}_4,
$$
so that off-diagonal contributions can be reconstructed from diagonal measurements on these superpositions. For a bipartite operator $O = O_1 \otimes O_2$, the expectation value becomes a sum of products of $N$-qubit expectation values on diagonal states and superposition states, rather than a direct $2N$-qubit measurement [2104.10220].

The chemistry-specific formulation in the 2025 Heron study makes this structure explicit for spin-separated electronic Hamiltonians. There the EF ansatz is
$$
|\Psi\rangle = \sum_\mu c_\mu |u_\mu\rangle \otimes |v_\mu\rangle,
$$
with $|u_\mu\rangle$ and $|v_\mu\rangle$ living on the $\alpha$ and $\beta$ partitions. For tensor-product operators $A\otimes B$, the expectation is reconstructed from diagonal contributions and from superposition states
$$
|u_{\mu\nu p}\rangle = \frac{|u_\mu\rangle + i^p |u_\nu\rangle}{U_{\mu\nu p}}, \qquad
|v_{\mu\nu r}\rangle = \frac{|v_\mu\rangle + i^r |v_\nu\rangle}{V_{\mu\nu r}},
$$
with normalization factors
$$
U_{\mu\nu p} = \sqrt{2 + 2\,\mathrm{Re}(i^p \langle u_\mu|u_\nu\rangle)}, \qquad
V_{\mu\nu r} = \sqrt{2 + 2\,\mathrm{Re}(i^r \langle v_\mu|v_\nu\rangle)}.
$$
This converts global expectation values into $\alpha$-only and $\beta$-only measurements plus a controlled set of superposition measurements [2508.08229].

For electronic structure, the same logic appears at the reduced-density-matrix level. With spin-resolved $1$-RDMs $D^\sigma$ and same-spin $2$-RDMs $\Gamma^\sigma$, the energy is written as
$$
E = \mathrm{Tr}(h D^\alpha) + \mathrm{Tr}(h D^\beta)
+ \frac{1}{2}\sum g^\alpha \Gamma^\alpha
+ \frac{1}{2}\sum g^\beta \Gamma^\beta
+ \sum J\, D^\alpha D^\beta.
$$
Under EF, the same-spin terms are reconstructed partition-wise, while the opposite-spin contribution is bilinear in $D^\alpha$ and $D^\beta$ and therefore requires only products of partition-wise $1$-RDM elements rather than any $2M$-qubit cross operator [2508.08229].

Several EF variants differ chiefly in what is factorized. Schrödinger-picture EF forges the state and reconstructs observables from local expectations. Heisenberg-picture forging instead decomposes observables and can provide constant-overhead estimators under additional symmetry assumptions, notably when $U_A = U_B$ and the problem admits the required Clifford decompositions [2104.10220]. This suggests a fundamental axis of variation in EF: state-side truncation versus operator-side decomposition.

## 3. Electronic-structure EF and qubit halving by spin separation

In conventional second-quantized encodings for chemistry, one qubit represents a spin-orbital, so a problem with $M$ spatial orbitals requires $2M$ qubits to represent the $\alpha$ and $\beta$ sectors. The 2025 quantum-centric hydrogen-abstraction study implements EF by mapping one qubit to one spatial orbital and handling the $\alpha$ and $\beta$ sectors on separate $M$-qubit circuits. Entanglement between these sectors is then “forged” classically from the measurement results of the two circuits [2508.08229].

The nonrelativistic electronic Hamiltonian is written in second quantization as
$$
H = \sum_{pq} h_{pq} a_p^\dagger a_q + \frac{1}{2}\sum_{pqrs} h_{pqrs} a_p^\dagger a_q^\dagger a_r a_s,
$$
with spin-separated contributions $H_\alpha$, $H_\beta$, and opposite-spin density-density terms $H_{\alpha\beta}$. EF is especially natural for this structure because the opposite-spin term can be assembled from products of partition-wise observables. The circuits in the Heron study use the standard Jordan–Wigner transformation, with the essential EF change being that the two spin sectors are implemented on separate spatial-orbital registers rather than on a single $2M$-qubit register [2508.08229].

That work also broadens the expressive family of EF states beyond earlier “bitstrings + hopgates” constructions. Starting from unitary coupled-cluster doubles,
$$
|\Psi_{\mathrm{uCCD}}\rangle = e^{T-T^\dagger}|x_{\mathrm{HF}}\rangle,
$$
the authors motivate EF states of the form
$$
|u_\mu\rangle = e^{T_{\alpha\alpha}-T_{\alpha\alpha}^\dagger}|d_{\mu\alpha}\rangle, \qquad
|v_\mu\rangle = e^{T_{\beta\beta}-T_{\beta\beta}^\dagger}|d_{\mu\beta}\rangle,
$$
where $|d_{\mu\sigma}\rangle$ are non-orthogonal Slater determinants from a resonating Hartree–Fock optimization and the unitary doubles are implemented by a local unitary cluster Jastrow ansatz. CCSD amplitudes parameterize the LUCJ layers, while ResHF provides the non-orthogonal determinants and EF coefficients [2508.08229].

The original chemistry EF demonstration on water used a more compact ansatz with five hop gates and a truncated Schmidt decomposition retaining three bitstrings,
$$
|b_1\rangle = |11100\rangle,\quad
|b_2\rangle = |01110\rangle,\quad
|b_3\rangle = |01101\rangle,
$$
on a five-qubit register representing one spin sector. The same work defined the two-qubit hop gate
$$
h(\varphi)=
\begin{bmatrix}
1 & 0 & 0 & 0\\
0 & \cos\varphi & -\sin\varphi & 0\\
0 & \sin\varphi & \cos\varphi & 0\\
0 & 0 & 0 & -1
\end{bmatrix},
$$
which was used as a hardware-efficient primitive for real-valued, fixed-particle-number wavefunctions [2104.10220].

The resource advantage of spin-separated EF is concrete. After transpilation to heavy-hex topology on IBM Heron, the 2025 study reports the following average resources per circuit:

| Active space | EF(ind) | EF(super) | Conventional LUCJ |
|---|---:|---:|---:|
| $(13e,13o)$ | 13 qubits, 248 CZ, depth 257 | 13+1 qubits, 450 CZ, depth 474 | 26+4 qubits, 1606 CZ, depth 631 |
| $(23e,23o)$ | 23 qubits, 800 CZ, depth 434 | 23+1 qubits, 1375 CZ, depth 834 | 46+6 qubits, 5167 CZ, depth 1222 |

The same study states that EF circuits used approximately $3\times$–$4\times$ fewer CZ gates and substantially shallower depths than the corresponding $2M$-qubit circuits, while also yielding a consistently larger fraction of bitstrings with correct $N_\alpha$ and $N_\beta$ across reactant, transition-state, and product geometries [2508.08229].

## 4. Measurement protocols, sampling, and integration with subspace methods

EF replaces a wide entangling circuit by a collection of smaller circuits plus classical stitching. The measurement burden therefore becomes central. In the 2025 Heron workflow, two classes of circuits are used. “EF ind” prepares the individual states $|u_\mu\rangle$ and $|v_\mu\rangle$ and measures the observables required for $D^\sigma$ and $\Gamma^\sigma$. “EF super” prepares the superposition states $|u_{\mu\nu p}\rangle$ and $|v_{\mu\nu r}\rangle$ using a controlled orbital rotation in a Hadamard-test-like construction with an ancilla initialized in
$$
|\phi_p\rangle = \frac{|0\rangle + i^p|1\rangle}{\sqrt{2}}.
$$
Measuring the ancilla in the $X$ basis probabilistically prepares the required superposition states, which capture the off-diagonal terms in the global expectation values [2508.08229].

The same workflow integrates EF with sample-based quantum diagonalization (SQD). SQD samples computational-basis configurations from a quantum circuit, projects the Schrödinger equation into the sampled subspace, and solves
$$
\sum_n H_{mn} c_n = E c_m, \qquad H_{mn} = \langle x_m|H|x_n\rangle,
$$
using Davidson because the sampled Slater determinants are orthonormal and therefore $S=I$. To handle noisy sampling and ansatz imperfections, SQD applies self-consistent configuration recovery: it postselects or repairs configurations to enforce correct particle numbers, forms randomized batches, diagonalizes each batch subspace, updates orbital occupations, and iterates to convergence [2508.08229].

EF contributes the joint distribution of $\alpha$ and $\beta$ bitstrings in the form
$$
p(x,y)=|\langle x,y|\Psi\rangle|^2 = \sum_I r_I\, p_I(x)\, q_I(y),
$$
but because the coefficients $\{r_I\}$ do not define a probability distribution, the study samples from the approximate compound distribution
$$
\tilde p(x,y)=\sum_I P_I\, p_I(x)\, q_I(y), \qquad
P_I = \frac{\max[0,\mathrm{Re}(r_I)]}{\sum_J \max[0,\mathrm{Re}(r_J)]}.
$$
Those paired configurations then seed the SQD recovery and diagonalization loop [2508.08229].

Outside chemistry, measurement and sampling are also the central challenge. In the 2021 formulation, the number of experiments required for additive error $\epsilon$ with $99\%$ confidence was given as
$$
S = \frac{200\,\|\mu\|_1^2}{\epsilon^2}, \qquad
\|\mu\|_1 = 1 + 4\left(\sum_i |\lambda_i|\right)^2,
$$
and the direct Schrödinger-forging sampling overhead scales as
$$
S \sim \frac{\|\vec\lambda\|_1^4}{\epsilon^2}.
$$
These formulas make explicit that EF is efficient when the bipartite entanglement is limited, since then $\|\vec\lambda\|_1$ remains small [2104.10220].

The generative-model EF literature addresses precisely this measurement bottleneck. The 2022 neural-network formulation uses an autoregressive neural network to model $p(\sigma)=\lambda_\sigma^2$ and the ratio
$$
R(\sigma,\sigma')=\frac{\lambda_{\sigma'}}{\lambda_\sigma},
$$
which enters the estimator for cross-register observables after a Clifford decomposition. The efficiency claim is not that EF eliminates measurement cost, but that the symmetry-restricted estimator remains polynomial in system size and measurement precision because the required conditional probabilities are sampled from single-register circuits [2205.00933].

## 5. Rank selection, entropy guidance, and neural augmentation

A central question in EF is how to choose the bipartition and how many forged terms to retain. “Entropy-driven entanglement forging” formalizes this using the Schmidt spectrum and subsystem entropy. For a pure state across $A|B$ with Schmidt coefficients $s_k$, the von Neumann entropy is
$$
S(\rho_A) = -\sum_k s_k^2 \log s_k^2,
$$
and for an equipartition of $N_q$ qubits the paper uses the bound
$$
2^S \le \chi \le 2^{N_q/2}.
$$
A heuristic forging rank is then $r \sim e^S$, refined by inspecting the Schmidt spectrum and symmetry-induced degeneracies [2409.04510].

That work shows how limited physical information can determine an effective EF structure. In the Fermi–Hubbard chain at half filling, a robust gap appears between keeping four and five Schmidt terms because one non-degenerate coefficient is followed by four symmetry-related degenerate ones; for $t_m=t$, a $1\%$ infidelity is reached by $r=5$. In neutron-rich nuclear shell-model systems, proton–neutron partitions exhibit low entanglement and fivefold degenerate subleading Schmidt sectors, motivating $r \approx 6$ and, in some cases, a second forging layer that quarters the qubit count per fragment [2409.04510].

The resource implications are substantial but strongly regime-dependent. In the reported benchmarks, one EF cut halves the qubits per fragment, and two cuts quarter them. For ${}^{28}\mathrm{Ne}$, the full ADAPT-VQE circuit required more than $10^4$ CNOTs by iteration $100$, whereas one-cut EDEF reduced the maximum CNOTs per circuit by an order of magnitude and two-cut EDEF by a further order of magnitude; for ${}^{60}\mathrm{Ti}$, two-cut EDEF kept the largest circuit within $O(10^2$–$10^3)$ CNOTs [2409.04510]. The same paper also identifies the principal failure mode: when the chosen cut carries strong entanglement, the fixed-rank approximation saturates and accuracy degrades unless $r$ is increased.

Neural approaches address the complementary problem of selecting the relevant forged basis states without enumerating exponentially many bitstrings. “Hybrid Ground-State Quantum Algorithms based on Neural Schrödinger Forging” introduces a generative autoregressive model that learns the most important forging bitstrings and restricts the EF sum to a resource-controlled subset $A$ of size $k$, using the truncated estimator
$$
E^{(k)} = \sum_{z\in A} w_z E_z,
$$
with truncation bias bounded by the omitted tail. The ARNN is trained with losses such as MMD,
$$
\mathcal{L}_{\mathrm{MMD}}(\theta)=\sum_{\sigma_1,\sigma_2\in\mathcal T}
\Big[q(\sigma_1)q(\sigma_2)-2q(\sigma_1)p_\theta(\sigma_2)+p_\theta(\sigma_1)p_\theta(\sigma_2)\Big]
K(\sigma_1,\sigma_2),
$$
or log-cosh regression on $\log p_\theta(\sigma)-\log q(\sigma)$, and the optimal restricted Schmidt coefficients are obtained by solving the eigenproblem
$$
C\lambda = \mu \lambda
$$
for the smallest

Source: https://www.emergentmind.com/topics/entanglement-forging-ef