---
title: Variational Compression of Trotter Terms
url: https://www.emergentmind.com/topics/variational-compression-of-trotter-terms
type: topic
---

# Variational Compression of Trotter Terms

Variational compression of Trotter terms denotes a family of hybrid quantum-classical methods in which product-formula circuits for real-time evolution are replaced, approximated, or refined by shallower parametrized circuits while retaining a controlled relation to the original Hamiltonian dynamics. In the basic setting, a Hamiltonian is decomposed as $H=\sum_\alpha H_\alpha$ and a first-order Trotter step is written as $U_{\mathrm{Trot}}(\Delta t)\coloneqq \prod_\alpha \exp[-i H_\alpha \Delta t]+O(\Delta t^2)$; compression then seeks a circuit $V(\theta)$ or $U_{\mathrm{var}}(\theta)$ that approximates either a Trotter-evolved state, an individual Trotter term, or a sequence of product-formula blocks with substantially lower depth [2112.12654, 2605.06122, 2604.26663]. Across recent work, this idea appears in state-compression schemes for many-body dynamics, constant-depth Hamiltonian variational ansätze for impurity solvers, Schur-decomposition-inspired compression of individual Trotter terms in molecular dynamics, Trotter-initialized variational compilation on hardware-native layouts, and operator-level variational product formulas derived from a global action principle [2508.10526, 2511.15124, 2311.06347].

## 1. Formal setting and conceptual scope

The common starting point is a decomposition of the propagator $U(t)=e^{-iHt}$ into local or blockwise exponentials. In the original variational Trotter compression procedure, one considers
$$
U_{\mathrm{Trot}}(t)=[U_{\mathrm{Trot}}(\Delta t)]^n,\qquad \Delta t=t/n,
$$
with each $\exp[-iH_\alpha \Delta t]$ compiled into a small network of single- and two-qubit gates when $H_\alpha$ acts locally [2112.12654]. The compression step is not a replacement of Hamiltonian simulation by an unrelated ansatz; rather, it is a variational approximation built directly from the same Hamiltonian terms or from a circuit structure mirroring those terms.

A second, closely related formulation compresses single Trotter factors rather than a full time-evolved state. For nonadiabatic molecular dynamics, each factor $e^{-iH_k\tau}$ in a first- or second-order product formula is identified as a “Trotter term” to be replaced by a compressed ansatz, with the target short-time step written as
$$
U(\tau)=e^{-iH\tau}\simeq \prod_k e^{-iH_k\tau}+O(\tau^2),
$$
and, for a second-order step,
$$
U_2(\tau)\simeq e^{-iT\tau/2}\cdot e^{-iV_{\mathrm{diag}}\tau/2}\cdot e^{-iV_{\mathrm{coup}}\tau}\cdot e^{-iV_{\mathrm{diag}}\tau/2}\cdot e^{-iT\tau/2}.
$$
This places variational compression at the level of individual product-formula components rather than only at the level of the full evolution operator [2605.06122].

A third formulation treats product formulas as synthesis primitives for compilation. In that setting, first-, second-, and fourth-order Suzuki blocks form a discrete library, and the variational stage refines a Trotter-initialized circuit after a greedy block-selection phase [2604.26663]. A still broader extension is the variational product formula
$$
U_{\mathrm{var}}(t)=\prod_{k=1}^K \exp[-i\lambda_k(t)H_{j_k}],
$$
whose coefficients are determined by Euler-Lagrange equations rather than fixed Suzuki coefficients, with the stated goal of preserving the unitary structure of time evolution while improving accuracy and gate efficiency [2511.15124].

Taken together, these formulations show that “compression” can refer to three technically distinct operations: compressing a Trotter-evolved state back into a shallow ansatz, compressing individual Trotter terms into trainable surrogates, and variationally refining a Trotter-inspired circuit template so that the final circuit is shallower or more hardware-compatible than a direct decomposition.

## 2. Core algorithmic patterns

In the iterative state-compression scheme, a variational state $|\psi(\theta_t)\rangle$ is first propagated by a short product-formula step and is then re-optimized so that a new shallow circuit reproduces the propagated state as closely as possible. The target overlap is
$$
F(\theta_{t+\tau})=\big|\langle \psi(\theta_{t+\tau})|U_{\mathrm{Trot}}(\tau)|\psi(\theta_t)\rangle\big|^2,
$$
with cost
$$
C(\theta_{t+\tau})=1-F(\theta_{t+\tau}),
$$
and the update $\theta_{t+\tau}^*=\arg\min_\theta C(\theta)$ is repeated step by step. The pseudocode in the original work makes the logic explicit: initialize an ansatz for the initial state, choose compression tolerance $\epsilon$, Trotter interval $\tau$, ansatz depth $\ell$, and maximum Trotter steps, then iterate propagation, compression, parameter update, fidelity recording, and stopping criteria. The stated outcome is simulations “to arbitrary times with an error controlled by the compression fidelity and a fixed Trotter step size” [2112.12654].

In the impurity-model setting, the same iterative logic is specialized to the states needed for the retarded impurity Green’s function in a DMFT loop. One writes
$$
V_n\equiv V(\theta_n)\approx [U_{\mathrm{trot}}(\Delta t)]^n,\qquad V_0=I,
$$
uses $V_{n-1}$ as a warm-start, and trains $\theta_n$ so that
$$
V_n\simeq U_{\mathrm{trot}}(\Delta t)\,V_{n-1}.
$$
The cost is not imposed on the full operator norm; it is imposed only on $|GS\rangle$ and $X_d|GS\rangle$, because those are “the only states entering the Hadamard-test circuits for $G^R_{\mathrm{imp}}(t)$” [2508.10526].

In the molecular-dynamics formulation, the workflow is termwise rather than stepwise. One chooses a target term $U=e^{-iH_k\tau}$, selects a variational ansatz $V(\gamma,\theta)$, trains it with the Local Hilbert-Schmidt Test, and then substitutes the optimized $V(\gamma^*,\theta^*;\tau)$ for the original Trotter term in the full circuit. Observable preservation is then assessed through the objective-qubit population $P_0(t)$ and the short-time slope $k=-dP_0/dt$ at $t\to 0$ [2605.06122].

In hardware-efficient compilation, the algorithm first chooses a nonuniform sequence of Suzuki blocks from a library by maximizing process fidelity with the target unitary, then promotes the fixed block angles to variational parameters and appends additional layers structured like one $S_2$ step. This produces a structure-aware approximate circuit rather than a generic synthesis of an arbitrary unitary [2604.26663]. By contrast, the global-action formulation replaces discrete re-optimization at each step by differential equations for the coefficients $\lambda_k(t)$, yielding state-independent parameters for a fixed Hamiltonian and ansatz structure [2511.15124].

## 3. Variational ansätze and circuit constructions

A notable feature of this literature is that the ansatz is typically Hamiltonian-variational rather than hardware-efficient in the narrow sense. For the Heisenberg chain, the variational Trotter compression work uses a layered “Hamiltonian-variational” brick-wall ansatz
$$
|\psi(\theta)\rangle=U(\theta)|\psi_0\rangle,\qquad
U(\theta)=\prod_{l=1}^{\ell}U_{\mathrm{even}}(\phi_l)\,U_{\mathrm{odd}}(\theta_l),
$$
with
$$
U_{\mathrm{odd}}(\theta_l)=\prod_{j\ \mathrm{odd}}\exp[-i\theta_{l,j}(X_jX_{j+1}+Y_jY_{j+1}+Z_jZ_{j+1})],
$$
and an analogous expression for $U_{\mathrm{even}}(\phi_l)$. For a nonintegrable model, an extra “$Z$–$Z$” layer
$$
U_Z(\gamma_l)=\prod_{j=1}^M \exp[-i\gamma_{l,j} Z_jZ_{j+2}]
$$
is inserted, “doubling the parameters per layer” [2112.12654].

For the single-impurity Anderson model, the Hamiltonian variational ansatz mirrors the Jordan-Wigner-mapped Pauli structure of the SIAM Hamiltonian. In compact form,
$$
U_{\mathrm{ans}}(\theta)=\prod_{\ell=1}^K\prod_{k=1}^M \exp[-i\theta_{\ell,k}h_k],
$$
where the $h_k$ include impurity, bath, and hybridization terms. The crucial design choice is that $K$ is kept independent of $t$: “one trains the set of parameters $\theta$ to approximate $U(t)$ at each desired time with only $K$ such ‘HVA layers’ rather than $N\to\infty$” [2508.10526].

The molecular nonadiabatic study adopts a Schur-decomposition-inspired “VFF ansatz”
$$
V(\gamma,\theta;\tau)\equiv W(\gamma)\,D(\theta;\tau)\,W(\gamma)^\dagger,
$$
where $W(\gamma)$ is a hardware-efficient eigenvector unitary made from alternating layers of single-qubit $R_x/R_z$ rotations and nearest-neighbor parity rotations, while
$$
D(\theta;\tau)=\prod_{j:|j|\le l}\exp\!\Big[-i\theta_j\bigotimes_{q=0}^{n-1} Z_q^{j_q}\Big]
$$
is diagonal in the computational basis and parameterized by a truncated Walsh expansion. In the “fully expressive case $(l\ge n)$” the explicit quadratic circuit is recovered exactly; for $l<n$, “the long-range $ZZ$’s are removed and the remaining $\theta_j$’s adjust to minimize the approximation error” [2605.06122].

Structure-aware compilation uses a Trotter-initialized variational ansatz in which the initial circuit is
$$
V_{\mathrm{initial}}=\prod_{k=1}^m S_2(\Delta t_k),
$$
and each subsequent variational layer is structured exactly like one $S_2$ step:
$$
V_\ell(\theta_\ell)=\Big(\prod_{\langle i,j\rangle}\prod_{\mu\in\{x,y,z\}} e^{-i\theta_{\ell,ij}^\mu \sigma_i^\mu\sigma_j^\mu}\Big)\cdot \Big(\prod_{k=1}^n e^{-i\theta_{\ell,k}^z Z_k}\Big).
$$
The same work emphasizes native placement on a linear nearest-neighbor chain and direct mapping of
$$
e^{-i\theta Z_iZ_j}=CX(i,j)\cdot R_z(2\theta)\cdot CX(i,j),
$$
with $XX$ and $YY$ terms obtained by conjugation into the $Z$ basis [2604.26663].

Constrained variational circuits constitute another compression mechanism. For the XXZ chain, the two-site gate
$$
U_{\mathrm{XXZ}}(\theta_\ell,\phi_\ell)=\exp\!\Big[i\theta_\ell(\sigma_i^x\sigma_{i+1}^x+\sigma_i^y\sigma_{i+1}^y)+i\phi_\ell \sigma_i^z\sigma_{i+1}^z\Big]
$$
encodes the global $U(1)$ symmetry directly, reducing the parameter count from $N_{\rm params}^{\rm TIVB}=12M+6$ in the translationally invariant brickwall circuit to $N_{\rm params}^{U(1)\mathrm{-block}}=2\tilde M$ [2311.06347]. This is compression by restriction of the variational manifold to symmetry-compatible circuits.

## 4. Cost functions, measurements, and optimization procedures

The cost function used in a variational compression protocol determines what is being preserved. In the original VTC work, the central quantity is the overlap fidelity between the time-evolved and variational states. Two hardware-compatible measurement protocols are given. The SWAP test with one ancilla prepares two registers in $|\Phi\rangle=U_{\mathrm{Trot}}(\tau)|\psi(\theta_t)\rangle$ and $|\Psi\rangle=U(\theta)|\psi_0\rangle$, and “the ancilla $Z$ expectation equals $|\langle \Psi|\Phi\rangle|^2$”; this requires $O(M)$ controlled-SWAPs and “$O(8M)$ CNOTs plus ancilla connectivity.” The double-time-contour or Loschmidt-echo circuit instead applies $U(\theta_t)$, then $U_{\mathrm{Trot}}(\tau)$, then $U^\dagger(\theta_{t+\tau})$ sequentially to $|\psi_0\rangle$, with the return probability to $|\psi_0\rangle$ equal to the required overlap. The same work lists readout-error calibration and inversion, symmetry-based postselection, Pauli twirling around CNOTs, and zero-noise extrapolation via gate-folding as error-mitigation strategies [2112.12654].

For DMFT, the cost function is tailored to the Green’s-function measurement task:
$$
C(\theta_n)=1-\mathrm{Re}\,[\langle GS|L_n^\dagger|GS\rangle\langle GS|K_n|GS\rangle],
$$
with
$$
L_n=V_{n-1}^\dagger U_{\mathrm{trot}}^\dagger(\Delta t)V_n,\qquad K_n=X_dL_nX_d.
$$
The Appendix result cited in the summary is that minimizing this cost “enforces $L_n|GS\rangle=|GS\rangle$ and $L_nX_d|GS\rangle=X_d|GS\rangle$ up to the same global phase for both states,” which is sufficient for the required overlaps in $G(t)$. Optimization is gradient-based and proceeds via the parameter-shift rule [2508.10526].

In the molecular setting, the comparison between target and compressed term is made with the Local Hilbert-Schmidt Test,
$$
C_{\mathrm{LHST}}(U,V)=1-\frac{1}{n}\sum_{j=1}^n F_e^{(j)},
$$
where $F_e^{(j)}$ is the entanglement fidelity on qubit pair $j$ measured through Bell-pair preparation and Bell-basis measurement. The stated faithfulness property is $C_{\mathrm{LHST}}=0$ iff $V=e^{i\phi}U$. Gradients are obtained from the parameter-shift rule
$$
\frac{\partial C}{\partial \alpha_k}=\frac{1}{2}\big[C(\alpha_k+\pi/2)-C(\alpha_k-\pi/2)\big],
$$
and parameters are updated with Adam until $C<C_{\mathrm{thresh}}$, with “$C_{\mathrm{thresh}}=0.01$” in the reported implementation [2605.06122].

Compilation-oriented compression uses process fidelity,
$$
F(U,V)=\frac{|\mathrm{Tr}(U^\dagger V)|^2}{2^{2n}},
$$
both for greedy block selection and for the variational cost $C(\theta)=1-F(U_{\mathrm{target}},U_{\mathrm{var}}(\theta))$. The reported optimizer is L-BFGS-B with finite-difference gradients, “ftol=$10^{-15}$, gtol=$10^{-10}$, maxiter=300,” and convergence in “$\sim100$–$300$ iterations, $\sim10^4$ total function calls, reaching $F>0.99$” [2604.26663]. Constrained-circuit compression instead minimizes the normalized Frobenius distance
$$
\epsilon(\{\theta_i\})=1-\frac{\mathrm{Tr}[C^\dagger(\{\theta_i\})U(t)]}{2^L},
$$
with gradients obtained by automatic differentiation of MPO representations and optimization by Adam [2311.06347].

## 5. Empirical performance across applications

In exact statevector simulations for the Heisenberg model with “$M$ up to $11$,” the original VTC study reports that “Trotter step-size error and compression error trade off $\to$ optimum $\tau$ and $n$ chosen so that both errors $\sim 5\times 10^{-3}$.” Under those conditions, “direct Trotter with depth $\sim 3\ell$ loses fidelity by $t\sim120\,J^{-1}$, whereas VTC retains $F\sim0.83$ out to $t=140\,J^{-1}$.” In ideal shot-noise simulations for “$M$ up to $6$,” “with $2^{14}$–$2^{16}$ shots per cost evaluation, VTC fidelities at large $t$ exceed direct Trotter (depth $3\ell$) fidelities,” and the “non-gradient optimizer (CMA-ES) [is] robust to shot noise; gradient methods struggle.” In noisy circuit simulation with IBM Santiago error rates for “$M=3$,” the combination of “readout mitigation + ZNE + twirling + conservation post-selection” yields “mean VTC fidelity $>0.9$ out to $t=20\,J^{-1}$, while direct Trotter (depth 6) fidelity vanishes by $t\sim16\,J^{-1}$.” On real IBM Santiago hardware, “VTC fidelity remains $\sim0.8$ at $t=20\,J^{-1}\gg$ device coherence limit” [2112.12654].

For impurity models, the DMFT-oriented compression work reports that, for “$B=2$ bath sites, $\Delta t=0.1$, and interaction $U=4\,v$,” compression with “$K=3$ layers yields fidelities $>0.999$ with the true Trotter evolution up to $t=50\,(1/v)$ (i.e. 500 Trotter steps).” After fitting the compressed $G(t)$ to “a few-pole Lehmann ansatz,” the authors “recover the Matsubara self-energy and quasiparticle weight $Z_{\mathrm{mats}}$ with $\Delta Z<10^{-3}$ using only $t_{\max}\sim10\,(1/v)$.” The full DMFT loop “converges in $O(10)$ iterations for the one-site Hubbard model at $U=4\,v$, $n=0.5$ on a Bethe lattice, reproducing the known DMFT quasiparticle weight $Z\approx0.68$” [2508.10526].

For nonadiabatic molecular dynamics, the reported optimization benchmarks show that, for the kinetic term $T$, “linear topology needs $l=4$ to reach $C<10^{-2}$ for all $n\le8$,” while “ring topology hits the same at $l=3$”; for the diabatic potentials $V_0/V_1$, “linear needs $l=6$, ring needs $l=4$.” On ibm_brisbane, the “$n=3$ compressed $T$ $(l=1)$” circuit attains “state fidelity $\ge0.80$ through $N=10$ steps at $\tau=32$ a.u.” For reaction-rate extraction, compressed circuits “reproduce[] peak position (normal/inverted Marcus regimes) but show[] quantitative deviations up to $\sim20\%$ in tails, due to combined Trotter + truncation error” [2605.06122].

For structure-aware compilation, the Heisenberg benchmarks at “$n=3$–$8$ qubits” and “$J=0.5,t=1.0$” show adaptive $S_2$ circuits with fidelities from “$0.9990$” at $n=3$ to “$0.9968$” at $n=8$, while black-box synthesis grows from “$30\to217\to1025\to4345$ CX at $n=3$–$6$.” On IBM Torino, the paper identifies a NISQ regime in which “shorter approximate circuits outperform deeper exact decompositions”: in “Scenario B,” a “27-CX” variational circuit with “$F_{\mathrm{sim}}=0.996$” achieves hardware fidelity “$F_{\mathrm{hw}}=0.987$” on $|++++\rangle$, whereas the “187-CX” exact black-box circuit reaches “$F_{\mathrm{hw}}=0.974$” [2604.26663].

## 6. Scalability, limitations, and broader formulations

The main limitations are not uniform across approaches. In the original VTC analysis, the required ansatz depth $\ell^*(t,\epsilon,M)$ “grows roughly linearly in $t$” at short times, but “for fixed final time $t_f\gg1$, $\ell^*(M)$ grows exponentially in $M$ for both integrable and nonintegrable Heisenberg chains.” The same work identifies “the difficulty of carrying out the optimization of the noisy cost function” as “the main bottleneck in going to larger system sizes” [2112.12654]. This makes clear that compression beyond the coherence time does not by itself imply favorable asymptotic scaling.

Constrained variational circuits reduce optimization cost dramatically, but the gain can come with an expressibility penalty. For the XXZ chain, encoding the symmetry reduces the parameter dimension “by factor $\sim20$” at comparable CNOT count and yields “one–to–two orders of magnitude” savings in classical optimization cost. However, in locally constrained models such as PXP, the blocked ansatz has a “restricted lightcone,” with $W_{\mathrm{PXP}}\approx M/4$, and “for times $t\gtrsim4$ the blocked ansatz loses expressibility and is overtaken by TIVB or even by standard Trotter in accuracy” [2311.06347]. A common misconception is therefore that stronger physical constraints always improve compressed simulation; the cited results show explicit exceptions.

Another important distinction concerns state dependence. The DMFT method deliberately optimizes only those states entering the measurement of $G^R_{\mathrm{imp}}(t)$, and the molecular scheme explicitly focuses on preserving reaction-rate coefficients rather than the full operator on all states [2508.10526, 2605.06122]. By contrast, the variational product-formula approach states that the optimized parameters “depend only on $H$ and the chosen ansatz structure, not on the initial state,” can be “re-used for any $|\psi_0\rangle$,” and effectively permit $\tau$ to be taken “$\sim2\times$ larger than in standard TS for the same accuracy,” leading to “$\sim2\times$ compression of Trotter layers” and “2×–5× error reductions” across the reported spin models [2511.15124]. This suggests that variational compression is not a single algorithmic template but a spectrum ranging from task-specific state compression to operator-level, state-independent product formulas.

Open problems are stated directly in the VTC work: “design more compact or adaptive ansätze,” “more efficient (possibly quantum-gradient) optimizers for noisy cost functions,” and “improved composite error-mitigation schemes (e.g. virtual distillation, probabilistic error cancellation) to push to larger system sizes” [2112.12654]. Within the current literature, the central technical tension remains the same across settings: depth reduction and hardware viability are obtained by introducing a variational approximation layer, and the quality of that layer is limited by ansatz expressibility, trainability, and the noise sensitivity of the chosen cost function.

Source: https://www.emergentmind.com/topics/variational-compression-of-trotter-terms