---
title: 'RL-CQE: Reinforcement Learning Quantum Eigensolver'
url: https://www.emergentmind.com/topics/reinforcement-learning-contracted-quantum-eigensolver-rl-cqe
type: topic
---

# RL-CQE: Reinforcement Learning Quantum Eigensolver

Searching arXiv for the cited RL-CQE papers and related context.
tool unavailable
Reinforcement Learning Contracted Quantum Eigensolver (RL-CQE) denotes a reinforcement-learning-based hybrid quantum–classical eigensolver framework that uses adaptive measurement, feedback, and operator selection to approximate eigenvectors or multi-state wavefunctions of Hermitian operators. In the cited literature, the framework appears in two closely connected forms: a 2019 measurement–feedback protocol for extracting approximate eigenvectors of an arbitrary Hermitian black-box operator, and a 2026 generalization that formulates RL-CQE through the anti-Hermitian contracted Schrödinger equation (ACSE) for electronic excited states and real-time dynamics. The original formulation emphasizes semi-autonomous learning with minimal memory and one measurement per iteration, whereas the later contracted formulation uses ACSE residuals, sign-free two-qubit operators, and a deep Q-network (DQN) agent to build compact ansätze for many-fermion problems [1906.06702], [2605.18569].

## 1. Origin, scope, and nomenclature

The original protocol introduced by Albarrán-Arriagada et al. is a reinforcement-learning procedure for obtaining an approximation of the eigenvectors of an arbitrary Hermitian quantum operator. It is built from two systems: a black box named environment and a quantum state named agent. The environment changes any quantum state by a unitary matrix
$$
\hat U_E(\tau)=\exp\big(-i\tau\hat{\mathcal O}_E\big),
\quad \tau\in\mathbb R,
$$
where $\hat{\mathcal O}_E$ is Hermitian. The agent adapts to some eigenvector of $\hat{\mathcal O}_E$ by repeated interactions with the environment, a feedback process, and semi-random rotations [1906.06702].

The later literature uses the name RL-CQE explicitly and states that it was “previously developed for ground-state problems.” In that formulation, RL-CQE is generalized to electronic excited states and real-time quantum dynamics. The generalization replaces the black-box eigenvector-learning viewpoint with a contracted-eigenvalue framework based on ACSE residual minimization, while retaining the reinforcement-learning principle that a policy adaptively selects the next transformation to apply. A key feature is that a DQN agent selects two-body operators at each iteration, yielding more compact ansätze and improved robustness with respect to critical hyperparameters [2605.18569].

A common source of confusion is that RL-CQE is not a single fixed circuit template. The 2019 construction is a measurement–feedback eigensolver for arbitrary Hermitian operators, while the 2026 construction is a contracted, many-body formulation targeted at electronic excited states and dynamics. The shared structural element is adaptive policy-driven operator selection rather than static variational optimization.

## 2. Original measurement–feedback formulation

In the original protocol, the environment operator is written in spectral form as
$$
\hat{\mathcal O}_E=\sum_{\ell=0}^{d-1}\lambda_\ell\,|v_\ell\rangle\langle v_\ell|,
$$
acting on a $d$-dimensional Hilbert space. An auxiliary agent system $A$ of the same Hilbert-space dimension is prepared in a fully known reference state $|\phi_{A,0}^{(1)}\rangle$, such as $|0\rangle$ for a single qubit or $|j\rangle$ in higher dimensions. After $k$ iterations, the agent state is denoted $|\phi_{A,0}^{(k)}\rangle$ [1906.06702].

The algorithm proceeds in discrete iterations. Before interaction, the agent is in $|\phi_{A,0}^{(k)}\rangle$. After applying the environment unitary, the state becomes
$$
|\bar\phi_{A,0}^{(k)}\rangle=\hat U_E\,|\phi_{A,0}^{(k)}\rangle.
$$
This state is then expressed in the current agent basis $\{|\phi_{A,j}^{(k)}\rangle\}_{j=0}^{d-1}$ through a diagonalization operator $\hat D^{(k)}$ satisfying
$$
\hat D^{(k)\dagger}:\{|\phi_{A,j}^{(k)}\rangle\}\longrightarrow \{|j\rangle\}.
$$
Operationally, the rotated state is measured in the computational basis. The outcome determines whether the current reference basis vector already approximates an eigenvector or whether a corrective update is required [1906.06702].

The role of the measurement is therefore not merely diagnostic. It defines the reinforcement signal that determines whether the current basis should be preserved or altered. In this sense, the eigenvector-learning task is cast as a closed-loop adaptive control problem on Hilbert space.

## 3. Policy update, search width, and convergence criterion

If the measurement outcome is $m=0$, the protocol assumes that $|\phi_{A,0}^{(k)}\rangle$ may already be an eigenvector, and no corrective rotation is applied. If the outcome is $m=j\neq 0$, the agent state is outside the desired eigen-subspace, and a small random rotation is applied within the two-dimensional subspace $\mathrm{span}\{|\phi_{A,0}^{(k)}\rangle,|\phi_{A,j}^{(k)}\rangle\}$:
$$
\hat{\mathcal U}_j^{(k)}
=\exp\big(-i\varphi_x\,\hat S_{x,j}\big)\,
 \exp\big(-i\varphi_z\,\hat S_{z,j}\big)\,
 \exp\big(-i\varphi_y\,\hat S_{y,j}\big),
$$
where, for example,
$$
\hat S_{x,j}=\tfrac12\big(|\phi_{A,0}^{(k)}\rangle\langle\phi_{A,j}^{(k)}|
+|\phi_{A,j}^{(k)}\rangle\langle\phi_{A,0}^{(k)}|\big),
$$
and each angle $\varphi_\alpha$ is chosen randomly in $[-w^{(k)}\pi,+w^{(k)}\pi]$ [1906.06702].

The reinforcement signal is encoded through a search width $w^{(k)}$ and two real scalars, a reward rate $0<r<1$ and a punishment rate $p>1$. The update rule is
$$
w^{(k+1)}=
\begin{cases}
r\,w^{(k)}, & m=0 \quad \text{(reward)}\\[4pt]
p\,w^{(k)}, & m\neq 0 \quad \text{(punishment)}~.
\end{cases}
$$
The diagonalizer evolves according to
$$
\hat D^{(k+1)}=
\begin{cases}
\hat D^{(k)}, & m=0,\\[4pt]
\hat{\mathcal U}_m^{(k)}\,\hat D^{(k)}, & m\neq 0~.
\end{cases}
$$
The agent is then reset to the post-measurement state in the computational basis before the next feedback step. One may take the value function as $V^{(k)}=w^{(k)}$, with convergence declared when $w^{(k)}\to 0$ [1906.06702].

The fidelity metric is defined by
$$
|\psi_{\rm agent}^{(\ell,k)}\rangle=\hat D^{(k)}|\ell\rangle,
\qquad
F_\ell(k)=\big|\langle v_\ell\,|\,\psi_{\rm agent}^{(\ell,k)}\rangle\big|^2.
$$
This makes the convergence target explicit: the protocol is designed to learn eigenvectors directly rather than to minimize an energy expectation value as an intermediate surrogate. That distinction is central to its later comparison with VQE [1906.06702].

## 4. Contracted formulation through ACSE residuals

The 2026 formulation casts RL-CQE as a multi-state contracted eigensolver for a fermionic Hamiltonian
$$
H = \sum_{pq} h_{pq}\,a_p^\dagger a_q
+\tfrac12\sum_{pqkl} g_{pq,kl}\,a_p^\dagger a_q^\dagger a_l a_k.
$$
The target is a set of $K$ eigenstates $\{|\Psi_v\rangle\}_{v=0\ldots K-1}$. The ACSE is written as
$$
[R_{pq,kl},H]\,|\Psi_v\rangle=0,
\qquad
R_{pq,kl}\equiv a_p^\dagger a_q^\dagger a_l a_k-\big(a_p^\dagger a_q^\dagger a_l a_k\big)^\dagger.
$$
At iteration $n$, the residual vector is
$$
r^{(n)}_{pq,kl}
=\sum_{v=0}^{K-1} w_v\,\langle\Psi_v^{(n)}|\,[R_{pq,kl},H]\,|\Psi_v^{(n)}\rangle,
$$
with fixed positive, descending weights $w_0\ge w_1\ge\cdots\ge w_{K-1}>0$ and $\sum_v w_v=1$. The cost function is the residual norm
$$
L^{(n)}=\bigl\lVert r^{(n)}\bigr\rVert_2^2
=\sum_{pqkl}\bigl(r^{(n)}_{pq,kl}\bigr)^2.
$$
The multi-state wavefunction is updated by a product of two-body unitaries,
$$
|\Psi_v^{(n+1)}\rangle
=\exp\!\bigl[\theta_k^{(n)}\,\hat O_k\bigr]\,|\Psi_v^{(n)}\rangle,
\quad v=0\ldots K-1,
$$
where one operator $\hat O_k$ is selected from an action pool at each step, and the scalar amplitude is obtained by the one-dimensional line search
$$
\theta_k^{(n)}=
\arg\min_\theta\;
\sum_{v=0}^{K-1} w_v\,
\bigl\langle\Psi_v^{(n)}\bigl|\,e^{-\theta O_k}\,H\,e^{+\theta O_k}\bigr|\Psi_v^{(n)}\bigr\rangle
$$
[2605.18569].

A distinguishing feature of this formulation is the use of sign-free qubit operators,
$$
O_{kl}=\sigma_k^+\sigma_l^+ + \sigma_k^-\sigma_l^-,
$$
with $\sigma^\pm=(\sigma^x\pm i\sigma^y)/2$. The paper states that the expectation value $\langle[a_p^\dagger a_q^\dagger a_l a_k,H]\rangle$ differs from $\langle[O_{kl},H]\rangle$ only by a known, state-independent phase that cancels in the residual norm, so that the iterative sequence of exponentials of $O_{kl}$ converges to the same unitary manifold as that generated by the full fermionic operators. This equivalence is important because it allows the RL action pool to be defined directly in terms of sign-free two-qubit mappings rather than the full set of non-local Jordan–Wigner strings [2605.18569].

## 5. Deep Q-network control and constant-scaling time evolution

In the contracted formulation, the RL state at step $n$ is the flattened residual vector $r^{(n)}_{pq,kl}$ over all orbital indices. Its dimension scales as $O(N_o^4)$ with the one-particle basis size $N_o$, but is independent of the number $K$ of targeted states. The action space is discrete, $a\in\{1,\ldots,N_{\rm act}\}$, where $N_{\rm act}$ is the number of distinct two-body operators in the pool. The immediate reward is
$$
R_n=-\bigl\|r^{(n+1)}\bigr\|_2+\lambda\,\Delta E_n,
$$
where $\Delta E_n$ is the ensemble energy drop and $\lambda\ge 0$ regularizes residual versus energy. In the reported benchmarks, the DQN is a feed-forward network with 8 linear layers, hidden width 512, GELU activations, replay buffer size $10^6$, discount $\gamma=0.99$, residual weight $\lambda=0.5$, AdamW optimization with learning rate $2\times 10^{-4}$, batch size $256$, and total episodes approximately $3000$; the maximum steps per episode are 5 for ground/excited-state calculations and up to 20 for H$_3$ time evolution [2605.18569].

The same work introduces a constant-scaling ansatz for time evolution. If the final excited eigenstates are prepared by shared unitaries,
$$
|\Psi_v\rangle = U_N\cdots U_1\,| \Phi_v^{\rm ref}\rangle,
$$
then any superposition $|\Psi(t)\rangle=\sum_v c_v(t)|\Psi_v\rangle$ can be written as
$$
|\Psi(t)\rangle
=U_N\cdots U_1
\Bigl(\sum_v c_v(t)\,|\Phi_v^{\rm ref}\rangle\Bigr)
\equiv U_N\cdots U_1\,|\Phi^{\rm ref}(t)\rangle.
$$
A time-dependent reference state is prepared from a fixed state $|\Phi_0\rangle$ by $N'$ additional unitaries $U_1'(t)\cdots U_{N'}'(t)$, so the full ansatz becomes
$$
U_N\cdots U_1\,U_{N'}'(t)\cdots U_1'(t)\,|\Phi_0\rangle,
$$
with total number of two-body exponentials $M=N+N'$ independent of simulation time $t$ [2605.18569].

For a time step $\Delta t$, the target state is defined as $|\Psi_{\rm target}\rangle=e^{-iH\Delta t}|\Psi(t)\rangle$ by classical eigendecomposition. RL-CQE is then run with reward equal to overlap fidelity until $\langle\Psi_{\rm target}|\Psi(n;t)\rangle\approx 1$, and the number of RL steps per $\Delta t$ is bounded by $M$, remaining constant in $t$. This suggests that the reinforcement-learning policy is used not only for state preparation but also as a mechanism for constraining temporal growth of the circuit ansatz [2605.18569].

## 6. Performance, resource profile, and relation to other eigensolvers

For the original black-box formulation, the reported single-qubit results for random $\hat{\mathcal O}_E$ use reward $r=0.9$ and punishment $p=2/r$. The average fidelity exceeds $0.90$ within fewer than about 10 iterations and exceeds $0.98$ within fewer than about 300 iterations. For two-qubit random operators, all four eigenvectors converge to fidelities above about $0.89$ in about 8,000 iterations for $r=0.9$ and $p\approx 2.22$. In the special Bell-basis case, when $\hat{\mathcal O}_E$ is block-diagonal in Bell states, the algorithm discovers each two-dimensional block independently and reaches $F>0.99$ in only about 1,000 iterations. The same description states that each new eigenvector is learned by a fresh copy of the loop in the subspace orthogonal to previously learned eigenvectors, and that in practice the total number of iterations scales roughly linearly with the dimension $d$, one block at a time, though the per-block cost grows with the subspace size [1906.06702].

The circuit form of one single-qubit loop consists of one black-box gate $e^{-i\tau\hat{\mathcal O}_E}$, one basis rotation $\hat D^{(k)\dagger}$, one computational-basis measurement, one corrective rotation $\hat{\mathcal U}_m$, and one re-preparation of $\hat D^{(k+1)}|0\rangle$. The description further states that circuit depth per iteration scales as $O(1)$ two-qubit and single-qubit gates. In its comparison with variational eigensolvers, the same source emphasizes that RL-CQE uses exactly one single-shot measurement per iteration, stores only the current diagonalizer $\hat D^{(k)}$, absorbs control noise into pseudo-random exploration, and converges toward eigenvectors directly, whereas VQE typically requires many shots to estimate expectation values or overlaps and maintains a classical optimizer history with gradient information [1906.06702].

For the contracted many-body formulation, the measurement cost per iteration is $O(N_o^4)$ because all residual components $r_{pq,kl}$ must be estimated. The ansatz size is typically much smaller than $N_o^4$: the benchmarks report 2–5 steps for H$_2$ and approximately 10–20 for H$_3$, with circuit depth scaling as $\sim M\times d_{\rm 2body}$. On H$_2$ in STO-6G with 4 qubits and 4 singlets over bond lengths $0.5$–$5.0$ Å, energies are within $10^{-3}$ Hartree of FCI across all $R$ using at most 5 RL steps; at $R\in\{0.5,0.7,1.2,3.0\}$ Å, even 2 steps give residual norms $\lesssim 10^{-8}$ and chemical accuracy. On linear H$_3$ in STO-6G with 6 qubits and 4 singlets over $R=1.0$–$3.0$ Å, the energies match exact diagonalization to $\lesssim 10^{-3}$ Hartree, and $\|r(n)\|$ decays to $\lesssim 10^{-8}$ within 5–10 steps. For real-time dynamics on $t\in[0,20]$ a.u. with $\Delta t=0.05$, the reported fidelity is at least $0.999$ in at most 5 RL steps for H$_2$ and at most 20 RL steps for H$_3$, constant for all $t$ [2605.18569].

These results clarify a second common misconception: RL-CQE is not uniformly “measurement-light” across all of its variants. The original semi-autonomous eigensolver uses one single-shot measurement per iteration, but the ACSE-based many-body version pays an $O(N_o^4)$ residual-estimation cost in exchange for compact operator sequences and constant-scaling time-evolution ansätze. The outlook described in the literature includes time-varying reward and punishment rates $r(k),p(k)$, integration of learned block diagonalizers into quantum simulation and quantum chemistry subroutines, continuous-time RL or policy-gradient updates using quantum amplitude estimation, larger molecules beyond STO-6G, hardware-efficient two-body decompositions, open quantum systems, and fully on-device quantum-RL loops for gradient-free autonomous calibration [1906.06702], [2605.18569].

Source: https://www.emergentmind.com/topics/reinforcement-learning-contracted-quantum-eigensolver-rl-cqe