---
title: Core-Optimized Orbitals (COO) in Spectroscopy and CI
url: https://www.emergentmind.com/topics/core-optimized-orbitals-coo
type: topic
---

# Core-Optimized Orbitals (COO) in Spectroscopy and CI

Core-Optimized Orbitals (COO) denotes an orbital optimization strategy in which the one-particle basis is variationally adapted to a targeted electronic structure problem rather than inherited unchanged from a reference Hartree–Fock or Kohn–Sham solution. In the arXiv literature, the acronym appears in two closely related but technically distinct settings. In core-level spectroscopy, COO refers to the orbital sets obtained by $\Delta$SCF or ROKS optimization for a $1s\!\to\!\text{virtual}$ excitation within an orbital-optimized DFT framework combined with the spin-free exact two-component (X2C) Hamiltonian [2111.08405]. In strongly correlated many-body calculations, COO denotes a co-optimization of sparse CI coefficients and orbital rotations, with the aim of absorbing a large fraction of the dynamical correlation directly into the single-particle basis [2605.22977]. In both usages, the central object is a unitary orbital rotation $U=e^\kappa$, and the central principle is that orbital flexibility can materially alter the compactness and accuracy of the many-electron ansatz.

## 1. Terminological scope and conceptual role

The term COO is used for two variational constructions that share a common mathematical mechanism but target different physical regimes. In the spectroscopy setting, one constructs a core-hole orbital manifold for K-edge calculations by enforcing a target occupancy pattern and optimizing the orbitals in the presence of a spin-free X2C one-electron operator [2111.08405]. In the sparse-CI setting, one seeks the ground state by varying both a selected determinant expansion and the orbital basis, so that the same many-body accuracy can be reached with substantially fewer determinants [2605.22977].

| Context | Optimization target | Principal outcome |
|---|---|---|
| OO-DFT/X2C spectroscopy | A $1s$ core-hole determinant or ROKS state | Accurate K-edge energies and spectra |
| TrimCI+COO many-body theory | Sparse CI coefficients and orbital rotations | Large compression in determinant count |

This dual usage is not merely terminological. In both cases, the orbital basis is treated as an active variational degree of freedom, and the orbital optimization is coupled to an explicitly constrained electronic ansatz. A plausible implication is that COO is best understood as a family of orbital-adaptive methods rather than a single formalism.

## 2. Variational structure in core-level spectroscopy

In the OO-DFT/X2C formulation for spectroscopy, the simplest $\Delta$SCF approach targets a single-determinant wavefunction $\Phi^*$ in which one or two $1s$ molecular orbitals on the chosen atom is emptied, and one electron is placed in a chosen virtual [2111.08405]. The total energy functional is
\[
E[\Phi^*]
=
\bigl\langle\Phi^*\bigl|\hat H_{\rm X2C}+\hat J[\rho^*]+\hat V_{\rm XC}[\rho^*]\bigr|\Phi^*\bigr\rangle,
\]
where $\hat H_{\rm X2C}$ is the spin-free X2C one-electron operator, $\hat J[\rho^*]$ is the Coulomb operator built from the density $\rho^*(\mathbf r)$, and $\hat V_{\rm XC}[\rho^*]$ is the chosen DFT exchange-correlation potential.

Because the orbitals $\{\phi_i\}$ must remain orthonormal and must carry the desired occupation pattern, the formulation introduces a matrix of Lagrange multipliers $\boldsymbol\varepsilon$ enforcing
\[
\langle\phi_p|\phi_q\rangle=\delta_{pq},
\]
together with additional constraints $C_\alpha[\{\phi\}]$ that enforce that a particular orbital on atom A remains empty. The corresponding Lagrangian is
\[
\mathcal L[\{\phi\},\boldsymbol\varepsilon,\{\lambda\}]
=
E[\Phi^*]
-
\sum_{pq}\varepsilon_{pq}\Bigl(\langle\phi_p|\phi_q\rangle-\delta_{pq}\Bigr)
-
\sum_\alpha \lambda_\alpha C_\alpha[\{\phi\}],
\]
and the stationarity conditions
\[
\frac{\partial \mathcal L}{\partial \phi_p}=0,\qquad
\frac{\partial \mathcal L}{\partial \varepsilon_{pq}}=0,\qquad
\frac{\partial \mathcal L}{\partial \lambda_\alpha}=0
\]
yield the usual DFT-like Fock equations with the desired core-hole occupancy.

This formulation makes explicit that the spectroscopy version of COO is not a perturbative excited-state correction layered onto a fixed ground-state calculation. Rather, it is a constrained variational optimization in which the target core-excited occupancy is embedded directly into the orbital equations.

## 3. Relativistic Hamiltonian, orbital optimization, and practical construction

For elements heavier than Ne, the spectroscopy framework combines orbital optimization with a scalar-relativistic treatment based on the spin-free X2C model [2111.08405]. The construction begins from the four-component one-electron Dirac equation in a restricted-kinetic-balance basis,
\[
\begin{pmatrix}
V & T\\[6pt]
T & \tfrac1{4c^2}W_{\rm SF}-T
\end{pmatrix}
\begin{pmatrix}
C_L\\ C_S
\end{pmatrix}
=
E
\begin{pmatrix}
S & 0\\[3pt]
0 & \tfrac1{2c^2}T
\end{pmatrix}
\begin{pmatrix}
C_L\\ C_S
\end{pmatrix},
\]
with $T$, $V$, and $S$ the usual kinetic, nuclear-attraction, and overlap integrals, and $W_{\rm SF}=\mathbf p\,V\,\mathbf p$ the spin-free piece of $(\sigma\!\cdot\!p)V(\sigma\!\cdot\!p)$. A one-step exact two-component unitary decoupling then yields an effective two-component Hamiltonian
\[
\hat h_{\rm X2C}=\hat T_{\rm X2C}+\hat V_{\rm X2C},
\]
with
\[
\begin{aligned}
T_{\rm X2C} &= R^\dagger\Bigl(TX+X^\dagger T-X^\dagger TX\Bigr)R,\\
V_{\rm X2C} &= R^\dagger\Bigl(V+\tfrac1{4c^2}X^\dagger W_{\rm SF}X\Bigr)R,
\end{aligned}
\]
where $X=C_S C_L^{-1}$, $\tilde S = S + \tfrac1{2c^2}X^\dagger T X$, and
\[
R=S^{-1/2}\Bigl(S^{-1/2}\tilde S S^{-1/2}\Bigr)^{-1/2}S^{1/2}.
\]

The orbital optimization itself is expressed through an infinitesimal unitary rotation generated by an anti-Hermitian matrix $\boldsymbol\kappa$,
\[
\phi_p^\prime=e^{\boldsymbol\kappa}\phi_p
=
\phi_p+\sum_q \kappa_{qp}\phi_p+\mathcal O(\kappa^2),
\]
together with the excited-state energy gradient
\[
g_{pq}=\frac{\partial E[\Phi^*]}{\partial \kappa_{pq}}\Bigg|_{\kappa=0}
\]
and, optionally, the orbital-rotation Hessian
\[
H_{pq,rs}
=
\frac{\partial^2 E[\Phi^*]}{\partial \kappa_{pq}\partial \kappa_{rs}}\Bigg|_{\kappa=0}.
\]
The workflow described for practical optimization uses quasi-Newton or trust-region methods rather than an explicit full Hessian. A convenient robust algorithm is square-gradient minimization, initialized with a guess that has the $1s$ core hole on the target atom, for example by the maximum-overlap method IMOM. The loop proceeds by building $\rho^*$, $J$, and $V_{\rm XC}$; forming
\[
F=h_{\rm X2C}+J+V_{\rm XC};
\]
computing the orbital-rotation gradient; proposing $\delta\kappa=-\mathbf B^{-1}g$ with an approximate Hessian such as BFGS; enforcing a trust region; optionally level-shifting the virtual block by $\Delta$ $(0.1$–$0.3\,{\rm a.u.})$ to avoid collapse to the ground state; and rotating the orbitals. The stated convergence controls are $\|g\|<\varepsilon$, for example $10^{-5}\,{\rm a.u.}$, and $\Delta E<10^{-7}\,{\rm a.u.}$.

The practical recipe is correspondingly specific: a radial grid of $99$ points and a Lebedev angular grid of $590$ points; a decontracted triple-$\zeta$ “pcX-n” or core-valence Dunning set on the core-hole atom, such as aug-pcX-2 or cc-pCVTZ; aug-pcseg-1 on all other atoms; and core-hole targeting by IMOM followed by freezing that occupation throughout the SCF loop, or equivalently by adding a small penalty functional $\lambda\,(n_{1s}-N_{\rm target})^2$ to the energy. Square-gradient minimization is stated to guarantee descent of $\|g\|^2$ and avoid “variational collapse.”

## 4. COO as a co-optimized sparse-CI orbital basis

In the many-body setting, COO is defined by simultaneous variation of a sparse CI expansion and the single-particle basis [2605.22977]. Starting from any initial orthonormal orbitals $\{\phi_p\}$, one applies a unitary rotation $U(\kappa)=e^\kappa$ generated by a real antisymmetric matrix $\kappa$, with $\kappa_{pq}=-\kappa_{qp}$ and $n(n-1)/2$ independent angles, to define
\[
\tilde \phi_p=\sum_q U_{qp}\phi_q,\qquad U=e^\kappa.
\]
The CI ansatz over a selected set of determinants $\mathcal V$ is then
\[
|\Psi(\kappa,c)\rangle
=
\sum_{I\in \mathcal V} c_I\,|D_I(\{\tilde \phi_p\})\rangle,
\]
and the variational objective is the Rayleigh quotient
\[
E[\kappa,c]
=
\frac{\langle \Psi(\kappa,c)|\hat H|\Psi(\kappa,c)\rangle}
{\langle \Psi(\kappa,c)|\Psi(\kappa,c)\rangle}
\longrightarrow \min_{\kappa,c},
\]
subject to $\|c\|=1$. Equivalently,
\[
E[\kappa,c]
=
\sum_{pq} h_{pq}(\kappa)\,D_{qp}
+
\tfrac12\sum_{pqrs} V_{pqrs}(\kappa)\,D_{rspq},
\]
where $D$ and $D_2$ are the $1$- and $2$-RDMs of $|\Psi(\kappa,c)\rangle$.

The baseline sparse ansatz is TrimCI, which builds a variational space $\mathcal V$ of $N_{\rm det}$ determinants by alternately “expanding” in large Hamiltonian couplings $|H_{IJ}c_J|>\theta$ and “trimming” by retaining the largest weight configurations after block-diagonal and global diagonalization. Its wavefunction is
\[
|\Psi_{\rm TrimCI}\rangle=\sum_{I\in \mathcal V} c_I\,|D_I\rangle,
\qquad
N_{\rm params}=|\mathcal V|\equiv N_{\rm det}.
\]
At fixed orbitals, stationarity with respect to the CI coefficients gives the projected eigenproblem
\[
\sum_{J\in \mathcal V} H_{IJ}(\kappa)c_J=E\,c_I,
\qquad
\frac{\partial E}{\partial c_I}=0.
\]
For the orbital variables, the generator
\[
\hat \kappa=\sum_{p<q}\kappa_{pq}(a_p^\dagger a_q-a_q^\dagger a_p)
\]
leads to the gradient
\[
g_{pq}
=
\frac{\partial E}{\partial \kappa_{pq}}
=
\langle \Psi|[a_p^\dagger a_q-a_q^\dagger a_p,H]|\Psi\rangle,
\]
which can be written in closed form from the $1$- and $2$-RDMs.

The joint update uses BFGS on $\kappa$, but every line-search trial step immediately re-optimizes $\{c\}$ via Davidson on the fixed core. The stated purpose is that the trial energy $E[\kappa+\alpha d,c(\kappa+\alpha d)]$ sees the full coupled response rather than a linear surrogate. The algorithm is staged: a global COO phase with repeated TrimCI core construction and BFGS orbital updates; a local-refinement phase in which the determinant count is grown gradually; and an expansion phase in which $\kappa$ is frozen while the CI space is enlarged further, optionally with semistochastic PT2 correction. Practical notes include using a small initial core $N_0\sim10^2$–$10^3$, seeding with HF and random determinants, and preferring BFGS over L-BFGS and CG for the orbital optimization.

## 5. Accuracy, compactness, and asymptotic scaling

For K-edge spectroscopy with OO-DFT/X2C, the reported benchmark is that over a broad set of third-row K-edges (Ne, Si–Cl, Ar), ROKS or $\Delta$SCF + X2C + SCAN yields RMS errors of $0.4$–$0.5\,{\rm eV}$ versus experiment, while conventional TDDFT routinely errs by $10$–$50\,{\rm eV}$ for heavy-atom K-edges and must be post-shifted [2111.08405]. The abstract further states that OO-DFT/X2C reproduces experimental spectra quite well, sans empirical shifts for alignment, and that all steps remain $O(N^3)$ or at worst $O(N^4)$, with only a modest prefactor overhead of $\sim 2\times$ relative to ground-state KS-DFT or TDDFT.

For the sparse-CI COO method, the central numerical claim is that the apparent parameter-heaviness of CI is strongly basis-dependent [2605.22977]. In a localized-MO basis, reaching $\sim1\,{\rm mHa}$ accuracy on $[\mathrm{Fe}_4\mathrm S_4]$ $(54e,36o)$ requires
\[
N_{\rm det}^{\rm LMO}\approx 3\times 10^{14},
\]
whereas in the COO basis the same accuracy is reached with
\[
N_{\rm det}^{\rm COO}\approx 10^9.
\]
The paper describes this as a factor $\sim 3\times 10^5$ compression in determinant count. At matched accuracy on the same cluster, TrimCI+COO is reported to be $8\times$ more compact than the largest unrestricted-DMRG benchmark and $25\times$ more compact with PT2. Across the iron-sulfur series, from $[\mathrm{Fe}_2\mathrm S_2]$ $(30e,20o)$ to the P-cluster $(114e,73o)$, TrimCI+COO is reported to be $10$–$100\times$ more compact than SU(2)-adapted DMRG with entanglement-minimized orbitals at matched accuracy.

| Setting | Metric | Reported result |
|---|---|---|
| OO-DFT/X2C K-edge spectroscopy | RMS error vs experiment | $0.4$–$0.5\,{\rm eV}$ |
| OO-DFT/X2C vs TDDFT | Heavy-atom K-edge error | TDDFT routinely errs by $10$–$50\,{\rm eV}$ |
| $[\mathrm{Fe}_4\mathrm S_4]$ TrimCI+COO | Determinants at matched accuracy | $\approx 10^9$ vs $\approx 3\times10^{14}$ in LMO basis |
| $[\mathrm{Fe}_4\mathrm S_4]$ vs UDMRG | Compactness at matched accuracy | $8\times$; $25\times$ with PT2 |
| Iron-sulfur series vs SU(2)-DMRG+EMO | Compactness at matched accuracy | $10$–$100\times$ |

The many-body paper also gives an asymptotic framing: if $E(N)-E_{\rm FCI}\sim aN^{-\alpha}$, then orbitals that increase $\alpha$ improve the asymptotic scaling of $N$ required for a fixed $\Delta E$. This suggests that COO is not only a finite-size compression device but also a means of changing the effective convergence law of sparse CI.

## 6. Factorization of gains, limitations, and open directions

A distinctive analysis in the many-body COO work is the separation of the total advantage into an orbital-basis gain and an ansatz gain using a tunable Hubbard-on-graph model on $L=8$ sites at half-filling with $U/t=4$ and topology parameter $\alpha\in[0,1]$ [2605.22977]. At fixed $\Delta E<0.1\,t$, the paper defines
\[
G_{\rm orb}=\frac{N_{\rm noCOO}}{N_{\rm COO}},
\qquad
G_{\rm ans}=\frac{N_{\rm DMRG}}{N_{\rm noCOO}},
\]
so that
\[
\frac{N_{\rm DMRG}}{N_{\rm COO}}=G_{\rm orb}\times G_{\rm ans}.
\]
As $\alpha$ grows from $0\to1$, $G_{\rm orb}$ rises from $\sim1.3\to3.3$ and $G_{\rm ans}$ from $\sim1\to3.5$, giving total $\mathrm{DMRG}/\mathrm{COO}$ from $\sim1\to12\times$. Within the paper’s interpretation, the ansatz gain captures multi-center entanglement that resists MPS localization.

The limitations are likewise explicit. In the many-body setting, accuracy remains variational; very large $N_{\rm det}$ expansions can be costly, though COO cuts $N_{\rm det}$ by $>10^3$; line-search re-diagonalization requires the core to remain small; selected-CI matvecs at $N_{\rm det}>10^9$ require distributed Davidson; and orbital rotations cost $O(n_{\rm orb}^4)$ per gradient step, which may dominate for very large active spaces. The proposed future directions are to apply COO to larger targets such as the FeMo cofactor and P-cluster refinement, combine COO with tensor networks, coupled-cluster, and neural quantum states, investigate hybrid schemes in which small COO cores seed VQE or quantum Monte Carlo, explore automated basin discovery, and extend the multi-node GPU Davidson to $N_{\rm det}\ge 10^{12}$.

In the spectroscopy setting, the paper reports that it also explored K and L edges of $3d$ transition metals to identify limitations of the OO-DFT/X2C approach in modeling the spectra of heavier atoms [2111.08405]. More broadly, the spectroscopy workflow makes clear that the method’s success depends on maintaining the target core occupancy during optimization and on controlling variational collapse through occupation freezing, penalty functionals, trust regions, or level shifts.

Taken together, these two lines of work establish COO as a general orbital-adaptive paradigm. In one application, it produces core-hole orbitals for relativistic K-edge spectroscopy at roughly ground-state DFT cost; in the other, it uses a small variational core to learn an orbital basis that can compress sparse CI by three to five orders of magnitude. A plausible unifying interpretation is that COO transfers correlation or excitation specificity from the many-electron expansion into the orbital manifold itself, thereby changing which degrees of freedom must be represented explicitly.

Source: https://www.emergentmind.com/topics/core-optimized-orbitals-coo