---
title: Wiener KL Divergence in Quantum Control
url: https://www.emergentmind.com/topics/wiener-kl-kl_w
type: topic
---

# Wiener KL Divergence in Quantum Control

to=arxiv_search.search  天天中彩票中大奖  彩神争霸官网  天天中彩票买json
{"query":"KL_W Wiener KL QMaxCal path-space regularization open quantum control arXiv 2606.19947", "max_results": 5}
to=arxiv_search.search  ахәыҷി  北京赛车女json
{"query":"Information Projection on Banach spaces with Applications to State Independent KL-Weighted Optimal Control Wiener space arXiv 2009.03504", "max_results": 5}
to=arxiv_search.search  福利彩票天天彩պեսjson
{"query":"A representation for the Kantorovich--Rubinstein distance on the abstract Wiener space arXiv 1608.07124", "max_results": 5}
Wiener KL, usually written \(KL_W\), denotes in the cited arXiv literature a Kullback–Leibler divergence defined on path space with a Wiener reference measure. In the exact notation \(KL_W\) used by QMaxCal, it is the KL divergence between the measurement-record distribution induced by a controlled stochastic Schrödinger evolution and standard Wiener measure, so it penalizes the observable trajectory-level consequences of control on decoherence channels rather than the control waveform itself [2606.19947]. In a broader Wiener-space context, this usage belongs to the standard relative-entropy framework for Gaussian and path measures; in particular, for Gaussian shift measures on classical Wiener space, the KL divergence reduces to a quadratic Cameron–Martin energy [2009.03504].

## 1. Definition in open quantum control

In QMaxCal, an open quantum system is continuously monitored under the diffusive or homodyne unravelling. Each decoherence channel \(k\) produces a measurement record
\[
dI_k(t)=\alpha_k(t)\,dt+dW_k(t),
\qquad
\alpha_k(t)=\langle \psi(t)\mid (L_k+L_k^\dagger)\mid \psi(t)\rangle,
\]
where the \(dW_k\) are independent Wiener increments. The drift \(\alpha_k(t)\) is the observable signature of the state’s exposure to channel \(k\) [2606.19947].

With \(P_\theta\) denoting the path measure induced by the controlled stochastic Schrödinger equation and \(P_W\) denoting standard Wiener measure, the Wiener KL is defined by
\[
\mathrm{KL}_W:=\mathrm{KL}(P_\theta\|P_W)
=
\frac12\sum_{k=1}^K
\mathbb{E}_{P_\theta}\!\left[\int_0^T \alpha_k^{(\theta)}(t)^2\,dt\right].
\]
Here \(P_W\) is the zero-drift reference process, \(dY_t=dW_t\). Accordingly, \(KL_W\) is literally the KL divergence between the control-induced record distribution and pure Brownian noise [2606.19947].

This definition makes the regularizer state- and trajectory-sensitive. It does not act on \(\{u_a^{(\theta)}(t)\}\) directly, but on the path-space distribution generated after the controls, Hamiltonian, Lindblad operators, and monitoring scheme have been fixed. The paper’s stated interpretation is that minimizing \(KL_W\) encourages the controlled trajectory to move into regions where decoherence has little or no effect, most importantly into a joint kernel or decoherence-free region when such a region is reachable [2606.19947].

## 2. Derivation from Girsanov’s theorem

The central derivation uses Girsanov’s theorem for diffusions with identical noise structure and different drifts. If two path measures \(P^{(1)}\) and \(P^{(2)}\) have drift difference \(\Delta\alpha_k(t)=\alpha_k^{(1)}(t)-\alpha_k^{(2)}(t)\), then the paper states the scalar change-of-measure formula as
\[
\log\frac{dP^{(1)}}{dP^{(2)}}=
\int_0^T \Delta a_t\, dB_t^{(2)}-\frac12\int_0^T (\Delta a_t)^2\,dt,
\]
and the \(K\)-channel quantum measurement-record analogue as
\[
\mathrm{KL}\!\left(P^{(1)}\|P^{(2)}\right)
=
\frac12\sum_{k=1}^K
\mathbb{E}_{P^{(1)}}\!\left[\int_0^T \Delta\alpha_k(t)^2\,dt\right].
\]
Choosing the reference drift to be zero immediately yields the Wiener KL formula above [2606.19947].

The derivation depends on a specific structural condition: the measurement record must be an Itô diffusion with unit diffusion coefficient independent of control. In the QMaxCal construction, two quantum evolutions share the same Lindblad operators \(\{L_k\}\), so the path measures differ only in the drift of the measurement record. Under exactly that hypothesis, the KL functional becomes a closed-form, differentiable expectation of a quadratic drift functional, and the paper emphasizes that it is straightforward to estimate by Monte Carlo over stochastic Schrödinger trajectories [2606.19947].

A plausible implication is that \(KL_W\) inherits the computational advantages typical of quadratic-energy path-space penalties while remaining tied to an explicitly observable signal, namely the monitored record itself rather than an auxiliary latent quantity.

## 3. Relation to Wiener-space KL on classical Wiener space

The QMaxCal construction is not an isolated use of KL on Wiener path space. On classical Wiener space \((\mathcal C_0[0,T],\mu_0)\), the Cameron–Martin theorem gives explicit Radon–Nikodym derivatives for Gaussian shift measures \(T_h^\ast(\mu_0)\), and the associated KL divergence is
\[
D_{KL}(\mu\|\mu_0):=
E_\mu\!\left[\log\!\left(\frac{d\mu}{d\mu_0}\right)\right].
\]
If \(\mu_1=T_{h_1}^\ast(\mu_0)\) and \(\mu_2=T_{h_2}^\ast(\mu_0)\) with \(h_1,h_2\in\mathcal H_{\mu_0}\), then
\[
D_{KL}(\mu_1\|\mu_2)=\frac12\|h_1-h_2\|_{\mu_0}^2.
\]
The same paper formulates an information projection problem over shift measures and shows that, in this Wiener-space setting, KL projection, KL-weighted optimal control, and minimization of an Onsager–Machlup function are equivalent formulations [2009.03504].

This places \(KL_W\) in a larger measure-theoretic lineage. The object is not a new divergence axiomatically; it is the ordinary KL divergence specialized to a Wiener reference structure. In the QMaxCal case the reference is standard Wiener measure on the monitored record; in the Banach- and Wiener-space control formulation, the reference is the law of standard Brownian motion and admissible measures are Cameron–Martin shifts [2009.03504].

A nearby but distinct result is the representation of the Kantorovich–Rubinstein distance on abstract Wiener space. There the relevant functional is not KL divergence but the \(1\)-Wasserstein distance \(W_1\), represented via the divergence or extended stochastic integral operator:
\[
W_1(v_0,v_1)=
\inf_{Iu=\frac{d(v_1-v_0)}{dp}}
\int_X |u(x)|\,p(dx).
\]
This is conceptually adjacent because it also turns a path-space measure discrepancy into a variational norm minimization, but it is not a Wiener KL [1608.07124].

## 4. Role inside QMaxCal and comparison with \(R_{\mathrm{DV}}\)

QMaxCal incorporates the Wiener KL into a fidelity-plus-regularization objective
\[
\mathcal L(\theta)
=
1-\mathbb{E}_{P_\theta}[F(\psi(T))]
+\lambda_W\,\mathrm{KL}_W
+\lambda_{DV}\,R_{\mathrm{DV}}
+\lambda_{\mathrm{flu}}\,\mathrm{Flu},
\]
where
\[
F(\psi(T))=\bigl|\langle \phi_{\mathrm{target}}\mid \psi(T)\rangle\bigr|^2,
\qquad
\mathrm{Flu}=\sum_a\int_0^T |u_a^{(\theta)}(t)|^2dt.
\]
The paper notes that experiments often use either \(KL_W\) or \(R_{\mathrm{DV}}\), not both simultaneously, depending on the benchmark [2606.19947].

The drift-variance regularizer \(R_{\mathrm{DV}}\) uses a different reference class. Instead of zero-drift Wiener measure, it chooses the closest constant-drift process:
\[
dY_t=c\,dt+dW_t,
\qquad c\in\mathbb R^K,
\]
and obtains
\[
R_{\mathrm{DV}}
=
\frac12\sum_{k=1}^K
\mathbb{E}_{P_\theta}\!\left[
\int_0^T
\left(\alpha_k^{(\theta)}(t)-\bar\alpha_k\right)^2 dt
\right],
\]
with
\[
\bar\alpha_k
=
\frac1T
\mathbb{E}_{P_\theta}\!\left[\int_0^T \alpha_k^{(\theta)}(t)\,dt\right].
\]
The paper’s distinction is explicit: \(KL_W\) penalizes drift magnitude relative to zero drift, whereas \(R_{\mathrm{DV}}\) penalizes fluctuations around a constant drift regardless of whether that constant is zero [2606.19947].

This leads to different inductive biases. \(KL_W\) has a stronger bias toward the joint kernel \(\bigcap_k \ker(L_k)\), where drift vanishes. By contrast, \(R_{\mathrm{DV}}\) vanishes on any decoherence-free subspace where the drift is constant in time and across realizations, even if that constant is nonzero. The paper therefore treats \(KL_W\) as the stronger but more specialized regularizer, and \(R_{\mathrm{DV}}\) as the more generally applicable one [2606.19947].

## 5. Benchmark behavior and reported performance

Across the reported single- and multi-qubit benchmarks, as well as a multi-qubit chain calibrated to a published snapshot of the IBM Kingston processor, the regularizers outperform unregularized gradient-based and reinforcement-learning baselines along final-state fidelity, robustness to mismatch in the assumed noise model, and occupation of forbidden states. The abstract reports gains growing from \(+17\) percentage points at training noise to \(+27\) percentage points under \(2.5\times\) noise mismatch, reduction of infidelity by up to \(50\%\), and approximately \(16\%\) gains on the calibrated IBM Kingston chain [2606.19947].

| Benchmark | Reported \(KL_W\) outcome | Structural interpretation |
|---|---|---|
| Single-qubit amplitude damping | At \(\gamma T=2\), baseline fidelity \(0.9715\), Wiener KL \(0.9810\), \(R_{\mathrm{DV}}\) \(0.9458\), PPO \(0.9684\) | Ground state is the kernel |
| STIRAP | At \(\gamma T=10\), fidelity stays essentially unchanged (\(0.9797\) baseline, \(0.9790\) Wiener KL), but peak \(|e\rangle\) population drops from \(0.097\) to \(0.043\) | Suppresses occupancy of the lossy intermediate state |
| Diamond system | At \(\gamma=2\), baseline \(F=0.665\), Wiener KL with \(\lambda_W=0.5\) gives \(0.810\), with \(\lambda_W=5\) gives \(0.834\), PPO gives \(0.605\) | Routes population through a safe kernel state |

The amplitude-damping example is the clearest demonstration of the mechanism. The Lindblad operator is \(L=\sqrt{\gamma}\,\sigma_-\), so the kernel is the ground state \(|0\rangle\). The paper reports that the trajectory-variance collapse is large: the time-integrated population variance falls from \(0.0321\) to \(0.0021\), approximately a \(15\times\) reduction, and the drift-squared integral \(\frac12\int \langle \alpha(t)^2\rangle\,dt\) drops from \(0.230\) to \(0.031\) at \(\gamma T=2\) [2606.19947].

The STIRAP benchmark illustrates a different regime. Here Wiener KL does not materially improve the final fidelity, but it reduces occupation of the lossy excited state. The paper gives a \(56\%\) reduction in peak \(|e\rangle\) population and a \(35\%\) reduction in time-integrated \(|e\rangle\) exposure at \(\gamma T=10\) [2606.19947]. This suggests that \(KL_W\) can improve protocol quality even when the terminal fidelity is already near saturation.

The diamond-system robustness test shows the strongest mismatch effect. Trained at \(\gamma_{\rm train}=2\) and evaluated at \(\gamma_{\rm test}=5\), the reported fidelities are \(0.393\) for the baseline, \(0.665\) for Wiener KL, and \(0.288\) for PPO, with the paper noting that the gain grows monotonically with test-noise strength and reaches \(+27\) percentage points over the baseline at the hardest mismatch setting [2606.19947].

## 6. Terminological scope and ambiguity of the notation

The notation \(KL_W\) is not uniform across arXiv literatures. In the QMaxCal usage, it denotes the Wiener KL regularizer just described [2606.19947]. In work on classical Wiener space, the operative object is still the ordinary KL divergence \(D_{KL}\) on Gaussian shift measures rather than a separately named \(KL_W\) invariant [2009.03504]. In abstract Wiener-space optimal transport, the relevant discrepancy is \(W_1\), not KL divergence [1608.07124].

The term “Wiener” is also heavily overloaded outside probability and control. In graph theory it usually denotes the Wiener index
\[
W(G)=\sum_{\{u,v\}\subseteq V(G)} d(u,v),
\]
or related invariants such as the terminal Wiener index or Wiener complexity \(C_W(G)\). Those papers explicitly do not define a KL-type quantity \(KL_W\) [1305.6196; 1905.01699; 2305.07405]. Separately, the shorthand \(KL_W\) may also refer to the KL divergence between two Weibull distributions,
\[
D_{KL}\!\left(\mathrm{Weibull}(k_1,l_1)\parallel \mathrm{Weibull}(k_2,l_2)\right),
\]
which is unrelated to Wiener measure despite the similar subscript notation [1310.3713].

Accordingly, the expression “Wiener KL” is context-dependent. As an exact named object, it refers most directly to the QMaxCal path-space regularizer
\[
\mathrm{KL}_W
=
\frac12\sum_{k=1}^K
\mathbb{E}_{P_\theta}\!\left[\int_0^T \alpha_k(t)^2\,dt\right],
\]
while its broader mathematical background is the standard KL divergence on Gaussian and Wiener path spaces rather than a separate divergence family [2606.19947].

Source: https://www.emergentmind.com/topics/wiener-kl-kl_w