---
title: 'Euclidean Path-Integral Control: Theory & Applications'
url: https://www.emergentmind.com/topics/euclidean-path-integral-control
type: topic
---

# Euclidean Path-Integral Control: Theory & Applications

Searching arXiv for the cited Euclidean/path-integral control papers to ground the article.
arxiv_search(query="2209.13733 Euclidean path integral control", max_results=5)
arxiv_search(query="2209.13733 Euclidean path integral control", max_results=5)
Euclidean Path-Integral Control is the class of stochastic optimal control problems in which the Hamilton–Jacobi–Bellman equation can be linearized by an exponential, or desirability, transform of the value function, yielding a backward Kolmogorov or heat-type PDE and an equivalent Feynman–Kac path integral with a real Euclidean action [1208.2523]. In the same literature, this class is identified with linearly-solvable stochastic optimal control, KL-control, and closely related path-integral formulations; its canonical output is an optimal feedback law expressed through the spatial gradient of the desirability function rather than through direct solution of a nonlinear value-function PDE [1406.4026]. The resulting framework has been extended from fully observed control-affine diffusions to generalized costs, state-dependent feedback synthesis, reproducing-kernel estimators, Koopman-based representations, partially observed belief-space control with controlled sensing, and nonlinear application domains including epidemiological control [1406.7869].

## 1. Linear-solvable structure

The standard Euclidean Path-Integral Control problem is posed for a controlled Itô diffusion
\[
d\mathbf{x}_t = \mathbf{f}(\mathbf{x}_t,t)\,dt + \mathbf{B}(\mathbf{x}_t,t)\,\mathbf{u}_t\,dt + \boldsymbol{\sigma}(\mathbf{x}_t,t)\,d\mathbf{W}_t,
\]
with finite-horizon cost
\[
J^{\mathbf{u}}(\mathbf{x}_t,t)=\mathbb{E}\left[\Phi(\mathbf{x}_T) + \int_t^T \Big(q(\mathbf{x}_s,s) + \tfrac12\mathbf{u}_s^\top \mathbf{R}(\mathbf{x}_s,s)\mathbf{u}_s\Big) ds\right].
\]
The defining structural assumption is the control–noise matching condition
\[
\boldsymbol{\Sigma}(\mathbf{x},t)\,\boldsymbol{\Sigma}(\mathbf{x},t)^\top
=
\lambda\,\mathbf{B}(\mathbf{x},t)\,\mathbf{R}^{-1}\,\mathbf{B}(\mathbf{x},t)^\top,
\]
or, in equivalent notation used elsewhere in the literature,
\[
B\,B^\top=\lambda\,U\,R^{-1}U^\top,
\qquad
\lambda I = R\nu.
\]
This coupling between diffusion, control authority, and quadratic control cost is the precise condition under which the nonlinear HJB becomes linear after exponentiation of the value function [1406.7869].

In matched formulations, control and noise act through the same channels. One consequence stated explicitly in the linearly-solvable literature is that no control is allowed in noiseless coordinates; if a coordinate is not directly diffused, it cannot be independently actuated without violating the matching condition [2012.05514]. This restriction is the price paid for linear solvability. It also clarifies why Euclidean Path-Integral Control is not a general method for arbitrary stochastic control problems, but a structurally specialized one.

The same basic architecture appears in several variants. In reproducing-kernel formulations, the dynamics are written as
\[
d x_t = f(x_t,t)\,dt + B(x_t,t)u_t\,dt + B(x_t,t)d\xi_t,
\]
again forcing control and noise through the same subspace [1208.2523]. In matched state-dependent feedback formulations, the diffusion is written as
\[
dX^u(t)=b(t,X^u(t))\,dt+\sigma(t,X^u(t))\big(u(t,X^u(t))\,dt+dW(t)\big),
\]
with the same alignment between control channels and noise channels [1406.4026]. Across these notational differences, the core object is unchanged: a stochastic control problem whose nonlinear optimization becomes a linear propagation problem after the appropriate transform.

## 2. Desirability transform, linear PDE, and Feynman–Kac representation

Starting from the stochastic HJB equation and minimizing over the quadratic control term gives the feedback minimizer
\[
\mathbf{u}^*(\mathbf{x},t)
=
-\,\mathbf{R}^{-1}\mathbf{B}(\mathbf{x},t)^\top \nabla V(\mathbf{x},t),
\]
with the nonlinear HJB reduced to
\[
-\partial_t V
=
q+\nabla V\cdot \mathbf{f}
+\tfrac12\mathrm{Tr}\big(\boldsymbol{\Sigma}\boldsymbol{\Sigma}^\top\nabla^2V\big)
-\tfrac12\nabla V^\top \mathbf{B}\mathbf{R}^{-1}\mathbf{B}^\top\nabla V
\]
under standard notation [1406.7869]. Defining the desirability function
\[
\psi(\mathbf{x},t)=\exp\!\big(-V(\mathbf{x},t)/\lambda\big)
\]
cancels the gradient-square term when the matching condition holds. The resulting PDE is linear:
\[
-\partial_t \psi
=
\frac{1}{\lambda}q(\mathbf{x},t)\,\psi
-\nabla\psi\cdot \mathbf{f}(\mathbf{x},t)
-\tfrac12\mathrm{Tr}\big(\boldsymbol{\Sigma}\boldsymbol{\Sigma}^\top\nabla^2\psi\big),
\qquad
\psi(\mathbf{x},T)=e^{-\Phi(\mathbf{x})/\lambda}.
\]
Equivalent sign conventions appear across the literature, but the substantive point is invariant: the HJB becomes a linear backward equation for \(\psi\) [1406.4026].

The optimal control is then recovered directly from the desirability gradient:
\[
\mathbf{u}^*(\mathbf{x},t)
=
\lambda\,\mathbf{R}^{-1}\mathbf{B}(\mathbf{x},t)^\top \nabla \log \psi(\mathbf{x},t).
\]
In notation used in other formulations, this appears as
\[
\bm{u}^*(\bm{x},t)
=
\lambda\,R^{-1}U^\top\nabla\log\psi(\bm{x},t),
\]
or, when the cost has already been scaled by \(\lambda\), without the prefactor [2012.05514].

The linear PDE admits a Feynman–Kac, hence Euclidean path-integral, representation under the uncontrolled or reference dynamics
\[
d\tilde{\mathbf{x}}_s
=
\mathbf{f}(\tilde{\mathbf{x}}_s,s)\,ds
+
\boldsymbol{\sigma}(\tilde{\mathbf{x}}_s,s)\,d\mathbf{W}_s.
\]
The desirability is
\[
\psi(\mathbf{x},t)
=
\mathbb{E}^{\mathrm{ref}}
\!\left[
\exp\!\Big(
-\tfrac{1}{\lambda}
\Big(
\Phi(\tilde{\mathbf{x}}_T)
+
\int_t^T q(\tilde{\mathbf{x}}_s,s)\,ds
\Big)
\Big)
\,\Big|\,\tilde{\mathbf{x}}_t=\mathbf{x}
\right].
\]
This is the sense in which the method is “Euclidean”: path weights are real, positive, and Boltzmann-like, \(\exp(-S/\lambda)\), rather than oscillatory amplitudes [1505.01874]. Several papers make the same point through the language of Wick rotation, Euclidean action, or linearly-solvable control [2002.09394].

## 3. Sampling, importance control, and feedback representations

The Feynman–Kac representation turns optimal control into a sampling problem. In its simplest form, one estimates \(\psi\) by Monte Carlo under the reference process and then differentiates \(\log\psi\) to recover \(\mathbf{u}^*\). A central refinement is to sample under an auxiliary control \(u\) and correct by Girsanov weights. In matched form, the resulting path cost includes the quadratic control term and a stochastic integral,
\[
S^u(t)
=
\int_t^T
\Big(
V(s,X^u(s))
+
\tfrac12 u(s,X^u(s))^\top u(s,X^u(s))
\Big)\,ds
+
\int_t^T u(s,X^u(s))^\top dW(s),
\]
and the desirability can be written as
\[
\psi(t)=\mathbb{E}\big[e^{-S^u(t)}\mid\mathcal{F}_t\big].
\]
This identity underlies importance sampling for path-integral control and explains why improving the sampling controller improves numerical efficiency [1406.4026].

A key result is that better control, in terms of control cost, yields more efficient importance sampling, in terms of effective sample size; the optimal control provides a zero-variance estimate [1406.4026]. In the same work, the weighted identity
\[
\big\langle (u^*(t)-u(t))\,f(t)^\top \big\rangle
=
\lim_{r\to t}
\Big\langle \frac{\int_t^r f(s)^\top dW(s)}{r-t}\Big\rangle
\]
supports direct estimation of state-dependent feedback laws from weighted trajectory ensembles. For parameterizations \(u^*(x,t)=A(t)h(x,t)\), this becomes a linear relation for the coefficient matrix \(A(t)\), making path-integral control compatible with explicit feedback synthesis rather than only open-loop control [1406.4026].

The same importance-sampling viewpoint is developed further through the Path Integral Cross Entropy method. There, the control problem is re-expressed as minimization of a reverse KL divergence between a parametrized controlled path measure and the optimal path measure
\[
p^*(\tau)=\frac{1}{\psi(t,x)}\,p_0(\tau)e^{-V(\tau)},
\]
leading to fixed-point equations or gradient updates for arbitrary controller parameterizations [1505.01874]. This reframes “how to compute” as “how to learn an effective importance sampler,” with state-feedback controllers as the preferred samplers.

Several non-sampling representations have also been developed. In the reproducing-kernel formulation, the discrete-time recursion
\[
\psi_i(x_{t_i})
=
\mathbb{E}_{X_{i+1}\mid x_{t_i}}
\big[
\Phi_i(x_{t_i},X_{i+1})\,\psi_{i+1}(X_{i+1})
\big]
\]
is embedded in RKHS covariance operators, producing a model-free, non-parametric estimator with a decomposition into invariant, dynamics-dependent operators and task-dependent cost factors. This decomposition enables sample reuse across tasks [1208.2523]. In the Koopman-based formulation, only the specific observable
\[
z(\bm{x})=\exp\big(-\phi(\bm{x})/\lambda\big)
\]
is propagated. A polynomial expansion of the extended observable dynamics then yields coupled linear ODEs for expansion coefficients, rather than direct Monte Carlo over trajectories [2012.05514]. These developments preserve the same desirability calculus while altering the numerical representation.

## 4. Generalizations of the Euclidean formulation

| Extension | Core construction | Source |
|---|---|---|
| Generalized costs | Augmented state \(y_t=\int_0^t V(s,\mathbf{x}_s)\,ds+\int_0^t c(s)\,d\mathbf{x}_s\) | [1406.7869] |
| Belief-space control | Controlled sensing enforces \(D(\Sigma,u)=\lambda B R_a^{-1}B^\top\) | [2604.18941] |
| Euclidean action in economics | Quantum Lagrangian and saddle-point condition \(\partial f/\partial u=0\) | [2002.09394] |

A major extension replaces the standard running cost by a generalized cost containing stochastic integral terms and linear control costs. The construction introduces an augmented state
\[
y_t
=
\int_0^t V(s,\mathbf{x}_s)\,ds
+
\int_0^t c(s)\,d\mathbf{x}_s,
\]
whose dynamics are
\[
dy_t
=
\big(V(t,\mathbf{x}_t)+c(t)(\mathbf{f}+\mathbf{B}\mathbf{u}_t)\big)\,dt
+
c(t)\boldsymbol{\sigma}\,d\mathbf{W}_t.
\]
The original problem is thereby converted into one with terminal cost \(\Phi(\mathbf{x}_T,y_T)\) and quadratic control effort in the augmented state space, and the same desirability transform applies after extending the matching condition to the augmented matrices \(\hat{\boldsymbol{\sigma}}\) and \(\hat{\mathbf{B}}\) [1406.7869]. This is the main mechanism by which Euclidean Path-Integral Control incorporates linear control costs and stochastic Itô terms without abandoning linear solvability.

A second extension addresses partial observation. For linear-Gaussian systems with controlled sensing, the Gaussian belief state \((\mu_t,\Sigma_t)\) obeys
\[
d\mu_t = \bigl(f(\mu_t)+Bu^a_t\bigr)\,dt + L(\Sigma_t,u^s_t)\,d\beta_t,
\qquad
\dot{\Sigma}_t = a(t,\Sigma_t,u^s_t),
\]
where the observation-driven diffusion satisfies
\[
L(\Sigma,u)L(\Sigma,u)^\top
=
D(\Sigma,u)
=
\Sigma C(u)^\top R_o^{-1}C(u)\Sigma.
\]
A fixed observation matrix generally cannot enforce the matching
\[
D(\Sigma_t,u^s_t)=\lambda B R_a^{-1}B^\top.
\]
By treating the observation matrix as a control variable and restricting sensing to a measurable selector from the matching set
\[
U(\Sigma)=\{u:\,D(\Sigma,u)=\lambda B R_a^{-1}B^\top\},
\]
the constrained belief-space HJB again linearizes under \(V=-\lambda\log\Psi\), yielding a linear PDE and Feynman–Kac representation on belief space [2604.18941]. Here the diffusion is not process-noise-driven but observation-driven, and the covariance evolves deterministically via a controlled Riccati equation.

A third line of work uses Euclidean action principles even when strict linearly-solvable assumptions are weakened. In a Walrasian or stochastic-volatility setting, the local Euclidean action density
\[
f(s,x,u)
=
\pi(s,x,u)
+
g(s,x)
+
\partial_s g
+
\mu(s,x,u)\,\partial_x g
+
\tfrac12 \sigma^2(s,x,u)\,\partial_{xx}g
\]
defines a Schrödinger- or Feynman–Kac-type evolution \(\partial_s\Psi=-f\Psi\), and optimal controls are extracted from local stationarity conditions such as \(\partial f/\partial u=0\) [2002.09394]. This is not identical to classical LSOC, but it preserves the Euclidean path-integral viewpoint.

## 5. Computational examples and empirical domains

The generalized-cost framework has been demonstrated on hierarchical electric load management. For a single thermostatically controlled load type, the path-integral method used 300 time nodes and 5 samples via implicit sampling, while grid-based HJB used 51, 101, and 401 state nodes; the path-integral solution matched the converged grid-based solution, and temperature trajectories differed by less than \(1\%\) against the finest grid [1406.7869]. For six TCL types, time discretization with 100 nodes and \(5^6=15625\) implicit samples was used; the dimensionality rendered grid-based methods impractical, while path-integral control remained effective [1406.7869]. These examples illustrate the standard claim that path-integral control avoids a global grid and is therefore attractive in moderate-to-large dimensions.

The Koopman-based Euclidean formulation has been tested on a nonlinear stochastic van der Pol system in which only the second state is actuated and the first-state noise is removed to satisfy the matching condition. With truncation \(n_1,n_2<60\), the polynomial method used \(60\times 60=3600\) coefficients, whereas the comparison HJB grid used \(400\times 400=160{,}000\) points; the reconstructed desirability agreed closely with the direct HJB solution [2012.05514]. In receding-horizon application, the resulting controller drove the noisy oscillator toward the desired target even though control acted only on one coordinate.

The control–inference duality has also been made explicit. In the inference interpretation, the posterior path law takes the same form as the optimal path measure,
\[
p^*(\tau)=\frac{1}{\psi}\,p_0(\tau)e^{-V(\tau)},
\]
with \(V\) interpreted as negative log-likelihood. In a two-dimensional firing-rate model with Gaussian observations on one neuron, the path-integral-based importance sampler achieved an effective sample size of about \(60\%\) with 22 iterations and 6000 particles per iteration, with computation time approximately \(35.1\) seconds; an open-loop control reduced ESS to about \(29\%\), while a forward-filter backward-smoother with 6000 forward and 3600 backward particles required about \(638\) seconds [1505.01874]. This places Euclidean Path-Integral Control at the boundary between stochastic control and latent-state inference rather than exclusively within either field.

## 6. Stochastic epidemiology and the SIR Euclidean path integral

A recent application extends the Euclidean construction to a stochastic SIR model with a non-linear incidence rate and two policy controls, lockdown intensity \(e(s)\) and vaccination rate \(v(s)\), under a COVID-19 setting [2209.13733]. The state is
\[
x(s)=[\beta(s),S(s),I(s),R(s)]^\top,
\qquad
u(s)=[e(s),v(s)]^\top,
\]
and the incidence term is saturated:
\[
F(S,I;e,v)
=
\beta(e,v)\,
\frac{SI}{[1+\rho I]+\eta N(s)}.
\]
The saturation is motivated by the claim that when the proportion of infected agents is very high, exposure is inevitable and the transmission rate responds slower than linearly to further increases in infections. In the simplified incidence \(b(S,I)=\beta SI/(1+\rho I)\), the derivatives satisfy
\[
\frac{\partial b}{\partial S}>0,\qquad
\frac{\partial b}{\partial I}>0,\qquad
\frac{\partial^2 b}{\partial I^2}<0,
\]
so the incidence is concave in \(I\) and grows sublinearly at high infection levels [2209.13733].

The controlled SDEs couple SIR compartments, environmental modifiers, and a stochastic infection-rate process \(\beta(s)\). The paper’s objective contains discounted quadratic control terms, linear terms in \(e\) and \(v\), and a transmission penalty:
\[
c[u(s),\mathbf X(s)]
=
\exp(-rs)\Big\{
S(s)\Big(\tfrac12\alpha_{11}v(s)^2+\alpha_{12}v(s)+\alpha_{13}\Big)
+
I(s)\Big(\tfrac12\alpha_{21}e(s)^2+\alpha_{22}e(s)+\alpha_{23}\Big)
\Big\}
+
\beta(e(s),v(s))S(s)I(s).
\]
Because the drift contains \(\beta(e,v)\), the control enters nonlinearly and the usual control-affine linear-solvable assumptions do not apply directly. The authors therefore use a Euclidean path-integral discretization based on a quantum Lagrangian
\[
\mathcal{L}(s,u,\mathbf X)
:=
\mathbb{E}_s\Big\{
c[u(s),\mathbf X(s)]
+
\lambda\big[\bm{\mu}(s,u,\mathbf X)\,ds+\bm{\sigma}(s,\mathbf X)\,d\mathbf B(s)-\Delta\mathbf X\big]
\Big\},
\]
and a transition function
\[
\Psi_{s,s+\varepsilon}(\mathbf X)
=
\frac{1}{L_\varepsilon}
\int_{\mathbb{R}^4}
\exp\big[-\varepsilon\,\mathcal{A}_{s,s+\varepsilon}(\mathbf X)\big]\,
\Psi_s(\mathbf X)\,d\mathbf X.
\]
Using Itô calculus, Taylor expansion, and Feynman–Kac, this yields a linear relation for \(\Psi\) and a Fokker–Plank type equation [2209.13733].

Optimality is imposed through the first-order condition
\[
-\frac{\partial}{\partial u}\tilde f(s,\bar Z)\,\Psi_s^\tau(\mathbf X)=0.
\]
With the specific choice
\[
g(s,\mathbf X)=\sum_{x\in\{\beta,S,I,R\}}[sx-1-\ln x],
\]
the paper derives explicit feedback policies \(e^*(s)\) and \(v^*(s)\) for \(\theta_1=\theta_2=2\), with
\[
e^*(s)=\frac{A_2+A_3}{A_1+A_2},
\qquad
v^*(s)=\frac{B_3}{B_1-B_2},
\]
assuming \(B_1>B_2\) [2209.13733]. Under perfect and complete information, the authors prove existence of a unique solution to the stochastic optimization and existence of a Brouwer fixed point for the quantum-Lagrangian mapping on a non-empty, convex, compact set \(\widetilde\Xi\subset\mathbb{R}^6\).

The numerical study uses \(N=100\), \(S(0)=99.8\), \(I(0)=0.1\), \(R(0)=0.1\), and reports two diffusion regimes, including a higher-noise case with \(\sigma_1=0.1\), \(\sigma_2=0.06\), and \(\sigma_3=0.12\). In that regime, susceptible and recovered curves maintain downward trends while the infected curve becomes ergodic, fluctuating around a level rather than converging monotonically [2209.13733]. The data analysis based on Office for National Statistics UK data for early 2021 sets \(S(0)=84.19\), \(I(0)=1.89\), \(R(0)=13.82\), \(e(0)=0.75\), and full vaccination \(v\approx 0.00557\); over the first 100 days of 2021, the controlled SIR paths reproduce high volatility in \(I\), downward but fluctuating \(S\) and \(R\), an initially declining optimal lockdown intensity with a spike around days 70–80, and a vaccination path without a clear upward trend [2209.13733].

## 7. Scope, limitations, and relations to adjacent frameworks

The central limitation of Euclidean Path-Integral Control is structural. The matching condition
\[
\Sigma\Sigma^\top=\lambda BR^{-1}B^\top
\]
or its equivalent forms is restrictive, and several papers identify it as the main barrier to broader applicability [1406.7869]. The partially observed extension makes this especially explicit: with a fixed observation matrix, the observation-driven belief diffusion generally cannot be made equal to \(\lambda B R_a^{-1}B^\top\); controlled sensing is introduced precisely to restore linear solvability, and even then existence of a measurable selector is an assumption tied to sensor richness and feasibility [2604.18941].

The computational burden is shifted rather than eliminated. Monte Carlo evaluation avoids global state-space grids, but sample efficiency depends strongly on the quality of the importance controller; the linearly-solvable literature therefore emphasizes adaptive importance sampling, ESS diagnostics, and feedback learning [1406.4026]. Deterministic alternatives such as RKHS embeddings and Koopman polynomial truncations introduce their own approximations: kernel choice and regularization in the former, truncation error and basis growth in the latter [1208.2523]. In Euclidean-action formulations outside strict LSOC, the choice of the auxiliary \(g\)-function becomes a modeling decision that affects tractability [2002.09394].

Euclidean Path-Integral Control is closely connected to KL-control and linearly-solvable MDPs, because all three rely on exponential transformations and linear backward equations [1208.2523]. It is also related to Schrödinger bridge formulations, which likewise lead to linear PDEs and admit optimal-transport interpretations. The stochastic SIR study explicitly notes, however, that its approach aligns with the Kappen-type Euclidean path-integral and Feynman–Kac grounding rather than with the entropic optimal-transport viewpoint [2209.13733]. A plausible implication is that Euclidean Path-Integral Control is best understood not as a single algorithm, but as a family of structurally linearizable stochastic-control representations whose practical effectiveness depends on how successfully a problem can be brought into, or approximated by, that linearizable form.

Source: https://www.emergentmind.com/topics/euclidean-path-integral-control