---
title: Stochastic Optimal Guaranteed Cost Control
url: https://www.emergentmind.com/topics/stochastic-optimal-guaranteed-cost-control
type: topic
---

# Stochastic Optimal Guaranteed Cost Control

Stochastic optimal guaranteed cost control is a family of stochastic control formulations in which a controller minimizes a performance criterion while certifying a bound, target, or risk level under uncertainty. In the literature, this certification takes several mathematically distinct forms: a worst-case bound under model uncertainty; a minimal achievable expected or average cost under stochastic dynamics; exact satisfaction of stochastic target or chance-type constraints; or certified upper and lower bounds on the optimal value function. Representative formulations include reflected stochastic differential systems with recursive backward-cost functionals, finite-horizon stochastic target problems, dynamic time-consistent risk-constrained Markov decision processes, robust control under \(G\)-Brownian motion, linearly solvable stochastic optimal control with semidefinite relaxations, stochastic PDE-constrained optimization with mean–variance objectives, and guaranteed-cost LQG control for uncertain linear quantum stochastic systems [1202.1412], [1511.06980], [0807.4619].

## 1. Conceptual scope and performance criteria

The core object is a stochastic control system, either continuous-time or discrete-time, together with a cost functional and a class of admissible controls. In continuous time, one standard form is
\[
J(u,x_0)=\mathbb{E}\Bigg[\int_0^T L(x_t^u,u_t,t)\,dt+h(x_T^u)\Bigg],
\]
with controlled SDE dynamics, while in discrete time one encounters both finite-horizon total-cost formulations and infinite-horizon average-cost criteria [2603.14310], [2010.06236]. In the average-cost setting for linear systems with multiplicative and additive Gaussian noise, the optimal long-run performance is
\[
\lambda=\lim_{N\to\infty}\frac{1}{N}\mathbb{E}\Big[\sum_{t=0}^{N-1}c(x_t,u_t)\Big],
\]
and the optimal gain minimizes \(\lambda\) over admissible linear policies [2010.06236].

What counts as “guaranteed cost” depends on the uncertainty model. In robust formulations under \(G\)-expectation, the cost already contains a worst-case operator,
\[
\mathbb{E}^G[X]=\sup_{P\in\mathcal P}E^P[X],
\]
so the value function is the minimal worst-case cost over a family of volatility scenarios [1606.01491]. In stochastic target problems, the guarantee is not a worst-case expectation but almost-sure satisfaction of a terminal target such as
\[
X_T^{x,u}\ge \mathbb{I}_{\{W_T>c\}} \quad \text{a.s.},
\]
with the controller minimizing effort among all controls that satisfy this hard stochastic constraint [1907.02429]. In value-function approximation and global optimization, the guarantee is again different: one constructs certified lower and upper bounds on the optimal stochastic cost, either through SOS/SDP approximations of a desirability PDE or through convex/concave relaxations of the expected terminal cost [1402.2763], [1711.08851].

This diversity suggests that stochastic optimal guaranteed cost control is not a single criterion but a collection of certification paradigms tied to how uncertainty, admissibility, and performance are modeled. A common thread is that the controller is not assessed only by nominal optimality; it is assessed by a mathematically explicit performance certificate.

## 2. Recursive costs, reflected dynamics, and nonlinear HJB characterizations

A particularly rich formulation arises for stochastic systems reflected in a domain, where the state is constrained to remain in a bounded convex set \(D=\{x\in\mathbb{R}^d:\phi(x)>0\}\) with reflection in the inward normal direction \(\nabla\phi(x)\) on \(\partial D\). The controlled reflected SDE is
\[
\begin{aligned}
X_s^{t,x;u}
&=x+\int_t^s b(r,X_r^{t,x;u},u_r)\,dr
+\int_t^s \sigma(r,X_r^{t,x;u},u_r)\,dB_r \\
&\qquad + \int_t^s \nabla\phi(X_r^{t,x;u})\,dK_r^{t,x;u},
\end{aligned}
\]
with \(K^{t,x;u}\) increasing only when the state lies on the boundary, thereby enforcing hard state constraints [1202.1412].

The cost is recursive and is defined by a generalized BSDE,
\[
\begin{aligned}
-dY_s^{t,x;u}
&=f(s,X_s^{t,x;u},Y_s^{t,x;u},Z_s^{t,x;u},u_s)\,ds \\
&\quad + g(s,X_s^{t,x;u},Y_s^{t,x;u})\,dK_s^{t,x;u}
- Z_s^{t,x;u}\,dB_s,
\qquad
Y_T^{t,x;u}=\Phi(X_T^{t,x;u}),
\end{aligned}
\]
so the cost process depends nonlinearly on its own continuation value. The control criterion is \(J(t,x;u):=Y_t^{t,x;u}\), and the value function is the essential supremum over admissible controls. The paper shows that this quantity is deterministic, continuous in time, locally Lipschitz in space, and satisfies a generalized dynamic programming principle through a backward stochastic semigroup [1202.1412].

The corresponding HJB equation is a fully nonlinear parabolic PDE with nonlinear Neumann boundary condition:
\[
\begin{cases}
\displaystyle \frac{\partial W}{\partial t}(t,x)+H\big(t,x,W,DW,D^2W\big)=0,
& (t,x)\in[0,T)\times D,\\[1ex]
\displaystyle \frac{\partial W}{\partial n}(t,x)+g\big(t,x,W(t,x)\big)=0,
& (t,x)\in[0,T)\times\partial D,\\[1ex]
W(T,x)=\Phi(x),
& x\in D,
\end{cases}
\]
with Hamiltonian
\[
H(t,x,y,p,A)=\sup_{u\in U}\left\{
\operatorname{tr}\big(\sigma\sigma^\top(t,x,u)A\big)
+p\cdot b(t,x,u)
+f\big(t,x,y,\sigma^\top(t,x,u)p,u\big)
\right\}.
\]
The reflection term in the state dynamics becomes the Neumann-type boundary condition, and the boundary generator \(g\) appears directly in the boundary cost term. The value function is the unique viscosity solution of this HJB system [1202.1412].

In guaranteed-cost language, this framework provides a precise mathematical object for the optimal performance level under state constraints and recursive, BSDE-based costs. The paper itself formulates the value as a maximization; under a minimizing sign convention, the same structure yields a minimal guaranteed cost. The recursive dependence on \(Y\) and \(Z\) also allows nonlinear valuation of future costs, including risk-sensitive or dynamically risk-adjusted structures.

## 3. Stochastic targets, dynamic risk constraints, and randomized policies

A different branch of the subject treats guarantees as explicit feasibility requirements. In the stochastic target formulation, the controlled state is
\[
X_t^{x,u}=x+\int_0^t u_s\,ds,
\]
and admissible controls must satisfy
\[
X_T^{x,u}\ge \mathbb{I}_{\{W_T>c\}} \quad \text{a.s.}
\]
The main power-cost problem is
\[
v(T,x,c)=\inf_{u\in U(T,x,c)}
\mathbb{E}\left[\int_0^T |u_t|^p\,dt\right], \qquad p>1.
\]
For this problem, the value function reduces, after scaling, to a one-dimensional function \(g\) that is the minimal positive solution of the semi-linear ODE
\[
h(y)\,g''(y)+(p-1)\Big(g(y)-g(y)^{\frac{p}{p-1}}\Big)=0,\qquad y\in(0,1),
\]
with boundary conditions \(\lim_{y\to0}g(y)=1\), \(\lim_{y\to1}g(y)=0\). The associated BSDE has singular terminal condition \(Y_T=\infty\,\mathbb{I}_{\{W_T>c\}}\), and the optimal control is given explicitly in feedback form through the conditional probability process \(M_t^{T,c}=\mathbb{P}(W_T<c\mid\mathcal F_t^W)\) [1907.02429]. For exponential running costs, by contrast, the optimal controller becomes deterministic and constant:
\[
u_t=\frac{(1-x)^+}{T},
\]
and the value is
\[
w(T,x,c)=T\left(e^{\frac{\lambda(1-x)^+}{T}}-1\right),
\]
which the paper describes as a “trivial” optimal control [1907.02429].

Time-consistent risk-constrained control introduces a different guarantee. In a finite-horizon finite-state MDP with stage-wise performance cost \(c\) and risk cost \(d\), the objective is to minimize
\[
J(x_0)=\mathbb{E}\Big[\sum_{k=0}^{N-1}c(x_k,u_k)\Big]
\]
subject to a dynamic, time-consistent risk constraint
\[
R^N(x_0):=P_{0,N}\big(d(x_0,u_0),\dots,d(x_{N-1},u_{N-1}),0\big)\le r_0.
\]
The dynamic risk measure is built recursively from coherent one-step conditional risk mappings, and the resulting dynamic program works on an augmented state \((x_k,r_k)\), where \(r_k\) is a risk budget carried forward through the Bellman operator [1511.06980]. This formulation replaces a single expected-cost constraint by a multistage guarantee on tail risk, with explicit time consistency.

Randomized control enters when the feasible performance–constraint set is nonconvex. For a finite-horizon stochastic optimal control problem with \(K\) stochastic constraints, initial randomization among at most \(K+1\) deterministic policies—“\(K\)-randomization”—is sufficient to attain the optimal mixed-strategy performance. The mixed strategy convexifies the achievable cost–constraint set, and the cost reduction relative to the optimal pure strategy is exactly the duality gap:
\[
c_{\rm M}^\star=c_{\rm P}^\star-\Delta.
\]
The paper gives necessary and sufficient optimality conditions for randomized solutions and a dual-optimization-based construction; for \(K=1\), the dual problem can be solved by root finding [1607.01478]. In this line of work, the guaranteed-cost interpretation is that the randomized controller achieves the minimal expected cost compatible with the stochastic constraints, whereas deterministic policies can be conservative when the underlying problem is nonconvex.

## 4. Robustness, worst-case formulations, and recurrence guarantees

Robust stochastic guaranteed cost control is explicit in the \(G\)-Brownian framework. The controlled state satisfies a \(G\)-SDE,
\[
X_s^{t,x,u}
=
x+\int_t^s b(X_r^{t,x,u},u_r)\,dr
+\int_t^s h_{ij}(X_r^{t,x,u},u_r)\,d\langle B^i,B^j\rangle_r
+\int_t^s \sigma(X_r^{t,x,u},u_r)\,dB_r,
\]
and the cost is defined by an infinite-horizon \(G\)-BSDE. Because
\[
\mathbb{E}^G[X]=\sup_{P\in\mathcal P}E^P[X],
\]
the control problem is intrinsically an \(\inf\sup\) problem and can be seen as a robust optimal control problem. The value function \(V(x)\) is deterministic, continuous, and the unique viscosity solution of the elliptic HJBI equation
\[
\inf_{u\in U}\Big\{
G\big(H(x,v,Dv,D^2v,u)\big)
+\langle Dv(x),b(x,u)\rangle
+f\big(x,v(x),Dv(x)\sigma(x,u),u\big)
\Big\}=0,
\]
where the operator \(G\) encodes the supremum over volatility matrices [1606.01491]. In this setting, the guaranteed cost is the minimal worst-case infinite-horizon cost under volatility uncertainty.

The quantum case gives a more classical guaranteed-cost statement. For uncertain linear quantum stochastic systems, the performance index is
\[
J(u(\cdot))
=
\mathbb{E}\int_0^{t_f}p^\top(t)p(t)\,dt
=
\mathbb{E}\int_0^{t_f}\big(x^\top(t)Rx(t)+u^\top(t)Gu(t)\big)\,dt.
\]
The uncertainty matrix \(A\) satisfies \(A^\top A\le I\). By establishing an exact correspondence with an auxiliary classical uncertain system, the problem reduces to a classical minimax LQG design. Under the Riccati conditions in the paper, the resulting dynamic output-feedback controller guarantees
\[
J(u(\cdot))\le V_T
\qquad
\forall A:\ A^\top A\le I,
\]
for all admissible uncertainties [0807.4619]. This is a direct robust guaranteed-cost result, with an explicit worst-case upper bound on expected quadratic cost.

A further performance guarantee appears in the recent discrete-time infinite-horizon discounted literature. For nonlinear stochastic systems
\[
x^+=f(x,u,v),
\]
with discounted value function
\[
V_\gamma(x)=\inf_h
\mathbb E\Bigg[\sum_{k=0}^\infty \gamma^k
\ell(\boldsymbol\varphi(k,x,h),h(\boldsymbol\varphi(k,x,h)))
\Bigg],
\]
the paper introduces stochastic cost-controllability and detectability conditions involving a state measure \(\sigma\) and proves uniform semi-global practical recurrence for closed-loop systems controlled by optimal or near-optimal inputs. Under additional continuity assumptions, this property is robust with respect to small strictly causal perturbations [2504.20705]. The guarantee here is not a min–max bound of the LQG type; it is a high-probability recurrence and boundedness certificate for closed-loop trajectories induced by discounted optimal control.

## 5. Certified computation, approximation, and learning

A substantial part of the literature studies how guaranteed-cost properties can be computed or certified numerically. One influential route begins from the linearly solvable stochastic optimal control transformation
\[
V(x,t)=-\lambda \log \Psi(x,t),
\]
under the matching condition
\[
\lambda\,G(x)R^{-1}G(x)^\top
=
B(x)\Sigma_\epsilon B(x)^\top.
\]
This converts the nonlinear HJB into a linear PDE for the desirability function \(\Psi\). Sum-of-squares relaxations of the PDE yield semidefinite programs whose solutions are guaranteed lower and upper bounds on \(\Psi^\ast\), hence guaranteed upper and lower bounds on the value function \(V^\ast\) through the logarithmic transform. The resulting hierarchy has monotonically decreasing residual bound \(\gamma\), and the approximate feedback law is obtained from
\[
\hat u(x,t)=\lambda R^{-1}G(x)^\top\frac{\nabla_x\hat\Psi(x,t)}{\hat\Psi(x,t)}.
\]
The main guarantee is a certified interval containing the optimal stochastic cost-to-go [1402.2763].

A complementary certification framework treats nonlinear stochastic optimal control with expected terminal cost
\[
\mathcal G(p)=\mathbb E[g(p,\omega,x(t_f,p,\omega))].
\]
Using dynamic convex/concave relaxations of the state and cost integrand on \(P\times\Omega\), combined with a partition \(\Phi=\{\Omega_i\}\) of the uncertainty space and Jensen’s inequality, the paper constructs finitely computable convex and concave relaxations
\[
\mathcal G_{P\times\Phi}^{cv}(p)
=
\sum_i \mathbb P[\Omega_i]\,G_{P\times\Omega_i}^{cv}(p,\omega_i),
\qquad
\mathcal G_{P\times\Phi}^{cc}(p)
=
\sum_i \mathbb P[\Omega_i]\,G_{P\times\Omega_i}^{cc}(p,\omega_i),
\]
with no sample-based approximation error. These give rigorous lower and upper bounds on the optimal objective value and are intended for use in spatial branch-and-bound algorithms [1711.08851].

For stochastic PDE-constrained control, Rosseel and Wells formulate optimization directly in stochastic function spaces and exploit the stochastic dimension to penalize statistical moments of the response. Two representative cost functionals are
\[
\mathcal J_1
=
\frac{\alpha}{2}\|z-\hat z\|^2_{L^2(D)\otimes L^2_\rho(\Gamma)}
+\frac{\beta}{2}\|\mathrm{std}(z)\|^2_{L^2(D)}
+\frac{\gamma}{2}\|u\|^2
+\frac{\delta}{2}\|g\|^2,
\]
and
\[
\mathcal J_2
=
\frac{\alpha}{2}\|\bar z-\hat z\|^2_{L^2(D)}
+\frac{\beta}{2}\|\mathrm{std}(z)\|^2_{L^2(D)}
+\frac{\gamma}{2}\|u\|^2
+\frac{\delta}{2}\|g\|^2.
\]
The optimal control can be decomposed as \(u(x,\omega)=\bar u(x)+u'(x,\omega)\), where \(\bar u\) is designed and \(u'\) is a known zero-mean stochastic component. One-shot stochastic finite element methods then solve the coupled state–adjoint system. A key computational observation is that stochastic collocation loses its usual decoupling advantage when deterministic controls or moment terms are present, because the collocation points become coupled [1107.3944].

Learning-based computation appears in discrete-time average-cost control. For linear systems with multiplicative and additive Gaussian noise, the value function under a fixed linear gain \(L\) is quadratic,
\[
J(x_k)=\mathbb E(x_k^\top P x_k)+\bar s,
\qquad
\lambda=\operatorname{tr}(PD),
\]
and the optimal gain satisfies a stochastic algebraic Riccati equation. The paper defines a quadratic Q-function with kernel matrix \(H\), proposes a model-free policy-evaluation/policy-improvement scheme using recursive least squares, and proves convergence of the learned kernel matrix and control gain to the optimal ones [2010.06236]. In that setting, the learned controller asymptotically achieves the minimal average cost among admissible linear policies.

More recently, Malliavin-calculus-based stochastic maximum principle methods have been used to avoid adjoint BSDE simulation altogether. For continuous-time SOCPs with cost
\[
J(u,x_0)=\mathbb E\Bigg[\int_0^T L(x_t^u,u_t,t)\,dt+h(x_T^u)\Bigg],
\]
the first variation is expressed in terms of Malliavin derivatives \(D_s x_t\) and the stochastic flow \(\boldsymbol\Gamma_{s,t}\), rather than a backward adjoint process. The resulting iterative algorithm updates piecewise-constant controls by projected gradient steps and proves convergence under Lipschitz, monotonicity, and consistency conditions:
\[
\|u^\ast-u^i\|\sim \mathcal O(\Delta t).
\]
The guarantee here concerns algorithmic convergence to the optimal control, rather than a robust worst-case performance bound [2603.14310].

## 6. Applications, interpretations, and limitations

The application range is broad. In finance, stochastic target control provides a model of super-hedging with transaction costs in the Bachelier model, where the terminal requirement is to hold at least one share when a call option is in the money [1907.02429]. In path planning, finite-state MDP control, and entry–descent–and–landing, randomized mixed strategies lower expected cost under stochastic constraints by exploiting nonconvexity in the feasible cost–constraint set [1607.01478]. In nonlinear continuous-time control, SOS relaxations and convex/concave expected-cost relaxations aim at globally valid certificates for the value function or optimal objective [1402.2763], [1711.08851]. In distributed-parameter systems, stochastic PDE control directly shapes both mean response and variance [1107.3944]. In quantum optics, guaranteed-cost LQG control is implemented through classical measurement and feedback for uncertain optical cavities [0807.4619].

The literature also imposes substantial structural assumptions. Reflected recursive control requires a \(C^2\) defining function for the domain, Lipschitz coefficients, and linear growth conditions [1202.1412]. The stochastic target ODE characterization leaves uniqueness of the positive solution as an open problem because the coefficient \(h(y)\) vanishes at \(y=0,1\) [1907.02429]. The linearly solvable SOS framework relies on the restrictive noise-control matching condition \(\lambda G R^{-1}G^\top=B\Sigma_\epsilon B^\top\) and polynomial data on compact semialgebraic domains [1402.2763]. The \(G\)-Brownian infinite-horizon theory requires dissipativity and monotonicity assumptions strong enough to make the \(G\)-BSDE well posed [1606.01491]. Dynamic programming with time-consistent risk constraints is developed for finite-state, finite-action MDPs and requires optimization over future risk-budget functions [1511.06980]. Stochastic PDE formulations inherit the high dimensionality of polynomial chaos and lose stochastic collocation non-intrusivity in the presence of deterministic controls or moment penalties [1107.3944]. Reinforcement-learning and Malliavin-gradient methods provide computational convergence results, but their guarantees remain tied to the assumed model class and regularity [2010.06236], [2603.14310].

A persistent misconception is that guaranteed cost control in stochastic systems must always mean a worst-case \(H_\infty\)-type bound. The literature shows a wider landscape. In some works the guarantee is a worst-case expectation over a set of models; in others it is a certified upper bound on the optimal value function; in others it is exact satisfaction of a stochastic target, a time-consistent risk budget, or asymptotic convergence of a learned average cost [1606.01491], [1711.08851], [1511.06980], [2010.06236]. A plausible implication is that the field is best organized by the type of certificate being enforced—worst-case, recursive, target-based, risk-based, or approximation-based—rather than by a single canonical performance index.

Taken together, these works establish stochastic optimal guaranteed cost control as a technically heterogeneous but conceptually coherent area: stochastic control under uncertainty with an explicit performance certificate. The certificate may be a viscosity-characterized value function, a risk budget, a target constraint, a Riccati-based bound, a convex relaxation envelope, a recurrent set reached with high probability, or a convergent computational approximation. The unifying problem is not merely to optimize, but to optimize with proof.

Source: https://www.emergentmind.com/topics/stochastic-optimal-guaranteed-cost-control