---
title: GU-HJBI Equation with Gradient Uncertainty
url: https://www.emergentmind.com/topics/hamilton-jacobi-bellman-isaacs-equation-with-gradient-uncertainty-gu-hjbi
type: topic
---

# GU-HJBI Equation with Gradient Uncertainty

Searching arXiv for recent and foundational papers on GU-HJBI, gradient constraints, and related HJBI structure.
The Hamilton–Jacobi–Bellman–Isaacs equation with Gradient Uncertainty (GU-HJBI) is a robust-control PDE introduced to model adversarial perturbations not only of system dynamics but also of the value function gradient itself. In the formulation developed in “Robust Control with Gradient Uncertainty” [2507.15082], the controller minimizes cost while an adversary perturbs both the drift through a standard robust-control channel and the local sensitivity \(\nabla V(x)\) through a bounded perturbation \(\delta\). This produces a zero-sum dynamic game with a fully nonlinear Isaacs operator and leads to a new PDE distinct from both the classical HJB equation and the standard robust HJBI equation. The topic sits at the intersection of robust stochastic control, differential games, viscosity solutions, and approximate dynamic programming, and it is especially motivated by applications in reinforcement learning, where value gradients are typically estimated rather than known exactly [2507.15082].

## 1. Definition and problem formulation

In the time-homogeneous setting of [2507.15082], the controlled state \(X_t\in\mathbb R^n\) satisfies
\[
dX_t = f(X_t,u_t)\,dt + \sigma(X_t,u_t)\,dB_t,
\]
with control \(u_t\in U\subset\mathbb R^k\), admissible controls \(u\in A\), running cost \(L(x,u)\), and discounted objective
\[
J(x;u) \coloneqq E_x\left[\int_0^\infty e^{-\rho t} L(X_t,u_t)\,dt\right],
\qquad
V(x) \coloneqq \inf_{u\in A} J(x;u).
\]
The classical HJB equation is
\[
\rho V(x)=\inf_{u\in U}\left\{L(x,u)+\mathcal{L}^u V(x)\right\},
\]
where
\[
\mathcal{L}^u \phi(x)\coloneqq \nabla \phi(x)^\top f(x,u)+\frac12 \operatorname{Tr}\!\left(\sigma(x,u)\sigma(x,u)^\top D^2\phi(x)\right).
\]
Standard robust control augments the drift by an adversarial distortion \(h_t\in\mathbb R^m\),
\[
dX_t = \big(f(X_t,u_t)+\sigma(X_t,u_t)h_t\big)\,dt+\sigma(X_t,u_t)\,dB_t,
\]
penalizing \(h\) quadratically and yielding the standard robust HJBI equation
\[
\begin{split}
\rho V(x) = \inf_{u \in U} \bigg\{ & L(x,u) + \nabla V(x)^T f(x,u) + \frac{1}{2}\operatorname{Tr}\!\left(\sigma(x,u)\sigma(x,u)^T D^2V (x)\right) \\
& + \frac{\eta}{2}\left\|\sigma(x,u)^T \nabla V(x)\right\|^2 \bigg\}.
\end{split}
\]

GU-HJBI adds a second adversarial channel by replacing the exact gradient with a perturbed gradient
\[
\nabla V(x)\leadsto \nabla V(x)+\delta,
\]
where \(\delta\) belongs to the uncertainty set
\[
\Delta_\epsilon \coloneqq \{\delta\in\mathbb{R}^n:\|\delta\|\le \epsilon\}.
\]
The full GU-HJBI PDE is
\[
\begin{split}
\rho V(x) = \inf_{u \in U} \sup_{h \in \mathbb{R}^m, \delta \in \Delta_\epsilon} \bigg\{ & L(x,u) + (\nabla V(x) + \delta)^T (f(x,u) + \sigma(x,u)h) \\
& - \frac{1}{2\eta} \|h\|^2 + \frac{1}{2} \operatorname{Tr}\left(\sigma(x,u)\sigma(x,u)^T D^2V(x)\right) \bigg\}.
\end{split}
\tag{GU-HJBI}
\]
This formulation is specific to [2507.15082]. It differs from singular-control HJBs with gradient constraints, such as
\[
\max\{Lu-f,\;H(Du)\}=0
\]
or
\[
\max\{q u-\Gamma u-h,\;|\nabla u|^2-1\}=0,
\]
because there the gradient term encodes a constraint arising from singular control rather than an uncertainty set over gradients [1102.1109], [1605.04993].

## 2. Reduced equation and Hamiltonian structure

For fixed \(x,u,p,\delta\), with \(p=\nabla V(x)\), the inner optimization over \(h\) in [2507.15082] is
\[
\sup_{h\in\mathbb{R}^m}\left\{(p+\delta)^\top \sigma(x,u) h - \frac{1}{2\eta}\|h\|^2\right\},
\]
whose maximizer is
\[
h^*(x,u,p,\delta)=\eta\,\sigma(x,u)^\top (p+\delta).
\]
Substituting this into the full game yields the reduced GU-HJBI equation
\[
\begin{split}
\rho V(x) = \inf_{u \in U} \bigg\{ & L(x,u) + \frac{1}{2} \operatorname{Tr}\left(\sigma\sigma^T D^2V(x)\right) \\
& + \sup_{\delta \in \Delta_\epsilon} \left[ (\nabla V(x) + \delta)^T f(x,u) + \frac{\eta}{2} \left\|\sigma(x,u)^T(\nabla V(x) + \delta)\right\|^2 \right] \bigg\}.
\end{split}
\tag{Reduced GU-HJBI}
\]
This is Proposition 3.1 of [2507.15082].

The corresponding Hamiltonian is
\[
\begin{aligned}
\mathcal{H}(x,p,X,u) \coloneqq &\, L(x,u) + \frac{1}{2}\operatorname{Tr}\left(\sigma(x,u)\sigma(x,u)^T X\right) \\
& + \sup_{\|\delta\|\le\epsilon} \left[ (p+\delta)^T f(x,u) + \frac{\eta}{2}\left\|\sigma(x,u)^T(p+\delta)\right\|^2 \right],
\end{aligned}
\]
and the PDE is written as
\[
F(x,r,p,X)\coloneqq \rho r - \inf_{u\in U}\mathcal H(x,p,X,u)=0.
\]
The paper states that this operator is proper / degenerate elliptic because increasing \(X\) increases \(\operatorname{Tr}(\sigma\sigma^\top X)\), hence decreases \(F\) [2507.15082].

This Isaacs structure is explicit. By contrast, graph HJBI formulations such as
\[
I(u,x)=\min_{\alpha}\max_{\beta}\left(f^{\alpha\beta}(x)+c^{\alpha\beta}(x)u(x)+\sum_yK^{\alpha\beta}(x,y)(u(y)-u(x))\right)
\]
encode uncertainty through coefficients acting on discrete gradients \(u(y)-u(x)\), which is structurally related but not identical to perturbing a continuum gradient variable directly [2511.07653]. Similarly, classical Bellman–Isaacs operators of the form
\[
\partial_t u + \min_{\beta\in B}\max_{\alpha\in A} \left\{ -\operatorname{tr}\!\big(a(x,\alpha,\beta)D^2u\big) + f(x,\alpha,\beta)\cdot Du + \ell(x,\alpha,\beta) \right\}=0
\]
capture adversarial uncertainty in the coefficient multiplying \(Du\), not in \(Du\) itself [1007.5445].

## 3. Interpretation of gradient uncertainty

The motivation of [2507.15082] is that in approximate dynamic programming and reinforcement learning, the value function is estimated from data, so \(\nabla V(x)\) is itself uncertain. The paper models this by allowing the adversary to perturb the controller’s local state valuation through \(\delta\in\Delta_\epsilon\). A central point is that, in standard robust control, the adversary’s optimal model distortion is proportional to \(\sigma^\top \nabla V\), so errors in \(\nabla V\) directly affect the robustness term [2507.15082].

This makes GU-HJBI conceptually different from several nearby PDE classes.

First, it is not a singular-control gradient-constraint equation. In the HJBs studied in [1102.1109] and [1605.04993], the gradient term enters as a constraint such as \(H(Du)\le 0\) or \(|\nabla u|\le 1\), derived from control costs or singular interventions, not from uncertainty in the gradient variable. The structural resemblance is real, but the semantics differ.

Second, it is not merely a Bellman–Isaacs equation with uncertain drift. In works such as [1007.5445] and [2104.14450], gradient dependence appears through affine terms like \(f^{\alpha\beta}(x)\cdot Du\) or \(b^{\alpha\beta}(x)\cdot \nabla u\). This means the adversary chooses a coefficient acting on the gradient. GU-HJBI goes further by perturbing the gradient argument itself [2507.15082].

Third, it is not equivalent to the recursive stochastic HJBI equations arising from BSDE games with random coefficients or non-Lipschitz generators. Those equations can involve gradient-sensitive quantities such as \(z+\sigma^\top p\) or \(p^\top \sigma\), but the uncertainty is attached to controls, coefficients, or the BSDE driver, rather than to a separate gradient ambiguity set [2004.05141], [2404.12129].

A plausible implication is that GU-HJBI should be viewed as an internal-uncertainty robust control model, complementing classical external model misspecification. That interpretation is explicit in [2507.15082].

## 4. Small-\(\epsilon\) asymptotics and induced nonlinearity

Let
\[
\mathcal{G}(x,u,p)\coloneqq \sup_{\|\delta\|\le \epsilon} \left\{(p+\delta)^T f(x,u)+\frac{\eta}{2}\left\|\sigma(x,u)^T(p+\delta)\right\|^2\right\}.
\]
Proposition 3.2 of [2507.15082] states that for small \(\epsilon\),
\[
\mathcal{G}(x,u,p) = \left( p^T f(x,u) + \frac{\eta}{2}\|\sigma(x,u)^T p\|^2 \right) + \epsilon \left\|f(x,u)+\eta \sigma(x,u)\sigma(x,u)^T p\right\| + O(\epsilon^2).
\]
Hence the approximate GU-HJBI equation is
\[
\begin{split}
\rho V(x) \approx \inf_{u \in U} \bigg\{ & L(x,u) + \nabla V(x)^T f(x,u) + \frac{1}{2}\operatorname{Tr}\left(\sigma\sigma^T D^2V(x)\right) \\
& + \frac{\eta}{2}\left\|\sigma(x,u)^T \nabla V(x)\right\|^2 + \epsilon \left\|f(x,u)+\eta \sigma(x,u)\sigma(x,u)^T \nabla V(x)\right\| \bigg\}.
\end{split}
\tag{Approximate GU-HJBI}
\]
The first-order correction is therefore a norm penalty in the vector
\[
v(x,u,p)\coloneqq f(x,u)+\eta \sigma(x,u)\sigma(x,u)^T p.
\]

This asymptotic structure is one of the key technical features of GU-HJBI. It shows that even arbitrarily small gradient uncertainty adds a non-polynomial first-order term. The paper further states that the geometry of the uncertainty set changes this penalty through dual norms. For small \(\epsilon\), if
\[
v=f+\eta \sigma\sigma^\top p,
\]
then:

| Uncertainty set | First-order correction |
|---|---|
| \(\ell_2\)-ball | \(\epsilon \|v\|_2\) |
| \(\ell_\infty\)-box | \(\epsilon \|v\|_1\) |
| \(\delta^\top M\delta\le \epsilon^2\) | \(\epsilon \sqrt{v^\top M^{-1}v}\) |

These are stated in Proposition 6.1 of [2507.15082]. This suggests that GU-HJBI is naturally sensitive to the dual geometry of the uncertainty model.

## 5. Well-posedness in viscosity sense

The paper [2507.15082] imposes the following assumptions. Assumption 2.1 requires continuity of \(f,\sigma,L\) in all arguments, uniform Lipschitz continuity in \(x\),
\[
\|f(x_1,u)-f(x_2,u)\|
+\|\sigma(x_1,u)-\sigma(x_2,u)\|_F
+|L(x_1,u)-L(x_2,u)|
\le K_L \|x_1-x_2\|,
\]
and linear growth
\[
\|f(x,u)\|^2 + \|\sigma(x,u)\|_F^2 + |L(x,u)|
\le K_G(1+\|x\|^2).
\]
Assumption 4.1 imposes uniform ellipticity:
\[
\sigma(x,u)\sigma(x,u)^T \ge \nu I, \qquad \forall (x,u)\in \mathbb{R}^n\times U,
\]
for some \(\nu>0\).

With the PDE written as
\[
F(x,r,p,X)=\rho r-\inf_{u\in U}\mathcal{H}(x,p,X,u),
\]
the viscosity definition is standard: a locally bounded \(V\) is a viscosity subsolution if for every \(\phi\in C^2\) touching from above at \(x_0\),
\[
F(x_0,V(x_0),\nabla \phi(x_0),D^2\phi(x_0))\le 0,
\]
and analogously for supersolutions [2507.15082].

The main result is Theorem 4.1: under Assumptions 2.1 and 4.1, if \(u\in \mathrm{USC}(\mathbb{R}^n)\) is a bounded viscosity subsolution and \(v\in \mathrm{LSC}(\mathbb{R}^n)\) is a bounded viscosity supersolution of the GU-HJBI equation, then
\[
u(x)\le v(x)\qquad \forall x\in \mathbb{R}^n.
\]
The paper then states uniqueness of bounded continuous viscosity solutions, and also states existence of a unique bounded and continuous viscosity solution under the same assumptions [2507.15082].

The proof uses the doubling-of-variables method with
\[
\Phi_{\alpha,\beta}(x,y) = u(x)-v(y) -\frac{\alpha}{2}\|x-y\|^2 -\frac{\beta}{2}(\|x\|^2+\|y\|^2),
\]
Ishii’s lemma, and continuity of the compactly maximized \(\delta\)-Hamiltonian [2507.15082]. The paper emphasizes that extension to the degenerate case \(\nu=0\) is open because the term
\[
\|\sigma^\top(p+\delta)\|^2
\]
couples adversarial gradient perturbations with diffusion geometry.

This well-posedness result places GU-HJBI alongside other viscosity-based Isaacs theories, but under a specific nonlinear Hamiltonian. Earlier Bellman–Isaacs comparison and continuous dependence results, such as those for parabolic HJBI operators in \(\mathbb R^n\), concern affine gradient dependence and do not directly address the GU term [1007.5445]. Likewise, stochastic HJBI equations with random coefficients establish viscosity characterizations in a random-field or BSPDE setting, but not this specific gradient-uncertainty Hamiltonian [2004.05141].

## 6. Linear-quadratic case and structural breakdown of Riccati theory

The LQ specialization in [2507.15082] uses
\[
dX_t = (A x_t + B u_t)\,dt + \Sigma\, dB_t,
\qquad
L(x,u)=x^\top Qx + u^\top R u,
\]
with \(Q\ge 0\), \(R>0\), and Assumption 5.1 that \((A,B)\) is stabilizable and \((A,Q^{1/2})\) is detectable.

At \(\epsilon=0\), one recovers standard robust LQ control. The quadratic ansatz
\[
V_0(x)=x^\top P_0 x + c_0
\]
leads to the robust algebraic Riccati equation
\[
A^T P_0 + P_0 A + Q - P_0 B R^{-1} B^T P_0 + 2\eta P_0 \Sigma\Sigma^T P_0 = \rho P_0,
\]
with optimal control
\[
u_0(x)=-R^{-1}B^\top P_0 x,
\qquad
c_0 = \frac{1}{\rho}\operatorname{Tr}(\Sigma\Sigma^\top P_0).
\]
Define
\[
A_{cl}\coloneqq A-BR^{-1}B^\top P_0,
\qquad
\tilde A \coloneqq A_{cl}+2\eta \Sigma\Sigma^\top P_0.
\]

The paper’s central LQ insight is Proposition 5.1: for any \(\epsilon>0\) and any non-degenerate problem data, the value function is not quadratic. The contradiction argument assumes
\[
V(x)=x^\top P x + c,
\quad
\nabla V(x)=2Px,
\quad
D^2V(x)=2P,
\]
and shows that the approximate GU-HJBI then contains the term
\[
\epsilon\|Mx\|,
\qquad
M = A - BR^{-1}B^\top P + 2\eta \Sigma\Sigma^\top P,
\]
so the PDE’s right-hand side is not a quadratic polynomial in \(x\) unless \(\epsilon=0\) [2507.15082].

This is a sharp structural difference from standard robust LQ control. A plausible implication is that GU-HJBI removes the finite-dimensional Riccati closure even in the simplest quadratic setting. That implication is directly supported by the paper’s statement that the classical quadratic value function assumption fails for any non-zero gradient uncertainty [2507.15082].

## 7. Perturbation theory, numerical evidence, and RL connection

The paper develops the expansion
\[
V(x)=V_0(x)+\epsilon V_1(x)+\epsilon^2 V_2(x)+O(\epsilon^3),
\qquad
u^*(x)=u_0(x)+\epsilon u_1(x)+\epsilon^2 u_2(x)+O(\epsilon^3).
\]
The first-order correction satisfies
\[
\rho V_1(x) = (\nabla V_1(x))^\top (\tilde A x) + \frac12 \operatorname{Tr}(\Sigma\Sigma^T D^2V_1(x)) + \|\tilde A x\|,
\]
with Feynman–Kac representation
\[
V_1(x) = E_x\left[\int_0^\infty e^{-\rho t}\|\tilde A Z_t\|\,dt\right],
\qquad
dZ_t = \tilde A Z_t\,dt + \Sigma\, dB_t,
\quad
Z_0=x.
\]
The main-text first-order control correction is
\[
u_1(x) = -R^{-1}B^\top \nabla V_1(x).
\]
The second-order correction solves
\[
\rho V_2(x)=L[V_2](x)+H_2(x),
\]
where
\[
L[W](x)= (\nabla W(x))^\top (\tilde A x)+\frac12 \operatorname{Tr}(\Sigma\Sigma^\top D^2W(x)).
\]
These formulas show that the non-polynomial structure propagates through the perturbation hierarchy [2507.15082].

The numerical section of [2507.15082] treats 1D and 2D LQ examples by solving the \(V_1\)-equation numerically. The paper reports that the approximate value function visibly deviates from the quadratic baseline, that the approximate control becomes nonlinear, and that in two dimensions the contour lines of \(V_1\) are not elliptical. These observations are presented as validation of the theoretical claim that gradient uncertainty destroys quadratic/linear structure [2507.15082].

The same paper proposes the Gradient-Uncertainty-Robust Actor-Critic algorithm, instantiated as GURAC-TD3. It uses the small-\(\epsilon\) GU-HJBI penalty
\[
\epsilon \| f(x,u)+\eta \sigma\sigma^\top \nabla V(x)\|
\]
as motivation for an actor regularizer based on critic state-gradients. With
\[
p_i = \nabla_x Q_{\theta_1}(x,a)\big|_{x=x_i,a=a_i},
\qquad
a_i=\pi_\phi(x_i),
\]
the actor loss is modified to
\[
L_{\text{actor}}
= \frac{1}{N}\sum_i \left( - Q_{\theta_1}(x_i,a_i) + \lambda_R \cdot \text{Penalty}_i \right),
\]
where
\[
\text{Penalty}_i = \left\|f_{\text{est}}(x_i,a_i)+\eta \sigma\sigma^\top p_i\right\|.
\]
On Pendulum-v1, the paper reports improved training stability relative to standard TD3, with smoother learning curves and tighter variance across seeds [2507.15082].

This suggests a concrete bridge from PDE-level GU-HJBI theory to robust approximate dynamic programming. By contrast, algebraic or idempotent approaches to HJB interpret dynamic programming operators through min-plus or max-plus linearity, but do not formulate a gradient-uncertainty Isaacs PDE of this kind [1203.0522].

## 8. Relation to adjacent theories and limitations

GU-HJBI belongs to a broader family of HJB, HJBI, and nonlocal or graph-based Isaacs equations, but it is not reducible to any one of them.

Gradient-constraint HJBs such as
\[
\max\{Lu-f,\;H(Du)\}=0
\]
and
\[
\max\{q u-\Gamma u-h,\;|\nabla u|^2-1\}=0
\]
provide rigorous techniques for convex gradient terms, penalization, and regularity, but they model singular-control constraints rather than uncertainty sets over \(\nabla V\) [1102.1109], [1605.04993]. This suggests that some analytic tools may transfer, but the game-theoretic interpretation does not.

Standard Bellman–Isaacs PDEs and their ergodic or homogenized versions incorporate adversarial drift and diffusion through min–max operators over linear coefficients acting on \(D^2u\) and \(Du\) [1007.5445], [2104.14450]. GU-HJBI extends this by inserting a second adversarial optimization over perturbations of the gradient argument itself [2507.15082].

Graph HJBI operators give a discrete analogue in which uncertainty acts through coefficients multiplying edge differences \(u(y)-u(x)\), that is, discrete gradients [2511.07653]. This is structurally close, but GU-HJBI is a continuum PDE with explicit uncertainty set \(\Delta_\epsilon\subset\mathbb R^n\).

Stochastic HJBI equations with random coefficients and BSPDE structure develop zero-sum differential games in random environments, including Hamiltonians depending on \(Du\) and \(z+\sigma^\top Du\), but they do not isolate an explicit gradient-uncertainty ambiguity set [2004.05141]. Recursive stochastic differential games with non-Lipschitz generators likewise give useful Isaacs and BSDE machinery, but their viscosity theory does not directly cover explicit gradient perturbation channels [2404.12129].

The main limitation stated in [2507.15082] is the reliance on uniform ellipticity. Extension to degenerate diffusions is open. The paper also notes the computational difficulty of solving GU-HJBI in high dimensions and leaves broader theory for other uncertainty models and solvers to future work. This suggests that the current theory is foundational rather than exhaustive.

## 9. Representative formulas

For reference, the principal formulas associated with GU-HJBI are collected below.

| Object | Formula |
|---|---|
| Gradient uncertainty set | \(\Delta_\epsilon = \{\delta \in \mathbb{R}^n \mid \|\delta\| \leq \epsilon \}\) |
| Full GU-HJBI | \(\rho V = \inf_u \sup_{h,\delta}\{\cdots\}\) as in (GU-HJBI) |
| Worst-case drift distortion | \(h^*(x,u,p,\delta)=\eta\,\sigma(x,u)^T(p+\delta)\) |
| Reduced GU-HJBI | \(\rho V = \inf_u \{ L + \frac12\operatorname{Tr}(\sigma\sigma^T D^2V) + \sup_{\delta\in\Delta_\epsilon}[ (\nabla V+\delta)^T f + \frac{\eta}{2}\|\sigma^T(\nabla V+\delta)\|^2]\}\) |
| Hamiltonian | \(\mathcal H(x,p,X,u)= L + \frac12\operatorname{Tr}(\sigma\sigma^T X) + \sup_{\|\delta\|\le\epsilon}[ (p+\delta)^T f + \frac{\eta}{2}\|\sigma^T(p+\delta)\|^2]\) |
| PDE operator | \(F(x,r,p,X)=\rho r-\inf_{u\in U}\mathcal H(x,p,X,u)\) |
| Small-\(\epsilon\) correction | \(\epsilon \|f(x,u)+\eta \sigma(x,u)\sigma(x,u)^T p\|\) |
| Robust ARE at \(\epsilon=0\) | \(A^T P_0 + P_0 A + Q - P_0 B R^{-1} B^T P_0 + 2\eta P_0 \Sigma\Sigma^T P_0 = \rho P_0\) |
| First-order correction PDE | \(\rho V_1 = (\nabla V_1)^\top (\tilde A x) + \frac12 \operatorname{Tr}(\Sigma\Sigma^T D^2V_1) + \|\tilde A x\|\) |

GU-HJBI therefore designates a specific class of Isaacs equations in which the adversary perturbs both the model and the controller’s gradient information. In the formulation currently available, it is characterized by an inner supremum over a compact gradient-uncertainty set, a reduced Hamiltonian after elimination of the drift distortion, a viscosity comparison principle under uniform ellipticity, and a perturbative structure that already invalidates classical quadratic LQ theory for any \(\epsilon>0\) [2507.15082].

Source: https://www.emergentmind.com/topics/hamilton-jacobi-bellman-isaacs-equation-with-gradient-uncertainty-gu-hjbi