---
title: Stochastic Stackelberg Game Models
url: https://www.emergentmind.com/topics/continuous-time-stochastic-stackelberg-game
type: topic
---

# Stochastic Stackelberg Game Models

Searching arXiv for the supplied papers and related continuous-time stochastic Stackelberg game work.
A continuous-time stochastic Stackelberg game is a hierarchical differential game on a finite horizon in which a leader commits to a control process or feedback law, a follower then optimizes in response, and the state evolves under stochastic dynamics driven by Brownian motion. In the formulations developed across the literature, the state equation is typically an Itô SDE, the equilibrium is defined through the follower’s best-response mapping and the leader’s anticipation of that mapping, and the resulting analysis is carried out through stochastic maximum principles, backward stochastic differential equations (BSDEs), and forward-backward stochastic differential equations (FBSDEs). In the linear-quadratic setting, this structure is often expressed through Riccati equations or, when conventional decoupling fails, through Riccati-free FBSDE characterizations and operator methods [2107.09315; 2303.07544; 2605.12950].

## 1. Formal structure and solution concepts

A basic two-player continuous-time stochastic Stackelberg differential game is specified by a controlled Itô SDE
\[
\begin{cases}
dx(t)=b\bigl(t,x(t),u(t),v(t)\bigr)\,dt+\sigma\bigl(t,x(t),u(t),v(t)\bigr)\,dW(t),\\
x(0)=x_0\in\mathbb{R}^n,
\end{cases}
\]
where \(u(\cdot)\) is the leader’s control and \(v(\cdot)\) is the follower’s control, both of which may enter both the drift and diffusion coefficients. The corresponding performance criteria are
\[
J_1(u,v)=\mathbb{E}\Bigl[\int_0^T \ell_1\bigl(t,x(t),u(t),v(t)\bigr)\,dt+\Phi_1(x(T))\Bigr],
\]
\[
J_2(u,v)=\mathbb{E}\Bigl[\int_0^T \ell_2\bigl(t,x(t),u(t),v(t)\bigr)\,dt+\Phi_2(x(T))\Bigr].
\]
In this notation, the leader minimizes \(J_1\) and the follower minimizes \(J_2\) [2107.09315].

The hierarchical structure is defined by the order of optimization. The leader first commits to a strategy on the whole horizon \([0,T]\), the follower then solves an optimization problem for that fixed leader strategy, and the leader finally minimizes along the follower’s reaction curve. In the terminology used for stochastic Stackelberg differential games with convex control constraints, a “global Stackelberg solution” means that the leader’s domination holds over the whole time interval \([0,T]\), so that the leader chooses a control process or feedback law anticipating the follower’s entire trajectory rather than only local or myopic reactions [2107.09315].

Two information patterns recur in the literature. In the adapted open-loop formulation, admissible controls are progressively measurable or adapted processes. In the adapted closed-loop memoryless formulation, the leader uses a feedback \(u(t,x(t))\) that depends only on the current state and is \(C^1\) in the state variable with bounded derivative. Other papers use open-loop formulations expressed through strategy-control pairs, including Elliott–Kalton strategies for the follower in zero-sum problems, or decentralized open-loop controls adapted to player-specific filtrations in large-population models [2107.09315; 2109.14893; 1905.09564].

The equilibrium conditions are correspondingly hierarchical. In the basic nonzero-sum two-player case, a pair \((u^*,v^*)\) satisfies
\[
J_2(u^*,v^*)\le J_2(u^*,v)\ \forall v,\qquad
J_1(u^*,v^*)\le J_1(u,v^*(u))\ \forall u,
\]
under the specified information structure. In zero-sum settings, the leader maximizes the follower’s value function, while in large-population settings the followers may themselves play a Nash game or a mean field game conditional on the leader’s announced strategy [2107.09315; 2109.14893; 1905.09564].

## 2. Follower’s optimization and stochastic maximum principles

Fixing the leader’s control reduces the follower’s problem to a stochastic optimal control problem with control constraints and, in many cases, control-dependent diffusion. For a convex control set \(V\), the follower’s Hamiltonian is
\[
H_2(t,x,u,v,p_2,q_2)=\langle p_2,b(t,x,u,v)\rangle+\langle q_2,\sigma(t,x,u,v)\rangle+\ell_2(t,x,u,v),
\]
and the stochastic maximum principle yields an adjoint BSDE of the form
\[
-dp_2(t)=H_{2x}\bigl(t,x(t),u(t),v^*(t),p_2(t),q_2(t)\bigr)\,dt-q_2(t)\,dW(t),
\]
with terminal condition \(p_2(T)=\Phi_{2x}(x(T))\), together with a variational inequality
\[
\big\langle H_{2v}(\cdots),\,v-v^*(t)\big\rangle\ge 0,\qquad \forall v\in V.
\]
This is the standard first-order characterization of the follower’s best response in the convex constrained stochastic Stackelberg framework [2107.09315].

In linear-quadratic models, the variational inequality becomes explicit. For dynamics
\[
dx(t)=\bigl[A(t)x(t)+B_1(t)u(t)+B_2(t)v(t)\bigr]dt+\bigl[C(t)x(t)+D_1(t)u(t)+D_2(t)v(t)\bigr]dW(t),
\]
and follower cost
\[
J_2(u,v)=\mathbb{E}\Bigl[\int_0^T \bigl(x^\top Q_2 x+v^\top R_2 v\bigr)dt+x(T)^\top\Phi_2 x(T)\Bigr],
\]
the optimal control under a closed convex constraint \(\Gamma\) is characterized by a projection:
\[
v^*(t)=\Pi_\Gamma^{R_2}\Bigl(-R_2^{-1}(t)\bigl(B_2(t)^\top p_2(t)+D_2(t)^\top q_2(t)\bigr)\Bigr).
\]
When \(\Gamma=\mathbb{R}^{m_2}\), the projection disappears and the follower’s control becomes linear in the adjoint variables [2107.09315].

Under partial or asymmetric information, the same optimality principle survives but the control law is written in terms of filtered adjoint or state-estimate processes. In the overlapping-information model, the follower’s admissible control is adapted to
\[
\mathcal G_t^1:=\sigma\{W_1(s),W_3(s);\,0\le s\le t\},
\]
and the optimal control is expressed through conditional expectations such as \(\hat\xi(t)=\mathbb E[\xi(t)\mid\mathcal G_t^1]\). The result is a state-estimate feedback representation derived from a Riccati equation plus filtering equations [1804.07466]. In the nested-observation model, the follower’s control is adapted to \(\mathcal Z_t^F=\sigma(Z_s^2:0\le s\le t)\), and the optimal control takes the form
\[
U_t^{F*}=-R_{FF}^{-1}B_F^\top \hat p_t^F,
\qquad
\hat p_t^F=\mathbb E[p_t^F\mid\mathcal Z_t^F],
\]
with Kalman–Bucy filtering used to represent the filtered adjoint in feedback form [2206.02699].

A plausible implication is that, even when the follower’s problem is formally “standard,” the information structure determines whether the best response is a projection in adjoint space, a state-estimate feedback, or a filtered-adjoint law.

## 3. Leader’s problem and fully coupled FBSDEs

Once the follower’s best response is substituted back into the state equation, the leader no longer faces a standard controlled SDE. In the convex constrained stochastic Stackelberg model, the follower’s best reply depends on \((x,p_2,q_2)\), so the leader’s effective state is the triple \((x,p_2,q_2)\), governed by a fully coupled FBSDE:
\[
\begin{cases}
dx(t)=b\bigl(t,x(t),u(t),v^*(t,x(t),u(t),p_2(t),q_2(t))\bigr)\,dt+\sigma(\cdots)\,dW(t),\\
-dp_2(t)=H_{2x}\bigl(t,x(t),u(t),v^*(t,x(t),u(t),p_2(t),q_2(t)),p_2(t),q_2(t)\bigr)\,dt-q_2(t)\,dW(t).
\end{cases}
\]
The forward component depends on the backward variables through \(v^*\), while the backward component depends on the forward state \(x\). This is the central structural feature of continuous-time stochastic Stackelberg games in which the follower’s maximum principle is embedded into the leader’s problem [2107.09315].

The leader’s Hamiltonian is correspondingly defined on this augmented state. In the adapted open-loop case, it has the form
\[
\begin{aligned}
H_1(t,u,x,k,p_1,p_2,q_1,q_2)
&=\langle p_1,b(t,x,u,v^*)\rangle+\langle q_1,\sigma(t,x,u,v^*)\rangle\\
&\quad+\ell_1(t,x,u,v^*)-\langle k,f_2(t,u,x,p_2,q_2)\rangle,
\end{aligned}
\]
where \(f_2=-H_{2x}\). The corresponding Pontryagin-type maximum principle introduces additional adjoint processes \((k,p_1,q_1)\), producing a triple-BSDE system and the optimality condition
\[
u^*(t)=\arg\min_{u\in U}H_1\bigl(t,u,x(t),k(t),p_1(t),p_2(t),q_1(t),q_2(t)\bigr)
\]
for convex \(U\) [2107.09315].

In adapted closed-loop memoryless formulations, the leader controls \(u(t,x)\), so the follower’s adjoint equation contains \(\partial_xu(t,x(t))\). This makes the leader’s problem nonstandard because the coefficients depend on both the control and its spatial derivative. One way to reduce the problem is to represent the feedback as an affine function
\[
u(t,x)=u_2(t)x+u_1(t),
\]
thereby converting it into a standard control problem in the new control pair \((u_1,u_2)\) with a state consisting of \((x,p_2)\) [2107.09315].

In zero-sum stochastic LQ Stackelberg games, the same hierarchy can be decomposed into a forward stochastic LQ problem for the follower and a backward stochastic LQ problem for the leader. The follower’s optimal response is
\[
\bar u_1(s)=\Theta(s)X(s)+v(s),
\]
with \(\Theta\) derived from a Riccati equation and \(v\) from a BSDE, and the leader then solves a backward SLQ problem whose cost depends only on \((Y,Z,u_2)\) rather than directly on the original forward state. This produces a second Riccati equation associated with the backward problem [2109.14893].

When the leader is deterministic and the follower is stochastic, the induced leader problem becomes a mean-field forward-backward stochastic differential equation. The leader’s optimality condition depends on \(\mathbb E[x_t^*],\mathbb E[y_t],\mathbb E[z_t],\mathbb E[p_t]\), and the control is represented as a functional of the expectation of the optimal state variable together with solutions to a two-point boundary value problem of ordinary differential equation [2004.00653].

## 4. Linear-quadratic structure, projections, and Riccati equations

The LQ specialization is the main tractable subclass of continuous-time stochastic Stackelberg games. In the basic constrained model, the state dynamics are
\[
dx(t)=[Ax+B_1u+B_2v]dt+[Cx+D_1u+D_2v]dW(t),
\]
the leader’s cost is
\[
J_1(u,v)=\mathbb E\Bigl[\int_0^T(x^\top Q_1x+u^\top R_1u)\,dt+x(T)^\top\Phi_1x(T)\Bigr],
\]
and the follower’s cost is defined analogously with \((Q_2,R_2,\Phi_2)\). Closed convex control constraints \(u(t)\in\Gamma_1\), \(v(t)\in\Gamma_2\) convert both players’ first-order conditions into projection formulas. The resulting Hamiltonian system is nonlinear because the projection operator enters directly into the state and adjoint equations [2107.09315].

The analysis of this constrained system relies on monotonicity of the associated FBSDE map. For the leader’s induced state equation, the map
\[
A=(x,p_2,q_2)\mapsto F(t,u,A)=(f_2(t,u,A),b(t,u,A),\sigma(t,u,A))
\]
is required to satisfy a Lipschitz condition and a Peng–Hu monotonicity condition
\[
\langle F(t,u,A_1)-F(t,u,A_2),A_1-A_2\rangle\le -c_2|A_1-A_2|^2,
\]
together with a similar condition for the terminal function. The projection operator contributes critically because, for a closed convex set \(\Gamma\),
\[
\langle \Pi_\Gamma(x)-\Pi_\Gamma(y),x-y\rangle\ge 0.
\]
Under these conditions, the fully coupled FBSDE has a unique adapted solution for each admissible leader control [2107.09315].

When the control domains are full space, the projection terms disappear and the Hamiltonian system becomes linear. The standard next step is to postulate a linear relation
\[
P(t)=R(t)X(t),
\qquad
R(t)=\Phi_1+\int_t^T\Lambda(s)\,ds-\int_t^T\Psi(s)\,dW(s),
\]
which yields a backward stochastic Riccati equation (BSRE). In the unconstrained Stackelberg model, the authors show that, after a suitable linear transformation and structural conditions such as
\[
Q_2=Q_1,\qquad B_2R_2^{-1}B_2^\top=B_1R_1^{-1}B_1^\top,
\]
the BSRE coincides with the standard form studied by Tang, and the optimal controls are linear in the augmented state [2107.09315].

In the zero-sum finite-horizon LQ case, two Riccati equations appear naturally: one for the follower’s forward SLQ problem and one for the leader’s induced backward SLQ problem. A central structural result is
\[
P=P_1-\Sigma,
\]
where \(P_1\) is the follower’s Riccati solution and \(\Sigma\) is the leader’s backward Riccati solution. The difference \(P_1-\Sigma\) solves the Riccati equation associated with the zero-sum Nash stochastic LQ differential game, which implies that the Stackelberg equilibrium and the Nash equilibrium are actually identical under the uniform convexity–concavity condition [2109.14893].

Closed-loop mean-field type Stackelberg games also lead to coupled Riccati systems. In the finite-horizon deterministic-coefficient mean-field model, the follower’s closed-loop optimal strategy is characterized by two coupled Riccati equations and a linear mean-field type backward stochastic differential equation, while the leader’s necessary conditions involve two coupled Riccati equations with another linear mean-field type backward stochastic differential equation. When the diffusion term does not contain the follower’s control, the solvability of the leader’s Riccati equations can be discussed more explicitly, and the leader’s value function is expressed via two BSDEs and two Lyapunov equations [2303.07544].

A plausible synthesis is that the role of Riccati equations in continuous-time stochastic Stackelberg games ranges from explicit decoupling devices in deterministic-coefficient LQ models to partial or unavailable tools in constrained, partially observed, or random-coefficient settings.

## 5. Information structures, mean-field interactions, and large populations

Beyond the two-player full-information model, the theory extends to partial observation, overlapping information, nested observations, and large-population mean-field interaction. In the large-population major-minor model with one major leader, minor leaders, and minor followers, the equilibrium notion is a Stackelberg–Nash–Cournot equilibrium. The mean-field methodology proceeds by solving representative follower and representative minor-leader problems, enforcing consistency conditions for the empirical means, and combining the resulting equations into a global fully coupled FBSDE. The resulting decentralized strategy profile is an \(\varepsilon\)-SNC equilibrium with
\[
\varepsilon(N)=O\!\left(\frac{1}{\sqrt{\min\{N_1,N_f\}}}\right),
\]
while empirical means converge to their mean-field limits at order \(O(1/N_1)\) and \(O(1/N_f)\) in the stated \(L^2\)-sup norms [1905.09564].

Partial observation changes the admissible filtrations. In one partially observed mean-field Stackelberg game, the leader’s information is generated by the observation process \(Y_0\), the follower’s information by \(Y_i\), and both state dynamics and costs contain mean-field terms. A state decomposition and backward separation principle are used to remove the control–observation circularity, yielding open-loop adapted decentralized strategies and feedback decentralized strategies. The constructed decentralized controls form an \(\varepsilon(N)\)-Stackelberg–Nash equilibrium with
\[
\varepsilon(N)=O(N^{-1/2}),
\]
and the approximation estimates for the state processes are of order \(O(1/N)\) in the stated second-moment sup norms [2503.15803].

A related mean-field model with partial information and common noise also yields decentralized strategies that form an \(\varepsilon\)-Stackelberg–Nash equilibrium with \(\varepsilon(N)=O(N^{-1/2})\). There the common Brownian motion enters all state equations but is not directly observable by the individual agents, so the mean field is conditional on the common-noise filtration and the analysis uses stochastic maximum principle with partial information and optimal filtering [2405.03102].

Overlapping-information games place leader and follower on incomparable filtrations. In the two-player version, the follower observes \((W_1,W_3)\), the leader observes \((W_2,W_3)\), and the optimal controls are expressed through filtered state estimates and filtered adjoints. In the four-player extension with two leaders and two followers, the followers first solve a partial-information nonzero-sum Nash game, after which the leaders solve another partial-information Nash game driven by a conditional mean-field type FBSDE. The equilibrium is represented in state-estimate feedback form through a coupled system of Riccati equations [1804.07466; 2401.08112].

In major-player mean-field Stackelberg problems with private information, the leader’s control path acts as a signaling device. The minor players observe the major’s control and update a belief process \(p_t\), and the leader optimizes anticipating the unique MFG equilibrium associated with that belief process. The relaxed formulation is stated on the space of càdlàg martingale laws for \(p_t\), and there exists a relaxed Stackelberg solution. The corresponding major-player strategy is approximately optimal in games with a large but finite number of small players [2311.05229].

These developments indicate that “continuous-time stochastic Stackelberg game” is not a single model class but a hierarchy of formulations in which the same leader–follower timing is combined with filtering, mean-field consistency, relaxed controls, or population limits.

## 6. Special cases, equivalences, and current technical frontiers

Several papers identify important special cases and structural equivalences. In the major-minor SNC framework, setting \(N_1=0\) and \(N_f=1\) removes the mean field and reduces the problem to a standard two-player LQ stochastic Stackelberg differential game. Setting \(N_f=0\) yields a major–minor mean-field game, and removing the major leader leads to a standard LQ mean-field game [1905.09564]. These embeddings show that many recent models are extensions of the classical two-player stochastic Stackelberg differential game rather than unrelated objects.

The zero-sum finite-horizon LQ model gives a particularly sharp equivalence result. Under bounded deterministic coefficients, uniform convexity in the follower’s control, and uniform concavity in the leader’s control, the Stackelberg equilibrium is exactly the unique open-loop saddle point of the corresponding Nash game. The value function is
\[
V(x)=\langle P(0)x,x\rangle=\langle(P_1(0)-\Sigma(0))x,x\rangle,
\]
so the Nash game can be solved in a leader-follower manner [2109.14893].

At the same time, several technically difficult points remain explicit in the literature. In the overlapping-information LQ game with controls in the diffusion, the follower’s Riccati equation is standard but the leader’s feedback representation requires a large four-matrix Riccati system, and the paper states that general solvability in the fully general control-in-diffusion case is nontrivial and open [1804.07466]. In the partially observed mean-field Stackelberg model, the coupled asymmetric Riccati equations for \(P_2,P_3\) and for \(\Sigma_1,\Sigma_2\) are assumed solvable rather than solved in full generality [2503.15803].

Random coefficients create a different obstruction. In the stochastic mean-field LQ Stackelberg game with random coefficients, the interaction between mean-field terms and random coefficients precludes the direct use of conventional decoupling techniques. The follower’s response is therefore derived through an extended Lagrange multiplier method, and the induced leader problem becomes a generalized stochastic LQ control problem with operator-valued coefficients. The equilibrium is characterized through a Riccati-free coupled FBSDE system, and the numerical side is addressed through a Deep FBSDE Picard Solver that preserves the Stackelberg order through follower-response learning, response-sensitivity extraction, leader optimization, and neural augmented Lagrangian enforcement of mean-field consistency constraints [2605.12950].

This suggests a broad division of the field. Deterministic-coefficient LQ models often admit Riccati-based closed-form or semi-closed-form feedback representations; convex constraints, partial observation, overlapping filtrations, and mean-field interaction preserve much of that structure but make the Riccati systems larger and more delicate; random coefficients, operator-valued dynamics, or signaling-based mean-field formulations may force a shift from Riccati decoupling to FBSDE, operator, relaxed-control, or learning-based methods.

Source: https://www.emergentmind.com/topics/continuous-time-stochastic-stackelberg-game