Stochastic Stackelberg Game Models
- Continuous-time stochastic Stackelberg games feature a leader committing to a control law and a follower reacting optimally, with state evolution governed by Itô SDEs.
- They employ stochastic maximum principles, BSDEs, and FBSDEs to characterize equilibrium conditions, often using Riccati equations or operator methods in linear-quadratic settings.
- They extend to models with filtering, partial observation, and mean-field interactions, offering robust tools for analyzing hierarchical decision-making under uncertainty.
Searching arXiv for the supplied papers and related continuous-time stochastic Stackelberg game work. A continuous-time stochastic Stackelberg game is a hierarchical differential game on a finite horizon in which a leader commits to a control process or feedback law, a follower then optimizes in response, and the state evolves under stochastic dynamics driven by Brownian motion. In the formulations developed across the literature, the state equation is typically an Itô SDE, the equilibrium is defined through the follower’s best-response mapping and the leader’s anticipation of that mapping, and the resulting analysis is carried out through stochastic maximum principles, backward stochastic differential equations (BSDEs), and forward-backward stochastic differential equations (FBSDEs). In the linear-quadratic setting, this structure is often expressed through Riccati equations or, when conventional decoupling fails, through Riccati-free FBSDE characterizations and operator methods (Zhang et al., 2021, Li et al., 2023, Yang et al., 13 May 2026).
1. Formal structure and solution concepts
A basic two-player continuous-time stochastic Stackelberg differential game is specified by a controlled Itô SDE
where is the leader’s control and is the follower’s control, both of which may enter both the drift and diffusion coefficients. The corresponding performance criteria are
In this notation, the leader minimizes and the follower minimizes (Zhang et al., 2021).
The hierarchical structure is defined by the order of optimization. The leader first commits to a strategy on the whole horizon , the follower then solves an optimization problem for that fixed leader strategy, and the leader finally minimizes along the follower’s reaction curve. In the terminology used for stochastic Stackelberg differential games with convex control constraints, a “global Stackelberg solution” means that the leader’s domination holds over the whole time interval , so that the leader chooses a control process or feedback law anticipating the follower’s entire trajectory rather than only local or myopic reactions (Zhang et al., 2021).
Two information patterns recur in the literature. In the adapted open-loop formulation, admissible controls are progressively measurable or adapted processes. In the adapted closed-loop memoryless formulation, the leader uses a feedback that depends only on the current state and is 0 in the state variable with bounded derivative. Other papers use open-loop formulations expressed through strategy-control pairs, including Elliott–Kalton strategies for the follower in zero-sum problems, or decentralized open-loop controls adapted to player-specific filtrations in large-population models (Zhang et al., 2021, Sun et al., 2021, Si et al., 2019).
The equilibrium conditions are correspondingly hierarchical. In the basic nonzero-sum two-player case, a pair 1 satisfies
2
under the specified information structure. In zero-sum settings, the leader maximizes the follower’s value function, while in large-population settings the followers may themselves play a Nash game or a mean field game conditional on the leader’s announced strategy (Zhang et al., 2021, Sun et al., 2021, Si et al., 2019).
2. Follower’s optimization and stochastic maximum principles
Fixing the leader’s control reduces the follower’s problem to a stochastic optimal control problem with control constraints and, in many cases, control-dependent diffusion. For a convex control set 3, the follower’s Hamiltonian is
4
and the stochastic maximum principle yields an adjoint BSDE of the form
5
with terminal condition 6, together with a variational inequality
7
This is the standard first-order characterization of the follower’s best response in the convex constrained stochastic Stackelberg framework (Zhang et al., 2021).
In linear-quadratic models, the variational inequality becomes explicit. For dynamics
8
and follower cost
9
the optimal control under a closed convex constraint 0 is characterized by a projection: 1 When 2, the projection disappears and the follower’s control becomes linear in the adjoint variables (Zhang et al., 2021).
Under partial or asymmetric information, the same optimality principle survives but the control law is written in terms of filtered adjoint or state-estimate processes. In the overlapping-information model, the follower’s admissible control is adapted to
3
and the optimal control is expressed through conditional expectations such as 4. The result is a state-estimate feedback representation derived from a Riccati equation plus filtering equations (Shi et al., 2018). In the nested-observation model, the follower’s control is adapted to 5, and the optimal control takes the form
6
with Kalman–Bucy filtering used to represent the filtered adjoint in feedback form (Li et al., 2022).
A plausible implication is that, even when the follower’s problem is formally “standard,” the information structure determines whether the best response is a projection in adjoint space, a state-estimate feedback, or a filtered-adjoint law.
3. Leader’s problem and fully coupled FBSDEs
Once the follower’s best response is substituted back into the state equation, the leader no longer faces a standard controlled SDE. In the convex constrained stochastic Stackelberg model, the follower’s best reply depends on 7, so the leader’s effective state is the triple 8, governed by a fully coupled FBSDE: 9 The forward component depends on the backward variables through 0, while the backward component depends on the forward state 1. This is the central structural feature of continuous-time stochastic Stackelberg games in which the follower’s maximum principle is embedded into the leader’s problem (Zhang et al., 2021).
The leader’s Hamiltonian is correspondingly defined on this augmented state. In the adapted open-loop case, it has the form
2
where 3. The corresponding Pontryagin-type maximum principle introduces additional adjoint processes 4, producing a triple-BSDE system and the optimality condition
5
for convex 6 (Zhang et al., 2021).
In adapted closed-loop memoryless formulations, the leader controls 7, so the follower’s adjoint equation contains 8. This makes the leader’s problem nonstandard because the coefficients depend on both the control and its spatial derivative. One way to reduce the problem is to represent the feedback as an affine function
9
thereby converting it into a standard control problem in the new control pair 0 with a state consisting of 1 (Zhang et al., 2021).
In zero-sum stochastic LQ Stackelberg games, the same hierarchy can be decomposed into a forward stochastic LQ problem for the follower and a backward stochastic LQ problem for the leader. The follower’s optimal response is
2
with 3 derived from a Riccati equation and 4 from a BSDE, and the leader then solves a backward SLQ problem whose cost depends only on 5 rather than directly on the original forward state. This produces a second Riccati equation associated with the backward problem (Sun et al., 2021).
When the leader is deterministic and the follower is stochastic, the induced leader problem becomes a mean-field forward-backward stochastic differential equation. The leader’s optimality condition depends on 6, and the control is represented as a functional of the expectation of the optimal state variable together with solutions to a two-point boundary value problem of ordinary differential equation (Shi et al., 2020).
4. Linear-quadratic structure, projections, and Riccati equations
The LQ specialization is the main tractable subclass of continuous-time stochastic Stackelberg games. In the basic constrained model, the state dynamics are
7
the leader’s cost is
8
and the follower’s cost is defined analogously with 9. Closed convex control constraints 0, 1 convert both players’ first-order conditions into projection formulas. The resulting Hamiltonian system is nonlinear because the projection operator enters directly into the state and adjoint equations (Zhang et al., 2021).
The analysis of this constrained system relies on monotonicity of the associated FBSDE map. For the leader’s induced state equation, the map
2
is required to satisfy a Lipschitz condition and a Peng–Hu monotonicity condition
3
together with a similar condition for the terminal function. The projection operator contributes critically because, for a closed convex set 4,
5
Under these conditions, the fully coupled FBSDE has a unique adapted solution for each admissible leader control (Zhang et al., 2021).
When the control domains are full space, the projection terms disappear and the Hamiltonian system becomes linear. The standard next step is to postulate a linear relation
6
which yields a backward stochastic Riccati equation (BSRE). In the unconstrained Stackelberg model, the authors show that, after a suitable linear transformation and structural conditions such as
7
the BSRE coincides with the standard form studied by Tang, and the optimal controls are linear in the augmented state (Zhang et al., 2021).
In the zero-sum finite-horizon LQ case, two Riccati equations appear naturally: one for the follower’s forward SLQ problem and one for the leader’s induced backward SLQ problem. A central structural result is
8
where 9 is the follower’s Riccati solution and 0 is the leader’s backward Riccati solution. The difference 1 solves the Riccati equation associated with the zero-sum Nash stochastic LQ differential game, which implies that the Stackelberg equilibrium and the Nash equilibrium are actually identical under the uniform convexity–concavity condition (Sun et al., 2021).
Closed-loop mean-field type Stackelberg games also lead to coupled Riccati systems. In the finite-horizon deterministic-coefficient mean-field model, the follower’s closed-loop optimal strategy is characterized by two coupled Riccati equations and a linear mean-field type backward stochastic differential equation, while the leader’s necessary conditions involve two coupled Riccati equations with another linear mean-field type backward stochastic differential equation. When the diffusion term does not contain the follower’s control, the solvability of the leader’s Riccati equations can be discussed more explicitly, and the leader’s value function is expressed via two BSDEs and two Lyapunov equations (Li et al., 2023).
A plausible synthesis is that the role of Riccati equations in continuous-time stochastic Stackelberg games ranges from explicit decoupling devices in deterministic-coefficient LQ models to partial or unavailable tools in constrained, partially observed, or random-coefficient settings.
5. Information structures, mean-field interactions, and large populations
Beyond the two-player full-information model, the theory extends to partial observation, overlapping information, nested observations, and large-population mean-field interaction. In the large-population major-minor model with one major leader, minor leaders, and minor followers, the equilibrium notion is a Stackelberg–Nash–Cournot equilibrium. The mean-field methodology proceeds by solving representative follower and representative minor-leader problems, enforcing consistency conditions for the empirical means, and combining the resulting equations into a global fully coupled FBSDE. The resulting decentralized strategy profile is an 2-SNC equilibrium with
3
while empirical means converge to their mean-field limits at order 4 and 5 in the stated 6-sup norms (Si et al., 2019).
Partial observation changes the admissible filtrations. In one partially observed mean-field Stackelberg game, the leader’s information is generated by the observation process 7, the follower’s information by 8, and both state dynamics and costs contain mean-field terms. A state decomposition and backward separation principle are used to remove the control–observation circularity, yielding open-loop adapted decentralized strategies and feedback decentralized strategies. The constructed decentralized controls form an 9-Stackelberg–Nash equilibrium with
0
and the approximation estimates for the state processes are of order 1 in the stated second-moment sup norms (Si et al., 20 Mar 2025).
A related mean-field model with partial information and common noise also yields decentralized strategies that form an 2-Stackelberg–Nash equilibrium with 3. There the common Brownian motion enters all state equations but is not directly observable by the individual agents, so the mean field is conditional on the common-noise filtration and the analysis uses stochastic maximum principle with partial information and optimal filtering (Si et al., 2024).
Overlapping-information games place leader and follower on incomparable filtrations. In the two-player version, the follower observes 4, the leader observes 5, and the optimal controls are expressed through filtered state estimates and filtered adjoints. In the four-player extension with two leaders and two followers, the followers first solve a partial-information nonzero-sum Nash game, after which the leaders solve another partial-information Nash game driven by a conditional mean-field type FBSDE. The equilibrium is represented in state-estimate feedback form through a coupled system of Riccati equations (Shi et al., 2018, Si et al., 2024).
In major-player mean-field Stackelberg problems with private information, the leader’s control path acts as a signaling device. The minor players observe the major’s control and update a belief process 6, and the leader optimizes anticipating the unique MFG equilibrium associated with that belief process. The relaxed formulation is stated on the space of càdlàg martingale laws for 7, and there exists a relaxed Stackelberg solution. The corresponding major-player strategy is approximately optimal in games with a large but finite number of small players (Bergault et al., 2023).
These developments indicate that “continuous-time stochastic Stackelberg game” is not a single model class but a hierarchy of formulations in which the same leader–follower timing is combined with filtering, mean-field consistency, relaxed controls, or population limits.
6. Special cases, equivalences, and current technical frontiers
Several papers identify important special cases and structural equivalences. In the major-minor SNC framework, setting 8 and 9 removes the mean field and reduces the problem to a standard two-player LQ stochastic Stackelberg differential game. Setting 0 yields a major–minor mean-field game, and removing the major leader leads to a standard LQ mean-field game (Si et al., 2019). These embeddings show that many recent models are extensions of the classical two-player stochastic Stackelberg differential game rather than unrelated objects.
The zero-sum finite-horizon LQ model gives a particularly sharp equivalence result. Under bounded deterministic coefficients, uniform convexity in the follower’s control, and uniform concavity in the leader’s control, the Stackelberg equilibrium is exactly the unique open-loop saddle point of the corresponding Nash game. The value function is
1
so the Nash game can be solved in a leader-follower manner (Sun et al., 2021).
At the same time, several technically difficult points remain explicit in the literature. In the overlapping-information LQ game with controls in the diffusion, the follower’s Riccati equation is standard but the leader’s feedback representation requires a large four-matrix Riccati system, and the paper states that general solvability in the fully general control-in-diffusion case is nontrivial and open (Shi et al., 2018). In the partially observed mean-field Stackelberg model, the coupled asymmetric Riccati equations for 2 and for 3 are assumed solvable rather than solved in full generality (Si et al., 20 Mar 2025).
Random coefficients create a different obstruction. In the stochastic mean-field LQ Stackelberg game with random coefficients, the interaction between mean-field terms and random coefficients precludes the direct use of conventional decoupling techniques. The follower’s response is therefore derived through an extended Lagrange multiplier method, and the induced leader problem becomes a generalized stochastic LQ control problem with operator-valued coefficients. The equilibrium is characterized through a Riccati-free coupled FBSDE system, and the numerical side is addressed through a Deep FBSDE Picard Solver that preserves the Stackelberg order through follower-response learning, response-sensitivity extraction, leader optimization, and neural augmented Lagrangian enforcement of mean-field consistency constraints (Yang et al., 13 May 2026).
This suggests a broad division of the field. Deterministic-coefficient LQ models often admit Riccati-based closed-form or semi-closed-form feedback representations; convex constraints, partial observation, overlapping filtrations, and mean-field interaction preserve much of that structure but make the Riccati systems larger and more delicate; random coefficients, operator-valued dynamics, or signaling-based mean-field formulations may force a shift from Riccati decoupling to FBSDE, operator, relaxed-control, or learning-based methods.