Papers
Topics
Authors
Recent
Search
2000 character limit reached

Stackelberg Markov Games: Models & Dynamics

Updated 12 July 2026
  • Stackelberg Markov games are stochastic dynamic leader–follower models where the leader commits to policies and the follower best responds, influencing rewards and state transitions.
  • They encompass various formulations—including finite-horizon, infinite-horizon, Bayesian, asymmetric-information, and mean-field approaches—to address strategic behavior under uncertainty.
  • Dynamic programming, Q-learning, and bilevel optimization methods are key to computing equilibria and learning in environments with unknown dynamics and follower utilities.

Stackelberg Markov games are stochastic dynamic leader–follower games in which a state variable evolves in a Markovian manner and strategic asymmetry is built into the move order: a leader commits to a policy, mixed action, or dynamic rule, followers respond after observing that commitment or the realized leader action, and rewards and transitions depend on the resulting joint behavior. In the literature, the term covers several related formulations rather than a single canonical model, including finite-horizon discrete-time dynamic Stackelberg games with per-state mixed-strategy commitment (Lauffer et al., 2022), discounted two-player general-sum Markov games with sequential action selection (Jeong et al., 6 Apr 2026), Bayesian typed formulations with uncertainty over attacker type (Sengupta et al., 2020), asymmetric-information models that become Markovian in a common-information belief state (Vasal, 2020), and mean-field extensions in which a leader interacts with a population summarized by a distributional state (Vasal et al., 2022).

1. Formal models and state representations

A common finite-horizon formulation is the discrete-time dynamic Stackelberg game

(S,A,B,r,u,P),(S,A,B,r,u,P),

where SS is the state space, AA the leader’s action set, BB the follower’s action set, rr the leader reward, uu the follower utility, and PP the transition kernel. In the episodic layered model, the state space is partitioned as S=S1SHS=S_1\cup\cdots\cup S_H with S1={s1}S_1=\{s_1\}, and transitions move only to the next layer. The leader uses a state-feedback policy π:SΔ(A)\pi:S\to\Delta(A), and the follower best-responds at each visited state. This formulation is Markovian because the current state, the leader’s mixed action, the follower response, and the transition kernel determine the next-state distribution and current payoff (Lauffer et al., 2022).

A different but equally standard formulation is the discounted infinite-horizon tabular two-player general-sum game with finite state space SS0, leader action space SS1, follower action space SS2, transition kernel SS3, rewards SS4, and discount factor SS5. The leader policy is SS6, while the follower policy conditions on both state and leader action, SS7. In this model the Stackelberg equilibrium is a bilevel object: SS8 and the associated fixed point can be represented through Stackelberg Q-functions (Jeong et al., 6 Apr 2026).

Typed and asymmetric-information variants enlarge the state. A Bayesian Stackelberg Markov game is written as

SS9

where AA0 is a collection of state-dependent distributions over follower types, AA1 contains the defender and type-contingent attacker action sets, AA2 depends on attacker type, and AA3 contains type-specific utilities. The game is discounted and sequential, with a defender policy committed in the presence of uncertainty over which follower type is realized (Sengupta et al., 2020). In partial-observation models, the strategically relevant Markov state is often a common belief AA4 over a private physical state rather than the physical state itself, so the “Markov” property is belief-state rather than fully observed state-state Markovianity (Mishra et al., 2020).

Mean-field models replace finitely many negligible followers by a population distribution. In discrete-time Stackelberg mean field games, the common state is typically a pair AA5, where AA6 is a common belief on the leader’s private state and AA7 is the mean-field distribution of follower states. The resulting equilibrium is a Stackelberg mean field equilibrium rather than a finite-player Markov perfect equilibrium (Vasal et al., 2022).

2. Commitment, observability, and move order

The defining structural feature is sequential commitment rather than simultaneous play. In the finite-horizon discrete-time model, at a visited state AA8 the leader chooses a mixed strategy AA9, the follower observes that mixed strategy and chooses

BB0

the leader observes BB1, and only then is a pure leader action BB2 sampled. The leader receives BB3 and the system transitions according to BB4. This ordering is explicitly different from simultaneous-move stochastic games and from formulations in which the leader’s pure action is chosen before the follower responds (Lauffer et al., 2022).

Other formulations use a realized-action version of the same asymmetry. In discounted tabular Stackelberg Q-value iteration, the leader chooses BB5, the follower observes BB6, chooses BB7, and the state transitions according to BB8. The follower’s policy therefore conditions on BB9, not just on rr0 (Jeong et al., 6 Apr 2026).

Asymmetric-information models change what is observed but preserve the leader–follower asymmetry. In one security formulation, the follower observes the physical state while the leader does not. The common information is the public action history

rr1

while the follower’s private history is

rr2

The common-information state is the belief

rr3

which evolves by a Bayesian filter

rr4

The leader acts from common history or belief, while the follower acts from common information plus private state (Mishra et al., 2020).

A common misconception is that Stackelberg Markov games differ from simultaneous-move Markov games only by equilibrium selection. The formulations above indicate a stronger distinction: the follower may observe the leader’s policy, mixed action, or realized action before moving; the follower’s feasible action set may depend on the leader’s move; and the relevant state can be a belief state rather than a publicly observed physical state. These differences change both the admissible policy classes and the form of the Bellman or fixed-point equations (Lauffer et al., 2022, Mishra et al., 2020, Goktas et al., 2024).

3. Equilibrium concepts

The baseline solution concept is Stackelberg equilibrium: the follower best-responds to the leader’s commitment, and the leader optimizes anticipating that response. In finite-horizon discrete-time models with a single follower, the follower best-response correspondence is often written as

rr5

which induces a leader reward

rr6

The leader’s benchmark is then the best policy in hindsight with respect to this induced reward (Lauffer et al., 2022).

Many papers adopt the strong Stackelberg equilibrium convention. If the follower has multiple best responses, ties are broken in the leader’s favor. In the discrete-time episodic model, if rr7 is the set of follower best responses, the follower is guaranteed to pick a rr8 maximizing the leader’s expected reward among all best replies. The Bayesian Stackelberg Markov game model likewise uses Strong Stackelberg Equilibrium, under which each follower type chooses a deterministic best-response policy and ties among type-wise best responses are resolved optimistically for the leader (Lauffer et al., 2022, Sengupta et al., 2020).

Once multiple followers are present, the lower level is no longer a single-agent best response. In general-sum episodic Markov games with one leader and rr9 followers, the relevant object is a Stackelberg-Nash equilibrium (SNE): for a fixed leader policy uu0, the followers play a Nash equilibrium uu1 of the stagewise follower game induced by uu2, and the leader optimizes against an equilibrium selection uu3 that is optimistic in the leader’s favor. The bilevel program is

uu4

In that paper, follower myopia is essential: followers maximize instantaneous rewards, not continuation values (Zhong et al., 2021).

Under partial observability, equilibrium concepts are belief-based. The security-game RL formulation uses Markov Perfect Stackelberg Equilibrium (MPSE): the leader commits to a policy over common-information beliefs, and the follower best-responds using the current private state and the same belief. In smart-grid demand response with private user storage, the paper introduces private Markovian strategies (PMS) and private Markovian equilibrium (PME), where admissible deviations are restricted to policies measurable with respect to each player’s private local state rather than the full system state. In mean-field settings the analogous concepts are Stackelberg mean field equilibrium (SMFE) and, with multiple leaders, SMFE-ML (Mishra et al., 2020, Huang et al., 6 Sep 2025, Vasal, 2022).

This proliferation of equilibrium notions reflects a substantive modeling point rather than mere terminological variation. A plausible implication is that “Stackelberg Markov equilibrium” denotes a family of bilevel fixed points indexed by the policy class, observation structure, and follower coordination model.

4. Dynamic programming, recursive structure, and computation

When the follower response model is known, Stackelberg Markov games often admit dynamic-programming structure on the leader side. In the finite-horizon discrete-time game with linearly parameterized follower utility, if the true follower parameter uu5 is known, the optimal mixed strategy uu6 at state uu7 can be obtained by a Bellman-style bilevel program and computed by backward induction. The resulting known-utility problem reduces to linear programs across state layers, even though the follower best-responds stagewise (Lauffer et al., 2022).

In asymmetric-information settings, recursive decomposition typically proceeds through a common-information state. In the stochastic Stackelberg game with one leader and multiple followers, a common agent chooses prescriptions

uu8

where uu9 is the vector of common marginal beliefs on private states. Beliefs evolve by a Bayes operator

PP0

and the equilibrium can be computed by backward recursion over PP1 rather than by a global fixed point over full histories. Discrete-time Stackelberg mean field games apply the same idea to the public state PP2, where the forward evolution of the mean field and the backward value recursion together form a “master equation” (Vasal, 2020, Vasal et al., 2022).

Model-based value-iteration-style methods also exist. Stackelberg Q-value iteration in a discounted two-player general-sum tabular game updates

PP3

PP4

where the follower’s greedy response is computed from PP5 and the leader then anticipates that response in PP6. The analysis does not prove exact convergence in full generality; it proves a finite-time sup-norm bound of the form

PP7

with an analogous bound for the follower, so the guaranteed limit object is an PP8-neighborhood unless stronger assumptions make PP9 (Jeong et al., 6 Apr 2026).

More structured subclasses admit stronger tractability. In convex-concave zero-sum Markov Stackelberg games, the policy-space problem

S=S1SHS=S_1\cup\cdots\cup S_H0

can be solved by nested stochastic gradient descent-ascent. Under convex-concavity, smoothness, Slater, and bounded-variance stochastic oracle assumptions, the averaged policy pair is an S=S1SHS=S_1\cup\cdots\cup S_H1-recursive Stackelberg equilibrium after S=S1SHS=S_1\cup\cdots\cup S_H2 gradient evaluations (Goktas et al., 2024). In the smart-grid private-information model, the lower-level follower game can be converted into an auxiliary Markov potential game and a pure PME can be computed in polynomial time, but that result depends on a specific aggregative and action-independent transition structure (Huang et al., 6 Sep 2025).

Computational difficulty remains central. In the no-regret learning model with unknown follower utility, the optimistic planning subproblem is a nonconvex quadratic program because the leader’s mixed strategy and the follower-utility parameter interact bilinearly in the margin constraints. The paper reports solving this subproblem with Gurobi. Thus theoretical no-regret guarantees and practical tractability need not coincide (Lauffer et al., 2022).

5. Learning with unknown dynamics, unknown follower utilities, or large populations

A substantial part of the literature concerns learning Stackelberg structure rather than assuming it known. One finite-horizon approach assumes that the leader knows transitions and its own reward but does not know the follower’s utility, except that it is linearly parameterized: S=S1SHS=S_1\cup\cdots\cup S_H3 Observed best responses induce linear inequalities in S=S1SHS=S_1\cup\cdots\cup S_H4, allowing an optimistic version-space algorithm. Before each episode the leader computes an optimistic S=S1SHS=S_1\cup\cdots\cup S_H5-conservative policy by backward induction over states and candidate follower responses, executes that policy, observes follower actions, and intersects the version space with the halfspaces implied by those observations. With high probability, the regret is sublinear in S=S1SHS=S_1\cup\cdots\cup S_H6, polynomial in S=S1SHS=S_1\cup\cdots\cup S_H7 and a condition parameter S=S1SHS=S_1\cup\cdots\cup S_H8, and independent of S=S1SHS=S_1\cup\cdots\cup S_H9; the dependence on time is S1={s1}S_1=\{s_1\}0, where S1={s1}S_1=\{s_1\}1 is the feature dimension (Lauffer et al., 2022).

When dynamics are unknown and state information is asymmetric, learning is often belief-based. In stochastic Stackelberg security games, a common-information reduction yields a belief state S1={s1}S_1=\{s_1\}2, and an Expected Sarsa procedure with particle filters approximates the belief recursion. The resulting learned policy pair is claimed to form an S1={s1}S_1=\{s_1\}3-MPSE, though the end-to-end RL convergence theory is partial and several proofs are omitted or sketched (Mishra et al., 2020).

Bayesian type uncertainty leads to a different learning architecture. In Bayesian Stackelberg Markov games for adaptive moving target defense, the statewise equilibrium is a Bayesian Strong Stackelberg Equilibrium computed from current Q-matrices. Bayesian Strong Stackelberg Q-learning (BSS-Q) combines Q-learning updates with repeated solution of a Bayesian Stackelberg stage game for each state. The paper states that BSS-Q converges to the BSS of a BSMG and uses the approach to learn movement policies under unknown rewards and transitions with state-dependent priors over attacker types (Sengupta et al., 2020).

For simultaneous-move general-sum Markov games with one leader and multiple myopic followers, optimistic and pessimistic least-squares value iteration yield sample-efficient online and offline learning. In the online case the regret satisfies

S1={s1}S_1=\{s_1\}4

while in the offline case the suboptimality scales as S1={s1}S_1=\{s_1\}5 under sufficient coverage. These guarantees rely on linear transition realizability, a statewise SNE oracle, and follower myopia, which reduces the lower level to stagewise normal-form equilibria (Zhong et al., 2021).

Large-population learning motivates mean-field approximation and regularization. In a stationary discounted Stackelberg Markov game, alternating best-response updates can be stabilized by a softmax approximation to exact best responses and by entropy-style regularization. The paper proves existence and uniqueness of stationary Stackelberg equilibrium under continuity, uniqueness, and Lipschitz best-response assumptions with S1={s1}S_1=\{s_1\}6, and extends the same logic to a Stackelberg–mean-field equilibrium. The softmax approximation yields an S1={s1}S_1=\{s_1\}7 error bound relative to the true equilibrium and is presented as enabling scalable and stable learning without full knowledge of follower objectives (He et al., 19 Sep 2025).

6. Applications, variants, and limitations

The application range is broad. Security scheduling, anti-poaching patrols, and dynamic resource allocation motivate finite-horizon leader–follower models with state-dependent leader rewards and actions (Lauffer et al., 2022). Adaptive moving target defense motivates Bayesian Stackelberg Markov games with uncertain attacker type and defender commitment (Sengupta et al., 2020). Smart-grid demand response motivates private-state Stackelberg Markov games in which an aggregator sets prices and users make storage and demand decisions based on private local states (Huang et al., 6 Sep 2025). Electricity tariff design motivates stationary Stackelberg Markov games and mean-field approximations in which a utility or commission sets time-varying rates for heterogeneous prosumers (He et al., 19 Sep 2025). Reach-avoid control is modeled as a zero-sum Markov Stackelberg game because the follower’s feasible action set depends on the leader’s move through a safety constraint (Goktas et al., 2024). Shared-autonomy and intelligent disobedience motivate a sequential leader–follower model with asymmetric safety information, although that work is presented more as a conceptual bridge than as a fully developed Stackelberg Markov game theory (Hornig et al., 22 Mar 2026).

Several nearby frameworks are related but not identical to standard Stackelberg Markov games. Continuous-time stochastic dynamic Stackelberg games with adapted open-loop controls treat the follower’s best response as an operator

S1={s1}S_1=\{s_1\}8

between spaces of stochastic processes, and show that attention-based neural operators can approximate that best-response operator uniformly on compact sets. This is directly about operator approximation in open-loop control space rather than about Markov perfect equilibrium in state-feedback policy space (Alvarez et al., 2024). Zero-sum stochastic linear-quadratic Stackelberg differential games with regime switching use a finite-state continuous-time Markov chain to modulate coefficients and derive explicit leader and follower feedback laws through coupled Riccati equations. That framework is Markovian, but it belongs to the regime-switching differential-games literature rather than the discrete-time stochastic-game literature (Wu et al., 2024).

The main limitations are structural. Many positive results require finite state and action spaces, known rewards or known transition kernels, exact observation of follower actions, strong Stackelberg tie-breaking, unique maximizers, myopic followers, or exact realizability of the follower model (Lauffer et al., 2022, Zhong et al., 2021). Belief-state and common-information approaches address asymmetric information, but they introduce enlarged state spaces and nontrivial fixed-point problems (Vasal, 2020, Mishra et al., 2020). Exact tractability in private-information multi-follower settings currently depends on restrictive structure such as action-independent public transitions, aggregative coupling, and elimination of dominated actions (Huang et al., 6 Sep 2025). Even when convergence is proved, the guarantee may be to an S1={s1}S_1=\{s_1\}9-neighborhood rather than an exact equilibrium (Jeong et al., 6 Apr 2026).

A second misconception is that followers in Stackelberg Markov games are always dynamically strategic. Several influential formulations instead assume stagewise or myopic follower response: the follower in the discrete-time dynamic Stackelberg game “aims to maximize its immediate expected utility,” and the followers in Stackelberg-Nash learning maximize instantaneous rewards only (Lauffer et al., 2022, Zhong et al., 2021). A forward-looking dynamic follower is therefore a modeling choice, not a defining property of the topic.

Taken together, the literature presents Stackelberg Markov games as a family of dynamic bilevel stochastic systems in which commitment, observation, and state evolution jointly determine the equilibrium object. The unifying theme is not a single formal tuple, but a recurring architecture: a Markov or belief-state environment, asymmetric move order, a lower-level best response or follower equilibrium, and a leader optimization problem defined over the induced response.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Stackelberg Markov Games.