Papers
Topics
Authors
Recent
Search
2000 character limit reached

Stackelberg-Nash Strategy in Hierarchical Games

Updated 10 July 2026
  • Stackelberg-Nash strategy is a hierarchical framework where leaders commit to actions first and followers respond via a Nash equilibrium, blending sequential decision-making with noncooperative interactions.
  • The approach underpins diverse applications—from control of stochastic PDEs and degenerate systems to multi-level security and inventory competition—by reducing complex interactions to tractable coupled optimality systems.
  • Analytical tools such as Carleman estimates, adjoint systems, and fixed-point theorems, along with computational techniques like gradient descent and convexification, are pivotal in solving these hierarchical control problems.

A Stackelberg-Nash strategy is a hierarchical game-theoretic or control-theoretic architecture in which one or more leaders move first and one or more followers respond through a Nash equilibrium, a generalized Nash equilibrium, or a closely related equilibrium object. In the cited literature, this structure appears in hierarchical null controllability for stochastic and deterministic PDEs, abstract stochastic evolution equations, security games, stochastic differential games, stopping games, inventory competition, mean-field systems, and multi-level equilibrium models. Its defining feature is the coexistence of a Stackelberg timing relation across levels and Nash interaction within at least one level, so the leader’s problem is solved against an endogenous equilibrium response rather than against passive or independently optimized followers (Oukdach et al., 2024, Elgrou et al., 8 Feb 2025, Zhang et al., 2022).

1. Structural definition and principal variants

In its most common form, the strategy has two layers. At the upper layer, the leader chooses a control or strategy first. At the lower layer, multiple followers observe that choice and then play a noncooperative game among themselves. This is the pattern used in hierarchical controllability papers where the leader aims at null controllability and the followers solve quadratic tracking problems, in algorithmic Stackelberg games with one leader and multiple followers, and in free-boundary or moving-domain control problems where follower responses are encoded by an optimality system (Oukdach et al., 2024, Wang et al., 2021, Carvalho et al., 1 May 2026).

The architecture is not restricted to a single leader. One cited formulation studies a Nash-Stackelberg-Nash game with multiple leaders at the upper level, a Stackelberg relation across stages, and multiple followers playing a generalized Nash game at the lower level. In that model, uncertainty itself is decision-dependent through a set-valued map W(x)\mathcal W(x), so the leaders affect not only payoffs but also the feasible range of uncertainty realizations (Zhang et al., 2022).

Other variants alter the hierarchy rather than abandon it. In the three-player leader-follower security model, the game is tri-level: a top-level leader chooses first, a middle follower responds, and a bottom follower responds last. In mixed major-minor stochastic games, the hierarchy combines a major leader, minor leaders, and minor followers, producing a Stackelberg-Nash-Cournot approximate equilibrium. In stopping games, the leader commits first to a stopping strategy and the follower responds optimally, which turns the leader’s problem into an optimal control problem over stopping rules (Xu et al., 2022, Si et al., 2019, Zhang et al., 26 Jul 2025).

These formulations show that “Stackelberg-Nash strategy” does not denote a single canonical model. It denotes a family of hierarchical constructions whose common element is sequential commitment at one level and equilibrium interaction at another.

2. Equilibrium objects and response maps

For fixed leader actions, the followers’ equilibrium is usually defined by unilateral optimality conditions. In the stochastic parabolic null controllability problem with two leaders and two followers, (v1,v2)(v_1^\star,v_2^\star) is a Nash equilibrium when each follower minimizes its own cost against the other follower’s choice. Because the costs are convex and differentiable, the equilibrium is characterized by first-order conditions, and for sufficiently large penalization parameters βi\beta_i there exists a unique Nash equilibrium with an a priori bound in the follower space (Oukdach et al., 2024).

A closely related characterization recurs in abstract stochastic evolution equations. For sufficiently large BiB_i, the followers admit a unique Nash equilibrium and the equilibrium controls have adjoint-state feedback form

vi=Ri1Bizi,v_i^\ast=-R_i^{-1}B_i^\ast z_i,

with ziz_i solving a suitable adjoint system. This converts the hierarchical problem into a coupled leader-follower system and is one of the main structural reductions used throughout the controllability literature (Elgrou et al., 8 Feb 2025).

When nonlinearities destroy convexity, the equilibrium notion is weakened. In the nonlinear coupled degenerate parabolic system and in the semilinear degenerate equation on a non-cylindrical domain, the papers introduce a Nash quasi-equilibrium, defined through first-order optimality conditions rather than global minimization. Both works then prove that sufficiently large penalization parameters restore convexity, so the quasi-equilibrium becomes a genuine Nash equilibrium (Djomegne et al., 2022, Gamboa et al., 30 Sep 2025).

The equilibrium object can also be more elaborate than an ordinary Nash profile. In decision-dependent uncertainty models, the equilibrium combines feasibility of leaders and uncertainty, leader-level Nash optimality, and a weak or strong Pareto condition on the uncertainty, which is treated as a “virtual player.” In stopping games, the cited work distinguishes precommitment strategies, which are optimal only at the initial time, from equilibrium strategies, which are time-consistent under future deviations. This shows that Stackelberg-Nash structures may require equilibrium refinements beyond the classical static Nash definition (Zhang et al., 2022, Zhang et al., 26 Jul 2025).

3. Control-theoretic realizations

The most technically developed realizations in the cited corpus are hierarchical controllability problems for PDEs and stochastic evolution equations. In these works, the leader’s objective is typically exact, null, or approximate controllability, while the followers pursue tracking of prescribed targets with control-energy penalties. A representative target is

y(T,)=0in G,P-a.s.,y(T,\cdot)=0 \quad \text{in } G,\qquad \mathbb P\text{-a.s.},

which appears in stochastic parabolic control and related null controllability results (Oukdach et al., 2024, Elgrou et al., 8 Feb 2025).

This structure has been implemented for a linear stochastic parabolic equation with gradient terms, where one leader control acts in the drift and another in the diffusion; for the anisotropic heat equation with general dynamic boundary conditions and drift terms; and for abstract forward and backward stochastic evolution equations, where the theory is developed for exact, null, and approximate controllability in the Stackelberg-Nash sense (Oukdach et al., 2024, Boutaayamou et al., 2021, Elgrou et al., 8 Feb 2025).

The same hierarchy extends to degenerate and moving-domain settings. One-dimensional semilinear degenerate equations in non-cylindrical domains, nonlinear coupled degenerate parabolic systems, degenerate elliptic equations, and degenerate parabolic equations controlled from measurable sets all admit Stackelberg-Nash formulations in which the degeneracy is handled by weighted Sobolev spaces, degenerate Carleman estimates, or compact embedding arguments. In the free-boundary setting, the Stefan problem combines one leader and two followers, and the original problem is reduced to the null controllability of a nonlinear optimality system coupled to the moving interface (t)\ell(t) (Gamboa et al., 30 Sep 2025, Djomegne et al., 2022, Liu et al., 2023, Liu et al., 2023, Carvalho et al., 1 May 2026).

A standard reduction in these control papers is the substitution of the followers’ equilibrium feedback into the state equation. Once the Nash response is fixed, the leader no longer faces a game in the original variables but a coupled forward-backward or backward-forward optimality system. The cited works repeatedly identify this reduction as the technical core of the Stackelberg-Nash program, because controllability of the original hierarchical model is then equivalent to controllability of the reduced coupled system (Oukdach et al., 2024, Elgrou et al., 8 Feb 2025, Boutaayamou et al., 2021).

4. Analytical and computational machinery

In controllability problems, the dominant analytic tools are adjoint systems, Carleman estimates, observability inequalities, and duality or HUM-type arguments. The stochastic parabolic paper with gradient terms derives a Carleman estimate for the coupled adjoint system and then an observability inequality that yields null controllability. The abstract stochastic evolution paper establishes a duality between Stackelberg-Nash controllability and observability properties, and its heat-equation applications rely on Carleman estimates, including a new Carleman estimate for the backward stochastic heat equation (Oukdach et al., 2024, Elgrou et al., 8 Feb 2025).

Degenerate and moving-boundary problems require adapted weights and localization procedures. In the non-cylindrical semilinear degenerate setting, the proof of local null controllability uses a Carleman inequality for degenerated non-autonomous systems and then Liusternik’s inverse function theorem. In the Stefan problem, observability is obtained through Carleman estimates adapted to the moving boundary, approximate controllability is first proved for a linearized system, and a Schauder fixed-point argument is used to recover the nonlinear free boundary (Gamboa et al., 30 Sep 2025, Carvalho et al., 1 May 2026).

Existence of equilibrium is often handled by fixed-point theorems. Kakutani’s fixed-point theorem is used in the degenerate elliptic Stackelberg-Nash game, in the measurable-set degenerate parabolic game, and in Nash-Stackelberg-Nash games under decision-dependent uncertainties. In the PDE papers, this is typically combined with well-posedness of the state equation, convexity of follower costs, and compactness of the state solution map (Liu et al., 2023, Liu et al., 2023, Zhang et al., 2022).

Outside controllability, the strategy is also treated algorithmically. One paper proposes end-to-end gradient descent for Stackelberg games with multiple followers by differentiating through the followers’ best responses. Its key theoretical device is the equilibrium flow, a vector field v(x,θ)v(x,\theta) satisfying

θp(x,θ)=x(p(x,θ)v(x,θ)),\frac{\partial}{\partial \theta} p(x,\theta) = -\nabla_x\cdot\big(p(x,\theta)v(x,\theta)\big),

which allows sampled Nash equilibria to provide unbiased gradient estimates under appropriate conditions (Wang et al., 2021).

Related computational ideas appear in equilibrium problems with equilibrium constraints. In Nash games among Stackelberg players, exact computation is carried out by full enumeration and convexification or by an inner approximation algorithm. In the Lottery Colonel Blotto game, the Stackelberg equilibrium is solved by a constructive method based on iterative game reductions, which reduces the possible follower-support patterns from (v1,v2)(v_1^\star,v_2^\star)0 to (v1,v2)(v_1^\star,v_2^\star)1 (Carvalho et al., 2019, Liu et al., 2024).

5. Relation to Nash, Stackelberg, and neighboring concepts

The Stackelberg-Nash strategy is distinct from plain Stackelberg control and from plain Nash equilibrium. In the Stefan controllability paper, this distinction is stated explicitly: plain Stackelberg control typically has one leader and one follower objective, whereas plain Nash control has multiple players but no hierarchy. The Stackelberg-Nash framework instead embeds a follower Nash response inside a leader-follower timing structure (Carvalho et al., 1 May 2026).

Its relation to Nash equilibrium is domain-dependent. In a finite-horizon zero-sum stochastic LQ differential game, the cited paper shows that the Stackelberg equilibrium and the Nash equilibrium are actually identical under the stated uniform convexity-concavity conditions. The crucial identity is (v1,v2)(v_1^\star,v_2^\star)2, where the difference between the follower and leader Riccati equations coincides with the Riccati equation of the zero-sum Nash game (Sun et al., 2021).

Security games exhibit both coincidence and divergence. In a general class of security games, Nash equilibria are interchangeable, and under the SSAS condition any Stackelberg strategy is also a Nash equilibrium strategy; in an important homogeneous-resource class, the resulting strategy is unique. By contrast, when the attacker can attack multiple targets, many of these properties no longer hold. A three-player leader-follower security model gives necessary and sufficient conditions for the case that every Stackelberg equilibrium is a Nash equilibrium, another condition for the case that at least one Stackelberg equilibrium is a Nash equilibrium, and a Hausdorff-distance upper bound when exact coincidence fails (Korzhyk et al., 2014, Xu et al., 2022).

Other papers sharpen the contrast. In the Lottery Colonel Blotto game, the Stackelberg equilibrium coincides with the Nash equilibrium only under a specific structural condition, namely when the battlefield value ratios have special form and the leader/follower budget ratio equals a calculable threshold. In Stackelberg stopping games, precommitment, equilibrium, and Nash are shown by example to be genuinely different notions. In multi-follower Stackelberg optimization, optimistic or pessimistic equilibrium-selection assumptions can be arbitrarily bad under stochastic equilibrium selection, which motivates equilibrium-flow-based learning against sampled equilibria rather than against a fixed tie-breaking rule (Liu et al., 2024, Zhang et al., 26 Jul 2025, Wang et al., 2021).

There is also a conceptual relation to extensive-form equilibrium. In two-player normal-form games, a Stackelberg mixed strategy can be represented as a subgame perfect Nash equilibrium of a corresponding extensive-form game. The cited paper nevertheless argues that the normal-form concept should be studied separately, because the extensive-form representation has a continuum of first-stage actions and obscures the linear-program structure, the link to maximin/minimax, and the comparison with correlated equilibrium (Conitzer, 2017).

6. Applications, scope, and current directions

The cited applications are unusually broad. In control theory, the strategy is used for stochastic heat equations, backward stochastic heat equations, stochastic parabolic equations with gradient terms, anisotropic heat equations with dynamic boundary conditions, semilinear degenerate equations in non-cylindrical domains, nonlinear coupled degenerate parabolic systems, and Stefan free-boundary problems. In these settings the leader typically aims at null controllability, while followers track distinct targets on possibly distinct observation regions (Elgrou et al., 8 Feb 2025, Oukdach et al., 2024, Boutaayamou et al., 2021, Gamboa et al., 30 Sep 2025, Djomegne et al., 2022, Carvalho et al., 1 May 2026).

In operations research and economics, the idea appears in multi-channel inventory competition, where a manufacturer acts as leader and retailers best-respond in inventory levels under substitution effects. Under specified parameter restrictions, the paper proves uniqueness of the Stackelberg equilibrium; in the special case (v1,v2)(v_1^\star,v_2^\star)3, the manufacturer carries more inventory than in the simultaneous-move game (0906.0151).

In strategic resource allocation and security, the framework appears in security games, Colonel Blotto, cybersecurity defense, and secure transmission. In data-driven or learning-based models, it appears in Stackelberg problems with multiple followers in normal-form games, security games with multiple defenders, and cyber insurance games with multiple customers. In energy systems and climate-aware regulation, Nash games among Stackelberg players model regulators whose own decision problems already contain lower-level Stackelberg subgames, and the resulting equilibrium problem is an EPEC (Korzhyk et al., 2014, Liu et al., 2024, Wang et al., 2021, Carvalho et al., 2019).

Current developments in the cited works push the framework in three directions. One direction is toward more difficult state dynamics: degeneracy, stochasticity, moving domains, measurable control sets, and free boundaries. A second direction is toward richer equilibrium notions: quasi-equilibrium, weak and strong equilibrium under decision-dependent uncertainty, and time-consistent equilibrium in stopping games. A third direction is toward computation and learning, including differentiated KKT systems, sampled-equilibrium optimization, convexification of complementarity-defined feasible regions, and exact algorithms for hierarchical equilibrium problems (Liu et al., 2023, Zhang et al., 2022, Zhang et al., 26 Jul 2025, Wang et al., 2021, Carvalho et al., 2019).

Taken together, these works indicate that the Stackelberg-Nash strategy is best understood not as a single theorem or algorithm but as a general hierarchical paradigm. Its central move is always the same: leaders optimize against a Nash-type response layer. What changes across domains is the equilibrium object, the regularity needed to justify the reduction, and the analytic or computational machinery used to solve the resulting problem.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (19)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Stackelberg-Nash Strategy.