Papers
Topics
Authors
Recent
Search
2000 character limit reached

Nash Equilibrium Control in Dynamic Systems

Updated 8 July 2026
  • Nash equilibrium control is a framework that designs feedback laws and adaptive protocols ensuring agents converge to equilibrium in multi-agent settings.
  • It employs methods like model-free extremum seeking, distributed consensus, and sliding-mode strategies to handle system constraints and uncertainties.
  • Applications include differential games, PDE control, and networked systems, demonstrating its versatility in dynamic and uncertain environments.

Nash equilibrium control denotes a family of control-theoretic constructions in which feedback laws, adaptive mechanisms, communication protocols, or causal interventions are designed so that interacting agents converge to a Nash equilibrium, remain within a quantified neighborhood of it, or are steered toward a selected equilibrium among several candidates. In the recent literature, the term spans model-free extremum-seeking schemes for quadratic noncooperative games, distributed primal-dual controllers for generalized Nash equilibrium problems, bounded and robust feedback for dynamical agents, hybrid and delayed differential games, PDE-governed bi-objective control, and equilibrium steering in evolutionary dynamics and LLMs (Rodrigues et al., 21 Jan 2025, Bianchi et al., 2019, Buda et al., 8 Dec 2025, Lekeas et al., 29 Apr 2026).

1. Problem classes and equilibrium notions

A common starting point is the quadratic noncooperative game. In one representative formulation, each of NN players has payoff

Ji(θ)=12j=1Nk=1NHjkiθjθk+j=1Nhjiθj+ci,J_i(\theta)=\frac{1}{2}\sum_{j=1}^N\sum_{k=1}^N H_{jk}^i \theta_j\theta_k+\sum_{j=1}^N h_j^i\theta_j+c^i,

with action vector θ=[θ1,,θN]T\theta=[\theta_1,\ldots,\theta_N]^T. The Nash equilibrium θ\theta^* satisfies

Hθ+h=0,H\theta^*+h=0,

and under assumptions such as strict diagonal dominance, or negative definiteness together with diagonal dominance in the averaged error dynamics, the equilibrium is unique, with θ=H1h\theta^*=-H^{-1}h (Rodrigues et al., 21 Jan 2025).

A second class is the generalized Nash equilibrium problem, in which agents optimize under shared constraints. One distributed non-model-based scheme addresses games with coupled equality constraints and local inequality constraints, using an exact penalty method for the local inequalities and distributed Lagrange multiplier estimates for the shared equality constraint. Another continuous-time framework studies strongly monotone games with convex separable coupling constraints and partial-decision information, and focuses on the variational GNE, characterized by a common dual multiplier and KKT inclusions of the form

0F(x)+xg(x)λ+NΩ(x),0g(x)+NR0m(λ).0\in F(x^*)+\nabla_x g(x^*)^\top \lambda^*+N_\Omega(x^*), \qquad 0\in -g(x^*)+N_{\mathbb{R}^m_{\ge 0}}(\lambda^*) .

These formulations place Nash equilibrium control squarely within projected dynamical systems and monotone operator methods (Xiao et al., 2022, Bianchi et al., 2019).

A third class is differential-game and optimal-control based. In a bi-objective fractional space-time parabolic PDE problem, a Nash equilibrium is a pair (uˉ1,uˉ2)(\bar u_1,\bar u_2) such that neither control can improve its own cost functional by unilateral deviation, and the equilibrium is characterized by a state equation, adjoint equations, and a variational inequality for each control (Buda et al., 8 Dec 2025). In a deterministic finite-horizon two-player nonzero-sum differential game with one ordinary-control player and one impulse-control player, feedback Nash equilibrium strategies are defined by unilateral optimality over feedback policies and are verified through coupled HJB and QVI conditions (Sadana et al., 2021).

2. Model-free extremum-seeking and averaging-based methods

A major research line uses extremum seeking to recover pseudogradient information from payoff measurements without explicit model knowledge. In Lie-bracket extremum seeking with bounded update rates, each player updates according to

θ˙i(t)=αiωicos(ωitkiJi(θ)).\dot{\theta}_i(t)=\sqrt{\alpha_i\omega_i}\cos(\omega_i t-k_i J_i(\theta)).

The update rate θ˙i(t)\dot\theta_i(t) is bounded regardless of payoff magnitude, and the fast dithers permit a Lie-bracket average system

Ji(θ)=12j=1Nk=1NHjkiθjθk+j=1Nhjiθj+ci,J_i(\theta)=\frac{1}{2}\sum_{j=1}^N\sum_{k=1}^N H_{jk}^i \theta_j\theta_k+\sum_{j=1}^N h_j^i\theta_j+c^i,0

For quadratic games, the averaged error dynamics are linear, Lyapunov analysis yields exponential decay of the averaged error, and the actual trajectory satisfies an ultimate bound of order Ji(θ)=12j=1Nk=1NHjkiθjθk+j=1Nhjiθj+ci,J_i(\theta)=\frac{1}{2}\sum_{j=1}^N\sum_{k=1}^N H_{jk}^i \theta_j\theta_k+\sum_{j=1}^N h_j^i\theta_j+c^i,1. The residual set vanishes as the dithering frequency increases and, in this scheme, is independent of the probing amplitude (Rodrigues et al., 21 Jan 2025).

Event-triggered extremum seeking adapts the same model-free logic to limited-bandwidth actuation. In a duopoly game, each player constructs a demodulated pseudogradient Ji(θ)=12j=1Nk=1NHjkiθjθk+j=1Nhjiθj+ci,J_i(\theta)=\frac{1}{2}\sum_{j=1}^N\sum_{k=1}^N H_{jk}^i \theta_j\theta_k+\sum_{j=1}^N h_j^i\theta_j+c^i,2 from a sinusoidally perturbed payoff signal and updates asynchronously using a static local trigger

Ji(θ)=12j=1Nk=1NHjkiθjθk+j=1Nhjiθj+ci,J_i(\theta)=\frac{1}{2}\sum_{j=1}^N\sum_{k=1}^N H_{jk}^i \theta_j\theta_k+\sum_{j=1}^N h_j^i\theta_j+c^i,3

The analysis combines time-scaling, Lyapunov’s direct method, and averaging theory for discontinuous systems, and establishes local convergence to a small residual set of size Ji(θ)=12j=1Nk=1NHjkiθjθk+j=1Nhjiθj+ci,J_i(\theta)=\frac{1}{2}\sum_{j=1}^N\sum_{k=1}^N H_{jk}^i \theta_j\theta_k+\sum_{j=1}^N h_j^i\theta_j+c^i,4 together with a strictly positive lower bound on inter-execution times, excluding Zeno behavior (Rodrigues et al., 2024).

Sliding-mode Nash equilibrium seeking replaces the proportional extremum-seeking feedback by a relay law,

Ji(θ)=12j=1Nk=1NHjkiθjθk+j=1Nhjiθj+ci,J_i(\theta)=\frac{1}{2}\sum_{j=1}^N\sum_{k=1}^N H_{jk}^i \theta_j\theta_k+\sum_{j=1}^N h_j^i\theta_j+c^i,5

again using sinusoidal perturbations for pseudogradient estimation. The averaged system converges in finite time under a strict diagonal dominance condition on Ji(θ)=12j=1Nk=1NHjkiθjθk+j=1Nhjiθj+ci,J_i(\theta)=\frac{1}{2}\sum_{j=1}^N\sum_{k=1}^N H_{jk}^i \theta_j\theta_k+\sum_{j=1}^N h_j^i\theta_j+c^i,6, and the original trajectories converge to a neighborhood of the Nash equilibrium whose size is Ji(θ)=12j=1Nk=1NHjkiθjθk+j=1Nhjiθj+ci,J_i(\theta)=\frac{1}{2}\sum_{j=1}^N\sum_{k=1}^N H_{jk}^i \theta_j\theta_k+\sum_{j=1}^N h_j^i\theta_j+c^i,7 (Rodrigues et al., 2024).

A further development targets constrained GNEs. A distributed non-model-based algorithm combines extremum seeking with learning dynamics, communicates only Lagrange multiplier estimates and an auxiliary dual variable, and employs a diminishing dither amplitude Ji(θ)=12j=1Nk=1NHjkiθjθk+j=1Nhjiθj+ci,J_i(\theta)=\frac{1}{2}\sum_{j=1}^N\sum_{k=1}^N H_{jk}^i \theta_j\theta_k+\sum_{j=1}^N h_j^i\theta_j+c^i,8 that vanishes near equilibrium. The convergence analysis is non-local and uses singular perturbation theory, averaging analysis, and Lyapunov stability theory. In contrast to classical fixed-amplitude ESC, the diminishing dither is designed to remove undesirable steady-state oscillations (Xiao et al., 2022).

Setting Control idea Guarantee
Quadratic Ji(θ)=12j=1Nk=1NHjkiθjθk+j=1Nhjiθj+ci,J_i(\theta)=\frac{1}{2}\sum_{j=1}^N\sum_{k=1}^N H_{jk}^i \theta_j\theta_k+\sum_{j=1}^N h_j^i\theta_j+c^i,9-player game Lie-bracket ES with bounded update rates Local convergence with θ=[θ1,,θN]T\theta=[\theta_1,\ldots,\theta_N]^T0 residual
Duopoly game Event-triggered ES with sinusoidal probing Local convergence with θ=[θ1,,θN]T\theta=[\theta_1,\ldots,\theta_N]^T1 residual
Duopoly game Sliding-mode ES with relay feedback Finite-time convergence of averaged system
Constrained GNE ESC with learning and diminishing dither Non-local convergence; oscillations removed

These methods share a structural theme: the controller does not require analytic payoff gradients, but extracts a gradient-like signal from measurements and then stabilizes the resulting slow dynamics. A recurring distinction is that the strongest guarantees are typically for an averaged system, while the original system is shown to track the average up to a residual term determined by dither frequency, dither amplitude, or triggering error.

3. Distributed feedback for dynamical players under constraints and uncertainty

When the agents are dynamical systems rather than static decision variables, Nash equilibrium control typically decomposes into an optimization layer and a regulation layer. For first-order and second-order integrator-type players with bounded control inputs, saturated gradient-play laws and consensus-based distributed estimators yield bounded inputs by construction. In the first-order distributed case, each agent maintains estimates θ=[θ1,,θN]T\theta=[\theta_1,\ldots,\theta_N]^T2 of all actions and updates

θ=[θ1,,θN]T\theta=[\theta_1,\ldots,\theta_N]^T3

together with a consensus protocol for θ=[θ1,,θN]T\theta=[\theta_1,\ldots,\theta_N]^T4. For second-order systems, auxiliary variables θ=[θ1,,θN]T\theta=[\theta_1,\ldots,\theta_N]^T5 and estimator states are added, and saturation is again inserted to enforce θ=[θ1,,θN]T\theta=[\theta_1,\ldots,\theta_N]^T6. Lyapunov analysis gives convergence, with semi-global convergence for the second-order distributed bounded-input case (Ye, 2019).

For higher-order players, one fully distributed strategy first applies a linear transformation to bring each agent into a controllable canonical form and then uses multiple saturation functions in a bounded controller. Consensus estimators with adaptive and time-varying gains allow the communication graph to be directed and avoid global parameter knowledge. Under globally Lipschitz gradients, strong monotonicity of the game Jacobian, and strong connectivity of the digraph, the paper proves

θ=[θ1,,θN]T\theta=[\theta_1,\ldots,\theta_N]^T7

while preserving bounded controls at all times (Ye et al., 2021).

Adaptive consensus design is another route to full distributivity. Node-based and edge-based adaptive laws update local consensus weights from instantaneous squared consensus errors, removing the need for a globally chosen singular perturbation gain. Under twice-continuously differentiable objectives, globally Lipschitz partial derivatives, strong monotonicity of the pseudo-gradient, and an undirected connected graph, LaSalle’s invariance principle yields global asymptotic stability of the Nash equilibrium. The edge-based method also extends to switching among undirected connected graphs (Ye et al., 2019).

Generalized equilibrium seeking with dynamical agents has been developed in continuous time through projected primal-dual dynamics. For single-integrator agents, controllers combine local projected gradients, consensus terms on strategy estimates, and dual consensus dynamics; an adaptive-weight variant replaces the fixed global consensus gain by uncoordinated integral adaptation. The same framework is then extended to aggregative games, heterogeneous multi-integrator agents, and nonlinear feedback-linearizable systems, with convergence to a variational equilibrium established through monotonicity properties and stability theory for projected dynamical systems (Bianchi et al., 2019).

Uncertain agent dynamics introduce a different control problem. With unknown control directions and parametric uncertainties, a modular design separates an optimization module,

θ=[θ1,,θN]T\theta=[\theta_1,\ldots,\theta_N]^T8

from a state-regulation module containing Nussbaum functions and adaptive laws. This permits asymptotic distributed Nash seeking without requiring homogeneity of the unknown control directions (Ye et al., 2020). With disturbances and unmodeled dynamics treated as extended states, PI-based observers drive actions to a small neighborhood of the Nash equilibrium, whereas a RISE-based observer yields asymptotic convergence to the equilibrium itself (Ye, 2019). For a special class of multi-agent linear systems, an integral Nash equilibrium seeking control law eliminates the proportional term used in earlier extremum-seeking designs; in the limited-information case it converges to a neighborhood of the equilibrium, and the reported simulations require smaller perturbation frequencies and amplitudes than the compared state-of-the-art method (Krilašević et al., 2019).

4. Differential games, delay systems, and PDE-governed Nash control

In deterministic state-delay systems, Nash equilibrium control has been developed for a two-player continuous-time LQ problem with generalized cross terms. The plant

θ=[θ1,,θN]T\theta=[\theta_1,\ldots,\theta_N]^T9

is paired with quadratic costs containing instantaneous, delayed, and cross terms. The candidate feedback has the form

θ\theta^*0

and Bellman functionals lead to coupled Riccati-type PDEs. The paper derives an explicit Nash feedback

θ\theta^*1

and uses Lyapunov-Krasovskii methods to show asymptotic stability of the closed loop. In a thermal prototype with an ESP32 micro-controller, the Nash strategy yields lower IAE, ITSE, and ITAE than the compared optimal PI controllers, although ISE is higher because of the initial transient (Castro et al., 6 Jun 2026).

Hybrid differential games with impulse control require a different equilibrium machinery. In a two-player nonzero-sum finite-horizon game with ordinary controls for Player 1 and impulse controls for Player 2, a verification theorem characterizes feedback Nash equilibrium by an HJB equation for the ordinary controller and a QVI for the impulse controller. The intervention operator

θ\theta^*2

defines continuation and intervention regions, and the equilibrium number of impulses is upper bounded by

θ\theta^*3

For a scalar linear-quadratic example, the paper gives an analytical characterization of threshold-based equilibrium policies (Sadana et al., 2021).

PDE-governed Nash control extends these ideas to infinite-dimensional systems. In a bi-objective optimal control problem for a fractional space-time parabolic PDE,

θ\theta^*4

each player minimizes

θ\theta^*5

Under convexity and coercivity assumptions, the Nash equilibrium is unique and is characterized by the coupled state-adjoint-variational inequality system. The solution is computed with conjugate gradient algorithms applied iteratively to the discretized control problems, and the numerical experiments agree with the theoretical estimates (Buda et al., 8 Dec 2025).

5. Networked and cyber-physical implementations

Several works treat Nash equilibrium control as an overview problem for networked physical systems. One distributed feedback controller steers a class of passive nonlinear second-order networks to a prescribed Cournot-Nash equilibrium. Production is locally controllable, demand is price-responsive, and each node updates a local price variable by

θ\theta^*6

With θ\theta^*7, the closed-loop system has a unique equilibrium corresponding to the optimal Cournot-Nash point, and Lyapunov plus LaSalle arguments give global asymptotic convergence using only reduced demand information (Persis et al., 2018).

Security-motivated control appears in CAN networks under bus-off attacks. The plant is modeled as

θ\theta^*8

where θ\theta^*9 is the transmitter decision and Hθ+h=0,H\theta^*+h=0,0 is the attacker decision. The problem is cast as a non-zero-sum LQG game between the controller-transmitter pair and the attacker. Under both closed-loop and open-loop attacker information structures, the attacker has a dominant strategy; under that strategy, the optimal control law is linear in the system state,

Hθ+h=0,H\theta^*+h=0,1

in finite horizon and Hθ+h=0,H\theta^*+h=0,2 in infinite horizon. A necessary and sufficient condition for bounded average cost is that the effective successful-transmission probability satisfies

Hθ+h=0,H\theta^*+h=0,3

where Hθ+h=0,H\theta^*+h=0,4 depends on the unstable eigenvalues of Hθ+h=0,H\theta^*+h=0,5 (Tang et al., 2021).

Large-population production control with sticky prices leads to a mean field formulation. Each firm’s output follows

Hθ+h=0,H\theta^*+h=0,6

while price obeys a controlled jump-diffusion with mean-field coupling. Solving the limiting control problem and a fixed-point problem yields an explicit feedback

Hθ+h=0,H\theta^*+h=0,7

and the resulting profile is an Hθ+h=0,H\theta^*+h=0,8-Nash equilibrium for the Hθ+h=0,H\theta^*+h=0,9-firm game, with θ=H1h\theta^*=-H^{-1}h0 as θ=H1h\theta^*=-H^{-1}h1 (Jiang et al., 2022).

These applications show that Nash equilibrium control is not confined to abstract game dynamics. It also functions as a design language for pricing layers in physical networks, attack-defense policies in networked control, and decentralized control laws in large stochastic markets.

6. Equilibrium selection, causal intervention, and recurring distinctions

A central distinction in the literature is between equilibrium seeking and equilibrium selection. In evolutionary game dynamics, equilibrium selection can itself be controlled by eigenvalue assignment. For replicator dynamics linearized at an equilibrium,

θ=H1h\theta^*=-H^{-1}h2

a feedback law modifies the Jacobian to

θ=H1h\theta^*=-H^{-1}h3

with a tax term chosen to preserve the equilibrium and budget balance. By assigning the closed-loop poles, one may destabilize one Nash equilibrium and strengthen the attraction of another. The cited study describes this as the first realization of control of equilibrium selection by design in the game dynamics theory paradigm (Zhijian, 2023).

A recent non-classical extension treats Nash behavior inside LLMs as a causal control problem. In the mechanistic study of four open-source models playing four canonical two-player games, opponent history is encoded with near-perfect fidelity at the first layer, while Nash action encoding remains weak throughout and no dedicated Nash module is found. The model privately favors the Nash action through most of its forward pass, but a prosocial override concentrated in the final layers reverses this. Injecting a learned Nash direction into the residual stream shifts behavior bidirectionally, and concept clamping confirms monotonic control over cooperation probabilities. This suggests a new sense of “Nash equilibrium control”: inference-time steering of equilibrium play rather than synthesis of a controller for a physical or multi-agent dynamical system (Lekeas et al., 29 Apr 2026).

Several recurring distinctions cut across the field. First, model-free does not imply communication-free: some extremum-seeking schemes require only each player’s own payoff measurements and no action sharing, whereas constrained GNE and consensus-based controllers exchange Lagrange multipliers, local estimates, or prices (Rodrigues et al., 21 Jan 2025, Xiao et al., 2022, Persis et al., 2018). Second, exact convergence is not universal: bounded-rate Lie-bracket ES, event-triggered ES, sliding-mode ES, PI-observer-based robust seeking, and integral NES often establish convergence to a small residual set or neighborhood, while diminishing-dither GNE seeking, RISE-observer-based designs, and several variational or optimal-control formulations establish asymptotic or exact equilibrium convergence under stronger assumptions (Rodrigues et al., 2024, Rodrigues et al., 2024, Ye, 2019, Buda et al., 8 Dec 2025). Third, convergence domains differ substantially: some schemes are explicitly local, especially in quadratic duopoly extremum-seeking settings, whereas others are non-local, global, or semi-global, depending on monotonicity, convexity, connectivity, and dynamic-order assumptions (Rodrigues et al., 21 Jan 2025, Xiao et al., 2022, Ye, 2019, Ye et al., 2019).

Taken together, the literature portrays Nash equilibrium control as a heterogeneous but coherent area at the intersection of game theory, nonlinear control, distributed optimization, and dynamical systems. Its unifying problem is not merely to compute an equilibrium offline, but to embed equilibrium behavior into feedback laws, communication protocols, hybrid intervention rules, or causal steering mechanisms so that strategic interaction and closed-loop dynamics become analytically tractable and, in many settings, implementable.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (19)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Nash Equilibrium Control.