---
title: Nash Equilibrium Control in Dynamic Systems
url: https://www.emergentmind.com/topics/nash-equilibrium-control
type: topic
---

# Nash Equilibrium Control in Dynamic Systems

Nash equilibrium control denotes a family of control-theoretic constructions in which feedback laws, adaptive mechanisms, communication protocols, or causal interventions are designed so that interacting agents converge to a Nash equilibrium, remain within a quantified neighborhood of it, or are steered toward a selected equilibrium among several candidates. In the recent literature, the term spans model-free extremum-seeking schemes for quadratic noncooperative games, distributed primal-dual controllers for generalized Nash equilibrium problems, bounded and robust feedback for dynamical agents, hybrid and delayed differential games, PDE-governed bi-objective control, and equilibrium steering in evolutionary dynamics and large language models [2501.12256][1911.12266][2512.07327][2604.27167].

## 1. Problem classes and equilibrium notions

A common starting point is the quadratic noncooperative game. In one representative formulation, each of \(N\) players has payoff
\[
J_i(\theta)=\frac{1}{2}\sum_{j=1}^N\sum_{k=1}^N H_{jk}^i \theta_j\theta_k+\sum_{j=1}^N h_j^i\theta_j+c^i,
\]
with action vector \(\theta=[\theta_1,\ldots,\theta_N]^T\). The Nash equilibrium \(\theta^*\) satisfies
\[
H\theta^*+h=0,
\]
and under assumptions such as strict diagonal dominance, or negative definiteness together with diagonal dominance in the averaged error dynamics, the equilibrium is unique, with \(\theta^*=-H^{-1}h\) [2501.12256].

A second class is the generalized Nash equilibrium problem, in which agents optimize under shared constraints. One distributed non-model-based scheme addresses games with coupled equality constraints and local inequality constraints, using an exact penalty method for the local inequalities and distributed Lagrange multiplier estimates for the shared equality constraint. Another continuous-time framework studies strongly monotone games with convex separable coupling constraints and partial-decision information, and focuses on the variational GNE, characterized by a common dual multiplier and KKT inclusions of the form
\[
0\in F(x^*)+\nabla_x g(x^*)^\top \lambda^*+N_\Omega(x^*), \qquad
0\in -g(x^*)+N_{\mathbb{R}^m_{\ge 0}}(\lambda^*) .
\]
These formulations place Nash equilibrium control squarely within projected dynamical systems and monotone operator methods [2201.11003][1911.12266].

A third class is differential-game and optimal-control based. In a bi-objective fractional space-time parabolic PDE problem, a Nash equilibrium is a pair \((\bar u_1,\bar u_2)\) such that neither control can improve its own cost functional by unilateral deviation, and the equilibrium is characterized by a state equation, adjoint equations, and a variational inequality for each control [2512.07327]. In a deterministic finite-horizon two-player nonzero-sum differential game with one ordinary-control player and one impulse-control player, feedback Nash equilibrium strategies are defined by unilateral optimality over feedback policies and are verified through coupled HJB and QVI conditions [2106.10706].

## 2. Model-free extremum-seeking and averaging-based methods

A major research line uses extremum seeking to recover pseudogradient information from payoff measurements without explicit model knowledge. In Lie-bracket extremum seeking with bounded update rates, each player updates according to
\[
\dot{\theta}_i(t)=\sqrt{\alpha_i\omega_i}\cos(\omega_i t-k_i J_i(\theta)).
\]
The update rate \(\dot\theta_i(t)\) is bounded regardless of payoff magnitude, and the fast dithers permit a Lie-bracket average system
\[
\dot{\bar{\theta}}(t)=\frac{1}{2}AK\nabla J(\bar{\theta}).
\]
For quadratic games, the averaged error dynamics are linear, Lyapunov analysis yields exponential decay of the averaged error, and the actual trajectory satisfies an ultimate bound of order \(\mathcal{O}(1/\tilde\omega)\). The residual set vanishes as the dithering frequency increases and, in this scheme, is independent of the probing amplitude [2501.12256].

Event-triggered extremum seeking adapts the same model-free logic to limited-bandwidth actuation. In a duopoly game, each player constructs a demodulated pseudogradient \(\hat G_i(t)\) from a sinusoidally perturbed payoff signal and updates asynchronously using a static local trigger
\[
t_{\kappa+1}^i=\inf\{t>t_\kappa^i:\ |e_i(t)|\ge \sigma_i |\hat G_i(t)|\}.
\]
The analysis combines time-scaling, Lyapunov’s direct method, and averaging theory for discontinuous systems, and establishes local convergence to a small residual set of size \(\mathcal{O}(a+1/\omega)\) together with a strictly positive lower bound on inter-execution times, excluding Zeno behavior [2404.07287].

Sliding-mode Nash equilibrium seeking replaces the proportional extremum-seeking feedback by a relay law,
\[
u_i(t)=K_i\,\operatorname{sgn}(\hat G_i(t)),
\]
again using sinusoidal perturbations for pseudogradient estimation. The averaged system converges in finite time under a strict diagonal dominance condition on \(HK\), and the original trajectories converge to a neighborhood of the Nash equilibrium whose size is \(\mathcal{O}(a+1/\omega)\) [2405.15762].

A further development targets constrained GNEs. A distributed non-model-based algorithm combines extremum seeking with learning dynamics, communicates only Lagrange multiplier estimates and an auxiliary dual variable, and employs a diminishing dither amplitude \(a_i\) that vanishes near equilibrium. The convergence analysis is non-local and uses singular perturbation theory, averaging analysis, and Lyapunov stability theory. In contrast to classical fixed-amplitude ESC, the diminishing dither is designed to remove undesirable steady-state oscillations [2201.11003].

| Setting | Control idea | Guarantee |
|---|---|---|
| Quadratic \(N\)-player game | Lie-bracket ES with bounded update rates | Local convergence with \(\mathcal{O}(1/\tilde\omega)\) residual |
| Duopoly game | Event-triggered ES with sinusoidal probing | Local convergence with \(\mathcal{O}(a+1/\omega)\) residual |
| Duopoly game | Sliding-mode ES with relay feedback | Finite-time convergence of averaged system |
| Constrained GNE | ESC with learning and diminishing dither | Non-local convergence; oscillations removed |

These methods share a structural theme: the controller does not require analytic payoff gradients, but extracts a gradient-like signal from measurements and then stabilizes the resulting slow dynamics. A recurring distinction is that the strongest guarantees are typically for an averaged system, while the original system is shown to track the average up to a residual term determined by dither frequency, dither amplitude, or triggering error.

## 3. Distributed feedback for dynamical players under constraints and uncertainty

When the agents are dynamical systems rather than static decision variables, Nash equilibrium control typically decomposes into an optimization layer and a regulation layer. For first-order and second-order integrator-type players with bounded control inputs, saturated gradient-play laws and consensus-based distributed estimators yield bounded inputs by construction. In the first-order distributed case, each agent maintains estimates \(y_{ij}\) of all actions and updates
\[
\dot{x}_i=-\sigma_{\bar U}\!\left(\partial_{x_i} f_i(\mathbf y_i)\right),
\]
together with a consensus protocol for \(\mathbf y_i\). For second-order systems, auxiliary variables \(z_i\) and estimator states are added, and saturation is again inserted to enforce \(|u_i|\le \bar U\). Lyapunov analysis gives convergence, with semi-global convergence for the second-order distributed bounded-input case [1901.09333].

For higher-order players, one fully distributed strategy first applies a linear transformation to bring each agent into a controllable canonical form and then uses multiple saturation functions in a bounded controller. Consensus estimators with adaptive and time-varying gains allow the communication graph to be directed and avoid global parameter knowledge. Under globally Lipschitz gradients, strong monotonicity of the game Jacobian, and strong connectivity of the digraph, the paper proves
\[
\lim_{t\to\infty}\|\mathbf y(t)-\mathbf y^*\|=0
\]
while preserving bounded controls at all times [2108.06573].

Adaptive consensus design is another route to full distributivity. Node-based and edge-based adaptive laws update local consensus weights from instantaneous squared consensus errors, removing the need for a globally chosen singular perturbation gain. Under twice-continuously differentiable objectives, globally Lipschitz partial derivatives, strong monotonicity of the pseudo-gradient, and an undirected connected graph, LaSalle’s invariance principle yields global asymptotic stability of the Nash equilibrium. The edge-based method also extends to switching among undirected connected graphs [1912.00415].

Generalized equilibrium seeking with dynamical agents has been developed in continuous time through projected primal-dual dynamics. For single-integrator agents, controllers combine local projected gradients, consensus terms on strategy estimates, and dual consensus dynamics; an adaptive-weight variant replaces the fixed global consensus gain by uncoordinated integral adaptation. The same framework is then extended to aggregative games, heterogeneous multi-integrator agents, and nonlinear feedback-linearizable systems, with convergence to a variational equilibrium established through monotonicity properties and stability theory for projected dynamical systems [1911.12266].

Uncertain agent dynamics introduce a different control problem. With unknown control directions and parametric uncertainties, a modular design separates an optimization module,
\[
\dot y_i=-\nabla_i f_i(\mathbf z_i),
\]
from a state-regulation module containing Nussbaum functions and adaptive laws. This permits asymptotic distributed Nash seeking without requiring homogeneity of the unknown control directions [2009.12748]. With disturbances and unmodeled dynamics treated as extended states, PI-based observers drive actions to a small neighborhood of the Nash equilibrium, whereas a RISE-based observer yields asymptotic convergence to the equilibrium itself [1902.00901]. For a special class of multi-agent linear systems, an integral Nash equilibrium seeking control law eliminates the proportional term used in earlier extremum-seeking designs; in the limited-information case it converges to a neighborhood of the equilibrium, and the reported simulations require smaller perturbation frequencies and amplitudes than the compared state-of-the-art method [1911.09409].

## 4. Differential games, delay systems, and PDE-governed Nash control

In deterministic state-delay systems, Nash equilibrium control has been developed for a two-player continuous-time LQ problem with generalized cross terms. The plant
\[
\dot x(t)=A_0x(t)+A_1x(t-h)+B_1u_1(t)+B_2u_2(t)
\]
is paired with quadratic costs containing instantaneous, delayed, and cross terms. The candidate feedback has the form
\[
u_i(x_t)=\Gamma_{i,0}x(t)+\int_{-h}^0 \Gamma_{i,1}(\theta)x(t+\theta)\,d\theta,
\]
and Bellman functionals lead to coupled Riccati-type PDEs. The paper derives an explicit Nash feedback
\[
u_i^*(x_t)=-R_{i,i}^{-1}B_i^\top\!\left[\Pi_{i,0}x(t)+\int_{-h}^0 \Pi_{i,1}(\theta)x(t+\theta)\,d\theta\right],
\]
and uses Lyapunov-Krasovskii methods to show asymptotic stability of the closed loop. In a thermal prototype with an ESP32 micro-controller, the Nash strategy yields lower IAE, ITSE, and ITAE than the compared optimal PI controllers, although ISE is higher because of the initial transient [2606.08344].

Hybrid differential games with impulse control require a different equilibrium machinery. In a two-player nonzero-sum finite-horizon game with ordinary controls for Player 1 and impulse controls for Player 2, a verification theorem characterizes feedback Nash equilibrium by an HJB equation for the ordinary controller and a QVI for the impulse controller. The intervention operator
\[
\mathcal RV_2(t,x)=\min_{\eta\in\Omega_2} V_2(t,x+g(x,\eta))+b_2(x,\eta)
\]
defines continuation and intervention regions, and the equilibrium number of impulses is upper bounded by
\[
K=\left\lceil \frac{2\left(T\|h_2\|_\infty+\|s_2\|_\infty\right)}{\mu}\right\rceil .
\]
For a scalar linear-quadratic example, the paper gives an analytical characterization of threshold-based equilibrium policies [2106.10706].

PDE-governed Nash control extends these ideas to infinite-dimensional systems. In a bi-objective optimal control problem for a fractional space-time parabolic PDE,
\[
\partial_t^\gamma w+(-\Delta)^s w = B_1u_1+B_2u_2+f,
\]
each player minimizes
\[
J_j(w,u_1,u_2)=\frac12\int_0^T\left(\|w-w_j\|_{L^2(\Omega)}^2+\mu_j\|B_ju_j\|_{L^2(\Omega)}^2\right)\,dt .
\]
Under convexity and coercivity assumptions, the Nash equilibrium is unique and is characterized by the coupled state-adjoint-variational inequality system. The solution is computed with conjugate gradient algorithms applied iteratively to the discretized control problems, and the numerical experiments agree with the theoretical estimates [2512.07327].

## 5. Networked and cyber-physical implementations

Several works treat Nash equilibrium control as a synthesis problem for networked physical systems. One distributed feedback controller steers a class of passive nonlinear second-order networks to a prescribed Cournot-Nash equilibrium. Production is locally controllable, demand is price-responsive, and each node updates a local price variable by
\[
\tau_i\dot p_i=-k_i y_i-Q_{di}^{-1}y_i-\sum_{j\in\mathcal N_i^c}\rho_{ij}(p_i-p_j).
\]
With \(k_i=(\alpha^*+Q_{gi})^{-1}\), the closed-loop system has a unique equilibrium corresponding to the optimal Cournot-Nash point, and Lyapunov plus LaSalle arguments give global asymptotic convergence using only reduced demand information [1803.03593].

Security-motivated control appears in CAN networks under bus-off attacks. The plant is modeled as
\[
x_{t+1}=Ax_t+(1-\beta_t)\alpha_t Bu_t+v_t,
\]
where \(\alpha_t\) is the transmitter decision and \(\beta_t\) is the attacker decision. The problem is cast as a non-zero-sum LQG game between the controller-transmitter pair and the attacker. Under both closed-loop and open-loop attacker information structures, the attacker has a dominant strategy; under that strategy, the optimal control law is linear in the system state,
\[
u_t^*=K_t x_t
\]
in finite horizon and \(u_t^*=K_\infty x_t\) in infinite horizon. A necessary and sufficient condition for bounded average cost is that the effective successful-transmission probability satisfies
\[
p(1-p)>\rho_{\min},
\]
where \(\rho_{\min}\) depends on the unstable eigenvalues of \(A\) [2111.04950].

Large-population production control with sticky prices leads to a mean field formulation. Each firm’s output follows
\[
dX_t^{i,u}=X_t^{i,u}(-\mu_i dt+\sigma_i dW_t^i)+u_t^i dt,
\]
while price obeys a controlled jump-diffusion with mean-field coupling. Solving the limiting control problem and a fixed-point problem yields an explicit feedback
\[
u_t^{*,i}=\frac{1}{2r_i}\bar g_t^i,
\]
and the resulting profile is an \(\epsilon_n\)-Nash equilibrium for the \(n\)-firm game, with \(\epsilon_n\to 0\) as \(n\to\infty\) [2204.03300].

These applications show that Nash equilibrium control is not confined to abstract game dynamics. It also functions as a design language for pricing layers in physical networks, attack-defense policies in networked control, and decentralized control laws in large stochastic markets.

## 6. Equilibrium selection, causal intervention, and recurring distinctions

A central distinction in the literature is between equilibrium seeking and equilibrium selection. In evolutionary game dynamics, equilibrium selection can itself be controlled by eigenvalue assignment. For replicator dynamics linearized at an equilibrium,
\[
\dot x \approx J^o(x-x^*),
\]
a feedback law modifies the Jacobian to
\[
J^c=J^o+BK+T,
\]
with a tax term chosen to preserve the equilibrium and budget balance. By assigning the closed-loop poles, one may destabilize one Nash equilibrium and strengthen the attraction of another. The cited study describes this as the first realization of control of equilibrium selection by design in the game dynamics theory paradigm [2302.09131].

A recent non-classical extension treats Nash behavior inside large language models as a causal control problem. In the mechanistic study of four open-source models playing four canonical two-player games, opponent history is encoded with near-perfect fidelity at the first layer, while Nash action encoding remains weak throughout and no dedicated Nash module is found. The model privately favors the Nash action through most of its forward pass, but a prosocial override concentrated in the final layers reverses this. Injecting a learned Nash direction into the residual stream shifts behavior bidirectionally, and concept clamping confirms monotonic control over cooperation probabilities. This suggests a new sense of “Nash equilibrium control”: inference-time steering of equilibrium play rather than synthesis of a controller for a physical or multi-agent dynamical system [2604.27167].

Several recurring distinctions cut across the field. First, model-free does not imply communication-free: some extremum-seeking schemes require only each player’s own payoff measurements and no action sharing, whereas constrained GNE and consensus-based controllers exchange Lagrange multipliers, local estimates, or prices [2501.12256][2201.11003][1803.03593]. Second, exact convergence is not universal: bounded-rate Lie-bracket ES, event-triggered ES, sliding-mode ES, PI-observer-based robust seeking, and integral NES often establish convergence to a small residual set or neighborhood, while diminishing-dither GNE seeking, RISE-observer-based designs, and several variational or optimal-control formulations establish asymptotic or exact equilibrium convergence under stronger assumptions [2404.07287][2405.15762][1902.00901][2512.07327]. Third, convergence domains differ substantially: some schemes are explicitly local, especially in quadratic duopoly extremum-seeking settings, whereas others are non-local, global, or semi-global, depending on monotonicity, convexity, connectivity, and dynamic-order assumptions [2501.12256][2201.11003][1901.09333][1912.00415].

Taken together, the literature portrays Nash equilibrium control as a heterogeneous but coherent area at the intersection of game theory, nonlinear control, distributed optimization, and dynamical systems. Its unifying problem is not merely to compute an equilibrium offline, but to embed equilibrium behavior into feedback laws, communication protocols, hybrid intervention rules, or causal steering mechanisms so that strategic interaction and closed-loop dynamics become analytically tractable and, in many settings, implementable.

Source: https://www.emergentmind.com/topics/nash-equilibrium-control