- The paper develops an estimation-based best-response algorithm that converges to the unique Nash equilibrium under multi-step delays when a Lyapunov–Krasovskii LMI is feasible.
- For one-step delays, the method achieves deterministic exponential convergence when the learning rate satisfies ξ < δ₁ and provably diverges when ξ exceeds δ₂.
- Simulations show that stable learning rates generally decrease as delay and network size grow, while the exact LMI-based stability boundary remains topology-dependent and conservative.
Problem setting and motivation
The paper studies distributed Nash equilibrium (NE) seeking in a static non-cooperative quadratic game in which each agent maximizes a private quadratic payoff Ji(s)=21s⊤Ais+bi⊤s+gi, with aiii<0. Under a strict diagonal-dominance condition on each Ai (Assumption 1), the NE exists, is unique, and equals (I−M)−1c for the best-response matrix M and offset c. The distinguishing feature of the setting is that agents communicate over an undirected connected graph and can only exchange τ-step-delayed strategy and estimation information: at stage t, agent i receives only sj(t−τ) and aiii<00 from neighbors. The authors motivate this both by communication latency in networked systems and by strategic privacy concerns—agents may deliberately disclose only outdated strategies to conceal their current decisions.
The core algorithmic idea is an estimation-based best-response scheme. Each agent maintains estimates aiii<01 of every other agent's current strategy and plays the best response to its own estimate vector. Estimates are updated by a consensus-plus-correction rule combining neighbor estimation disagreement aiii<02 with the innovation aiii<03, scaled by a learning rate aiii<04. Stacking all quantities yields a linear delay system of dimension aiii<05 driven by matrices built from the Laplacian aiii<06, the adjacency-derived diagonal aiii<07, and the best-response structure aiii<08.
Convergence under multi-step delays (aiii<09)
For Ai0, the closed-loop dynamics take the form Ai1 after differencing. Theorem 1 establishes global asymptotic convergence of the strategy profile to the NE provided there exist positive definite matrices Ai2 satisfying a linear-matrix-inequality-type condition Ai3 of dimension Ai4. The proof constructs a Lyapunov–Krasovskii functional with three terms—a quadratic term on an augmented state, a delay-window sum, and a summation-inequality-based term following Seuret, Gouaisbaut, and Fridman—and shows the forward difference is negative definite. Equilibrium analysis then confirms that the unique equilibrium of the augmented system coincides with the augmented NE state, using irreducible diagonal dominance of Ai5.
Two caveats deserve emphasis. First, the Lyapunov–Krasovskii construction is explicitly valid only for Ai6: at Ai7 the summation limits invert and the analysis breaks down, which is why the one-step case is treated separately. Second, the condition Ai8 is sufficient but not constructive—the paper concedes that the structural complexity of Ai9 prevents a theoretical characterization of admissible learning rates, and feasibility must be checked numerically via LMI solvers. In the five-agent wheel-graph example, feasible solutions exist up to (I−M)−1c0 with (I−M)−1c1 but not for (I−M)−1c2, where simulations show divergence; the authors note this "may imply" divergence rather than proving it, so the failure of the LMI at (I−M)−1c3 should not be read as a necessary instability condition.
Exponential convergence and an instability threshold for (I−M)−1c4
For one-step-delay exchange, the estimation update degenerates into a delay-free recursion, and the augmented system becomes (I−M)−1c5. Theorem 2 proves global exponential convergence when
(I−M)−1c6
via a Gershgorin disk argument: disks associated with the first block lie inside the unit circle by Assumption 1, while disks of the second block are inscribed in the unit circle whenever (I−M)−1c7. A contradiction argument rules out (I−M)−1c8 as an eigenvalue, giving (I−M)−1c9. This is a stronger result than related work: compared with asynchronous algorithms requiring an auxiliary interference graph that guarantee only almost-sure convergence [Li2025], or sub-linearly convergent delayed seeking schemes [LiuJ2024], this algorithm is synchronous, uses only the communication graph, and achieves deterministic exponential convergence.
Theorem 3 complements this with a lower bound for instability: if M0, then the mean eigenvalue of M1 falls below M2, forcing M3 and divergence. Together, Theorems 2 and 3 bracket the admissible learning rate from above and below, though the interval M4 may be empty or the bounds conservative. Indeed, the paper acknowledges that M5 is conservative—convergence is observed empirically for some M6—and numerical sensitivity analysis shows M7 tracks the true maximal convergent learning rate M8 well for complete graphs (both tending to zero as M9 grows) but poorly for ring graphs, where c0 while c1. A notable topological finding is that the guaranteed upper bound for ring graphs is always c2 regardless of agent count, making rings the most robust topology for large games among those analyzed.
Numerical evidence and scalability
Simulations corroborate the theory across three examples. In the five-agent wheel graph with c3 and c4, strategy errors decay exponentially; sweeping c5 from 0.05 to 1 reveals a U-shaped terminal-stage curve, with instability beyond c6 at c7, consistent with c8 being a valid but loose instability bound. A 20-agent example demonstrates scalability: convergence holds in four of five tested configurations, while the divergent case (c9, τ0, ring graph) violates the LMI condition of Theorem 1. The most practically relevant empirical observation is an inverse relationship between the maximal stable learning rate and both the delay step τ1 and the number of agents τ2: longer delays and larger networks require smaller learning rates.
Limitations and open questions
Several limitations are stated or evident. The LMI condition of Theorem 1 provides no analytic recipe for choosing τ3 as a function of τ4 and τ5, and no existence conditions for τ6 are given; the gap between LMI feasibility and actual stability (as at τ7) is unresolved. The bounds τ8 and τ9 are provably conservative for some topologies, and the exact stability boundary t0 lacks closed-form characterization. The framework assumes noiseless, lossless, synchronous communication over a fixed undirected connected graph, deterministic quadratic payoffs satisfying strict diagonal dominance, and scalar strategies per agent—all restrictive relative to realistic deployments. The authors list as future work: existence conditions for the LMI variables, more general payoff structures, quantitative analysis of how t1 and t2 affect convergence speed, and extensions to noise, packet dropout, and asynchronous updates.
Conclusion
This paper delivers a complete convergence picture for estimation-based best-response NE seeking under delayed information exchange in quadratic games: asymptotic convergence via a Lyapunov–Krasovskii LMI condition for multi-step delays, exponential convergence under an explicit learning-rate upper bound for one-step delays, and an explicit learning-rate threshold above which the dynamics provably diverge. The main practical takeaway—that permissible learning rates shrink inversely with delay length and network size—is supported both theoretically and numerically, though the conservatism of the derived bounds and the lack of constructive rate selection remain open problems.