The paper introduces sequential policy iteration that preserves mean-square stability at every player update and whose fixed points are stabilizing linear feedback Nash equilibria.
The paper shows that local convergence depends on the spectral radius of the iteration derivative and, for three or more players, update order can determine whether an equilibrium attracts or repels iterates.
The paper develops a homotopy initialization with finite-step termination under mean-square stabilizability and validates the method numerically, including a case where sequential updates converge while decoupled updates fail.
Problem setting and motivation
The paper addresses the computation of stabilizing linear feedback Nash equilibria in infinite-horizon N-player nonzero-sum LQ stochastic differential games whose dynamics carry multiplicative noise depending on both the state and the controls. The state evolves under an SDE with driftAx+∑i​Bi​ui​ and diffusion Cx+∑i​Di​ui​, and each player minimizes an infinite-horizon quadratic cost with Qi​≻0 and Rii​≻0. Because the diffusion does not vanish, mean-square stability (MSS) replaces asymptotic stability as the relevant closed-loop property, and the coupled equilibrium conditions become rational matrix equations rather than polynomial ones. This combination—arbitrary player count, control-dependent noise, infinite horizon—is precisely the gap identified in a recent survey as open for Riccati-type solvability and explicit feedback implementation (2608.17940). Existing iterative methods cover either two-player formulations, purely state-dependent noise, or finite horizons, and none of their convergence arguments carries over, since the game's equilibrium conditions form a fixed-point problem to which monotonicity arguments do not apply.
Sequential policy iteration and stability preservation
The algorithm alternates between policy evaluation, which solves a generalized Lyapunov equation for each player's value matrix Pi​(K) at the current partially updated profile Kk,i, and policy improvement, which enforces the fixed-value stationarity condition quadratic in Ki​. Players are updated one at a time in a fixed order, each improvement evaluated at a profile already incorporating preceding updates.
The central well-posedness result is that every player-wise update preserves MSS: completing the square on the Lyapunov operator shows that after updating player i,
so the strict MSS Lyapunov inequality holds at every intermediate profile by induction over players. Consequently the entire iteration remains on Ax+∑i​Bi​ui​0, and every iterate is implementable as a stabilizing feedback—a property the decoupled scheme does not enjoy, as the numerics later confirm. A fixed-point characterization then establishes that Ax+∑i​Bi​ui​1 satisfies Ax+∑i​Bi​ui​2 if and only if it solves the coupled equilibrium equations, i.e., is a stabilizing linear feedback Nash equilibrium. Together these results resolve the well-posedness problem for this class of games; convergence itself remains local and equilibrium-dependent, since multiple stabilizing equilibria can exist even deterministically.
Two structural results organize the analysis. First, expanding the product Ax+∑i​Bi​ui​6 separates the derivative into an order-invariant part—which Proposition 4 identifies as the derivative of the decoupled (simultaneous-update) map—and an order-dependent term Ax+∑i​Bi​ui​7 collecting products of length at least two. Second, a triangular block representation shows that the permuted sequential derivative is exactly the block Gauss–Seidel iteration matrix of the decoupled derivative's splitting, placing the ordering phenomenon within classical matrix-splitting theory (2608.17940). This gives the first structural explanation of the empirically reported faster convergence of sequential updates.
The sharpest claims concern order dependence:
Cyclic invariance: cyclic shifts of an update order leave the spectrum unchanged.
Order-dependent attraction: for Ax+∑i​Bi​ui​8, there exist games and equilibria where one order has Ax+∑i​Bi​ui​9 while another has Cx+∑i​Di​ui​0. The explicit three-player scalar example exhibits Cx+∑i​Di​ui​1 versus Cx+∑i​Di​ui​2 across the two cyclic classes. For Cx+∑i​Di​ui​3 the effect vanishes: both orders share the same spectral radius, and in fact Cx+∑i​Di​ui​4, so one sequential iteration contracts as much as two decoupled iterations wherever they converge together.
A weighted block-Cx+∑i​Di​ui​5 comparison theorem guarantees that whenever the decoupled map is non-expansive in the transported norm, the sequential map contracts at least as strongly—an explicit, checkable guarantee supplementing the existence-of-norm statement from the spectral criterion. An explicit sufficient contraction criterion follows from computable coefficients Cx+∑i​Di​ui​6 capturing cross-player diffusion coupling Cx+∑i​Di​ui​7 and policy-evaluation sensitivity through the inverse Lyapunov-operator norm. Notably, the dependence on a player's own diffusion Cx+∑i​Di​ui​8 is not monotone: larger Cx+∑i​Di​ui​9 may tighten or relax the criterion depending on how it enters the premultiplier and Qi​≻00.
Homotopy initialization
Since policy evaluation presupposes a stabilizing profile, the paper constructs one via homotopy continuation on an auxiliary single-controller problem obtained by stacking all inputs. A uniform negative drift shift Qi​≻01 renders the trivial gain Qi​≻02 stabilizing; single-controller PI steps preserve MSS at each level, and admissible step sizes Qi​≻03, with Qi​≻04 derived from the value matrix and running-cost weights, ensure stability along the path. A safe-overshoot lemma permits terminating once Qi​≻05. Under mean-square stabilizability, finite termination holds with the explicit bound
Qi​≻06
continuation steps, established via a monotonicity argument bounding Qi​≻07 of the value matrices uniformly along the path. Unlike prior homotopy initializations for deterministic settings, this construction requires no individual player to stabilize the system alone and uses the correct MSS notion.
Numerical validation
A three-player, two-state example with scaled diffusions Qi​≻08, Qi​≻09 exercises all analytical predictions. The initialization returned a stabilizing profile after two continuation levels and 16 PI steps, within the theoretical bound of 5, and the main iteration converged in nine further iterations. Along the sweep from Rii​≻00 to Rii​≻01:
Scheme
Behavior
Sequential order Rii​≻02
Contractive throughout tested range (Rii​≻03 at Rii​≻04)
Sequential order Rii​≻05
Loses contraction near Rii​≻06
Decoupled
Loses contraction near Rii​≻07
At Rii​≻08, running all schemes from a common admissible profile confirms the spectral predictions: Rii​≻09 converges at its predicted rate Pi​(K)0 to machine precision in roughly 30 iterations, Pi​(K)1 is repelled at its predicted rate Pi​(K)2, and the decoupled scheme leaves Pi​(K)3 after 33 iterations, at which point its value equations cease to be solvable—while both sequential orders remain admissible throughout, as guaranteed. The preferred order also crosses near Pi​(K)4, so the locally best order is not constant along the parameter path. The weighted block-norm comparison quantity crosses 1 near Pi​(K)5 for the decoupled core while both sequential quantities remain below it, consistent with the guaranteed comparison.
Limitations and open questions
The convergence theory is inherently local and equilibrium-dependent: global convergence to a distinguished equilibrium cannot be expected without strong additional assumptions, and the paper provides only sufficient conditions—the spectral criterion guarantees existence of a contracting norm but not an explicit one, and the weighted-norm comparison requires the decoupled non-expansiveness assumption. The finite-termination bound for the homotopy depends on Pi​(K)6, which is not known a priori and may be conservative. The analysis assumes constant linear feedback, a one-dimensional Brownian motion, and strictly positive definite state weights; extension to degenerate or correlated noise is not addressed. Whether the order-dependence phenomenon admits a constructive characterization—i.e., how to select a contracting order for a given equilibrium beyond enumerating cyclic classes—remains open, as does global convergence behavior of the sequential scheme.
“Emergent Mind helps me see which AI papers have caught fire online.”
Philip
Creator, AI Explained on YouTube
Sign up for free to explore the frontiers of research
Discover trending papers, chat with arXiv, and track the latest research shaping the future of science and technology.Discover trending papers, chat with arXiv, and more.