Papers
Topics
Authors
Recent
Search
2000 character limit reached

Policy Iteration for Linear-Quadratic Stochastic Differential Games with State- and Control-Dependent Noise

Published 18 Aug 2026 in eess.SY | (2608.17940v1)

Abstract: This paper presents a novel sequential policy iteration (PI) method for stochastic differential games with state- and control-dependent noise. The updates preserve mean-square stability, so that the iteration is well posed. We further derive a closed-form expression for the Fréchet derivative of the sequential PI map at a Nash equilibrium. The resulting characterization reveals how control-dependent noise, policy-evaluation sensitivity, and update ordering govern local error propagation, and yields explicit sufficient conditions for local linear convergence. Since finding an initial stabilizing solution is a major challenge in policy iteration, we also propose a homotopy-based initialization that ensures a valid starting point. The effectiveness of the proposed PI algorithm and the analytical results are verified through a numerical example.

Summary

  • The paper introduces sequential policy iteration that preserves mean-square stability at every player update and whose fixed points are stabilizing linear feedback Nash equilibria.
  • The paper shows that local convergence depends on the spectral radius of the iteration derivative and, for three or more players, update order can determine whether an equilibrium attracts or repels iterates.
  • The paper develops a homotopy initialization with finite-step termination under mean-square stabilizability and validates the method numerically, including a case where sequential updates converge while decoupled updates fail.

Problem setting and motivation

The paper addresses the computation of stabilizing linear feedback Nash equilibria in infinite-horizon NN-player nonzero-sum LQ stochastic differential games whose dynamics carry multiplicative noise depending on both the state and the controls. The state evolves under an SDE with drift Ax+∑iBiuiAx + \sum_i B_i u_i and diffusion Cx+∑iDiuiCx + \sum_i D_i u_i, and each player minimizes an infinite-horizon quadratic cost with Qi≻0Q_i \succ 0 and Rii≻0R_{ii} \succ 0. Because the diffusion does not vanish, mean-square stability (MSS) replaces asymptotic stability as the relevant closed-loop property, and the coupled equilibrium conditions become rational matrix equations rather than polynomial ones. This combination—arbitrary player count, control-dependent noise, infinite horizon—is precisely the gap identified in a recent survey as open for Riccati-type solvability and explicit feedback implementation (2608.17940). Existing iterative methods cover either two-player formulations, purely state-dependent noise, or finite horizons, and none of their convergence arguments carries over, since the game's equilibrium conditions form a fixed-point problem to which monotonicity arguments do not apply.

The paper makes three contributions: (1) a sequential policy-iteration (PI) scheme whose player-wise updates preserve MSS, with fixed points exactly the stabilizing feedback Nash equilibria; (2) a local convergence analysis via the Fréchet derivative of the iteration map, showing that update ordering can decide whether an equilibrium is locally attracting at all; and (3) a homotopy-based initialization producing a stabilizing gain profile in finitely many steps whenever the game is mean-square stabilizable.

Sequential policy iteration and stability preservation

The algorithm alternates between policy evaluation, which solves a generalized Lyapunov equation for each player's value matrix Pi(K)P_i(K) at the current partially updated profile K~k,i\widetilde K^{k,i}, and policy improvement, which enforces the fixed-value stationarity condition quadratic in KiK_i. Players are updated one at a time in a fixed order, each improvement evaluated at a profile already incorporating preceding updates.

The central well-posedness result is that every player-wise update preserves MSS: completing the square on the Lyapunov operator shows that after updating player ii,

LK~k,i+1(Pik)=−Sik,i+1−ΔKi⊤(Rii+Di⊤PikDi)ΔKi≺0,\mathcal L_{\widetilde K^{k,i+1}}(P_i^k) = -S_i^{k,i+1} - \Delta K_i^\top\bigl(R_{ii}+D_i^\top P_i^k D_i\bigr)\Delta K_i \prec 0,

so the strict MSS Lyapunov inequality holds at every intermediate profile by induction over players. Consequently the entire iteration remains on Ax+∑iBiuiAx + \sum_i B_i u_i0, and every iterate is implementable as a stabilizing feedback—a property the decoupled scheme does not enjoy, as the numerics later confirm. A fixed-point characterization then establishes that Ax+∑iBiuiAx + \sum_i B_i u_i1 satisfies Ax+∑iBiuiAx + \sum_i B_i u_i2 if and only if it solves the coupled equilibrium equations, i.e., is a stabilizing linear feedback Nash equilibrium. Together these results resolve the well-posedness problem for this class of games; convergence itself remains local and equilibrium-dependent, since multiple stabilizing equilibria can exist even deterministically.

Local convergence and the role of update order

The local analysis derives the closed-form Fréchet derivative of the sequential PI map at an equilibrium via the implicit function theorem applied to the evaluation residual Ax+∑iBiuiAx + \sum_i B_i u_i3 and stationarity residual Ax+∑iBiuiAx + \sum_i B_i u_i4. The spectral condition Ax+∑iBiuiAx + \sum_i B_i u_i5 yields local linear convergence in an equivalent norm, with superlinear convergence when the derivative vanishes.

Two structural results organize the analysis. First, expanding the product Ax+∑iBiuiAx + \sum_i B_i u_i6 separates the derivative into an order-invariant part—which Proposition 4 identifies as the derivative of the decoupled (simultaneous-update) map—and an order-dependent term Ax+∑iBiuiAx + \sum_i B_i u_i7 collecting products of length at least two. Second, a triangular block representation shows that the permuted sequential derivative is exactly the block Gauss–Seidel iteration matrix of the decoupled derivative's splitting, placing the ordering phenomenon within classical matrix-splitting theory (2608.17940). This gives the first structural explanation of the empirically reported faster convergence of sequential updates.

The sharpest claims concern order dependence:

  • Cyclic invariance: cyclic shifts of an update order leave the spectrum unchanged.
  • Order-dependent attraction: for Ax+∑iBiuiAx + \sum_i B_i u_i8, there exist games and equilibria where one order has Ax+∑iBiuiAx + \sum_i B_i u_i9 while another has Cx+∑iDiuiCx + \sum_i D_i u_i0. The explicit three-player scalar example exhibits Cx+∑iDiuiCx + \sum_i D_i u_i1 versus Cx+∑iDiuiCx + \sum_i D_i u_i2 across the two cyclic classes. For Cx+∑iDiuiCx + \sum_i D_i u_i3 the effect vanishes: both orders share the same spectral radius, and in fact Cx+∑iDiuiCx + \sum_i D_i u_i4, so one sequential iteration contracts as much as two decoupled iterations wherever they converge together.

A weighted block-Cx+∑iDiuiCx + \sum_i D_i u_i5 comparison theorem guarantees that whenever the decoupled map is non-expansive in the transported norm, the sequential map contracts at least as strongly—an explicit, checkable guarantee supplementing the existence-of-norm statement from the spectral criterion. An explicit sufficient contraction criterion follows from computable coefficients Cx+∑iDiuiCx + \sum_i D_i u_i6 capturing cross-player diffusion coupling Cx+∑iDiuiCx + \sum_i D_i u_i7 and policy-evaluation sensitivity through the inverse Lyapunov-operator norm. Notably, the dependence on a player's own diffusion Cx+∑iDiuiCx + \sum_i D_i u_i8 is not monotone: larger Cx+∑iDiuiCx + \sum_i D_i u_i9 may tighten or relax the criterion depending on how it enters the premultiplier and Qi≻0Q_i \succ 00.

Homotopy initialization

Since policy evaluation presupposes a stabilizing profile, the paper constructs one via homotopy continuation on an auxiliary single-controller problem obtained by stacking all inputs. A uniform negative drift shift Qi≻0Q_i \succ 01 renders the trivial gain Qi≻0Q_i \succ 02 stabilizing; single-controller PI steps preserve MSS at each level, and admissible step sizes Qi≻0Q_i \succ 03, with Qi≻0Q_i \succ 04 derived from the value matrix and running-cost weights, ensure stability along the path. A safe-overshoot lemma permits terminating once Qi≻0Q_i \succ 05. Under mean-square stabilizability, finite termination holds with the explicit bound

Qi≻0Q_i \succ 06

continuation steps, established via a monotonicity argument bounding Qi≻0Q_i \succ 07 of the value matrices uniformly along the path. Unlike prior homotopy initializations for deterministic settings, this construction requires no individual player to stabilize the system alone and uses the correct MSS notion.

Numerical validation

A three-player, two-state example with scaled diffusions Qi≻0Q_i \succ 08, Qi≻0Q_i \succ 09 exercises all analytical predictions. The initialization returned a stabilizing profile after two continuation levels and 16 PI steps, within the theoretical bound of 5, and the main iteration converged in nine further iterations. Along the sweep from Rii≻0R_{ii} \succ 00 to Rii≻0R_{ii} \succ 01:

Scheme Behavior
Sequential order Rii≻0R_{ii} \succ 02 Contractive throughout tested range (Rii≻0R_{ii} \succ 03 at Rii≻0R_{ii} \succ 04)
Sequential order Rii≻0R_{ii} \succ 05 Loses contraction near Rii≻0R_{ii} \succ 06
Decoupled Loses contraction near Rii≻0R_{ii} \succ 07

At Rii≻0R_{ii} \succ 08, running all schemes from a common admissible profile confirms the spectral predictions: Rii≻0R_{ii} \succ 09 converges at its predicted rate Pi(K)P_i(K)0 to machine precision in roughly 30 iterations, Pi(K)P_i(K)1 is repelled at its predicted rate Pi(K)P_i(K)2, and the decoupled scheme leaves Pi(K)P_i(K)3 after 33 iterations, at which point its value equations cease to be solvable—while both sequential orders remain admissible throughout, as guaranteed. The preferred order also crosses near Pi(K)P_i(K)4, so the locally best order is not constant along the parameter path. The weighted block-norm comparison quantity crosses 1 near Pi(K)P_i(K)5 for the decoupled core while both sequential quantities remain below it, consistent with the guaranteed comparison.

Limitations and open questions

The convergence theory is inherently local and equilibrium-dependent: global convergence to a distinguished equilibrium cannot be expected without strong additional assumptions, and the paper provides only sufficient conditions—the spectral criterion guarantees existence of a contracting norm but not an explicit one, and the weighted-norm comparison requires the decoupled non-expansiveness assumption. The finite-termination bound for the homotopy depends on Pi(K)P_i(K)6, which is not known a priori and may be conservative. The analysis assumes constant linear feedback, a one-dimensional Brownian motion, and strictly positive definite state weights; extension to degenerate or correlated noise is not addressed. Whether the order-dependence phenomenon admits a constructive characterization—i.e., how to select a contracting order for a given equilibrium beyond enumerating cyclic classes—remains open, as does global convergence behavior of the sequential scheme.

Conclusion

The paper delivers a well-posed sequential policy-iteration method for Pi(K)P_i(K)7-player LQ stochastic differential games with state- and control-dependent noise, preserving mean-square stability at every intermediate update and fixing the stabilizing feedback Nash equilibria as its fixed points. Its Fréchet-derivative analysis explains sequential-versus-decoupled convergence as Gauss–Seidel versus Jacobi structure, proves a guaranteed norm comparison favoring sequential updates, and demonstrates—analytically and numerically—that with three or more players the update order can determine whether an equilibrium is locally attracting at all. The homotopy initialization closes the stabilizing-start requirement with a finite-step guarantee under mean-square stabilizability.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.