Papers
Topics
Authors
Recent
Search
2000 character limit reached

Mean-Field Stochastic LQR Controller

Updated 11 December 2025
  • Mean-Field Stochastic LQR Controller is a control strategy that extends classical LQR by incorporating both individual-state and population-averaged interactions under stochastic uncertainty.
  • It employs a Lagrangian dual formulation to decouple mean and fluctuation dynamics, solving coupled Riccati equations for robust, scalable state-feedback synthesis.
  • Applications in microgrid frequency control and large-scale power networks demonstrate its ability to reduce overshoots and manage risk through variance-based constraints.

A mean-field stochastic linear quadratic regulator (MF-SLQR) controller generalizes the classical LQR approach to systems with both individual-state and mean-field (population-averaged) interactions under stochastic uncertainty, and further constrains state fluctuations to address risk. Such controllers are crucial in high-dimensional multi-agent networks, power grids, and large-scale coupled stochastic systems, particularly when low-probability, high-impact events must be systematically attenuated rather than averaged away as in the risk-neutral setting.

1. Problem Formulation: Dynamics, Cost, and Risk Constraint

Consider nn exchangeable agents indexed by ii with individual state xti∈Rdxx_t^i \in \mathbb{R}^{d_x} and control uti∈Rduu_t^i \in \mathbb{R}^{d_u}. Each agent's dynamics incorporate both local and mean-field coupling: xt+1i=A xti+B uti+Aˉ xˉt+Bˉ uˉt+wti,x_{t+1}^i = A\,x_t^i + B\,u_t^i + \bar{A}\,\bar{x}_t + \bar{B}\,\bar{u}_t + w_t^i, where xˉt=1n∑j=1nxtj\bar{x}_t = \frac{1}{n}\sum_{j=1}^n x_t^j, uˉt=1n∑j=1nutj\bar{u}_t = \frac{1}{n}\sum_{j=1}^n u_t^j, and wtiw_t^i is an i.i.d. zero-mean noise sequence. In the n→∞n\to\infty mean-field limit, agent states decompose into fluctuation and mean components: x~t+1i=A x~ti+B u~ti+w~ti, xˉt+1=(A+Aˉ) xˉt+(B+Bˉ) uˉt+wˉt,\begin{aligned} &\tilde{x}^i_{t+1} = A\,\tilde{x}_t^i + B\,\tilde{u}_t^i + \tilde{w}_t^i, \ &\bar{x}_{t+1} = (A+\bar{A})\,\bar{x}_t + (B+\bar{B})\,\bar{u}_t + \bar{w}_t, \end{aligned} where ii0, ii1.

The infinite-horizon average quadratic cost per agent (risk-neutral) is

ii2

To control rare but impactful fluctuations, a variance-type risk constraint is imposed. Define

ii3

where ii4 denotes agent ii5's history. The per-player time-averaged variance is

ii6

The MF-SLQR controller seeks to

ii7

with risk budget ii8 (Roudneshin et al., 2023).

2. Lagrangian Dual Formulation and Decomposition

A Lagrange multiplier ii9 combines the nominal and risk costs: xti∈Rdxx_t^i \in \mathbb{R}^{d_x}0 Owing to orthogonal mean–fluctuation decomposition, xti∈Rdxx_t^i \in \mathbb{R}^{d_x}1 decouples into two independent infinite-horizon LQR objectives: xti∈Rdxx_t^i \in \mathbb{R}^{d_x}2 where

xti∈Rdxx_t^i \in \mathbb{R}^{d_x}3

and xti∈Rdxx_t^i \in \mathbb{R}^{d_x}4. The dependence on xti∈Rdxx_t^i \in \mathbb{R}^{d_x}5 drives risk sensitivity (Roudneshin et al., 2023).

3. Coupled Riccati Equations and Controller Synthesis

Both mean and fluctuation subsystems admit closed-form LQR solutions. The optimal value of xti∈Rdxx_t^i \in \mathbb{R}^{d_x}6 is enforced via a primal–dual algorithm.

Riccati Equations

xti∈Rdxx_t^i \in \mathbb{R}^{d_x}7

These equations admit a unique positive semidefinite stabilizing solution under standard MF-LQR stabilizability/detectability criteria (Roudneshin et al., 2023).

Controller Law

The optimal state-feedback is affine in the agent's own state and in the population mean: xti∈Rdxx_t^i \in \mathbb{R}^{d_x}8 with

xti∈Rdxx_t^i \in \mathbb{R}^{d_x}9

and uti∈Rduu_t^i \in \mathbb{R}^{d_u}0 arises if the noise is nonzero-mean in the dual formulation but is zero otherwise (Roudneshin et al., 2023). The feedback structure matches that of classical mean-field LQR, but the gain matrices internalize the variance penalty via the modified cost matrices.

4. Risk Constraint Enforcement and Computational Aspects

The optimal Lagrange multiplier uti∈Rduu_t^i \in \mathbb{R}^{d_u}1 is computed via a primal–dual loop. At each iteration:

  1. Update uti∈Rduu_t^i \in \mathbb{R}^{d_u}2, uti∈Rduu_t^i \in \mathbb{R}^{d_u}3;
  2. Solve the two Riccati equations for uti∈Rduu_t^i \in \mathbb{R}^{d_u}4, uti∈Rduu_t^i \in \mathbb{R}^{d_u}5;
  3. Form uti∈Rduu_t^i \in \mathbb{R}^{d_u}6, uti∈Rduu_t^i \in \mathbb{R}^{d_u}7, uti∈Rduu_t^i \in \mathbb{R}^{d_u}8;
  4. Simulate closed-loop dynamics and evaluate uti∈Rduu_t^i \in \mathbb{R}^{d_u}9;
  5. Update xt+1i=A xti+B uti+Aˉ xˉt+Bˉ uˉt+wti,x_{t+1}^i = A\,x_t^i + B\,u_t^i + \bar{A}\,\bar{x}_t + \bar{B}\,\bar{u}_t + w_t^i,0 by a projected subgradient step: xt+1i=A xti+B uti+Aˉ xˉt+Bˉ uˉt+wti,x_{t+1}^i = A\,x_t^i + B\,u_t^i + \bar{A}\,\bar{x}_t + \bar{B}\,\bar{u}_t + w_t^i,1 Each iteration requires xt+1i=A xti+B uti+Aˉ xˉt+Bˉ uˉt+wti,x_{t+1}^i = A\,x_t^i + B\,u_t^i + \bar{A}\,\bar{x}_t + \bar{B}\,\bar{u}_t + w_t^i,2 operations (Riccati/Lyapunov equations) and is independent of the agent population xt+1i=A xti+B uti+Aˉ xˉt+Bˉ uˉt+wti,x_{t+1}^i = A\,x_t^i + B\,u_t^i + \bar{A}\,\bar{x}_t + \bar{B}\,\bar{u}_t + w_t^i,3. The resulting gains do not depend on xt+1i=A xti+B uti+Aˉ xˉt+Bˉ uˉt+wti,x_{t+1}^i = A\,x_t^i + B\,u_t^i + \bar{A}\,\bar{x}_t + \bar{B}\,\bar{u}_t + w_t^i,4, which ensures scalability (Roudneshin et al., 2023).

5. Structural Properties and Solution Comparison

  • Independence from Number of Players: All Riccati equations and controller gains depend only on local and mean-field coupling (xt+1i=A xti+B uti+Aˉ xˉt+Bˉ uˉt+wti,x_{t+1}^i = A\,x_t^i + B\,u_t^i + \bar{A}\,\bar{x}_t + \bar{B}\,\bar{u}_t + w_t^i,5) and cost weights, not on xt+1i=A xti+B uti+Aˉ xˉt+Bˉ uˉt+wti,x_{t+1}^i = A\,x_t^i + B\,u_t^i + \bar{A}\,\bar{x}_t + \bar{B}\,\bar{u}_t + w_t^i,6. This makes the approach viable for large-scale networks (Roudneshin et al., 2023).
  • Existence and Uniqueness: Unique stabilizing Riccati solutions and dual multipliers xt+1i=A xti+B uti+Aˉ xˉt+Bˉ uˉt+wti,x_{t+1}^i = A\,x_t^i + B\,u_t^i + \bar{A}\,\bar{x}_t + \bar{B}\,\bar{u}_t + w_t^i,7 are guaranteed by standard LQR system-theoretic conditions and strong duality (Roudneshin et al., 2023).
  • Risk Parameter Influence: As xt+1i=A xti+B uti+Aˉ xˉt+Bˉ uˉt+wti,x_{t+1}^i = A\,x_t^i + B\,u_t^i + \bar{A}\,\bar{x}_t + \bar{B}\,\bar{u}_t + w_t^i,8 is decreased (tighter variance constraint), xt+1i=A xti+B uti+Aˉ xˉt+Bˉ uˉt+wti,x_{t+1}^i = A\,x_t^i + B\,u_t^i + \bar{A}\,\bar{x}_t + \bar{B}\,\bar{u}_t + w_t^i,9 increases, which modifies xˉt=1n∑j=1nxtj\bar{x}_t = \frac{1}{n}\sum_{j=1}^n x_t^j0, leading to more conservative xˉt=1n∑j=1nxtj\bar{x}_t = \frac{1}{n}\sum_{j=1}^n x_t^j1 (higher-magnitude feedback gains). This reduces the amplitude of risky excursions (e.g., overshoots) at the cost of modestly increased average cost xˉt=1n∑j=1nxtj\bar{x}_t = \frac{1}{n}\sum_{j=1}^n x_t^j2 (Roudneshin et al., 2023).
  • Comparison to Risk-Neutral MF-LQR: Setting xˉt=1n∑j=1nxtj\bar{x}_t = \frac{1}{n}\sum_{j=1}^n x_t^j3 recovers the standard mean-field LQR controller, which ignores state variance (Roudneshin et al., 2023).

6. Practical Example and Performance Analysis

A microgrid frequency control scenario is used as a high-dimensional case study. Each area (agent) maintains a xˉt=1n∑j=1nxtj\bar{x}_t = \frac{1}{n}\sum_{j=1}^n x_t^j4-dimensional state including local frequency, generation, tie-line flow, and area control error integral. System matrices xˉt=1n∑j=1nxtj\bar{x}_t = \frac{1}{n}\sum_{j=1}^n x_t^j5 are drawn from standard load frequency control (LFC) parameters and quadratic costs xˉt=1n∑j=1nxtj\bar{x}_t = \frac{1}{n}\sum_{j=1}^n x_t^j6 are assigned.

Table: Impact of Variance Constraint on Controller Behavior | xˉt=1n∑j=1nxtj\bar{x}_t = \frac{1}{n}\sum_{j=1}^n x_t^j7 (risk tolerance) | xˉt=1n∑j=1nxtj\bar{x}_t = \frac{1}{n}\sum_{j=1}^n x_t^j8 (dual) | Gain Magnitude | Overshoot | Average Cost xˉt=1n∑j=1nxtj\bar{x}_t = \frac{1}{n}\sum_{j=1}^n x_t^j9 | |--------------------------|--------------------|---------------|-----------|------------------| | High | uˉt=1n∑j=1nutj\bar{u}_t = \frac{1}{n}\sum_{j=1}^n u_t^j0 | Baseline | Large | Baseline | | Moderate | Moderate | Increased | Reduced | Slightly higher | | Low | Large | Largest | Minimal | Highest |

Reducing uˉt=1n∑j=1nutj\bar{u}_t = \frac{1}{n}\sum_{j=1}^n u_t^j1 leads to increased uˉt=1n∑j=1nutj\bar{u}_t = \frac{1}{n}\sum_{j=1}^n u_t^j2 and higher feedback gain magnitude, yielding faster disturbance damping and less vulnerability to rare high-variance events (Roudneshin et al., 2023).

7. Generalizations and Methodological Significance

Risk-constrained MF-SLQR controllers extend classic LQR and mean-field LQR by systematically regulating not only expected cost but also rare-event risk, via variance-type constraints. The affine law and dual-based synthesis parallel but augment the classical mean–fluctuation separation and are compatible with standard Riccati-based implementation. Scalability and algorithmic simplicity are preserved, and the general methodology readily extends to other risk proxies and limit regimes (Roudneshin et al., 2023).

These contributions are directly relevant for modern applications requiring explicit robustness to stochastic volatility and population-level coupling, such as large-scale energy systems and networked autonomous agents.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Mean-Field Stochastic LQR (MF-SLQR) Controller.