Distributionally Robust Controller Synthesis
- Distributionally robust controller synthesis is the design of control laws that optimize worst-case performance over an ambiguity set of disturbance distributions rather than a fixed law.
- It leverages various ambiguity sets—moment-based, Wasserstein, and KL-divergence—to statistically calibrate uncertainty, ensuring robust stability, safety, and reachability.
- Computational reformulations transform the infinite-dimensional minimax problem into tractable convex programs and frequency-domain approximations that balance rigor with practical implementation.
Distributionally robust controller synthesis is the design of control laws against uncertainty modeled not as a single disturbance distribution, nor only as deterministic bounded perturbations, but as an ambiguity set of probability measures. Across the literature, the synthesized controller is chosen to optimize a worst-case criterion—typically expected cost, regret, or specification satisfaction probability—over all distributions in that set. The resulting paradigm spans infinite-horizon and finite-horizon stochastic control, linear-quadratic regulation with additive or multiplicative noise, partially observed LQG, formal safety and reach-avoid synthesis, System Level Synthesis (SLS), path-integral control, CLF-CBF safety filters, and learning-based or reinforcement-learning settings (Coppens et al., 2019, Micheli et al., 2024, Chen et al., 6 Jan 2025).
1. Problem classes and the meaning of distributional robustness
Distributionally robust controller synthesis arises when the uncertainty entering a dynamical system is stochastic, but its law is not known exactly. In the linear-quadratic setting, this uncertainty may appear as multiplicative noise,
with
or as additive disturbance,
The synthesis objective is then to construct a feedback law—static state feedback, causal finite-horizon state feedback, affine disturbance feedback, dynamic output feedback, or optimization-induced feedback—that minimizes a worst-case performance functional over an ambiguity set of admissible disturbance distributions (Coppens et al., 2019, Taha et al., 11 Dec 2025, Gramlich et al., 27 Sep 2025).
The criterion optimized depends on the application class. In stochastic LQR and LQG, the objective is worst-case expected quadratic cost. In regret-optimal control, the objective is worst-case expected regret relative to a clairvoyant noncausal benchmark,
with the optimal noncausal controller (Taha et al., 11 Dec 2025, Kargin et al., 2024). In formal methods, the objective is worst-case probability of satisfying safety or reach-avoid specifications under adversarial distribution selection at each stage (Chen et al., 6 Jan 2025, Gracia et al., 2022). In CLF-CBF filtering, the goal is to enforce stability and safety constraints with high probability under all distributions in a Wasserstein ambiguity set (Long et al., 2022). In learning-based settings, robustness is expressed as a worst-case expected rollout loss over deployment distributions lying near the training distribution, or as robustness to scenario-weighted exogenous demand shifts (Herceg et al., 12 Apr 2026, Pei et al., 21 Dec 2025).
A defining distinction from classical robust control is that the adversary acts on probability laws rather than directly on pointwise disturbance realizations. A defining distinction from nominal stochastic control is that the disturbance law is not fixed and known. A plausible implication is that distributionally robust synthesis interpolates between the over-confidence of plug-in stochastic design and the pessimism of worst-case signal robustness. This interpolation is explicit in several formulations: the infinite-horizon Wasserstein DR-LQR tends to nominal /LQR behavior as the ambiguity radius tends to zero and to -type behavior as the radius grows (Hajar et al., 2024), while the KL-based Distributionally Robust Path Integral formulation approaches risk-neutral control as the ambiguity radius vanishes (Park et al., 2023).
2. Ambiguity sets and their statistical calibration
The ambiguity set is the central modeling object. Several classes appear repeatedly.
A first class is the moment-based ambiguity set. In data-driven multiplicative-noise LQR, the unknown disturbance law is represented by first and second moments estimated from i.i.d. samples,
The ambiguity set is of Delage-Ye type, constraining the mean to an ellipsoid around and the covariance to lie below a scaled empirical covariance: Under a sub-Gaussian assumption on 0, the paper derives explicit concentration bounds for 1 and 2, and shows that the true distribution lies in the ambiguity set with probability at least 3 (Coppens et al., 2019). Closely related moment ambiguity appears in finite-horizon regret-optimal control, where the ambiguity set is
4
with covariance uncertainty measured in a Schatten 5-norm ball (Taha et al., 11 Dec 2025).
A second class is Wasserstein ambiguity. In finite-horizon SLS-based data-driven control, the ambiguity set is a Wasserstein ball centered at the predictive empirical closed-loop distribution,
6
and its radius is decision dependent because the controller changes the predictive-to-actual distribution shift (Micheli et al., 2024). In infinite-horizon LQR and regret-optimal control, the ambiguity is a Wasserstein-2 ball around a nominal disturbance law or nominal covariance operator, with radius scaling as 7 so that the steady-state criterion is well defined (Hajar et al., 2024, Kargin et al., 2024). In formal safety and reach-avoid synthesis, the disturbance law at each stage is chosen from
8
a Wasserstein ball around a nominal distribution 9, often with finite support (Chen et al., 6 Jan 2025). Wasserstein balls also underpin robust MDP abstractions for switched stochastic systems (Gracia et al., 2022), STL-constrained interacting-agent control (Kordabad et al., 12 Mar 2025), CLF-CBF safety filtering (Long et al., 2022), PAC-Bayesian deployment-shift certification (Herceg et al., 12 Apr 2026), and end-to-end metric learning for finite-horizon DRC (Wu et al., 11 Oct 2025).
A third class is KL-divergence ambiguity. In continuous-time path-integral control, the ambiguity set is
0
a KL ball around an empirical nominal path law 1 built from disturbance trajectories (Park et al., 2023). In finite-horizon partially observed LQG, the ambiguity is specified stagewise by KL balls around nominal Gaussian laws for 2, 3, and 4, preserving temporal independence of exogenous variables under every admissible law (Fochesato et al., 13 May 2025).
These ambiguity sets are often calibrated statistically. The moment-based approach uses explicit finite-sample concentration inequalities under sub-Gaussian assumptions (Coppens et al., 2019). Wasserstein-based approaches use concentration inequalities of Fournier and Guillin under light-tail assumptions to select a radius that contains the true law with prescribed confidence (Micheli et al., 2024, Kordabad et al., 12 Mar 2025). KL-based path-integral control derives a deterministic-drift finite-sample guarantee,
5
so that the true diffusion law lies in the KL ball with probability at least 6 (Park et al., 2023). In end-to-end anisotropic metric learning, the ambiguity geometry is changed through a learned SPD matrix 7, but the radius is rescaled as
8
to preserve the same finite-sample Wasserstein coverage guarantee (Wu et al., 11 Oct 2025).
3. Synthesis formulations and computational reformulations
A recurring theme is that the original minimax problem is infinite-dimensional—because the adversary chooses a distribution—but admits exact or approximate finite-dimensional reformulations.
In moment-based multiplicative-noise LQR, the nominal infinite-horizon Riccati equation is equivalent to an SDP. When the disturbance mean is known and only the covariance is ambiguous, the distributionally robust controller is obtained exactly by replacing 9 with 0 in the nominal Riccati or SDP. When both mean and covariance are uncertain, the exact minimax problem is intractable, so the paper constructs a quadratic upper-bound value function 1, introduces 2 and 3, and derives an LMI program whose feasibility implies robust Lyapunov decrease (Coppens et al., 2019).
In SLS-based finite-horizon data-driven control, the closed-loop responses 4 satisfy the affine achievability condition
5
and the predictive empirical closed-loop distribution depends linearly on 6. For piecewise affine cost and constraint functions, Wasserstein DRO duality yields finite-dimensional worst-case expectation formulas. A tractable LP approximation is then obtained through a small-gain upper bound on the decision-dependent Wasserstein radius (Micheli et al., 2024).
In finite-horizon moment-robust regret-optimal control, the expected regret depends only on the first two moments of the disturbance trajectory. The exact minimax problem is reformulated as the convex program
7
with
8
This exposes mean ambiguity as spectral-norm regularization and covariance ambiguity as dual Schatten-norm regularization. The formulation admits an SDP representation, but the paper develops a dual projected subgradient method with projection onto Schatten norm balls, reducing the dominant cost to eigendecomposition and making horizons up to 9 practically accessible (Taha et al., 11 Dec 2025).
In infinite-horizon Wasserstein DR-LQR and infinite-horizon DR regret-optimal control, the synthesis becomes a saddle-point problem over causal operators and disturbance covariance operators. The optimal controller is characterized in the frequency domain through Wiener–Hopf factorization. In DR-LQR, the controller has the form
0
with 1 the spectral factor of the worst-case disturbance spectrum, and is generally non-rational (Hajar et al., 2024). In infinite-horizon DR regret-optimal control, the optimal saddle point is unique and similarly characterized, but the resulting controller also lacks a finite-order state-space realization. Both papers therefore compute the non-rational controller on a frequency grid and then construct a rational approximation by solving a convex 2-norm approximation problem for the spectrum, yielding a realizable state-space controller (Hajar et al., 2024, Kargin et al., 2024).
In distributionally robust LMI synthesis for LTI systems, the worst-case asymptotic average expected quadratic output energy under a Wasserstein ambiguity set is first rewritten via an exact Wasserstein-coupling formulation, then reduced through an exact stationary second-moment relaxation, then dualized into robust-analysis LMIs, and finally convexified through Scherer’s dynamic output-feedback parameterization. The final synthesis conditions are exact convex SDPs for both correlated and independent disturbance classes, and they expose the problem as an instance of robust 3 synthesis (Gramlich et al., 27 Sep 2025).
Other architectures use different computational reductions. KL-based path-integral control dualizes the robust min-max problem into
4
so controller synthesis becomes an outer one-dimensional optimization over 5 wrapped around a risk-sensitive path-integral control subproblem (Park et al., 2023). Distributionally robust LQG retains linear output-feedback optimality and solves the resulting Gaussian zero-sum game by regularized best-response dynamics, with closed-form controller best responses and blockwise covariance updates for the adversary (Fochesato et al., 13 May 2025). Distributionally robust CLF-CBF filtering turns Wasserstein DR chance constraints into SOCPs that can be solved online at each state (Long et al., 2022).
4. Stability, safety, and formal specifications
Distributionally robust controller synthesis is not confined to expected-cost optimality. A substantial part of the literature replaces or complements cost minimization with stability, safety, and logic-based specification guarantees.
For multiplicative-noise systems, the relevant notion is mean-square stability. The Lyapunov characterization
6
is equivalent to mean-square stability of the autonomous system, and its distributionally robust counterpart,
7
implies exponential mean-square stability with probability at least 8 whenever the true law lies in the ambiguity set (Coppens et al., 2019).
In finite-horizon SLS-based DRO, the guarantee is not asymptotic stability but out-of-sample performance and risk-feasibility transfer. If the ambiguity radius is chosen according to the predictive-to-actual distribution-shift bound, then the true closed-loop distribution lies in the Wasserstein ball with probability at least 9, and the robust objective upper-bounds actual cost while the CVaR constraint remains satisfied under the true dynamics (Micheli et al., 2024). In PAC-Bayesian finite-horizon control, the certified quantity is expected rollout loss under any deployment distribution within Wasserstein radius 0 of the training distribution. The robust PAC-Bayes certificate takes the form
1
with the same operator norm controlling both Wasserstein sensitivity and the sub-Gaussian proxy for unbounded loss (Herceg et al., 12 Apr 2026).
For reach-avoid and safety specifications, the dynamic game formulation is explicit. In finite-horizon stochastic systems with disturbance distributions selected adversarially from a Wasserstein ball at each stage, the optimal worst-case reach-avoid and safety probabilities are characterized by Bellman recursions of the form
2
and
3
with measurable optimal Markov controllers and measurable 4-optimal adversary selectors (Chen et al., 6 Jan 2025). For polynomial systems under safety specifications, the same paper introduces distributionally robust control barrier certificates satisfying a robust Bellman-like decrease inequality and turns their search into an SOS program, producing a time-invariant polynomial feedback law with a certified lower bound on worst-case safety probability (Chen et al., 6 Jan 2025).
For switched stochastic systems, the formal-specification route proceeds through abstraction. A continuous-state switched system with additive noise and Wasserstein ambiguity is abstracted into a robust MDP whose uncertainty sets 5 are transport-based liftings of disturbance ambiguity. Robust Bellman updates compute worst-case reach-avoid probabilities for finite or infinite horizons, and the resulting abstract strategy is refined to a switching strategy for the original system, with lower and upper probability bounds bracketing the true satisfaction probability for all distributions in the ambiguity set (Gracia et al., 2022).
In CLF-CBF filtering, safety and stability are encoded by chance constraints over uncertain CLF and CBF inequalities. The distributionally robust version requires these inequalities to hold with probability at least 6 for every law in a Wasserstein ball around the empirical uncertainty distribution. A CVaR-based reformulation then yields online SOCPs; the adaptive cruise control experiments illustrate the intended tradeoff between the over-conservatism of robust support-based methods and the over-confidence of nominal Gaussian chance constraints under distribution shift (Long et al., 2022).
5. Data-driven, learning-based, and end-to-end variants
Recent work extends distributionally robust controller synthesis beyond classical convex control design into learning-based pipelines, but the underlying uncertainty model remains distributional.
One line uses data-driven robustification of explicit control programs. In interacting-agent control under STL constraints, the original problem is a chance-constrained program requiring
7
It is converted into an expectation-constrained program either by concentration of measure,
8
or by CVaR. Since the disturbance law is unknown, each expectation is then replaced by a worst-case expectation over a Wasserstein ball around the empirical distribution, yielding a robust sample-based program with a second confidence level that depends on the number of samples (Kordabad et al., 12 Mar 2025). A closely related pattern appears in CLF-CBF filtering, where only a finite set of model-uncertainty samples is available and the ambiguity set is again centered on the empirical distribution (Long et al., 2022).
A second line learns both controller and certificate. In distributionally robust policy and Lyapunov-certificate learning, the uncertain continuous-time system
9
is stabilized by jointly learning a neural state-feedback controller and a Lyapunov function. The nominal derivative decrease condition is replaced by a distributionally robust chance constraint over the uncertainty law, and a deterministic sufficient condition,
0
is embedded into the training loss. The paper then argues that global asymptotic stability of the equilibrium can be certified with high confidence, even under out-of-distribution model uncertainties (Long et al., 2024).
A third line integrates ambiguity-set design with downstream control. End-to-end statistically guaranteed metric learning for finite-horizon Wasserstein DRC treats the Wasserstein metric itself as a learnable SPD matrix 1. The inner problem is a convex reformulation of finite-horizon affine disturbance-feedback DRC with CVaR safety constraints, while the outer problem updates 2 using closed-loop rollout performance over a distribution of initial conditions. The ambiguity radius is adjusted by
3
so that the learned anisotropic ambiguity set preserves the same finite-sample statistical guarantee as the isotropic one (Wu et al., 11 Oct 2025). This suggests a controller-oriented alternative to task-agnostic ambiguity geometry.
A fourth line treats distributional robustness in reinforcement learning as scenario robustness. In multi-agent traffic signal control on a 4 Athens grid calibrated from pNEUMA, robustness is not expressed through Wasserstein or moment balls but through a finite family of eight origin-destination scenarios and their convex mixtures. A contextual-bandit worst-case estimator learns context-dependent scenario weights, and a baseline PPO-based MARL controller is fine-tuned under these adversarial mixtures. The result is a scenario-based DRO-like synthesis pipeline with empirical improvements on worst-case queues and speeds, including an unseen Sioux Falls validation network (Pei et al., 21 Dec 2025).
6. Conservatism, structural limitations, and recurring open directions
The literature repeatedly emphasizes that distributionally robust controller synthesis is a tradeoff between protection against model misspecification and computational or structural conservatism.
Several sources of conservatism are explicit. Moment-based robustification can become loose when uncertain means are handled through ellipsoidal upper bounds and LMI slack variables (Coppens et al., 2019). Wasserstein finite-horizon SLS control inherits the curse of dimensionality through the rate
5
and its tractable LP relies on a small-gain bound, triangle-inequality decomposition, and conservative handling of CVaR (Micheli et al., 2024). Safety and reach-avoid dynamic programming with stagewise Wasserstein adversaries is exact for a dynamically varying adversarial environment but conservative if the true disturbance law is fixed and unknown over the horizon (Chen et al., 6 Jan 2025). CVaR approximations in CLF-CBF filtering and STL control are sufficient but conservative relative to the original chance constraints (Long et al., 2022, Kordabad et al., 12 Mar 2025). Scenario-based MARL robustness is confined to the convex hull of a hand-designed finite scenario family and does not provide theorem-level worst-case guarantees (Pei et al., 21 Dec 2025).
A second recurring limitation is horizon and realizability. Many formulations are finite horizon only: SLS-based doubly robust control, moment-robust regret control, KL-robust LQG, PAC-Bayesian robust control, STL chance-constrained control, and anisotropic metric learning all work over a fixed finite horizon (Micheli et al., 2024, Taha et al., 11 Dec 2025, Fochesato et al., 13 May 2025, Herceg et al., 12 Apr 2026, Kordabad et al., 12 Mar 2025, Wu et al., 11 Oct 2025). Infinite-horizon Wasserstein DR-LQR and DR regret-optimal control address this limitation but produce generally non-rational optimal controllers, requiring rational approximation for implementation (Hajar et al., 2024, Kargin et al., 2024).
A third limitation concerns observability and policy class. Some methods assume full state measurement and static or affine state feedback (Coppens et al., 2019, Micheli et al., 2024), while others require disturbance feedback or full-information disturbance access (Taha et al., 11 Dec 2025, Kargin et al., 2024, Hajar et al., 2024). Partially observed synthesis is substantially harder; the finite-horizon KL-robust LQG result is notable precisely because it preserves linear output-feedback optimality under distributional uncertainty (Fochesato et al., 13 May 2025). The PAC-Bayesian framework certifies randomized controller posteriors rather than deterministic controllers directly (Herceg et al., 12 Apr 2026).
A fourth limitation is model specificity. Some approaches require linearly solvable path-integral structure and KL ambiguity on path space (Park et al., 2023); others require piecewise affine costs and constraints (Micheli et al., 2024), or polynomial dynamics and semialgebraic sets for SOS synthesis (Chen et al., 6 Jan 2025). The robust 6 LMI formulation is exact but specialized to LTI systems with Gaussian nominal disturbance law and Wasserstein geometry (Gramlich et al., 27 Sep 2025). The learning-based methods can relax analytical structure but shift complexity into nonconvex optimization and finite-sample generalization questions (Long et al., 2024, Wu et al., 11 Oct 2025).
A plausible synthesis of these trends is that the field is converging on three complementary regimes. The first is exact convex synthesis for narrowly structured linear problems, where ambiguity sets and performance functionals are chosen to preserve SDP, SOCP, LP, or frequency-domain tractability. The second is exact or near-exact formal synthesis for safety and reachability, where dynamic programming, robust MDP abstraction, or SOS certificates provide verifiable guarantees. The third is learning-based robustification, where ambiguity geometry, controller parameters, and certificates are adapted from data, but theorem-level guarantees are typically finite horizon, local to the chosen architecture, or restricted to high-probability performance bounds rather than exact optimality. Across all three regimes, the unifying idea remains the same: controller synthesis is performed against a set of plausible probability laws rather than a single estimated law, so that robustness is calibrated to statistical uncertainty instead of being imposed purely as deterministic worst-case protection.