---
title: 'Dynamic System Optimum: Principles & Applications'
url: https://www.emergentmind.com/topics/dynamic-system-optimum
type: topic
---

# Dynamic System Optimum: Principles & Applications

Searching arXiv for recent and foundational uses of “dynamic system optimum” across online optimization, traffic assignment, and dynamical systems.
Dynamic system optimum denotes a class of optimization benchmarks in which the reference object is the best admissible evolution of a system over time, rather than the best static configuration. Across the literature, the term appears in several technically distinct settings. In online convex optimization, it refers to a time-varying comparator sequence constrained by a path-variation budget, and optimality is assessed through dynamic regret [1810.03594]. In dynamic traffic assignment, it denotes a socially optimal time-dependent allocation of flows, routes, or departure decisions that minimizes total system cost under congestion dynamics [2101.00116], [2102.01899], [2508.18392]. In parameter-dependent dynamical systems, it denotes the choice of fixed parameters that optimize global stability properties such as maximal domain of attraction or minimal absorption time [1111.0495]. A related recent usage defines a dynamically optimal projection onto a prescribed slow spectral manifold by minimizing integrated future trajectory mismatch [2503.18021]. Another closely aligned formulation treats policy tuning itself as optimization of an induced autonomous Markov system, calling this “Dynamical System Optimization” [2506.08340]. These usages differ in state space, decision variables, and optimality criteria, but they share a common principle: optimality is assigned to system evolution under temporal constraints, rather than to a static snapshot.

## 1. Dynamic optimum as a moving benchmark in online optimization

In online convex optimization, dynamic system optimum arises through **dynamic regret**, which replaces the classical fixed comparator by a time-varying comparator sequence. The setting considered in “Proximal Online Gradient is Optimum for Dynamic Regret” uses composite losses
\[
f_t(x)=F_t(x)+H(x),
\]
with \(F_t\) and \(H\) convex and closed, a compact convex domain \(\mathcal X\subset\mathbb R^d\), bounded diameter
\[
\|x-y\|_2^2 \le R,\qquad \forall x,y\in\mathcal X,
\]
and bounded subgradients
\[
\|G_t(x)\|_2 \le G,\qquad \forall x\in\mathcal X,\quad G_t(x)\in \partial F_t(x)
\]
[1810.03594].

The dynamic comparator class is constrained by a weighted path-variation budget,
\[
\mathcal L_{D_\beta} := \left\{ \{y_t\}_{t=1}^T : \sum_{t=1}^{T-1} t^\beta \|y_{t+1}-y_t\| \le D_\beta \right\}, \qquad 0\le \beta <1,
\]
which generalizes the classical path budget \(\sum_{t=1}^{T-1}\|y_{t+1}-y_t\|\le D_0\) obtained when \(\beta=0\) [1810.03594]. Dynamic regret is then
\[
\mathcal R_T^A := \sum_{t=1}^T f_t(x_t) - \min_{\{y_t\}_{t=1}^T\in \mathcal L_{D_\beta}} \sum_{t=1}^T f_t(y_t).
\]
Here the “dynamic optimum” is the best sequence of actions in hindsight subject to the variation budget, not an arbitrary clairvoyant sequence [1810.03594].

The algorithmic result is that **Proximal Online Gradient (POG)**,
\[
x_{t+1} = \operatorname{prox}_{H,\eta_t}\bigl(x_t-\eta_t G_t(x_t)\bigr),
\]
with
\[
\operatorname{prox}_{H,\eta_t}(x') := \arg\min_{x\in\mathcal X} \left\{ H(x)+\frac{1}{2\eta_t}\|x-x'\|_2^2 \right\},
\]
achieves the minimax-optimal order of dynamic regret under these assumptions [1810.03594]. The upper bound has the form
\[
\sup_{\{f_t\}_{t=1}^T\in \mathcal F^T} \mathcal R_T^{\mathrm{POG}}
\;\lesssim\;
\sqrt{R}\max_{t\in[T]}\left\{\frac{t^\beta}{\eta_t}\right\}D_\beta
+ \frac{R}{2\eta_T}
+ \frac{G^2}{2}\sum_{t=1}^T \eta_t
+ H(x_1)-H(x_{T+1}),
\]
and with tuned non-increasing steps \(\eta_t=t^{-\gamma}\sigma_1\), \(\gamma\in(\beta,1)\), this becomes
\[
\sup_{\{f_t\}_{t=1}^T\in \mathcal F^T} \mathcal R_T^{\mathrm{POG}}
\;\lesssim\;
\sqrt{D_\beta\,T^{1-\beta}}+\sqrt{T}
\]
[1810.03594]. A matching lower bound,
\[
\inf_{A\in\mathcal A}\sup_{\{f_t\}_{t=1}^T\in\mathcal F^T}\mathcal R_T^A
\;\gtrsim\;
\sqrt{D_\beta\,T^{1-\beta}}+\sqrt{T},
\]
shows that no online algorithm can do better in order [1810.03594]. In this sense, the dynamic system optimum is a constrained moving benchmark, and POG is optimal relative to it.

A closely related interpretation appears in shifting regret. When the comparator is allowed only a limited number of switches, the same framework yields
\[
\mathcal R_T^{\mathrm{POG}} \lesssim \sqrt{MT}+\sqrt T,
\]
so the moving optimum may also be parameterized by a switching budget rather than a path-length budget [1810.03594]. This suggests that, in online learning, “dynamic system optimum” is best understood as a feasible nonstationary reference trajectory endowed with explicit temporal regularity.

## 2. Dynamic system optimum in traffic assignment

In traffic assignment, dynamic system optimum usually denotes a socially optimal time-dependent traffic state that minimizes total system travel cost under congestion propagation. The term is used in both microscopic atomic-user models and macroscopic aggregate-flow models, with the common feature that users interact through time-dependent congestion rather than static link costs [2101.00116], [2102.01899], [2508.18392].

In “Dynamic system optimal traffic assignment with atomic users: Convergence and stability,” the problem is formulated on general many-to-many networks with **fixed departure times**, **atomic users**, and **route choice only** [2101.00116]. Each user \(i\in\mathcal P\) has origin \(o_i\), destination \(d_i\), fixed departure time \(s_i\), and feasible route set \(\mathcal R_i=\mathcal R(o_i,d_i)\) augmented by a null strategy \(\phi_i\). For a route profile \(\mathbf r\), the dynamic loading model uniquely determines each user’s travel time \(C_i(\mathbf r)\), under assumptions including FIFO on each link, causality, node rules consistent with realistic macroscopic node-model requirements, and specified merge priority to ensure unique trajectories [2101.00116].

The DSO objective is
\[
TC(\mathbf r)=\sum_{i\in\mathcal P} C_i(\mathbf r),
\qquad
\min_{\mathbf r\in\mathcal R} TC(\mathbf r),
\]
so optimality means minimizing total travel time of all users in the dynamic network state [2101.00116]. The paper defines the external cost imposed by user \(i\) as
\[
E_i(r_i,\mathbf r_{-i}) =
\sum_{i' \in \mathcal P\setminus\{i\}}
\left\{
C_{i'}(r_i,\mathbf r_{-i}) - C_{i'}(\phi_i,\mathbf r_{-i})
\right\},
\]
and the marginal social cost as \(C_i(r_i,\mathbf r_{-i})+E_i(r_i,\mathbf r_{-i})\) [2101.00116]. With utility
\[
U_i(r_i,\mathbf r_{-i}) = - C_i(r_i,\mathbf r_{-i}) - E_i(r_i,\mathbf r_{-i}),
\]
the authors show that the resulting **DSO game** is an exact potential game with potential
\[
\Pi(\mathbf r)=-TC(\mathbf r).
\]
For any unilateral deviation, the utility change equals the potential change, so Nash equilibria coincide with local minima of total system cost in the discrete route-profile space [2101.00116].

This game-theoretic formulation yields several dynamical results. Under **better response dynamics**, from any initial profile, the route profile converges almost surely to a Nash equilibrium state, that is, a locally or globally optimal state, and the final total cost is lower than the initial total cost [2101.00116]. Under **best response dynamics**, closed communication classes consist of Nash equilibrium states with equal total costs, and the process converges almost surely to a set of such equilibrium states [2101.00116]. Under **logit response dynamics**
\[
p_i^\beta(r,\mathbf r^\tau) =
\frac{\exp\!\left(\beta U_i(r,\mathbf r_{-i}^\tau)\right)}
{\sum_{r'\in\mathcal R_i}\exp\!\left(\beta U_i(r',\mathbf r_{-i}^\tau)\right)},
\]
the stochastically stable state is the globally optimal state minimizing total cost [2101.00116]. Thus, in this formulation, dynamic system optimum is a dynamic route assignment minimizing total travel time, and stochastic stability selects the global optimizer rather than merely a local equilibrium.

The paper further interprets the DSO game as an **evolutionary implementation scheme** for marginal cost pricing. Since each user’s utility is negative marginal social cost, the dynamics emulate exact state-dependent internalization of congestion externalities. The authors contrast this with a fixed implementation scheme using tolls \(T_i(r_i)=E_i(r_i,\mathbf r_{-i}^*)\) defined at a target state \(\mathbf r^*\) [2101.00116]. The theoretical comparison yields two properties stated in the abstract: first, “the total travel time decreases smoother to an efficient traffic state as congestion externalities are perfectly internalised”; second, “a traffic state would reach a more efficient state as the globally optimal state is stabilised” [2101.00116]. Numerical experiments suggest that this evolutionary scheme is robust in the sense that it prevents the process from visiting worse traffic states with high total travel times as often [2101.00116].

A different but related traffic meaning appears in “Dynamic traffic assignment in a corridor network: Optimum versus Equilibrium,” where DSO is defined as a **queue-free dynamic assignment** of departure or arrival times in a tandem-bottleneck corridor [2102.01899]. For the morning commute, the objective is
\[
\min_{\mathbf q\ge \mathbf 0}
\sum_{i\in \mathcal N}\int_t \left(s(t)+c_i\right)q_i(t)\,dt,
\]
subject to queue-free bottleneck capacity constraints
\[
\sum_{j=i}^{N} q_j(t)\le \mu_i,\qquad \forall i\in\mathcal N,\forall t,
\]
and demand conservation
\[
\int_t q_i(t)\,dt = Q_i,\qquad \forall i\in\mathcal N
\]
[2102.01899]. Here \(q_i(t)\) is destination arrival flow for class \(i\), \(s(t)\) is schedule delay, \(c_i\) free-flow time, and \(\mu_i\) bottleneck capacity [2102.01899]. The KKT multipliers \(p_i(t)\) serve as optimal dynamic tolls, and after eliminating “false bottlenecks,” the solution has a closed form:
\[
q_i(t)=
\begin{cases}
\hat\mu_i & \text{if}\quad t\in\mathcal T_i,\\
0 & \text{otherwise},
\end{cases}
\qquad
\hat\mu_i=
\begin{cases}
\mu_i-\mu_{i+1} & i\in\mathcal N\setminus\{N\},\\
\mu_N & i=N.
\end{cases}
\]
The active time windows are nested,
\[
\mathcal T_i\subset \mathcal T_{i+1},\qquad \forall i\in \mathcal N\setminus\{N\},
\]
their lengths satisfy
\[
T_i=\frac{Q_i}{\hat\mu_i},
\]
and class equilibrium costs are
\[
\rho_i=\bar s(T_i)+c_i
\]
[2102.01899]. Thus the corridor DSO is a layered peak-spreading solution in which each class uses a constant-rate schedule over a nested time window.

The same paper proves a precise relation between DSO and DUE under conditions on the schedule-delay function: queueing delay at bottleneck \(i\) in DUE equals the optimal toll at bottleneck \(i\) in DSO,
\[
w_i^E(t)=p_i(t),
\]
and class costs also coincide,
\[
\rho_i^E=\rho_i
\]
[2102.01899]. The broader implication is that the DSO toll is exactly the queueing delay needed to decentralize the queue-free optimum, extending the single-bottleneck intuition to a multi-bottleneck corridor.

## 3. Regional macroscopic formulations and desired arrival times

A more recent traffic formulation appears in “Dynamic System Optimum: A Projection-based Framework for Macroscopic Traffic Models,” which defines DSO on a **regional** network represented by Macroscopic Fundamental Diagram dynamics [2508.18392]. The regional graph is
\[
\mathcal G=(R,U),
\]
with regions \(r\in R\), directed adjacency \(u=(i,j)\in U\), successors \(\Gamma^+(i)\), predecessors \(\Gamma^-(i)\), and region classes \(\mathcal O,\mathcal D,\mathcal J\) for origins, destinations, and intermediate regions [2508.18392]. Demand is indexed by desired arrival time \(t_a\in\mathcal T_a\):
\[
\Delta^{OD,t_a}(t),
\qquad
\Delta^{OD,t_a}=\int_{\mathcal T_d}\Delta^{OD,t_a}(t)\,dt.
\]
The state variable is the class-specific regional accumulation
\[
N_i^{D,t_a}(t),
\qquad
N_i(t)=\sum_{D\in\mathcal D}\sum_{t_a\in\mathcal T_a}N_i^{D,t_a}(t)
\]
[2508.18392].

Regional demand and supply are induced by MFD functions,
\[
\delta_i(t)=\Delta_i(N_i(t)),
\qquad
\sigma_i(t)=\Sigma_i(N_i(t)),
\]
while the control variable
\[
\gamma_{i\ell}^{D,t_a}(t)
\]
specifies the fraction of class \((D,t_a)\) travelers in region \(i\) sent to successor \(\ell\in\Gamma^+(i)\) [2508.18392]. Under the STRADA-style split-supply model,
\[
\sigma_{ij}(t)=\beta_{ij}\sigma_j(t),
\qquad
q_{ij}(t)=\min\{\sigma_{ij}(t),\delta_{ij}(t)\},
\]
where the partial demand \(\delta_{ij}(t)\) is determined by turning fractions and class composition [2508.18392]. An alternative invariant merge model defines flows by a local optimization problem with KKT solution
\[
q_{\ell i}(t)=P_{[0,\delta_{\ell i}(t)]}\!\left(q_{\ell i,x}(1-\zeta_i(t))\right)
\]
[2508.18392].

The state evolution is governed by regional conservation laws. For intermediate regions,
\[
\dot N_i^{D,t_a}(t)=\sum_{\ell\in\Gamma^-(i)} q_{\ell i}^{D,t_a}(t)-\sum_{j\in\Gamma^+(i)} q_{ij}^{D,t_a}(t),
\]
and for origins,
\[
\dot N_O^{D,t_a}(t)=\Delta_O^{D,t_a}(t)-\sum_{i\in\Gamma^+(O)} q_{Oi}^{D,t_a}(t)
\]
[2508.18392]. In compact form,
\[
\dot N(t)=F(N(t),\Delta(t),\gamma(t)).
\]

The DSO is then formulated as an optimal control problem over time-dependent departure controls \(\Delta(t)\) and routing splits \(\gamma(t)\). The feasible set for \(\gamma\) is simplex-valued:
\[
\sum_{j\in\Gamma^+(i)}\gamma_{ij}^{D,t_a}(t)=1,
\qquad
\gamma_{ij}^{D,t_a}(t)\ge 0,
\]
while each departure profile \(\Delta_O^{D,t_a}(\cdot)\) must satisfy a fixed total mass constraint [2508.18392]. The objective includes total time spent,
\[
TTS=\sum_{i\in\mathcal J\cup\mathcal O}\int_0^{t_f} N_i(t)\,dt,
\]
arrival-time penalty,
\[
TAC=
\sum_{D\in\mathcal D}\sum_{i\in\Gamma^-(D)}\sum_{t_a\in\mathcal T_a}
\int_0^{t_f} q_{iD}^{D,t_a}(t)\,\mathcal L(t_a,t)\,dt,
\]
and a terminal penalty
\[
TC=\frac12 N(t_f)\cdot \mu \cdot N(t_f)
\]
[2508.18392]. The full objective is
\[
\min
\int_0^{t_f}
\left[
c\cdot N(t) + \sum_{t_a\in\mathcal T_a} s\cdot Q^{t_a}(N(t),\gamma(t))\,\mathcal L(t_a,t)
\right]dt
+\frac12 N(t_f)\cdot\mu\cdot N(t_f),
\]
subject to the regional dynamics and feasible-set constraints [2508.18392].

The methodological novelty is a **projected gradient algorithm** that computes gradients directly from the discretized MFD dynamics rather than from approximations of marginal path travel times. With the discretization
\[
N_{k+1}=E(N_k,\beta_k)=N_k+\Delta t\,F(N_k,\beta_k),
\qquad
\beta_k=(\gamma_k,\Delta_k),
\]
the stage cost is
\[
g_k(N_k,\beta_k)=
\left[
c\cdot N_k + \sum_{t_a\in\mathcal T_a} s\cdot Q^{t_a}(N_k,\beta_k)\,\mathcal L(t_a,k\Delta t)
\right]\Delta t,
\]
and the adjoint recursion is
\[
\Upsilon_K=\partial_{N_K}g_K,
\qquad
\Upsilon_k=\partial_{N_k}g_k+\Upsilon_{k+1}\cdot \partial_{N_k}E_k.
\]
The control gradient is
\[
\partial_{\beta_k}J = \partial_{\beta_k}g_k+\Upsilon_{k+1}\cdot \partial_{\beta_k}E_k,
\]
and the update is
\[
\beta^{\tau+1} = \mathcal P_{\mathcal K}\!\left[\beta^\tau-\alpha^\tau \partial_\beta J(\beta^\tau)\right]
\]
[2508.18392]. Projections onto routing and departure simplices are given explicitly through KKT multipliers, with sorting-based complexity \(\mathcal O(F\log F)\) for route fractions [2508.18392].

On the reported 8-region network, the paper compares the “SO \(\beta\)-controller” against MSA and a gap-based method using total travel cost and average cost. The reported totals are \(536{,}908\) for the SO controller, \(624{,}333\) for MSA, and \(611{,}752\) for the gap method, corresponding to about \(14\%\) reduction versus MSA and about \(12\%\) versus the gap-based method [2508.18392]. In the central region \(R5\), the average speed is reported as \(6.37\) m/s for the SO controller, \(5.60\) m/s for MSA, and \(5.11\) m/s for the gap method [2508.18392]. This suggests that, in regional MFD models, dynamic system optimum is most naturally interpreted as a constrained dynamic optimal control problem over aggregate traffic states, with schedule-delay terms and direct adjoint-based optimization.

## 4. System-level optimum in continuous dynamical systems

Outside traffic and online learning, the term also denotes optimization of the long-term behavior of a continuous-time dynamical system. In “Optimizing the stable behavior of parameter-dependent dynamical systems — maximal domains of attraction, minimal absorption times,” the system is
\[
\dot x = v(x;b),
\]
with state \(x(t)\in\mathbb R^d\), fixed parameter \(b\in\mathbb R^r\), compact state space \(\mathcal X\subset\mathbb R^d\), and target region \(\mathcal T\subset\mathcal X\) [1111.0495]. Trajectories leaving \(\mathcal X\) are treated as terminated. The domain of attraction is
\[
\mathcal D(b):=
\left\{
x\in \mathcal X\,\big|\,
\phi^t x\in \mathcal T\text{ for some }t\ge 0
\text{ and }
\phi^s x\in \mathcal X\ \forall s\in[0,t]
\right\},
\]
and the absorption time is
\[
\tau(x;b):=
\left\{
\begin{array}{ll}
\inf\{t\ge 0\mid \phi^t x\in\mathcal T\}, & x\in\mathcal D(b),\\[1mm]
\infty, & \text{otherwise}.
\end{array}
\right.
\]
Two optimization objectives are considered:
\[
f(b):=m(\mathcal D(b))-\alpha |b|^2,
\qquad
f(b):=\int_{\mathcal D_0}\tau(x;b)\,dx+\alpha |b|^2
\]
[1111.0495]. The first maximizes basin volume penalized by parameter magnitude; the second minimizes average time to reach the target over a prescribed region \(\mathcal D_0\) [1111.0495].

The paper’s central approximation replaces the deterministic system by a finite-state **Markov jump process (MJP)** induced by a partition \(\mathcal X_1,\dots,\mathcal X_n\) of the state space. The generator is defined by fluxes across cell boundaries:
\[
G_{n,ij} =
\begin{cases}
\displaystyle
\frac{1}{m(\mathcal X_j)}
\int_{\partial \mathcal X_i\cap \partial \mathcal X_j}
\left(v(x)\cdot n_j(x)\right)^+\,dm_{d-1}(x), & i\neq j,\\[3mm]
\displaystyle
-\frac{1}{m(\mathcal X_i)}
\int_{\partial \mathcal X_i}
\left(v(x)\cdot n_i(x)\right)^+\,dm_{d-1}(x), & i=j.
\end{cases}
\]
Off-diagonal entries are nonnegative, column sums are nonpositive, and strict negativity corresponds to leakage out of \(\mathcal X\) into a fictive outside state \(\omega\) [1111.0495].

Absorption probabilities \(p_i\) solve
\[
\sum_{j\in \mathcal Y\setminus \mathcal T} p_j G_{ji}
=
-\sum_{j\in\mathcal T} G_{ji},
\qquad
i\in\mathcal Y\setminus\mathcal T,
\]
equivalently
\[
\hat p=-\widehat G^{-T}q,
\]
while expected termination times \(t_i\) solve
\[
\sum_{j\in \mathcal Y\setminus \mathcal T} t_j G_{ji}=-1,
\qquad
i\in\mathcal Y\setminus\mathcal T,
\]
equivalently
\[
\hat t=-\widehat G^{-T}e
\]
[1111.0495]. The deterministic objectives are approximated by
\[
f_n(b):=\sum_{i=1}^n m(\mathcal X_i)\,p_{n,i}-\alpha|b|^2
\]
for domain-of-attraction maximization, and
\[
f_n(b):=\sum_{i\in \mathcal D_0} m(\mathcal X_i)\,t_{n,i}+\alpha|b|^2
\]
for absorption-time minimization [1111.0495].

A key advantage is that the chain
\[
b\mapsto v(\cdot;b)\mapsto G_n\mapsto p_n\mapsto f_n(b)
\]
is differentiable under mild assumptions, allowing analytical or semi-analytical gradients. The derivative of the generator is expressed through boundary integrals,
\[
DG(v)_{ij}\cdot \delta v=
\begin{cases}
\displaystyle
\frac{1}{m(\mathcal X_j)}
\int_{\mathcal X_{ij}^+}\delta v(x)\cdot n_j(x)\,dm_{d-1}(x), & i\neq j,\\[3mm]
\displaystyle
-\frac{1}{m(\mathcal X_i)}
\int_{\mathcal X_{ii}^+}\delta v(x)\cdot n_i(x)\,dm_{d-1}(x), & i=j,
\end{cases}
\]
and the resulting optimization is performed by gradient ascent or descent, with optional projection steps to preserve feasibility of \(\mathcal D_0\subseteq\mathcal D(b)\) [1111.0495].

In this usage, dynamic system optimum is not a moving trajectory benchmark or social traffic state. It is a parameter choice that optimizes the **global stability landscape** of the system: how much of state space reaches the target and how rapidly it does so. A plausible implication is that this usage is closest to system design or control synthesis, rather than online decision making.

## 5. Dynamically optimal projection and autonomous-system optimization

Two recent papers extend the phrase into model reduction and autonomous Markov-system optimization. In “Dynamically Optimal Projection onto Slow Spectral Manifolds for Linear Systems,” the system is the linear dissipative evolution equation
\[
\frac{dx}{dt}=Lx
\]
on a complex Hilbert space \(H\), with isolated slow eigenvalues \(\{\lambda_j\}_{1\le j\le n}\) above the essential spectrum and slow invariant manifold
\[
\mathcal M_{\rm slow}^n=\operatorname{span}\{\hat x_1,\dots,\hat x_n\}
\]
[2503.18021]. The novelty is not the manifold itself, which is assumed known spectrally, but the choice of representative point on that manifold for a general initial condition \(x_0\).

The criterion is the integrated future trajectory misfit,
\[
\mathcal E(x_0,\xi)=
\frac12\int_0^\infty
\|e^{tL}x_0-e^{tL}x_{\rm slow}(\xi)\|^2\,dt,
\qquad
x_{\rm slow}(\xi)=\sum_{j=1}^n \xi_j\hat x_j.
\]
Minimizing this over \(\xi\in\mathbb C^n\) yields the **dynamically optimal projection** [2503.18021]. The solution is explicit:
\[
\xi_i^{\rm min}(x_0)=
\sum_{j=1}^n
\left[\left(G^T\right)^{-1}\right]_{ij}
\langle (L+\lambda_j^*)^{-1}x_0,\hat x_j\rangle,
\]
where the spectrally weighted Gramian is
\[
G_{jk}=\frac{\langle \hat x_j,\hat x_k\rangle}{\lambda_j+\lambda_k^*}.
\]
The induced projection operator is
\[
\mathbb P_{\rm DOP}x=
\sum_{i=1}^n\sum_{j=1}^n
\hat x_i
\left[\left(G^T\right)^{-1}\right]_{ij}
\langle (L+\lambda_j^*)^{-1}x,\hat x_j\rangle,
\]
and satisfies
\[
\mathbb P_{\rm DOP}^2=\mathbb P_{\rm DOP}
\]
[2503.18021].

This construction differs from orthogonal projection and from the Riesz projection when \(L\) is non-normal. For normal \(L\), however, DOP reduces to the canonical orthogonal projection
\[
\mathbb P_{\rm DOP}x=\sum_{j=1}^n \langle x,\hat x_j\rangle \hat x_j
\]
[2503.18021]. In the two-dimensional shear example
\[
L=\begin{pmatrix}-1 & \gamma\\ 0 & -\alpha\end{pmatrix},
\qquad \alpha>1,\ \gamma>0,
\]
the DOP onto the slow eigenspace becomes
\[
\mathbb P_{\rm DOP}x=
\left(x_1+\frac{\gamma}{1+\alpha}x_2\right)
\begin{pmatrix}1\\0\end{pmatrix},
\]
which incorporates the transient influence of the fast variable \(x_2\) on the future slow trajectory [2503.18021]. Here “dynamic optimum” means best reproduction of future evolution, integrated over time, rather than best instantaneous geometric fit.

A more expansive autonomous-system formulation appears in “Dynamical System Optimization,” where a parameterized policy is viewed as inducing an autonomous Markov chain with transition law \(P(x'|x,\theta)\) and step cost \(L(x,\theta)\) [2506.08340]. The central claim is that once a policy is fixed, “control authority is transferred to the policy,” so optimization should be posed directly over the parameters of the induced autonomous system rather than through action-level dynamic programming machinery [2506.08340]. The generic objective is
\[
\min_\theta J(\theta),
\]
with discounted, average-cost, or finite-horizon variants. In the discounted setting,
\[
J(\theta)=\mathbb E_{x\sim P_0(\cdot)}[V(x,\theta)],
\qquad
V(x,\theta)=
\mathbb E\!\left[
\sum_t \gamma^t L(x_t,\theta)
\right],
\]
and the Bellman equation is
\[
V(x,\theta)=L(x,\theta)+\gamma \mathbb E_{x'\sim P(\cdot|x,\theta)}[V(x',\theta)]
\]
[2506.08340].

The paper derives the gradient theorem
\[
\nabla_\theta J(\theta)=
\mathbb E_{x\sim \rho(\cdot,\theta)}
\left[
\nabla_\theta L(x,\theta)
+\gamma \int \nabla_\theta P(x'|x,\theta)\,V(x',\theta)\,dx'
\right],
\]
and its score-function form
\[
\nabla_\theta J(\theta)=
\mathbb E_{x\sim \rho}
\left[
\nabla_\theta L(x,\theta)
+\gamma \mathbb E_{x'\sim P(\cdot|x,\theta)}
\left[
\nabla_\theta \ln P(x'|x,\theta)\,V(x',\theta)
\right]
\right]
\]
[2506.08340]. It also defines a DSO Fisher matrix
\[
F(\theta)=
\mathbb E_{x\sim \rho,\;x'\sim P(\cdot|x,\theta)}
\left[
\nabla_\theta \ln P(x'|x,\theta)
\nabla_\theta \ln P(x'|x,\theta)^T
\right]
\]
for natural-gradient updates [2506.08340]. A surrogate objective \(S(\theta,\alpha)\) is introduced such that
\[
\nabla_\alpha S(\theta,\alpha)\big|_{\alpha=0}=\nabla_\theta J(\theta),
\]
leading to chain-iteration and proximal-chain analogs of policy iteration and PPO [2506.08340].

The paper also identifies a special linearly-solvable setting in which optimization over all chains yields a global optimum:
\[
P^*(x'|x)=
\frac{\bar p(x'|x) Z^\gamma(x')}{G[Z^\gamma](x)},
\qquad
Z(x)=\exp(-r(x))\,G[Z^\gamma](x)
\]
[2506.08340]. This usage broadens “dynamic system optimum” into an autonomous-system view of learning, system identification, mechanism design, estimator tuning, and related tasks [2506.08340].

## 6. Structural themes, contrasts, and scope conditions

Across these literatures, dynamic system optimum is not a single formal object but a family of temporally structured optimality concepts. The commonality is optimization over system evolution; the differences lie in what is allowed to vary and what notion of feasibility restricts the optimum.

In online optimization, the admissible comparator sequence is constrained by a weighted path budget \(D_\beta\), and optimality is inherently minimax: the benchmark is the best moving comparator in hindsight, and algorithmic performance is quantified by regret against it [1810.03594]. In atomic traffic assignment, the optimum is a route profile minimizing total travel time, and game-theoretic dynamics distinguish local optima from globally stochastically stable states [2101.00116]. In corridor traffic assignment, the optimum is queue-free and constrained by capacity, producing nested time windows and constant-rate flow patterns [2102.01899]. In regional MFD models, the optimum becomes a dynamic optimal control problem over aggregate accumulations, departure releases, and routing fractions, augmented by desired arrival times and terminal penalties [2508.18392]. In parameter-dependent ODEs, the optimum is a parameter value maximizing domain of attraction or minimizing average absorption time [1111.0495]. In DOP, the optimum is a projection minimizing integrated future trajectory mismatch [2503.18021]. In autonomous Markov-system optimization, it is a parameterization minimizing cumulative cost of the induced closed-loop chain [2506.08340].

Several recurring structural devices appear. One is the use of **potential or value functions** to turn dynamic optimality into an exact characterization. In the DSO traffic game, the potential is \(-TC(\mathbf r)\) [2101.00116]. In online learning, one-step proximal inequalities and telescoping bounds reduce dynamic regret to comparator movement, gradient accumulation, and diameter terms [1810.03594]. In DOP, a convex quadratic functional over future trajectories yields a closed-form oblique projection [2503.18021]. In dynamical-system optimization via MJP approximation, absorption probabilities and termination times solve linear systems, so global stability criteria become differentiable surrogates amenable to local optimization [1111.0495]. In autonomous Markov-chain optimization, Bellman equations on the induced chain yield gradient, Hessian, and natural-gradient formulas without introducing action-value functions as primary objects [2506.08340].

Several misconceptions are also clarified by the literature. Dynamic system optimum is not necessarily the instantaneous minimizer of a time-varying objective. The online-learning paper explicitly notes that POG does **not** compute \(x_t^*=\arg\min_x f_t(x)\) at each round; instead it controls cumulative loss relative to the best feasible moving comparator sequence [1810.03594]. In traffic, DSO is not equivalent to dynamic user equilibrium: under certain conditions DUE queueing delay equals DSO toll, but the corresponding flow patterns may still differ, and equivalence can fail when schedule-delay slope conditions are violated [2102.01899]. In the atomic-user traffic game, deterministic better or best response dynamics converge only to local optima; global optimality appears only through logit-based stochastic stability or a slowly increasing inverse-noise schedule \(\beta(\tau)=\ln(\tau+1)/\lceil |\mathcal P|/2\rceil\) [2101.00116]. In dynamical-system optimization via MJP approximation, the method is global in criterion but still local in parameter search, since the proposed optimization uses gradient ascent or descent and can have multiple local optima [1111.0495]. In DOP, the manifold is assumed known; what is optimized is the projection fiber, not the invariant manifold itself [2503.18021].

The assumptions under which each notion is valid are correspondingly specific. The dynamic-regret optimality result requires convexity, closedness, compact domain, bounded subgradients, and non-increasing positive step sizes, but does not require smoothness or strong convexity [1810.03594]. The atomic-user DSO game depends on exact finite-user externalities, fixed departure times, route choice only, and a dynamic loading model with unique travel times [2101.00116]. The corridor results rely on homogeneous travelers, common desired time, deterministic bottlenecks, strict quasi-convexity of schedule delay, and corridor topology [2102.01899]. The MFD framework assumes regional homogeneity, deterministic demand, predefined MFDs, and aggregate rather than microscopic FIFO fidelity [2508.18392]. The MJP stability-design method is subject to discretization error, leakage, and lack of full convergence theorems for absorption quantities in the paper itself [1111.0495]. DOP is restricted to linear systems with a prescribed slow spectral manifold [2503.18021]. Autonomous Markov-system optimization generally assumes access to \(P(x'|x,\theta)\), \(L(x,\theta)\), and their gradients, and does not provide general convergence theorems for its stochastic optimization algorithms [2506.08340].

Taken together, these works indicate that “dynamic system optimum” is best treated as a context-dependent technical term. In each context, it specifies the best admissible temporal organization of a system under explicit structural constraints: comparator regularity in online learning [1810.03594], socially efficient traffic evolution under congestion dynamics [2101.00116], [2102.01899], [2508.18392], global stability design in nonlinear ODEs [1111.0495], time-domain trajectory fit in spectral model reduction [2503.18021], or parameter-optimal autonomous stochastic dynamics [2506.08340]. The unifying interpretation is that optimality is assigned to dynamics at the system level, not merely to isolated actions, states, or equilibrium points.

Source: https://www.emergentmind.com/topics/dynamic-system-optimum