---
title: Sparse Optimal Control
url: https://www.emergentmind.com/topics/sparse-optimal-control
type: topic
---

# Sparse Optimal Control

Sparse optimal control denotes a family of optimal control formulations in which some notion of sparsity is optimized together with stability or performance. In the literature, the sparse object may be the time support of the control, the set of active actuator or sensor channels, the spatial or spatio-temporal support of a distributed control, or the frequency support of a time–frequency representation. Correspondingly, sparse optimal control appears as minimum-support or “maximum hands-off” control, as $\ell^1$- or group-regularized controller synthesis, as PDE-constrained control with sparsity-inducing penalties, and as finite-actuator control of continuum or mean-field models [1412.5707] [2409.09596] [1507.00768].

## 1. Foundational formulations and equivalence principles

A classical finite-dimensional formulation considers the linear time-invariant single-input system
\[
\frac{d\mathbf{x}(t)}{dt} = A \mathbf{x}(t) + B u(t),
\]
with fixed horizon $T>0$, terminal condition $\mathbf{x}(T)=\mathbf{0}$, and amplitude constraint $\|u\|_\infty \le 1$. The sparse objective is the Lebesgue measure of the control support,
\[
\|u\|_0 \triangleq m(\mathrm{supp}(u)),
\]
leading to
\[
P_0:\quad \min_{u \in U(\boldsymbol{\xi})} \ \|u\|_0.
\]
The associated minimum-fuel problem is
\[
P_1:\quad \min_{u \in U(\boldsymbol{\xi})} \ \|u\|_1.
\]
Under the normality assumption, the two problems coincide: if $P_1$ is normal and admits an optimal control, then
\[
U_0^*(\boldsymbol{\xi}) = U_1^*(\boldsymbol{\xi}),
\qquad
\|u_0\|_0 = \|u_1\|_1,
\]
and therefore $V_0(\boldsymbol{\xi}) = V_1(\boldsymbol{\xi})$ on the reachable set $R(T)$. In that regime, sparse optimal controls can be computed through the convex $L^1$ problem rather than directly through the nonconvex $L^0$ problem. The corresponding Pontryagin structure is bang-off-bang:
\[
u^*(t) = -\mathrm{dez}\bigl(B^{\top}\mathbf{p}^*(t)\bigr),
\]
where the switching function is $B^\top \mathbf{p}^*(t)$ and $\mathrm{dez}$ is the dead-zone map. Under controllability and nonsingular $A$, the value function $V_0$ is continuous on $R(T)$ [1412.5707].

Infinite-horizon sparse optimal control extends this viewpoint to discounted problems with running cost
\[
J(x,u)=\int_0^\infty e^{-\lambda s}\Big(\ell_1(y(s))+\gamma\|u(s)\|_p^p\Big)\,ds,
\qquad 0<p\leq 1.
\]
For $p=1$, the Hamiltonian minimization yields hard thresholding. For $0<p<1$, the singular derivative of $|u|^p$ near the origin strengthens sparsity and produces bang-off-bang behavior. For box constraints $U=\prod_{i=1}^m[-\rho_i,\rho_i]$, the threshold condition is explicit:
\[
\bar u_i(s)=0 \quad \text{if } \rho_i^{1-p}|c_i(s)|<1,
\]
while saturation occurs when the threshold is exceeded. In the one-dimensional Eikonal case, the optimal control is
\[
\bar u(t)=
\begin{cases}
-\rho\,\operatorname{sgn}(x),& 0\le t<\tfrac{|x|-\lambda\gamma}{\rho},\\
0,& t\ge \tfrac{|x|-\lambda\gamma}{\rho},
\end{cases}
\]
whenever $|x|>\lambda\gamma$ [1602.08931].

## 2. Convex co-design of sparse controllers, actuators, and sensors

A distinct line of work treats sparse optimal control as a controller-synthesis problem for continuous-time LTI plants with prescribed $\mathcal H_2$ or $\mathcal H_\infty$ performance. The plant is written as
\[
\dot{x} = A x + B u + B_1 w,\qquad
y = C_y x + D_{yw} w,\qquad
z = C_z x + D_z u + D_{zw} w,
\]
and two controller classes are considered: static full-state feedback $u=Kx$ and dynamic output feedback
\[
\dot{x}_K = A_K x_K + B_K y,\qquad u = C_K x_K + D_K y.
\]
Sparsity is induced in two ways. Channel-wise actuator sparsity is promoted by minimizing a weighted $\ell_1$ norm of bounds on the disturbance-to-control $\mathcal H_2$ norms of individual channels, while structural sparsity is induced by $\ell_1/\ell_2$ penalties on rows of $[\hat C_K\ \hat D_K]$ and columns of $[\hat B_K\ \hat D_K]$. In the static full-state case, the change of variables $X \succ 0$, $W:=KX$ converts bilinear synthesis constraints into semidefinite programs. In the output-feedback case, Scherer’s transformed variables $\hat A_K,\hat B_K,\hat C_K,\hat D_K$ and Lyapunov partitions are used, together with realizability constraints such as $D_K D_{yw}=0$ and, for $\mathcal H_2$, $D_z=0$ when $D_{yw}\neq 0$ [2409.09596].

The crucial structural result is sparsity preservation: if $[\hat C_K\ \hat D_K]$ is row-sparse and $[\hat B_K;\hat D_K]$ is column-sparse, then the reconstructed controller matrices $[C_K\ D_K]$ and $[B_K;D_K]$ inherit the same actuator and sensor sparsity. The resulting problems are convex SDPs, so the synthesized solutions are globally optimal within the relaxed formulation class. The paper’s tensegrity-wing study makes the architecture implications explicit. For $\gamma_0=0.1$, sparse full-state and sparse output-feedback designs select primarily cables $3$, $4$, and $7$, with slight use of cable $6$ in output feedback. Under structural sparsity, the $\mathcal H_2$ design selects actuators $3,4,6,7$ and sensors $4,6,10,12$, whereas the $\mathcal H_\infty$ design selects actuators $3,4,6,7$ and sensors $6,10,11,12$. Tightening channel bounds $\gamma_i \le \gamma_{i,\max}$ forces additional actuators into the design, exposing the performance–magnitude–sparsity trade-off [2409.09596].

## 3. PDE-constrained sparse control

In PDE-constrained settings, sparsity is usually enforced by an $L^1$ term in the control cost. For fractional diffusion with the spectral fractional operator $\mathcal L^s$, the cost
\[
J(u,z) = \frac{1}{2}\,\|u - u_d\|_{L^2(\Omega)}^2 + \frac{\sigma}{2}\,\|z\|_{L^2(\Omega)}^2 + \nu\,\|z\|_{L^1(\Omega)}
\]
leads to the pointwise optimality condition
\[
\bar z(x') = \operatorname{Proj}_{[a,b]}\!\left( -\frac{1}{\sigma}\,S_{\nu}\big(\bar p(x')\big) \right),
\qquad
\bar z(x')=0 \iff |\bar p(x')|\le \nu.
\]
Using the Caffarelli–Silvestre extension on the truncated cylinder and piecewise-constant controls with tensor-product finite elements for the state, the a priori control and state errors scale like
\[
|\log(\#\mathscr{T}_Y)|^{\,2s}\,(\#\mathscr{T}_Y)^{-1/(n+1)}
\]
up to the norms specified in the analysis [1704.01058].

For semilinear elliptic equations with the integral fractional Laplacian, the same thresholding structure persists. With
\[
J(y,u) = \int_\Omega L(x,y)\,dx + \frac{\alpha}{2}\|u\|_{L^2(\Omega)}^2 + \beta\|u\|_{L^1(\Omega)},
\]
the first-order condition becomes
\[
u^*(x) = \Pi_{[a,b]}\!\left(-\frac{1}{\alpha}\big(p^*(x)+\beta\,\eta^*(x)\big)\right),
\qquad
\eta^*(x) = \Pi_{[-1,1]}\!\left(-\frac{1}{\beta}p^*(x)\right),
\]
hence
\[
u^*(x)=0 \iff |p^*(x)|\le \beta.
\]
The analysis includes second-order conditions on a critical cone and finite-element error estimates for both fully discrete and semidiscrete schemes [2312.08335].

Other distributed PDE models exhibit more structured sparsity. For the viscous Camassa–Holm equations, three convex sparsity functionals are treated:
\[
j_1(u)=\|u\|_{L^1(Q)^3},\qquad
j_2(u)=\|u\|_{L^2(0,T;L^1(\Omega)^3)},\qquad
j_3(u)=\|u\|_{L^1(\Omega;L^2(0,T)^3)}.
\]
They induce, respectively, pointwise spatio-temporal sparsity, time-sparse thresholds, and spatial group sparsity. The corresponding control laws imply
\[
|\bar\lambda_i(x,t)|\le \kappa \Rightarrow \bar u_i(x,t)=0
\]
for $j_1$,
\[
|\bar\lambda_i(x,t)|\le \kappa\,\bar\sigma_i(t) \Rightarrow \bar u_i(x,t)=0
\]
for $j_2$, and
\[
\|\bar\lambda_i(x,\cdot)\|_{L^2(0,T)}\le \kappa
\iff
\|\bar u_i(x)\|_{L^2(0,T)}=0
\]
for $j_3$. As $\kappa\to 0$, the optimal controls converge to those of the nonsparse problem, with rates $O(\kappa)$ for $j_1$ and $j_3$, and $O(\kappa^{2/3})$ for $j_2$ [2604.02724].

For viscous Cahn–Hilliard systems with logarithmic potential, sparsity enters through
\[
\mathcal{J}(\varphi,u)=\frac{b_1}{2}\int_Q|\varphi-\varphi_Q|^2\,dx\,dt
+\frac{b_2}{2}\int_\Omega|\varphi(T)-\varphi_\Omega|^2\,dx
+\frac{b_3}{2}\int_Q|u|^2\,dx\,dt
+K\,\|u\|_{L^1(Q)}.
\]
The first-order condition is
\[
\int_Q\big(r^*+b_3 u^*+K\,\chi^*\big)\,(u-u^*)\,dx\,dt \ge 0,
\]
with $\chi^*\in \partial \|u^*\|_{L^1(Q)}$. If $0\in[\underline u,\overline u]$, then
\[
|r^*(x,t)|<K \Rightarrow u^*(x,t)=0.
\]
Second-order sufficient conditions are formulated on the critical cone and yield quadratic growth [2402.18506].

Under parametric uncertainty, sparse PDE control may target shared support rather than pointwise sparsity. For linear PDEs with Gaussian random parameters, the stochastic control formulation uses
\[
\int_D \left(\int_\Omega u(\omega,x)^2\,d\mu\right)^{1/2} dx,
\]
which enforces the same spatial support for all realizations. The norm-reweighting variable
\[
\nu(x)=\big(\|u\|_\Omega^2(x)+\varepsilon^2\big)^{-1/2}
\]
depends on physical space only, so the algorithm avoids approximation of the random space by samples or quadrature. In the reported experiments, the Newton variant clearly outperforms the IRLS method [1804.05678].

## 4. Stochastic, noisy, and random extensions

For continuous-time stochastic systems,
\[
d x_s = f(x_s,u_s)\, ds + \sigma(x_s,u_s)\, d w_s,
\]
sparse control can be posed with the time-support penalty
\[
\|u\|_0 := \sum_{j=1}^m \mu_L(\{ s \in [t,T] : u_s^{(j)} \neq 0 \})
= \int_t^T \psi_0(u_s)\,ds,
\qquad
\psi_0(a)=\sum_{j=1}^m |a^{(j)}|^0.
\]
The value function
\[
V^s(t,x)=\inf_{u\in\mathbb U^s[t,T]} \mathbb E\!\left[\int_t^T \psi_0(u_s)\,ds + g(x_T)\right]
\]
is, in general, not differentiable, and is characterized as a viscosity solution of
\[
- v_t(t,x) + H^s(x,D_x v,D_x^2 v)=0,
\]
with
\[
H^s(x,p,M)=\sup_{u\in U}\left\{- f(x,u)\cdot p - \frac12 \operatorname{tr}(\sigma \sigma^\top(x,u) M) - \psi_0(u)\right\}.
\]
For control-affine systems with $U=[-1,1]^m$, the $L^0$ and $L^1$ value functions coincide, and optimal controls are bang-off-bang, with each component taking values in $\{-1,0,1\}$ according to the threshold rule built from $f_j(x)\cdot D_x V^s$ [2109.07716].

In discrete time, multiplicative-noise LQR leads to a different sparse-control problem: the sparse object is the feedback matrix $K$ rather than the open-loop signal. The system is
\[
x_{t+1} = \left(A  + \sum_{i=1}^p \delta_{it} A_i\right) x_t
+ \left(B + \sum_{j=1}^q \gamma_{jt} B_j\right) u_t,
\]
and the regularized objective is
\[
J(\pi) + \gamma J_{\text{reg}}(\pi).
\]
The paper studies entrywise $\ell^1$, row and column max norms, group LASSO, and sparse group LASSO penalties, and computes analytic policy gradients through generalized Lyapunov equations. On a $51$-node Erdős–Rényi network, $\ell^1$ regularization under low multiplicative noise achieved $75.5\%$ sparsity at $\gamma=10$ and $94.3\%$ sparsity at $\gamma=320$, while row group LASSO achieved $18\%$ sparsity at $\gamma=400$ and $94\%$ sparsity at $\gamma=18102$. The reported algorithms converged to performant sparse mean-square stabilizing controllers [1905.13548].

## 5. Mean-field, Wasserstein, and graphon limits

Sparse optimal control also arises in large-population limits where only finitely many agents are directly actuated. In leader–follower systems, the finite-dimensional dynamics couple controlled leaders and uncontrolled followers through an interaction kernel $H$, and the finite-dimensional cost is
\[
J_N(u) := \int_0^T \left\{ L\big(y(t),w(t),\mu_N(t)\big) + \frac{1}{m}\sum_{k=1}^m |u_k(t)| \right\}\,dt.
\]
As the number of followers tends to infinity, the empirical measure $\mu_N$ converges to a probability measure solving a Vlasov-type PDE coupled with ODEs for the leaders. The finite-dimensional functionals $\Gamma$-converge to the mean-field functional
\[
J_\infty(u) := \int_0^T \left\{ L\big(y(t),w(t),\mu(t)\big) + \frac{1}{m}\sum_{k=1}^m |u_k(t)| \right\}\,dt,
\]
in weak $L^1([0,T];\mathcal U)$, and weakly convergent subsequences of finite-dimensional minimizers converge to minimizers of the mean-field problem [1402.5657].

A more recent Wasserstein-space formulation treats sparse control as structural sparsity: a finite set of controllable agents steers a continuum distribution. The coupled PDE–ODE system is
\[
\partial_t \mu(t) + \nabla_x \cdot \bigl(F[\mu(t)](\cdot,y(t))\,\mu(t)\bigr) = 0,
\qquad
\dot y(t) = G\bigl[\mu(t)\bigr](y(t)) + u(t),
\]
and the terminal objective can be
\[
\psi(\mu)=\tfrac12 W_2(\mu,\hat\mu)^2.
\]
Under Wasserstein differentiability of the terminal cost, the adjoint-based first variation yields
\[
\psi\bigl(\mu^\varepsilon(T)\bigr)=\psi\bigl(\mu^\star(T)\bigr)+\varepsilon\int_0^T q^\star(t)\cdot \delta u(t)\,dt+o(\varepsilon),
\]
and a Pontryagin-type stationarity condition
\[
q^\star(t)\cdot u^\star(t)=\min_{\omega\in U}q^\star(t)\cdot\omega
\quad \text{for a.e. } t\in[0,T].
\]
In the paper’s distribution-splitting experiment, the mean-field terminal cost was approximately $0.023$, while the same control applied to a finite-agent system with $N=500$ achieved terminal cost approximately $0.050$, averaged over $50$ seeds with standard deviation $0.027$ [2603.00373].

Infinite-dimensional linear systems and graphon control furnish another limit regime. On $L^2_{[0,1]}$, the sparse problem is
\[
\min_{u\in\mathcal U}\ \|u\|_0 + \lambda \|x(T)-x_f\|_2^2,
\qquad
\mathcal U := \{u \in L^\infty_{[0,T]}(\mathbb R^m): \|u\|_\infty \le 1\}.
\]
Under Assumption \ref{ass:BOB}, any optimal solution of the $L^1$ relaxation is bang-off-bang and is also an optimal solution of the $L^0$ problem. A nonconvex penalty class satisfying $\psi_j(0)=0$, $|u_j|<\psi_j(u_j)\le 1$ on $(-1,1)\setminus\{0\}$, and $\psi_j(\pm1)=1$ yields the same property whenever an optimizer exists. In the graphon application, sparse optimal controls for large finite networks are approximated by those of the limit graphon system under cut-norm convergence, and in one example the sparsity rate was
\[
1 - \|\check{u}\|_0/(mT) = 0.7837
\]
for $\lambda=10^6$ [2507.18030].

## 6. Broader interpretations and computational exploitation of sparsity

A recurrent misconception is that sparse optimal control always refers to sparse time-domain input signals. In quantum control, sparsity is instead imposed in the frequency variable. Controls are modeled as measures on a frequency set $\Omega$ with values in a Hilbert space $U$, and the cost is
\[
J(\psi,u) = \frac{1}{2}\,\langle \psi(T),O\,\psi(T)\rangle_H + \alpha\,\|u\|_{\mathcal M(\Omega;U)}.
\]
The optimality system yields the support condition
\[
\operatorname{supp}|\,\bar u\,|
\subset
\{\omega\in\Omega:\ \|(B^*q)(\omega)\|_U=\alpha\},
\]
and for two-scale, Gabor-$L^2$, and Fourier syntheses the support of the optimal control is finite. In the three-level example, the sparse cost selected exactly two Bohr frequencies, $\omega=3$ and $\omega=4$, with more than $99.999\%$ target population [1507.00768].

Other works use sparsity to mean exploitation of structure in the dynamics or in the numerical representation. For structurally sparse systems, the sparse object is the coupling graph itself: subsystem interactions are local, and the optimal control problem is cast as inference on a constrained factor graph. Variable elimination then yields linear time complexity in the state and control dimensions and in the time horizon under bounded-degree coupling [2104.02945]. For PDE-discretized systems, the exact solutions of generalized Lyapunov equations have banded dominant patterns, and approximate generalized Riccati solutions can be computed as banded matrices. This leads to banded feedback laws that preserve sparsity in large FE or FD models [1801.05194].

The terminology is even broader in some neighboring literatures. In neuro-dynamic programming, “sparse optimal control” can denote sparse coding of sensory inputs for value-function approximation rather than sparsity of the control inputs themselves [2006.11968]. In uncertainty quantification for PDE-constrained control, “sparse” may refer to sparse Hermite polynomial chaos truncations or sparse quadrature in parameter space, again without implying sparse control supports [1903.05547]. This suggests that the field is unified less by a single penalty than by a structural principle: sparsity is introduced where it matches the physical, architectural, or computational bottleneck of the control problem.

Source: https://www.emergentmind.com/topics/sparse-optimal-control