Papers
Topics
Authors
Recent
Search
2000 character limit reached

Optimal Control with LL^\infty and Integral Cost Functionals

Published 14 Aug 2026 in math.OC | (2608.14316v1)

Abstract: Many control problems are classically formulated using integral costs that capture the cumulative performance of a system. Peak or worst-case behavior is captured via supremum or L<sup>L<sup>\infty-costs in another variety of control problems. When both cumulative and peak performance are important, it is natural to consider objective functions that combine the two costs. Although each criterion is well studied, their combination has not been explored extensively and we precisely work on such control problems. Towards establishing the existence, we first consider the relaxed framework, where control is considered using probability distributions. Using the well-known compactness and convexity properties of such control spaces, we establish the existence of an optimal relaxed control---we eventually establish the existence of an εε-optimal pure (or classical) control, for every $ε&gt; 0.$ Despite these existence results, computing optimal policies remains challenging due to the non-smoothness introduced by the supremum term, and it is not clear whether the dynamic programming principle holds for our combined problem. To address this, we introduce a family of smooth approximations that yield standard control problems with well-defined optimal (pure) solutions. Using Maximum Theorem, we establish that the solutions of the smooth problems among pure controls form εε-optimal for the original problem, with εε tending to zero as the smoothing parameter converges to zero. Finally using the methods proposed in this paper, we study a queueing problem to illustrate (among others) that the required trade-off between peak congestion levels and cumulative performance can be achieved.

Summary

  • The paper establishes relaxed-control existence and pure-control ε-optimality by proving compactness, trajectory stability, and continuity of running maxima under broad regularity assumptions.
  • The paper introduces a mollified maximum-dynamics approximation with a uniform error bound of δ, proving convergence of smooth optimal values and ε-optimality of computed solutions for the original nonsmooth problem.
  • The paper’s queueing experiments show that peak penalties can reduce maximum congestion by up to 27% while increasing cumulative congestion costs by about 20%, revealing important multi-objective trade-offs.

Problem formulation and motivation

This paper studies deterministic optimal control problems whose objective combines a classical integral (running) reward with an LL^\infty-type peak penalty. For dynamics x˙(s)=f(s,x(s),u(s))\dot{x}(s)=f(s,x(s),u(s)), x(0)=x0x(0)=x_0, on a finite horizon T=[0,t1]T=[0,t_1] with compact control set UU, the authors maximize

J(x0;u)=0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))supsT{Lmx(s,x(s))}.J(x_0;{\bm u})=\int_0^{t_1}L_r(s,x(s),u(s))\,ds+\Psi(x(t_1))-\sup_{s\in T}\{L_{mx}(s,x(s))\}.

This differs fundamentally from Barron's formulation, which optimizes supt{tt1LrdsLmx(t,x(t))}\sup_t\{\int_t^{t_1}L_r\,ds - L_{mx}(t,x(t))\} and is amenable to a standard HJB equation. Barron's structure optimizes a running maximum of a quantity mixing instantaneous penalty with future cumulative reward; the present objective instead adds accumulated rewards to a negated running maximum over the whole horizon, which is what multi-objective applications require. The paper also distinguishes the problem from hard state-constraint formulations: there, a Lagrange multiplier must be tuned to enforce supsLmx(s,x(s))κ\sup_s L_{mx}(s,x(s))\le\kappa, whereas here the trade-off weight is prescribed exogenously.

Motivating applications include queueing systems (peak congestion determines waiting-room capacity and QoS), inventory control (peak stock level dictates storage capacity design), and epidemic control under SIS dynamics coupled with information spread on online social networks, where peak infection burden on healthcare infrastructure matters alongside cumulative infection and vaccination costs.

Existence via relaxed controls

The first contribution establishes existence of solutions. Since it is unclear whether dynamic programming applies to the combined objective, the authors take a direct route: continuity of the objective plus compactness of the domain. They embed pure controls into the space S\mathcal{S} of relaxed controls—measurable maps TP(U)T\to\mathcal{P}(U) endowed with the r-weak topology—which is compact, convex, and metrizable, with pure controls dense in it (Warga's classical results).

Under assumptions A.0–A.2 (compact x˙(s)=f(s,x(s),u(s))\dot{x}(s)=f(s,x(s),u(s))0; x˙(s)=f(s,x(s),u(s))\dot{x}(s)=f(s,x(s),u(s))1 continuous, Lipschitz in x˙(s)=f(s,x(s),u(s))\dot{x}(s)=f(s,x(s),u(s))2 uniformly in x˙(s)=f(s,x(s),u(s))\dot{x}(s)=f(s,x(s),u(s))3 with linear growth; x˙(s)=f(s,x(s),u(s))\dot{x}(s)=f(s,x(s),u(s))4, x˙(s)=f(s,x(s),u(s))\dot{x}(s)=f(s,x(s),u(s))5, x˙(s)=f(s,x(s),u(s))\dot{x}(s)=f(s,x(s),u(s))6 continuous with x˙(s)=f(s,x(s),u(s))\dot{x}(s)=f(s,x(s),u(s))7 Lipschitz), the paper proves:

  • Compactness and joint stability: the set of relaxed trajectory–control–initial-condition triples is sequentially compact, and trajectories are continuous in uniform topology with respect to both relaxed controls and initial conditions.
  • Continuity of max-trajectories: the running maximum x˙(s)=f(s,x(s),u(s))\dot{x}(s)=f(s,x(s),u(s))8 is continuous in uniform topology as a function of x˙(s)=f(s,x(s),u(s))\dot{x}(s)=f(s,x(s),u(s))9; this is identified as the first substantial result and is instrumental for everything that follows.
  • Existence: an optimal relaxed control exists for the relaxed problem.
  • x(0)=x0x(0)=x_00-optimality among pure controls: by denseness, for every x(0)=x0x(0)=x_01 there exists a measurable pure control that is x(0)=x0x(0)=x_02-optimal for the original problem.

The continuity results with respect to initial condition are noted as independently useful for stability analysis and for any future dynamic programming investigation.

Smooth approximation and convergence

Direct computation remains obstructed because the max-trajectory ODE,

x(0)=x0x(0)=x_03

has a state-dependent discontinuity (Filippov-type analysis would be required). The authors replace the indicator with a continuous mollifier x(0)=x0x(0)=x_04, yielding a family of smooth problems parametrized by x(0)=x0x(0)=x_05. Under additional assumptions A.3–A.5 (x(0)=x0x(0)=x_06 structure and growth bounds on x(0)=x0x(0)=x_07; convexity of x(0)=x0x(0)=x_08; Lipschitz/growth conditions on x(0)=x0x(0)=x_09), they show:

  • Well-posedness and existence among pure controls for each fixed T=[0,t1]T=[0,t_1]0: the joint T=[0,t1]T=[0,t_1]1 ODE has a unique solution continuous in the relaxed control, and—an important structural fact—an optimal solution exists within the pure controls T=[0,t1]T=[0,t_1]2, so standard numerical machinery applies without relaxation.
  • Uniform approximation bound: for every relaxed control, T=[0,t1]T=[0,t_1]3, proved by induction over the alternating active/inactive intervals of the smoothed dynamics. This uniform-in-T=[0,t1]T=[0,t_1]4 bound is the key quantitative estimate.
  • Convergence of values and T=[0,t1]T=[0,t_1]5-optimality: defining T=[0,t1]T=[0,t_1]6 so that T=[0,t1]T=[0,t_1]7 interpolates between the smooth value (T=[0,t1]T=[0,t_1]8) and the original value (T=[0,t1]T=[0,t_1]9), joint sequential continuity of UU0 on UU1 follows from the trajectory estimates; the Maximum Theorem then gives UU2, and consequently any sequence of smooth-problem optimizers UU3 satisfies UU4. Thus smooth-problem solutions are UU5-optimal for the original combined problem once UU6 is small enough.

Numerical solution and HJB characterization

For fixed UU7, the value function of the smooth augmented-state problem is the unique viscosity solution of an HJB equation on the extended state UU8, with Hamiltonian containing the term UU9 and terminal condition J(x0;u)=0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))supsT{Lmx(s,x(s))}.J(x_0;{\bm u})=\int_0^{t_1}L_r(s,x(s),u(s))\,ds+\Psi(x(t_1))-\sup_{s\in T}\{L_{mx}(s,x(s))\}.0; verification theorems apply when the data are J(x0;u)=0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))supsT{Lmx(s,x(s))}.J(x_0;{\bm u})=\int_0^{t_1}L_r(s,x(s),u(s))\,ds+\Psi(x(t_1))-\sup_{s\in T}\{L_{mx}(s,x(s))\}.1. Alternatively, Pontryagin's maximum principle yields two-point boundary-value co-state equations with terminal conditions J(x0;u)=0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))supsT{Lmx(s,x(s))}.J(x_0;{\bm u})=\int_0^{t_1}L_r(s,x(s),u(s))\,ds+\Psi(x(t_1))-\sup_{s\in T}\{L_{mx}(s,x(s))\}.2, J(x0;u)=0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))supsT{Lmx(s,x(s))}.J(x_0;{\bm u})=\int_0^{t_1}L_r(s,x(s),u(s))\,ds+\Psi(x(t_1))-\sup_{s\in T}\{L_{mx}(s,x(s))\}.3, and the optimizer maximizes the Hamiltonian pointwise. Both routes produce policies that are J(x0;u)=0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))supsT{Lmx(s,x(s))}.J(x_0;{\bm u})=\int_0^{t_1}L_r(s,x(s),u(s))\,ds+\Psi(x(t_1))-\sup_{s\in T}\{L_{mx}(s,x(s))\}.4-optimal for the original non-smooth problem.

Queueing application

The framework is applied to a fluid-limit single-server queue with arrival rate J(x0;u)=0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))supsT{Lmx(s,x(s))}.J(x_0;{\bm u})=\int_0^{t_1}L_r(s,x(s),u(s))\,ds+\Psi(x(t_1))-\sup_{s\in T}\{L_{mx}(s,x(s))\}.5 and controlled service rate J(x0;u)=0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))supsT{Lmx(s,x(s))}.J(x_0;{\bm u})=\int_0^{t_1}L_r(s,x(s),u(s))\,ds+\Psi(x(t_1))-\sup_{s\in T}\{L_{mx}(s,x(s))\}.6, minimizing a weighted combination of peak occupancy (J(x0;u)=0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))supsT{Lmx(s,x(s))}.J(x_0;{\bm u})=\int_0^{t_1}L_r(s,x(s),u(s))\,ds+\Psi(x(t_1))-\sup_{s\in T}\{L_{mx}(s,x(s))\}.7), linear congestion cost J(x0;u)=0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))supsT{Lmx(s,x(s))}.J(x_0;{\bm u})=\int_0^{t_1}L_r(s,x(s),u(s))\,ds+\Psi(x(t_1))-\sup_{s\in T}\{L_{mx}(s,x(s))\}.8, quadratic server-utilization deviation from an ideal rate J(x0;u)=0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))supsT{Lmx(s,x(s))}.J(x_0;{\bm u})=\int_0^{t_1}L_r(s,x(s),u(s))\,ds+\Psi(x(t_1))-\sup_{s\in T}\{L_{mx}(s,x(s))\}.9, and terminal costs. The Pareto analysis compares two frontiers: peak congestion versus utilization cost (varying supt{tt1LrdsLmx(t,x(t))}\sup_t\{\int_t^{t_1}L_r\,ds - L_{mx}(t,x(t))\}0 with supt{tt1LrdsLmx(t,x(t))}\sup_t\{\int_t^{t_1}L_r\,ds - L_{mx}(t,x(t))\}1), and cumulative congestion versus utilization cost (varying supt{tt1LrdsLmx(t,x(t))}\sup_t\{\int_t^{t_1}L_r\,ds - L_{mx}(t,x(t))\}2 with supt{tt1LrdsLmx(t,x(t))}\sup_t\{\int_t^{t_1}L_r\,ds - L_{mx}(t,x(t))\}3).

The headline numerical findings are strong and somewhat counterintuitive:

  • Explicitly penalizing peak congestion reduces peak levels by up to 27% relative to designs ignoring it—at the expense of roughly 20% degradation in cumulative congestion cost.
  • When cumulative utilization cost is either very low or very high, the difference between designs is negligible; the divergence is largest in the mid-range, described as typical of practical operating regimes.
  • Without a peak term in the objective, significant congestion can accumulate near the end of the horizon, since late-horizon congestion contributes negligibly to the cumulative integral—a failure mode invisible to purely integral formulations.

Convergence experiments with supt{tt1LrdsLmx(t,x(t))}\sup_t\{\int_t^{t_1}L_r\,ds - L_{mx}(t,x(t))\}4 show state trajectories and policies becoming nearly indistinguishable, with diminishing changes in the terminal peak estimate, consistent with the theoretical convergence result.

Limitations and open questions

Several caveats bear directly on the strength of the results. First, only supt{tt1LrdsLmx(t,x(t))}\sup_t\{\int_t^{t_1}L_r\,ds - L_{mx}(t,x(t))\}5-optimal pure controls are guaranteed for the original problem; exact optimality among pure controls is not established, and whether an optimal pure control exists at all remains open. Second, the applicability of the dynamic programming principle—and hence a direct HJB-PDE characterization—for the combined problem is unresolved; the authors derive dynamic programming relations in integral form but state plainly that it is unclear whether a corresponding HJB PDE characterizes the value function, suggesting Filippov-type solution frameworks as a possible route. Third, the approximation guarantee degrades gracefully but is only asymptotic: choosing supt{tt1LrdsLmx(t,x(t))}\sup_t\{\int_t^{t_1}L_r\,ds - L_{mx}(t,x(t))\}6 for a target accuracy relies on the qualitative convergence rather than sharp rates beyond the uniform supt{tt1LrdsLmx(t,x(t))}\sup_t\{\int_t^{t_1}L_r\,ds - L_{mx}(t,x(t))\}7-bound on the max-trajectory. Fourth, the numerical study is confined to a specific queueing instance with particular cost structures satisfying the regularity assumptions; transfer of the observed 27%/20% trade-off magnitudes to other systems is not claimed.

Conclusion

The paper delivers a complete analytical pipeline for optimal control problems combining integral and supt{tt1LrdsLmx(t,x(t))}\sup_t\{\int_t^{t_1}L_r\,ds - L_{mx}(t,x(t))\}8 criteria: existence via relaxed controls, supt{tt1LrdsLmx(t,x(t))}\sup_t\{\int_t^{t_1}L_r\,ds - L_{mx}(t,x(t))\}9-optimality of pure controls, a uniformly convergent family of smooth approximations justified through the Maximum Theorem, and computable HJB/PMP-based solvers for each smoothing level. The queueing study demonstrates that the framework yields actionable three-objective trade-off curves, quantifying how much cumulative performance must be sacrificed to control peak behavior. The principal open theoretical question is whether a viscosity-solution/HJB theory can be developed directly for the non-smooth combined problem, bypassing the smoothing argument.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.