- The paper establishes relaxed-control existence and pure-control ε-optimality by proving compactness, trajectory stability, and continuity of running maxima under broad regularity assumptions.
- The paper introduces a mollified maximum-dynamics approximation with a uniform error bound of δ, proving convergence of smooth optimal values and ε-optimality of computed solutions for the original nonsmooth problem.
- The paper’s queueing experiments show that peak penalties can reduce maximum congestion by up to 27% while increasing cumulative congestion costs by about 20%, revealing important multi-objective trade-offs.
This paper studies deterministic optimal control problems whose objective combines a classical integral (running) reward with an L∞-type peak penalty. For dynamics x˙(s)=f(s,x(s),u(s)), x(0)=x0, on a finite horizon T=[0,t1] with compact control set U, the authors maximize
J(x0;u)=∫0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))−s∈Tsup{Lmx(s,x(s))}.
This differs fundamentally from Barron's formulation, which optimizes supt{∫tt1Lrds−Lmx(t,x(t))} and is amenable to a standard HJB equation. Barron's structure optimizes a running maximum of a quantity mixing instantaneous penalty with future cumulative reward; the present objective instead adds accumulated rewards to a negated running maximum over the whole horizon, which is what multi-objective applications require. The paper also distinguishes the problem from hard state-constraint formulations: there, a Lagrange multiplier must be tuned to enforce supsLmx(s,x(s))≤κ, whereas here the trade-off weight is prescribed exogenously.
Motivating applications include queueing systems (peak congestion determines waiting-room capacity and QoS), inventory control (peak stock level dictates storage capacity design), and epidemic control under SIS dynamics coupled with information spread on online social networks, where peak infection burden on healthcare infrastructure matters alongside cumulative infection and vaccination costs.
Existence via relaxed controls
The first contribution establishes existence of solutions. Since it is unclear whether dynamic programming applies to the combined objective, the authors take a direct route: continuity of the objective plus compactness of the domain. They embed pure controls into the space S of relaxed controls—measurable maps T→P(U) endowed with the r-weak topology—which is compact, convex, and metrizable, with pure controls dense in it (Warga's classical results).
Under assumptions A.0–A.2 (compact x˙(s)=f(s,x(s),u(s))0; x˙(s)=f(s,x(s),u(s))1 continuous, Lipschitz in x˙(s)=f(s,x(s),u(s))2 uniformly in x˙(s)=f(s,x(s),u(s))3 with linear growth; x˙(s)=f(s,x(s),u(s))4, x˙(s)=f(s,x(s),u(s))5, x˙(s)=f(s,x(s),u(s))6 continuous with x˙(s)=f(s,x(s),u(s))7 Lipschitz), the paper proves:
- Compactness and joint stability: the set of relaxed trajectory–control–initial-condition triples is sequentially compact, and trajectories are continuous in uniform topology with respect to both relaxed controls and initial conditions.
- Continuity of max-trajectories: the running maximum x˙(s)=f(s,x(s),u(s))8 is continuous in uniform topology as a function of x˙(s)=f(s,x(s),u(s))9; this is identified as the first substantial result and is instrumental for everything that follows.
- Existence: an optimal relaxed control exists for the relaxed problem.
- x(0)=x00-optimality among pure controls: by denseness, for every x(0)=x01 there exists a measurable pure control that is x(0)=x02-optimal for the original problem.
The continuity results with respect to initial condition are noted as independently useful for stability analysis and for any future dynamic programming investigation.
Smooth approximation and convergence
Direct computation remains obstructed because the max-trajectory ODE,
x(0)=x03
has a state-dependent discontinuity (Filippov-type analysis would be required). The authors replace the indicator with a continuous mollifier x(0)=x04, yielding a family of smooth problems parametrized by x(0)=x05. Under additional assumptions A.3–A.5 (x(0)=x06 structure and growth bounds on x(0)=x07; convexity of x(0)=x08; Lipschitz/growth conditions on x(0)=x09), they show:
- Well-posedness and existence among pure controls for each fixed T=[0,t1]0: the joint T=[0,t1]1 ODE has a unique solution continuous in the relaxed control, and—an important structural fact—an optimal solution exists within the pure controls T=[0,t1]2, so standard numerical machinery applies without relaxation.
- Uniform approximation bound: for every relaxed control, T=[0,t1]3, proved by induction over the alternating active/inactive intervals of the smoothed dynamics. This uniform-in-T=[0,t1]4 bound is the key quantitative estimate.
- Convergence of values and T=[0,t1]5-optimality: defining T=[0,t1]6 so that T=[0,t1]7 interpolates between the smooth value (T=[0,t1]8) and the original value (T=[0,t1]9), joint sequential continuity of U0 on U1 follows from the trajectory estimates; the Maximum Theorem then gives U2, and consequently any sequence of smooth-problem optimizers U3 satisfies U4. Thus smooth-problem solutions are U5-optimal for the original combined problem once U6 is small enough.
Numerical solution and HJB characterization
For fixed U7, the value function of the smooth augmented-state problem is the unique viscosity solution of an HJB equation on the extended state U8, with Hamiltonian containing the term U9 and terminal condition J(x0;u)=∫0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))−s∈Tsup{Lmx(s,x(s))}.0; verification theorems apply when the data are J(x0;u)=∫0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))−s∈Tsup{Lmx(s,x(s))}.1. Alternatively, Pontryagin's maximum principle yields two-point boundary-value co-state equations with terminal conditions J(x0;u)=∫0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))−s∈Tsup{Lmx(s,x(s))}.2, J(x0;u)=∫0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))−s∈Tsup{Lmx(s,x(s))}.3, and the optimizer maximizes the Hamiltonian pointwise. Both routes produce policies that are J(x0;u)=∫0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))−s∈Tsup{Lmx(s,x(s))}.4-optimal for the original non-smooth problem.
Queueing application
The framework is applied to a fluid-limit single-server queue with arrival rate J(x0;u)=∫0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))−s∈Tsup{Lmx(s,x(s))}.5 and controlled service rate J(x0;u)=∫0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))−s∈Tsup{Lmx(s,x(s))}.6, minimizing a weighted combination of peak occupancy (J(x0;u)=∫0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))−s∈Tsup{Lmx(s,x(s))}.7), linear congestion cost J(x0;u)=∫0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))−s∈Tsup{Lmx(s,x(s))}.8, quadratic server-utilization deviation from an ideal rate J(x0;u)=∫0t1Lr(s,x(s),u(s))ds+Ψ(x(t1))−s∈Tsup{Lmx(s,x(s))}.9, and terminal costs. The Pareto analysis compares two frontiers: peak congestion versus utilization cost (varying supt{∫tt1Lrds−Lmx(t,x(t))}0 with supt{∫tt1Lrds−Lmx(t,x(t))}1), and cumulative congestion versus utilization cost (varying supt{∫tt1Lrds−Lmx(t,x(t))}2 with supt{∫tt1Lrds−Lmx(t,x(t))}3).
The headline numerical findings are strong and somewhat counterintuitive:
- Explicitly penalizing peak congestion reduces peak levels by up to 27% relative to designs ignoring it—at the expense of roughly 20% degradation in cumulative congestion cost.
- When cumulative utilization cost is either very low or very high, the difference between designs is negligible; the divergence is largest in the mid-range, described as typical of practical operating regimes.
- Without a peak term in the objective, significant congestion can accumulate near the end of the horizon, since late-horizon congestion contributes negligibly to the cumulative integral—a failure mode invisible to purely integral formulations.
Convergence experiments with supt{∫tt1Lrds−Lmx(t,x(t))}4 show state trajectories and policies becoming nearly indistinguishable, with diminishing changes in the terminal peak estimate, consistent with the theoretical convergence result.
Limitations and open questions
Several caveats bear directly on the strength of the results. First, only supt{∫tt1Lrds−Lmx(t,x(t))}5-optimal pure controls are guaranteed for the original problem; exact optimality among pure controls is not established, and whether an optimal pure control exists at all remains open. Second, the applicability of the dynamic programming principle—and hence a direct HJB-PDE characterization—for the combined problem is unresolved; the authors derive dynamic programming relations in integral form but state plainly that it is unclear whether a corresponding HJB PDE characterizes the value function, suggesting Filippov-type solution frameworks as a possible route. Third, the approximation guarantee degrades gracefully but is only asymptotic: choosing supt{∫tt1Lrds−Lmx(t,x(t))}6 for a target accuracy relies on the qualitative convergence rather than sharp rates beyond the uniform supt{∫tt1Lrds−Lmx(t,x(t))}7-bound on the max-trajectory. Fourth, the numerical study is confined to a specific queueing instance with particular cost structures satisfying the regularity assumptions; transfer of the observed 27%/20% trade-off magnitudes to other systems is not claimed.
Conclusion
The paper delivers a complete analytical pipeline for optimal control problems combining integral and supt{∫tt1Lrds−Lmx(t,x(t))}8 criteria: existence via relaxed controls, supt{∫tt1Lrds−Lmx(t,x(t))}9-optimality of pure controls, a uniformly convergent family of smooth approximations justified through the Maximum Theorem, and computable HJB/PMP-based solvers for each smoothing level. The queueing study demonstrates that the framework yields actionable three-objective trade-off curves, quantifying how much cumulative performance must be sacrificed to control peak behavior. The principal open theoretical question is whether a viscosity-solution/HJB theory can be developed directly for the non-smooth combined problem, bypassing the smoothing argument.