Tracking of the Optimal Trajectory (TOT)
- TOT is a family of optimal control formulations that balance state mismatch and control actuation across various plant classes and constraints.
- It encompasses diverse methodologies such as nonlinear affine, augmented-state, regional, and hybrid formulations tailored to specific system dynamics.
- TOT research integrates learning-based adaptations, safety constraints, and planning–tracking pipelines to achieve robust, application-oriented performance.
Searching arXiv for recent and foundational papers on Tracking of the Optimal Trajectory. Tracking of the Optimal Trajectory (TOT) designates optimal-control formulations in which a controlled system is required to follow a prescribed desired or reference trajectory while minimizing a criterion that penalizes tracking error and actuation effort. In the cited literature, the problem appears in finite-horizon affine nonlinear systems, infinite-horizon augmented-state formulations, regional tracking for time-fractional diffusion equations, nonholonomic mechanical systems, hybrid systems with state jumps, and spatio-temporal planning-and-tracking pipelines (Löber, 2015, Kamalapurkar et al., 2013, Ge et al., 2021, Colombo et al., 2020, Saccon et al., 2014, Solano-Castellanos, 2023). Taken together, these formulations suggest that TOT is best understood not as a single algorithm, but as a family of optimal tracking problems whose mathematical structure is dictated by the plant class, admissible constraints, and the meaning assigned to “optimality.”
1. Canonical problem statements
A standard TOT formulation starts from a control system and a desired trajectory, then optimizes a performance index that balances state mismatch and control effort. In the nonlinear affine setting, the plant is written as
with tracking objective
where is a Tikhonov regularization parameter. This formulation makes the reference trajectory explicit and renders the problem well posed by suppressing unbounded controls (Löber, 2015).
A second canonical form converts tracking into stationary optimal control on an augmented state. For the nonlinear control-affine system
with tracking error and desired trajectory generated by , the augmented state
satisfies
with local cost
This removes explicit time dependence from the optimal tracking problem and enables Hamilton–Jacobi–Bellman and actor–critic constructions on rather than on 0 alone (Kamalapurkar et al., 2013).
A third form arises in distributed-parameter systems, where tracking may be imposed only on a spatial subregion. For the time-fractional diffusion system, the objective is not global exact tracking on all of 1, but regional tracking on 2 over 3, with both trajectory error and terminal error penalized in 4 (Ge et al., 2021).
| Formulation | Dynamics | Objective |
|---|---|---|
| Nonlinear affine finite horizon | 5 | 6 |
| Augmented-state infinite horizon | 7 | 8 |
| Regional time-fractional diffusion | 9 | regional tracking on 0 plus terminal penalty |
| Nonholonomic geometric tracking | dynamics on 1 | position and velocity mismatch plus control effort |
2. Exact realizability, regularization, and singular structure
A distinctive line of TOT research studies when a desired trajectory can be tracked exactly rather than approximately. In the nonlinear affine framework, the input matrix is assumed full rank, and two complementary projectors are introduced,
2
Multiplication of the state equation by 3 removes the control term and yields the constraint equation
4
A desired trajectory is exactly realizable if it satisfies the corresponding constraint equation together with the initial condition, in which case the control can be written explicitly as
5
This exposes the controlled and uncontrolled directions separately and yields an open-loop exact tracking law expressed directly in terms of the desired trajectory (Löber, 2015).
The same work interprets the regularization parameter 6 as a singular perturbation parameter. In the optimality system, the stationarity condition
7
implies that the limit 8 is singular. The resulting structure consists of an outer problem, inner boundary-layer problems near 9 and 0, and a composite solution. For vanishing regularization, the state trajectory may become discontinuous and the control may diverge, with delta-like boundary impulses in distributional form. This is a central warning against interpreting unregularized optimal tracking as a benign limit (Löber, 2015).
Hybrid systems add a different singularity: the event time itself varies under perturbation. For state-triggered jumps defined by 1, the perturbed event occurs at 2. The first-order jump-time sensitivity,
3
and the linearized jump relation,
4
with saltation-like matrix
5
show that local TOT for hybrid references cannot be reduced to a naive smooth linearization. The proposed local tracking law switches between extended ante-event and post-event reference branches, and the associated hybrid LQR Riccati equation carries a jump condition at the event (Saccon et al., 2014). This suggests that in hybrid TOT, the relevant tracking error is branch-dependent rather than the raw state difference.
3. Geometric, nonholonomic, and safety-constrained tracking
For nonholonomic mechanical systems, TOT is formulated directly on the constraint distribution rather than on unconstrained error dynamics. The system is modeled by a triple 6, where 7 is a Riemannian metric, 8 a potential, and 9 a regular nonintegrable distribution. A reference trajectory
0
is given, and an admissible curve 1 is sought so as to minimize
2
In adapted coordinates on 3, admissibility is 4, while the controlled reduced dynamics are
5
Pontryagin minimum conditions then give state equations, costate equations, and the stationarity relation that yields 6 in terms of the costates (Nayak et al., 2019).
A later geometric treatment recasts the same problem as a constrained variational problem on a second-order manifold 7. The variational formulation is equivalent to the PMP description when 8, but it also supports structure-preserving variational integrators through a discretized action. The resulting discrete Euler–Lagrange system preserves symplecticity, momentum behavior, good long-time energy behavior, and accurate nonholonomic constraint evolution more faithfully than generic ODE solvers (Colombo et al., 2020). In this line of work, TOT is explicitly motivated as an alternative to direct error stabilization, which is obstructed in general by Brockett-type conditions.
Safety-constrained TOT can also be formulated through barrier transformations and temporal logic. For the control-affine system
9
the tracking error relative to a selected trajectory 0 is 1, and safety is imposed through a polytope
2
Each constrained coordinate is mapped by the barrier
3
which produces an unconstrained transformed system 4. The tracking-and-mission problem is decomposed by a finite state automaton derived from co-safe and safe LTL formulas, and the optimal policy in transformed coordinates takes the bounded-input form
5
Under the event-trigger condition given in the paper, the closed-loop tracking error system is asymptotically stable when 6 and ultimately uniformly bounded when 7, while Zeno behavior is excluded (Kanellopoulos et al., 2021).
4. Regional and distributed-parameter TOT
A distributed-parameter realization of TOT is given by optimal regional tracking for time-fractional diffusion systems under Neumann boundary conditions. The controlled plant is
8
with Caputo derivative order 9. The regional cost functional is
0
where 1 and 2 is the target subdomain. The system admits a unique mild solution with spectral representation based on eigenpairs of 3 and Mittag-Leffler kernels 4 and 5 (Ge et al., 2021).
The optimal control construction uses the Hilbert Uniqueness Method. Because 6 is strictly convex and coercive on the admissible set, the minimizer satisfies a variational inequality involving an adjoint state 7. For unconstrained controls, 8, the optimality condition reduces to the explicit feedback law
9
The adjoint itself is represented spectrally through the adjoint kernel 0, so the controller is written directly in terms of the tracking error on 1 over the time horizon and at terminal time. A major point of the construction is that global controllability on 2 is not required; the method is explicitly intended for systems that may be uncontrollable on the whole domain, for situations in which only a critical subregion needs regulation, or where computation should be reduced (Ge et al., 2021).
The same paper emphasizes that the result remains novel even in the classical limit 3, where the Caputo derivative becomes the usual time derivative and the model reduces to standard diffusion. In the numerical example with 4, 5, and large tracking weights, the final tracking error on the target interval satisfies
6
while the control energy remains finite and moderate (Ge et al., 2021).
5. Learning-based and data-driven TOT
Approximate dynamic programming provides one route to nonlinear infinite-horizon TOT. In the augmented-state setting, the optimal value function satisfies the HJB equation
7
with optimal policy
8
A neural-network approximation
9
supports actor–critic adaptation laws based on the Bellman error. The main result is not exact asymptotic optimal tracking, but ultimate boundedness of the tracking error and ultimate boundedness of the policy estimation error under the stated gain conditions and persistent excitation assumption (Kamalapurkar et al., 2013).
A hardware-oriented variant appears in the large-scale ball-on-plate system, where the unknown discrete-time dynamics are controlled by an off-policy LSPI algorithm. Instead of setpoint tracking, the desired trajectory is locally approximated by
0
and in the experiments by the quadratic basis
1
The learned Q-function is quadratic in an augmented vector built from state, input, and reference parameters, which yields an analytic policy-improvement step and a final controller of the form
2
The constant term automatically compensates static asymmetry of the real apparatus. The reported training set uses 1200 tuples, about 48 s of excitation data, and the experiments show lower accumulated cost than both setpoint controllers and a model-based optimal trajectory controller (Köpf et al., 2020).
Model-free reinforcement learning also appears in time-optimal path tracking for industrial manipulators. TOPTO-SARSA learns first a safe trajectory in the 3 phase plane using kinematic feasibility, then refines it through interaction with the real robot until measured torques remain within limits. The method modifies SARSA by selecting the greatest feasible pseudo-velocity during exploitation and by propagating penalties backward more aggressively after failed episodes. On the reported experiments, execution time improves from 0.8004 s to 0.7806 s on a line path and from 1.3648 s to 1.3065 s on a cosine path, while measured torques satisfy the specified bounds after enough interactions (Xiao et al., 2019).
In microswimmer control, TOT is formulated as a constrained optimal control problem over a fixed horizon and solved by B-spline control parametrization together with Scalable Constrained Bayesian Optimization. The objective penalizes trajectory mismatch over time and terminal error, while the decision variables are the spline control points. The same optimization pattern is applied across a flagellated magnetic swimmer ODE model and a three-sphere swimmer PDE model, including wall-induced hydrodynamic effects. The reported behavior includes accurate tracking of straight, elliptical, and non-planar ellipsoidal trajectories, as well as wall-compensating figure-eight patterns in the three-sphere swimmer (Palazzolo et al., 10 Feb 2026).
A different but related data-driven insight is that state-estimation optimality need not coincide with tracking optimality. For a nanoAUV using USBL, IMU, depth sensor, and either MAG or AHRS, Bayesian optimization is applied to tune UKF covariance parameters under open-loop and closed-loop objectives. The key conclusion is explicit: the filter that is best for state estimation is not necessarily best for tracking. With BO-tuned navigation parameters, the abstract reports that the median tracking error is reduced by up to 50% compared to default parametrization, and the closed-loop tuning yields more successful goal reaches in extensive Monte Carlo simulations (Nitsch et al., 2023).
6. Planning–tracking architectures and spatio-temporal optimization
Several works separate TOT into trajectory generation and trajectory tracking. For on-road autonomous vehicles, the planner supplies a reference trajectory, a strictly convex quadratic program generates a smoother nominal trajectory and feedforward input sequence, and a TVLQR controller tracks that nominal motion. Strict convexity yields a unique optimizer, eliminating nondeterminism in the control synthesis. In the reported 5-second optimization benchmark, the QP has 100 decision variables and 200 inequality constraints, and the average MATLAB quadprog solve time is 43.65 ms over 150 runs (Liu et al., 2018).
A related two-stage pipeline is used for Dubins-car time-optimal control. Pontryagin’s Maximum Principle produces a two-point boundary value problem, Galerkin’s Weighted Residuals Method discretizes state and adjoint trajectories, Sequential Convex Programming refines the nonlinear algebraic system, and a receding-horizon MPC then tracks the nominal open-loop optimum. The paper explicitly motivates the MPC stage as a way to reject disturbances because time-optimal control yields an open-loop controller (Solano-Castellanos, 2023).
For agile quadrotors, time-optimal planning and tracking are coupled differently. The planner fixes waypoint-node assignments and uses separate segment sampling intervals 4, thereby minimizing total flight time without the combinatorial waypoint-allocation burden of earlier methods. The discrete solution is then converted into a continuous reference trajectory tracked by time-adaptive MPC, which optimizes an initial timing offset 5 online. This temporal alignment improves robustness to disturbances and model mismatch. In dynamic waypoint replanning experiments, the reported replanning time is 0.12 s, the maximum flight speed is 10.2 m/s, and the average throttle is about 21 m/s6, close to the theoretical maximum 23 m/s7 (Zhou et al., 2023).
Spatio-temporal target tracking in cluttered environments leads to yet another architecture. Elastic Tracker predicts target waypoints, finds an occlusion-aware guiding path, builds a safe flight corridor as a sequence of convex polytopes, defines sector-shaped visible regions, and finally optimizes a MINCO polynomial trajectory under smoothness, time, dynamic-feasibility, corridor, visibility, and distance-keeping costs. The base objective is
8
augmented by analytical occlusion and distance penalties and by sampled integral penalties for velocity, acceleration, and corridor violations. In the benchmark timings, total computation is reported as 3.184 ms for 9 m/s and 5.2 ms for 0 m/s, both lower than the compared baselines (Ji et al., 2021).
Space-object tracking places the same planning–tracking logic into an orbital optimization setting. A chaser spacecraft tracks a known target trajectory while remaining within the annular corridor
1
with discrete on/off thrusters modeled by binary variables and exact inertial-frame dynamics discretized by RK4. The mixed-integer nonlinear program is tightened by a perspective reformulation and a rounding heuristic. In the reported case study, 2 km, 3 km, the horizon is 1 hour with 4, and the total pipeline time is about 38.2 s including model construction and data loading (Kazi et al., 28 Jun 2025).
7. Common themes, misconceptions, and limitations
A recurrent misconception is that TOT necessarily means exact global coincidence between state and reference. The cited literature shows otherwise. Regional TOT explicitly tracks only on 5 rather than on all of 6 (Ge et al., 2021). Space-object tracking is defined by remaining inside a prescribed proximity corridor rather than matching the target position exactly (Kazi et al., 28 Jun 2025). In target-tracking theory, the “optimal trajectory” may denote the best-fitting trajectory function of time obtained from the regularized criterion
7
rather than a control law at all; the corresponding polynomial T-FoT is estimated by either order-limiting ORLS or 8-regularized hybrid Newton iterations (Li et al., 22 Feb 2025).
A second misconception is that improving a subsystem in isolation necessarily improves TOT. The AUV study shows explicitly that the best navigation filter in open-loop state-estimation terms is not necessarily best for closed-loop tracking (Nitsch et al., 2023). The nonlinear affine theory shows that sending the control penalty to zero may yield discontinuous states and diverging controls rather than an innocuous “better” tracker (Löber, 2015). The hybrid-systems analysis shows that using the raw difference 9 near a jump can be misleading because perturbed and nominal trajectories jump at different times; extended ante-event and post-event branches are required for a meaningful first-order error notion (Saccon et al., 2014).
Across the literature, several structural devices recur: projector decompositions in affine systems, augmented-state embeddings for HJB and ADP, geometric reduction to the constraint distribution in nonholonomic mechanics, barrier maps for safe tracking, spline or polynomial parametrizations for expensive optimal-control problems, and two-stage planning–tracking decompositions with MPC or LQR wrappers. This suggests that TOT is less a singular doctrine than a unifying optimization viewpoint: the reference trajectory, admissible dynamics, and tracking architecture are jointly designed so that trajectory fidelity is measured relative to the correct geometry, the correct domain, and the correct performance criterion.