---
title: Closed-Loop Terminal-Set-Based Evasion (TSE)
url: https://www.emergentmind.com/topics/closed-loop-terminal-set-based-evasion-tse
type: topic
---

# Closed-Loop Terminal-Set-Based Evasion (TSE)

Searching arXiv for the cited papers to ground the article in current sources.
arXiv search: 2511.21633 "Bang-Bang Evasion: Its Stochastic Optimality and a Terminal-Set-Based Implementation"
Closed-Loop Terminal-Set-Based Evasion (TSE) is a practical feedback policy for stochastic endgame evasion in which the current command is restricted to the two admissible extremal accelerations and selected online by comparing analytically predicted terminal outcomes. In "Bang-Bang Evasion: Its Stochastic Optimality and a Terminal-Set-Based Implementation" [2511.21633], TSE is introduced after a structural result showing that, under linear-affine state evolution in the target input, bounded scalar target control, a convex and continuous terminal objective, finite horizon, and existence of the expectation, there exists at least one optimal control sequence and at least one optimal sequence is bang-bang. The resulting controller is receding-horizon, posterior-dependent, and terminal-set-based in the paper’s specific sense: the sign of the current maximum-magnitude input is chosen from posterior-state information, uncertainty over terminal time, and uncertainty over pursuer guidance mode, using terminal affine predictions rather than a full stochastic dynamic program.

## 1. Endgame formulation and information structure

The TSE formulation concerns a planar lateral endgame engagement between an interceptor missile and an evading target, with geometry linearized along the initial line of sight and both vehicles treated as point masses [2511.21633]. The interceptor is denoted \(M\), the target \(T\), with speeds \(V_M,V_T\), path angles \(\gamma_M,\gamma_T\), normal accelerations \(a_M,a_T\), slant range \(\rho\), and LOS angle \(\lambda\). The accelerations normal to the initial LOS are
\[
\grave{a}_{M} = a_{M} \cos (\gamma_{M}^{0} - \lambda^{0}), \qquad 
\grave{a}_{T} = a_{T} \cos (\gamma_{T}^{0} + \lambda^{0}).
\]

The discrete-time state is
\[
\mathbf{x} = \begin{bmatrix} \xi & \dot{\xi} & \mathbf{q}_{M}^{\top} & \mathbf{q}_{T}^{\top} \end{bmatrix}^{\top}\in\mathbb{R}^{n_x},
\]
where \(\xi,\dot\xi\) are relative lateral states and \(\mathbf{q}_M,\mathbf{q}_T\) are internal pursuer and evader dynamics. The state evolves according to
\[
\mathbf{x}^{k+1} = \mathbf{F}^{k}(\eta)\mathbf{x}^{k} + \mathbf{g}_{M}^{k}(\eta)u_{M}^{k} + \mathbf{g}_{T}^{k}u_{T}^{k} + \omega^{k},
\]
with random interceptor parameters \(\eta\), process noise \(\omega^k\), and target acceleration bound
\[
|u_T^k|\le u_T^{\max}.
\]

The target does not have perfect information about the engagement state from the missile’s standpoint. Instead, it observes through noise,
\[
y^{k}=h(\mathbf{x}^{k})+\nu^{k},
\]
and maintains a posterior distribution over the state. The interceptor is assumed to use a linear feedback guidance law covering PN, APN, OGL, LQDG, and more general delayed linear state feedback laws. The paper also models terminal-time uncertainty explicitly: the terminal step \(f\) is random with PMF \(p_f(i)\), and the corresponding time-to-go at time \(t^k\) has PMF
\[
p_{t_{go}^k}(i)=\frac{p_f(i/\Delta t + k)}{\sum_{i=0}^{\infty} p_f(i/\Delta t + k)}.
\]

The performance objective is stochastic. The general criterion is
\[
J\triangleq \mathbb{E}\!\left[g(\mathbf{x}^f)\right],
\]
and in the TSE instantiation the terminal objective is quadratic in a selected terminal output:
\[
g(\mathbf{x}^{f})=\|\mathbf{C}\mathbf{x}^{f}\|^2.
\]
This makes the problem a stochastic optimal control problem with imperfect information, uncertain interceptor parameters, process noise, and random terminal time.

## 2. Stochastic optimal-control problem and sufficient statistic

At decision time \(t^n\), the evasion problem is posed over the remaining horizon as maximization of expected terminal performance subject to the stochastic dynamics, uncertain guidance law, uncertain terminal time, and input bounds [2511.21633]. The target’s control must be based on posterior-state information rather than the true state. A key conceptual commitment is to the generalized separation theorem: the estimator may be designed separately, but the controller should depend on the posterior distribution \(p_x^k\), not merely on a point estimate.

The paper does not write a full Bellman recursion or an HJB equation. Instead, the sufficient statistic for control is the posterior PDF \(p_x^k\), or in the Gaussian and Gaussian-mixture instantiations used for TSE, the associated posterior means, posterior covariances, guidance-mode probabilities \(P_j^n\), and terminal-time PMF \(p_f(i)\). The missile guidance model may be classical,
\[
u_M = \frac{\mathcal N_p z_p}{t_{go}^2},
\]
with example terminal variables \(z_{\text{PN}}, z_{\text{APN}}, z_{\text{OGL}}, z_{\text{LQDG}}\), or more general delayed linear state feedback,
\[
u_M^k = \sum_{i=0}^{d}\mathbf{K}^{k-i}(\eta,t_{go}^{k-i})\mathbf{x}^{k-i}.
\]
Because the target does not know the interceptor’s exact information state or exact mode, it uses a modeled interceptor acceleration estimate
\[
\hat{u}_M^k
=
\sum_{j=1}^{p^{\max}}P_j^k
\sum_{i=0}^{d}\mathbf{K}^{k-i}(\eta,t_{go}^{k-i})\hat{\mathbf{x}}^{k-i}.
\]

This information pattern distinguishes TSE from deterministic bang-bang evasion. Classical deterministic results usually assume exact state, exact final time, and exact opponent model. Here, the control law is belief-dependent, terminal-time uncertainty is explicit, and the terminal criterion is an expectation under uncertainty. A frequent misconception is that TSE is certainty-equivalent bang-bang guidance with noisy inputs; the formulation is stricter than that. The controller is explicitly posterior-dependent, and the terminal evaluation averages over terminal-time and mode uncertainty.

## 3. Bang-bang optimality in the stochastic setting

The theoretical foundation of TSE is Theorem 1 of the paper: at each time \(t^n\), there exists at least one optimal control sequence for the evasion problem, and among the optimal sequences there is at least one with bang-bang structure,
\[
\overset{\star}{u}_{T}^{k}\in\{-u_T^{\max},u_T^{\max}\}
\quad
\text{for all }k=n,\dots,k'-1
\]
[2511.21633]. The assumptions are linear-affine state evolution in the target input, bounded scalar target control, convex and continuous terminal objective, finite horizon, and existence of the expectation.

The proof does not proceed through Pontryagin switching functions. Instead, the paper uses a convex-analysis argument based on extreme points. By recursively eliminating state and estimated-guidance variables, the terminal state can be written as an affine function of the target command sequence,
\[
\mathbf{x}^f(\mathbf{u}) = A^f\mathbf{u}+\mathbf{b}^f,
\]
where \(\mathbf{u}=[u_T^n,\dots,u_T^{k'-1}]^\top\). The feasible control set is the hypercube
\[
[-u_T^{\max},u_T^{\max}]^{k'-n}.
\]
Since \(g(\cdot)\) is convex and continuous, \(g(A^f\mathbf{u}+\mathbf{b}^f)\) is convex and continuous in \(\mathbf{u}\), and expectation preserves convexity and continuity. By the cited theorem from Beck, a maximizer of a convex continuous function over a convex compact set exists at an extreme point of that set. The extreme points of the hypercube are exactly the bang-bang sequences.

This result is structurally important. It reduces the continuum-valued control search to a combinatorial one over \(2^{\bar H}\) bang sequences over an average remaining horizon \(\bar H\). The paper is careful not to overstate this reduction: it removes one layer of intractability but does not remove the curse of dimensionality. Brute-force bang-bang MPC still scales exponentially in horizon length. TSE is introduced precisely to avoid that exhaustive search while preserving the bang-bang structure.

## 4. Terminal-set semantics and derivation of the TSE law

TSE is terminal-set-based because, at stage \(n\), the terminal state under each candidate terminal index \(i\) and each pursuer mode \(j\in\{1,\dots,p^{\max}\}\) is unrolled as an affine function of the current target command,
\[
\mathbf{x}_i^f=\mathbf{a}^n(i,j)\,u_T^n+\mathbf{z}_{i,j}^n,
\]
with
\[
\mathbf{a}^n(i,j)\triangleq \mathbf{\Phi}(i,n+1;p_j)\mathbf{g}_T,
\qquad
\mathbf{\Phi}(i,\ell;p_j)
=
\mathbf{F}^{i-1}(\eta,p_j)\cdots \mathbf{F}^{\ell}(\eta,p_j),
\]
and \(\mathbf{z}_{i,j}^n\) collecting all terms independent of the current command [2511.21633]. The random vector \(\mathbf{z}_{i,j}^n\) captures propagated process noise, measurement noise, and posterior-state uncertainty, with mean and covariance
\[
\boldsymbol{\mu}_{i,j}^n\triangleq \mathbb E[\mathbf{z}_{i,j}^n],
\qquad
\boldsymbol{\Sigma}_{i,j}^n\triangleq \operatorname{Var}(\mathbf{z}_{i,j}^n).
\]

The paper states explicitly that “terminal-set” is not a strict set-valued reachability object in the HJ sense. Rather, it refers to the pair of terminal outcome collections generated by the two admissible current controls \(u_T^n=\pm u_T^{\max}\), across all candidate terminal indices and pursuer modes. This is one of the defining semantic features of TSE.

The expected terminal cost for current action \(u_T^n\) is
\[
J(u_T^n)
=
\sum_i p_f(i)\sum_{j=1}^{p^{\max}} P_j^n\,
\mathbb E\!\left[\|\mathbf{C}\mathbf{x}_i^f\|^2\,\middle|\, f=i,p_j\right].
\]
Using the affine decomposition,
\[
\mathbb E\!\left[\|\mathbf{C}(\mathbf{a}^n(i,j)u_T^n+\mathbf{z}_{i,j}^n)\|^2\right]
=
\|\mathbf{C}(\mathbf{a}^n(i,j)u_T^n+\boldsymbol{\mu}_{i,j}^n)\|^2
+
\operatorname{tr}(\mathbf{C}\boldsymbol{\Sigma}_{i,j}^n\mathbf{C}^\top).
\]
Because the covariance term is independent of the current sign, comparing the two admissible current bang commands reduces to comparing the mean terms. The paper therefore defines
\[
S_+^n
=
\sum_i p_f(i)\sum_{j=1}^{p^{\max}} P_j^n
\,
\|\mathbf{C}(\mathbf{a}^n(i,j)u_T^{\max}+\boldsymbol{\mu}_{i,j}^n)\|^2,
\]
\[
S_-^n
=
\sum_i p_f(i)\sum_{j=1}^{p^{\max}} P_j^n
\,
\|\mathbf{C}(-\mathbf{a}^n(i,j)u_T^{\max}+\boldsymbol{\mu}_{i,j}^n)\|^2.
\]
Letting \(\tilde{\mathbf a}_{i,j}^n=\mathbf C\mathbf a^n(i,j)\) and \(\tilde{\boldsymbol\mu}_{i,j}^n=\mathbf C\boldsymbol\mu_{i,j}^n\), the difference satisfies
\[
S_+^n-S_-^n
=
4u_T^{\max}
\sum_i p_f(i)\sum_{j=1}^{p^{\max}} P_j^n
\langle \tilde{\mathbf a}_{i,j}^n,\tilde{\boldsymbol\mu}_{i,j}^n\rangle.
\]
This motivates the shaping scalar
\[
S_n
=
\sum_i p_f(i)\sum_{j=1}^{p^{\max}} P_j^n
\langle \tilde{\mathbf a}_{i,j}^n,\tilde{\boldsymbol\mu}_{i,j}^n\rangle,
\]
and the TSE law
\[
\overset{\star}{u}_T^n=u_T^{\max}\,\operatorname{sign}(S_n).
\]

The controller is closed-loop because \(S_n\) is recomputed at each decision time from the current posterior moments, current guidance-mode probabilities, current terminal-time PMF, and current mode-conditioned transition gains. The paper specifies the sign rule exactly: if \(S_n>0\), choose \(+u_T^{\max}\); if \(S_n<0\), choose \(-u_T^{\max}\); if \(S_n=0\), the formula is indifferent and no tie-breaking rule is specified. The paper does not provide pseudocode, but it reconstructs the implementation as posterior update, mode/time-indexed terminal prediction, scalar shaping evaluation, sign selection, and repetition.

## 5. Numerical behavior and comparative performance

The numerical study specializes to a planar lateral engagement with a single guidance mode and simplified lateral state
\[
\mathbf{x}^k=\begin{bmatrix}\xi^k & \dot{\xi}^k\end{bmatrix}^{\!\top},
\qquad
\mathbf{C}=\begin{bmatrix}1&0\end{bmatrix},
\]
with
\[
\mathbf{x}^{k+1}=\mathbf{F}\mathbf{x}^k+\mathbf{g}_T u_T^k+\mathbf{g}_M u_M^k,
\quad
\mathbf{F}=
\begin{bmatrix}
1&\Delta t\\
0&1
\end{bmatrix},
\quad
\mathbf{g}_T=\mathbf{g}_M=
\begin{bmatrix}
(\Delta t)^2/2\\
\Delta t
\end{bmatrix}
\]
[2511.21633]. The reported parameters are \(\Delta t=0.01~\mathrm{s}\), \(V_c=400~\mathrm{m/s}\), \(u_T^{\max}=9g\), \(u_M^{\max}=27g\), PN guidance with \(N=3\), and terminal step uncertainty
\[
f\sim \mathrm{U}\{295,\dots,305\},\qquad N_f=11.
\]

To illustrate extremality, future evader commands \(u_T^k\), \(k>n\), are modeled as i.i.d. uniform on \([-u_T^{\max},u_T^{\max}]\), so
\[
\operatorname{Var}(u_T^k)=\frac{(u_T^{\max})^2}{3}.
\]
Under this model, the resulting \(J(u_T^n)\) is quadratic and strictly convex in the current control, with maxima at the bounds \(\pm u_T^{\max}\). The paper reports that the corresponding figure confirms bang-bang extremality numerically.

The Monte Carlo study uses \(N_{\mathrm{MC}}=10{,}000\) trials, initial state \(x^0\sim\mathcal N(0,P^0)\) with \(P^0=\operatorname{diag}(100,4)\), measurement model \(y^k=\mathbf C\mathbf x^k+\nu^k\), LOS-angle jitter \(\sigma_\lambda=5~\mathrm{mrad}\), estimator noise model
\[
R_k = (\sigma_\lambda V_c t_{go}^{k,\mathrm{nom}})^2,
\]
and process noise covariance
\[
Q_k=(u_T^{\max})^2\mathbf g_T\mathbf g_T^\top.
\]
Both sides use KFs, and the pursuer receives an information advantage through
\[
P_M^0=\beta P_T^0,\qquad \beta=0.25.
\]

TSE is compared against random telegraph switching, Singer acceleration, and weaving. In one representative engagement, the reported miss distances are \(6.29~\mathrm{m}\) for TSE, \(5.04~\mathrm{m}\) for RTS, \(1.93~\mathrm{m}\) for Singer, and \(0.63~\mathrm{m}\) for weaving. The paper attributes the difference mainly to switch timing: TSE uses strategically timed bang-bang reversals rather than random or periodic switching.

The Monte Carlo CDF of \(|\xi(t^f)|\) shows that TSE and RTS outperform the smoother stochastic maneuvers, and TSE dominates RTS across nearly the whole distribution. For a warhead lethality radius of \(1~\mathrm{m}\), the approximate single-shot kill probabilities are reported as \(\approx 0.9\) for Singer, \(\approx 0.9\) for weaving, \(\approx 0.4\) for RTS, and \(\approx 0.2\) for TSE. The reported summary statistics for \(|\xi(t^f)|\) are also consistent with that ordering: TSE has mean \(2.55\), median \(2.4\), \(P5=0.13\), \(P20=1.02\), \(P80=4.16\), and \(P95=5.98\), all in meters. RTS yields mean \(1.75\) and median \(1.26\); Singer yields mean \(0.4\) and median \(0.24\); weaving yields mean \(0.51\) and median \(0.42\). The paper therefore reports the largest mean and median miss distance for TSE, together with the broadest upper-tail evasion performance.

## 6. Interpretation, neighboring terminal-set frameworks, and limitations

TSE occupies a specific position within the broader terminal-set literature. Its terminal-set object is not a robustly invariant polytope or ellipsoid, and not a strict HJ reachability set. Instead, it is a mode- and terminal-time-indexed family of terminal affine images associated with the two admissible current bang commands [2511.21633]. This distinguishes it from terminal ingredients in robust predictive control, where the terminal object is usually an invariant set under a local feedback law. For example, "Robust Data-Driven Tube-Based Zonotopic Predictive Control with Closed-Loop Guarantees" uses terminal ingredients together with a tube and tightened constraints to prove recursive feasibility and robust exponential stability around a terminal set/tube [2409.14366]. Likewise, "Data-Driven Tube-Based Zonotopic Predictive Control With Nonconvex Layered Terminal Sets" separates a small contractive region, a larger nonconvex terminal region, and an outer screening region under a fixed feedback law [2605.12467]. This suggests a useful contrast: TSE makes the current sign decision by comparing analytically predicted terminal outcomes, whereas those predictive-control papers use terminal sets primarily to certify recursive feasibility, invariance, and stability.

The paper also differs from feedback synthesis work that starts from open-loop differential-game solutions and then learns a state-feedback implementation. "From open-loop representations to closed-loop feedback implementations in differential games: A numerical case study" emphasizes learning value, gradient, and feedback jointly because learning only the scalar value and differentiating it did not lead to satisfying results [2605.04768]. A plausible implication is that TSE and such learned-feedback approaches address complementary regimes: TSE exploits an explicit affine terminal-state decomposition and a closed-form sign test, while learned approaches become attractive when no comparable analytic reduction is available.

The stated limitations of TSE are substantial and precise. The derivation is for a planar linearized endgame, the pursuer is assumed to use a linear guidance law, the structural bang-bang theorem relies on a convex terminal cost, and the implemented TSE law optimizes only the current command while modeling future evader inputs stochastically rather than solving the full finite-dimensional bang-bang problem [2511.21633]. The implementation is moment-based, exploiting Gaussian or Gaussian-mixture posterior summaries, and the paper does not solve an integrated estimator-controller dual-control problem. It also does not extend the method to full \(3\)D nonlinear geometry.

These limitations clarify both the scope and the significance of TSE. Its theoretical contribution is the extension of bang-bang optimality to a stochastic imperfect-information setting under the generalized separation theorem. Its practical contribution is a low-complexity closed-loop sign-selection law that replaces exponential bang-sequence enumeration by a single analytically computable scalar test. Within those assumptions, TSE is a rigorous terminal-outcome-based evasion policy rather than a heuristic maneuver generator, and its empirical comparison with RTS, Singer, and weaving is consistent with that distinction.

Source: https://www.emergentmind.com/topics/closed-loop-terminal-set-based-evasion-tse