---
title: 'Most Likely Trajectory: Optimality Under Uncertainty'
url: https://www.emergentmind.com/topics/most-likely-trajectory
type: topic
---

# Most Likely Trajectory: Optimality Under Uncertainty

Searching arXiv for recent and directly relevant papers on “most likely trajectory” across domains.
“Most likely trajectory” is a domain-dependent term for a trajectory, path, or evolution selected by an explicit optimality principle under uncertainty. In the research literature, the phrase does not denote a single universal construction. Instead, it refers to distinct but structurally related objects: a minimum-relative-entropy path-space law in Schrödinger bridge theory for diffusing and vanishing particles [2108.02879]; a minimum-action transition path in small-noise large deviations for distribution-dependent stochastic systems [2111.06030]; a maximum-probability path on a Markov transition graph for ocean drifters [2002.07774]; a best-fit physically constrained trajectory under weighted least squares for meteoroid dynamics [1911.00816] and coronal ejecta tracking [2009.09323]; a deterministic nominal continuation inferred from online state estimation in interception and maneuver prediction [2010.02512], [2501.04273]; a highest-ranked trajectory proposal in multimodal forecasting systems [2307.14788], [2605.20388]; and, in monitored quantum dynamics, the dominant measurement record or noise-conditioned path obtained from a trajectory probability functional [2509.24520], [2209.13164]. Across these settings, the common idea is selection of one representative evolution from many admissible possibilities, but the selection criterion varies sharply with the underlying probabilistic, variational, or physical model.

## 1. Relative-entropy formulations on path space

In one major line of work, “most likely trajectory” is defined in the original Schrödinger sense: among all path-space laws consistent with observed endpoint marginals, choose the one that is closest to a prior stochastic law in relative entropy [2108.02879]. For a classical Schrödinger bridge with diffusion prior
\[
dX_t = b(t,X_t)dt + \sigma(t,X_t)dW_t,
\]
with diffusion matrix
\[
a(t,x)=\sigma(t,x)\sigma(t,x)',
\]
the bridge solves
\[
P^\star := {\rm arg}\min_{P\in \mathcal P(\Omega)} \left\{H(P\mid R)\mid P_0=\rho_0,\;P_1=\rho_1\right\}.
\]
Here the “most likely” evolution is the minimum-relative-entropy path law relative to the prior law \(R\) [2108.02879].

The unbalanced extension for diffusing and vanishing particles generalizes this definition to priors with killing. The prior Fokker–Planck equation becomes
\[
\partial_t R_t+ \nabla\cdot(b R_t) + V R_t = \frac{1}{2}\sum_{i,j=1}^n \frac{\partial^2(a_{ij}R_t)}{\partial x_i\partial x_j},
\]
or equivalently
\[
\partial_t R_t = -\nabla\cdot(bR_t) + \frac12 \sum_{i,j}\partial_{x_i x_j}^2(a_{ij}R_t) - VR_t.
\]
Because killing destroys mass, the endpoint marginals are unbalanced. The construction therefore augments the state space with a coffin state
\[
\mathcal X = \mathbb R^n\cup\{\mathfrak c\},
\]
and works on path space
\[
\mathbf\Omega = D([0,1],\mathcal X).
\]
This converts a lossy transport problem into a probability-constrained path-space problem on an augmented space [2108.02879].

The resulting generalized Schrödinger bridge problem is
\[
\mu^\star := {\rm arg}\min_{\mu\in \mathcal P(\mathbf \Omega)} \left\{H(\mu\mid R)~\mid~ \mu_0 = p_0, \mu_1 = p_1\right\},
\]
with augmented marginals
\[
p_0=(\rho_0(\cdot),0), \qquad
p_1=\left(\rho_1(\cdot),\,1-\int_{\mathbb R^n}\rho_1(x)\,dx\right).
\]
In this formulation, the most likely evolution describes not only transport of surviving particles but also when and where particles most likely vanish [2108.02879].

The optimizer has the multiplicative endpoint form
\[
\mu^\star = f(\mathbf X_0)g(\mathbf X_1)R,
\qquad
\frac{d\mu^\star}{dR} = f(\mathbf X_0)g(\mathbf X_1),
\]
and is justified by a large-deviation principle:
\[
\exp(-N H(\mu\mid R)).
\]
This gives a direct probabilistic meaning to “most likely”: among empirical path laws compatible with the observed marginals, the entropy minimizer is exponentially most probable [2108.02879].

## 2. Variational trajectories from large deviations and action minimization

A second major usage appears in Freidlin–Wentzell theory, where the most likely trajectory is the minimum-action path of a rare transition. For the McKean–Vlasov stochastic system
\[
d X_{t}^{\epsilon}=V\left(X_{t}^{\epsilon}\right) d t- F*u_t^\epsilon(X_t^\epsilon)d t+\sqrt{\epsilon}\, d B_t,
\]
with \(u_t^\epsilon=\mathcal L(X_t^\epsilon)\), the most likely transition path from \(x_1\) to \(x_2\) over fixed time \(T\) is defined as the minimizer of the action
\[
I_T(\varphi^*) =  \min_{\varphi} I_T(\varphi),
\]
subject to
\[
\varphi(0)=x_1,\qquad \varphi(T)=x_2,\qquad \varphi\in C([0,T]).
\]
The associated quasipotential is
\[
\tilde{V}(x_1,x_2) = \inf_{T>0}\inf_{\varphi \in \bar{C}_{x_1}^{x_2}([0,T])} I_T(\varphi)
\]
[2111.06030].

For the class treated there, the action functional is
\[
I^{x_1}_T(\varphi)=\frac{1}{2} \int_0^T \left|\dot{\varphi}(s)-V(\varphi(s))+F(\varphi(s)-\eta(s))\right|^2 ds,
\]
where the skeleton equation is
\[
\dot{\eta}(t)=V(\eta(t))-F(0), \qquad \eta(0)=x_1.
\]
A key reduction occurs when \(F(0)=0\) and \(x_1\) is an equilibrium of \(V\). Then \(\eta(t)\equiv x_1\), and the action reduces to that of an ordinary SDE with modified drift
\[
\bar b(x)=V(x)-F(x-x_1).
\]
In that case, the distribution-dependent problem can be computed through the non-distribution-dependent SDE
\[
dZ^\epsilon_t=\left(V(Z^\epsilon_t)-F(Z^\epsilon_t-x_1)\right)dt+\sqrt{\epsilon}dB_t
\]
[2111.06030].

This notion of “most likely” is explicitly distinct from deterministic drift following. The deterministic reduced system
\[
\dot{\xi}(t)=V(\xi(t))-F(\xi(t)-x_1)
\]
has zero action along its own orbits, but the most likely transition path is the path minimizing the action among rare trajectories that connect metastable states against the drift [2111.06030]. The paper derives the corresponding Euler–Lagrange equation
\[
\left\{ \begin{array}{l}
\varphi_{tt}-\nabla_x b(\varphi,\eta) \varphi_t+\big(\nabla_x b(\varphi,\eta)\big)^\top (\varphi_t-b(\varphi,\eta))-\nabla_y b(\varphi,\eta) b(\eta,\eta)=0, \\
\varphi(0)=x_1,\quad \varphi(T)=x_2,
\end{array}\right.
\]
and computes solutions with the adaptive minimum action method [2111.06030].

A related but formally different variational use appears in unscented trajectory optimization, where the paper does not define a statistical mode path and instead optimizes what it calls the “most desirable tychastic trajectory” under uncertainty. The uncertain dynamics are
\[
\dot x = f(x,u,t;\xi),
\]
and the optimization may penalize endpoint covariance
\[
\operatorname{tr}\,\operatorname{Cov}[x(t_f,\xi)]
\]
or chance-constraint risk. This suggests a different but nearby interpretation: in some control settings, the representative trajectory is not the most probable one but the one that is most desirable under a chosen risk-sensitive criterion [2405.02753].

## 3. Dynamic laws, posterior corrections, and control interpretations

In the unbalanced Schrödinger bridge with killing, the optimal law remains Markovian and changes both drift and killing rate [2108.02879]. The generalized Schrödinger system is
\[
\partial_t \hat\varphi = - \nabla\cdot(b \hat\varphi)-V\hat\varphi +\frac{1}{2} \sum_{i,j=1}^n \frac{\partial^2 (a_{ij} \hat \varphi)}{\partial x_i\partial x_j},
\]
\[
\frac{d\hat\psi}{dt} = \int V\hat\varphi(t,x)\,dx,
\]
\[
\partial_t\varphi = - b \cdot\nabla\varphi+V\varphi - \frac{1}{2} \sum_{i,j=1}^n a_{ij}\frac{\partial^2 \varphi}{\partial x_i\partial x_j} - V\psi,
\]
\[
\frac{d\psi}{dt}=0,
\]
with boundary couplings
\[
\rho_0 = \varphi(0,\cdot)\hat\varphi(0,\cdot), \qquad
\rho_1 = \varphi(1,\cdot)\hat\varphi(1,\cdot),
\]
\[
\hat\psi(0)=0, \qquad
\psi(1)\hat\psi(1)=1-\int \rho_1.
\]
The one-time marginals are
\[
P_t(x)=\varphi(t,x)\hat\varphi(t,x), \qquad q_t=\psi(t)\hat\psi(t).
\]
The optimal diffusion law is
\[
dX_t = (b(t,X_t) +a(t,X_t)\nabla\log\varphi(t,X_t))dt + \sigma(t,X_t) dW_t
\]
with posterior killing rate
\[
\frac{\psi(t)}{\varphi(t,x)}\,V(t,x).
\]
Thus the most likely trajectory is not merely conditioned survival; it is a full posterior law over both surviving and vanished particles, with an inferred state-dependent killing mechanism [2108.02879].

The same object admits a fluid-dynamic and control-theoretic characterization:
\[
\min_{P, u,\alpha} \int_0^1 \int_{\mathbb R^n} \left[ \frac{1}{2} \|u(t,x)\|^2 P_t +(\alpha \log \alpha -\alpha +1)VP_t \right]dx\,dt
\]
subject to
\[
\partial_t  P_t+ \nabla\cdot((b+\sigma u) P_t) +\alpha VP_t - \frac{1}{2} \sum_{i,j=1}^n \frac{\partial^2 (a_{ij} P_t)}{\partial x_i\partial x_j}=0,
\]
\[
P_0 = \rho_0,\quad P_1=\rho_1.
\]
The stationarity conditions give
\[
u^\star(t,x)=\sigma(t,x)'\nabla\log\varphi(t,x), \qquad
\alpha^\star(t,x)=\frac{\psi(t)}{\varphi(t,x)}.
\]
This suggests a general pattern: in several fields, “most likely trajectory” can be equivalently described either as a probabilistic posterior object or as a solution to a deterministic control problem with an action or entropy cost [2108.02879].

A different control-centered use appears in free-flight aircraft convergence. There the trajectory judged most appropriate is the minimum-time, bounded-turn, safe-separation convergence path toward a leader’s route. The Hamiltonian is affine in the turn-rate control, producing bang-bang structure and regular turn-straight-turn trajectories, later approximated online by a neural network [1302.4858]. That work does not use a probabilistic definition, but it shares the same structural idea: a trajectory is selected by an explicit optimality principle under constraints.

## 4. Data-driven maximum-probability paths and best-fit trajectories

In data-driven ocean transport, the “most likely path” is defined directly as a maximum-probability path on a Markov chain inferred from drifter data [2002.07774]. Continuous positions are discretized by
\[
f : \mathbb{R}^2 \rightarrow \mathcal{S},
\]
and an empirical transition matrix is estimated:
\[
T_{s,p} =\dfrac{\sum_{j=1}^N\sum_{i=1}^{n_{j}-4\mathcal{T}_L} \mathbb{I}[g_{i+4\mathcal{T}_L,j}=p]\mathbb{I}(g_{i,j}=s)}{ \sum_{j=1}^N\sum_{i=1}^{n_{j}-4\mathcal{T}_L} \mathbb{I}[g_{i,j}=s]}.
\]
For a path \(\mathbf{p}=(p_0,\dots,p_n)\),
\[
P(\mathbf{p}) = \prod_{i=0}^{n-1} T_{p_i,p_{i+1}}.
\]
The most likely path is
\[
\hat{\mathbf{p}} = \argmax_{\mathbf{p}\in \mathcal{P}_{o,d}} \left\{ \prod_{i=0}^{n-1}T_{p_i,p_{i+1}} \right\}.
\]
Taking logarithms gives a shortest-path problem with weights
\[
w(e_{i,j})=-\log T_{i,j},
\]
so the maximum-probability path is exactly a shortest path in the transformed graph [2002.07774]. This notion differs sharply from minimum-time or mean-path descriptions: the path can be circuitous in Euclidean space if dominant currents make that route more probable [2002.07774].

The same literature also estimates an expected travel time conditional on following the inferred path. For each step \(p_i\to p_{i+1}\),
\[
\mathbb{E}[k_i] = \dfrac{T_{p_i,p_i}}{T_{p_i,p_{i+1}}}+1,
\]
and
\[
\mathbb{E}[\mathbf{k}] = \sum_{i=0}^{n-1}\mathbb{E}[k_i].
\]
This shows that “most likely path” and “fastest path” are not the same object [2002.07774].

A conceptually different data-driven maximum-selection procedure appears in vessel trajectory reconstruction from AIS data [2007.13712]. There the next point of a vessel is chosen as the “best possible next point” using a local 3D feature space of time difference, scaled forward/backward prediction error, and space-time angle. For moving pairs, the score is
\[
d_{ij}=\frac{1}{2}(d^+ + d^-),
\]
with angle screening by
\[
\theta_{ij}\le \cos^{-1}(0.1)\approx84.26^\circ.
\]
Points are then linked by their selected next points to reconstruct complete trajectories [2007.13712]. A plausible implication is that in reconstruction problems without explicit global probabilities, “most likely trajectory” may be operationalized through repeated local best-match decisions.

In observational astronomy and heliophysics, the term often collapses into a best-fit trajectory under a physical model. For fireballs, the Dynamic Trajectory Fit treats the most plausible trajectory as the bounded weighted least-squares optimum of a 3D physical flight model fitted directly to multi-sensor lines of sight [1911.00816]. The state includes position, velocity, ballistic coefficient, ablation coefficient, fragmentation parameters, timing offsets, and missing times, and the fit minimizes weighted along-track and cross-track angular residuals [1911.00816]. For coronal ejecta observed by WISPR, the preferred 3D trajectory is the radial constant-speed parameter set
\[
(r_{20},V,\phi_2,\delta_2)
\]
that best matches the measured angle tracks \(\gamma(t),\beta(t)\) under weighted nonlinear least squares [2009.09323]. In both cases, “most likely” is effectively “best-fitting physically admissible trajectory,” not a path-space entropy or action minimizer.

## 5. Machine learning, ranking, and nominal extrapolation

In multimodal trajectory forecasting, “most likely trajectory” often denotes the highest-ranked candidate among generated proposals rather than a closed-form mode of a density. In the context-free clustering-based forecaster of [2307.14788], the pipeline is: displacement-space transformation, clustering, cluster-conditioned proposal generation, and proposal ranking. One trajectory proposal is generated per cluster:
\[
\{\hat{\boldsymbol{dy}_i}\}_{i=1}^K.
\]
Probabilities are assigned post hoc by
\[
\hat{p}_{c_i} = \frac{\exp(\frac{1}{m_{c_i} / \tau})}{\sum_{j}^K \exp(\frac{1}{m_{c_j} / \tau})},
\]
and the implied top-ranked trajectory is
\[
i^* = \arg\max_i \hat{p}_{c_i}, \qquad \hat{\boldsymbol{y}}_{ML} = \hat{\boldsymbol{y}}_{i^*}.
\]
The paper is explicit that this top-ranked generated proposal is its operational “most likely trajectory” [2307.14788].

A recent egocentric prediction system uses a closely related top-1 selection idea, but with future camera motion as the latent intent variable [2605.20388]. The trajectory representation is a “16-knot relative 6-DoF trajectory”
\[
u \in \mathbb{R}^{16 \times 6},
\]
with each knot storing
\[
[\Delta x,\Delta y,\Delta z,\Delta r_x,\Delta r_y,\Delta r_z].
\]
At test time, the system retrieves a bank of plausible future trajectories, evaluates them with a gate-then-rank scorer, and either falls back to a no-trajectory predictor or uses the top-ranked candidate. In that setup, the most likely trajectory is the highest-ranked retrieved trajectory candidate conditioned on the current egocentric context [2605.20388].

By contrast, some forecasting papers use “most likely trajectory” only in a deterministic nominal sense. In high-speed interception, an extended Kalman filter estimates current position from visual observations, then a fitted local circular-motion model is rolled forward recursively. The resulting future path is best interpreted as the nominal continuation implied by the current filtered state and motion model, not as a trajectory distribution [2010.02512]. Similarly, AISE/FS trajectory prediction extrapolates a local Frenet–Serret geometry with estimated speed, curvature, and torsion:
\[
\hat{R}_{k+l} = R_k [ \Gamma_0(\omega_k T_s) ]^l,
\]
\[
\hat{p}_{k+l} = p_k + T_s   \bigg[R_k + \sum_{i=1}^{l-1} \hat{R}_{k+i}\bigg]\Gamma_1(\omega_k T_s) \begin{bmatrix} u_{k} & 0 & 0 \end{bmatrix}^{T}.
\]
That work explicitly does not define a probabilistic “most likely trajectory”; it produces a single deterministic short-horizon continuation [2501.04273].

A probabilistic but still geometric variant appears in lateral aircraft ground-track prediction. There the future path is parameterized by piecewise-linear control points
\[
P=[\boldsymbol{p}_0,\boldsymbol{p}_1,\dots,\boldsymbol{p}_d],
\]
with normalized times
\[
t_i=t_{i-1}+\frac{\mathcal{D}(\boldsymbol{p}_i,\boldsymbol{p}_{i-1})}{\sum_{i=1}^d \mathcal{D}(\boldsymbol{p}_i,\boldsymbol{p}_{i-1})}
\]
and interpolation
\[
\boldsymbol{x}(t)=\boldsymbol{x}(t_i)+\frac{t-t_i}{t_{i+1}-t_i}(\boldsymbol{p}_{i+1}-\boldsymbol{p}_i).
\]
The learned object is a context-conditioned distribution over control points \(\boldsymbol{y}\in\mathbb{R}^{2d}\). The paper does not define a formal most-likely decoder, but a plausible reading is that it would correspond to the most probable or central control-point configuration under the predictive model [2309.14957]. This suggests a general caveat: in modern generative forecasting, “most likely trajectory” may be a downstream decoding choice, not an intrinsic object of the model.

## 6. Quantum and scientific uses of dominant trajectories

In monitored quantum dynamics, “most likely trajectory” acquires a literal path-probability meaning. For continuously monitored bosonic systems, the single-step readout probability under Gaussian weak measurement is
\[
P(r|\psi)=\left(\frac{2\gamma\delta t}{\pi}\right)^{1/2}\int dR\, e^{-2\gamma\delta t (r-R)^2}|\psi(R)|^2.
\]
Approximating \( |\psi(R)|^2 \) by a sharply peaked distribution yields
\[
r^*=\langle \hat R\rangle_\psi,
\]
and, in continuous time,
\[
r^*(t)=\langle \hat R\rangle_{\psi^*(t)}.
\]
Equivalently, in Wiener-process form, the most likely trajectory is the one with
\[
dW^*=0 \qquad \forall t.
\]
The resulting conditioned density matrix evolves by a deterministic nonlinear equation rather than a stochastic ensemble average [2509.24520].

The same paper defines the global trajectory probability functional
\[
P[r|\psi_0;t]=\mathrm{Tr}\bigl(|\tilde\psi_r\rangle\langle \tilde\psi_r|\bigr)
\]
and a fictitious action
\[
\mathcal S[\mathbf r;t]=\log P[\mathbf r;t].
\]
The most likely trajectory is then the saddle point satisfying
\[
\frac{\delta \mathcal S[\mathbf r;t]}{\delta r_j(\tau)}\bigg|_{\mathbf r=\mathbf r^*}=0.
\]
For monitored free bosons, this construction is exact; for the interacting Sine-Gordon model, it is combined with a self-consistent time-dependent harmonic approximation and reveals an entanglement phase transition from area-law to logarithmic-law scaling [2509.24520]. This is a case where “most likely trajectory” is neither an action-minimizing rare-event path nor a best-fit curve, but the dominant measurement record in a monitored quantum ensemble.

A related quantum-control use appears in state preparation under dephasing noise [2209.13164]. There a stochastic path integral is built over noise histories, and the most-likely path is the state trajectory and associated noise realization that maximize path likelihood subject to fixed initial and final states. In the qubit proof of concept, the extremal equations imply a most-likely noise path \(\xi_t^*=0\), yielding analytic Rabi controls for arbitrary target states [2209.13164]. This sharply contrasts with mean-path Lindblad control, which optimizes average fidelity over all noise realizations.

Outside mathematical trajectory inference, the phrase can also mean an inferred physical evolutionary path. In GW170817, the “most likely trajectory” refers to the post-merger evolution of the compact remnant, with the authors concluding that the system most likely formed a black hole well before day 107 rather than remaining a long-lived neutron star [1712.03240]. A plausible implication is that, in astrophysical discourse, the term may designate the most plausible physical scenario consistent with multi-epoch observations rather than a formal mathematical path.

## 7. Misconceptions, distinctions, and cross-domain synthesis

A recurring misconception is that “most likely trajectory” always means the deterministic trajectory of the noiseless system. The large-deviation formulation for McKean–Vlasov systems explicitly rejects this: the most likely transition path is the minimum-action rare path, not the zero-noise orbit [2111.06030]. In Schrödinger bridge theory, the object is a posterior path law on path space, not a single sample trajectory [2108.02879]. In monitored quantum systems, it is the dominant readout history, not the measurement-averaged state [2509.24520]. In machine learning forecasting, it may be merely the top-ranked proposal among finitely many generated futures [2307.14788], [2605.20388].

Another common confusion is between “most likely,” “mean,” and “best fit.” The drifter path literature explicitly distinguishes the most likely path from the fastest path and from geometric shortest paths [2002.07774]. Fireball and coronal ejecta studies identify a best-fit trajectory under measurement models, which is likelihood-based only through observational residuals and model assumptions [1911.00816], [2009.09323]. Interception and local motion extrapolation papers output a single deterministic continuation without a full posterior over futures [2010.02512], [2501.04273].

The literature also shows that pathwise likelihood can live at different levels of abstraction:

| Setting | Object selected as “most likely” | Selection principle |
|---|---|---|
| Schrödinger bridge with killing | Path-space law | Relative entropy minimization [2108.02879] |
| Small-noise transition theory | Continuous path | Action minimization [2111.06030] |
| Ocean drifters | Discrete state path | Maximum Markov path probability [2002.07774] |
| Proposal-based forecasting | Generated future proposal | Highest ranking score/probability [2307.14788], [2605.20388] |
| Monitored quantum dynamics | Measurement record / conditioned state path | Saddle point of trajectory probability functional [2509.24520] |
| Best-fit observational dynamics | Physical trajectory parameters | Weighted least squares under a physical model [1911.00816], [2009.09323] |

This suggests that “most likely trajectory” is best treated as an umbrella term whose meaning is fixed by the measure on trajectories, the admissible class, and the optimization or inference rule. Where the underlying object is a path-space law, the phrase refers to a posterior distribution. Where the underlying object is a rare event, it refers to an action minimizer. Where the underlying object is a discrete candidate set, it refers to a highest-ranked proposal. Where the model is observational and deterministic, it may simply mean the trajectory that best matches data under a chosen physical parameterization.

The broad research trajectory of the concept therefore runs from classical large deviations and entropy projection [2108.02879], [2111.06030], through data-derived graph and filtering formulations [2002.07774], [2010.02512], to modern proposal-ranking and latent-intent systems [2307.14788], [2605.20388], and to trajectory-probability methods in monitored quantum many-body dynamics [2509.24520]. This suggests a unifying editor’s term, “selection principle,” for the rule that converts a set or distribution of possible evolutions into one designated trajectory. The literature does not support a single universal definition beyond that abstraction.

Source: https://www.emergentmind.com/topics/most-likely-trajectory