---
title: Continuous-Time Maximum Likelihood Estimation
url: https://www.emergentmind.com/topics/continuous-time-maximum-likelihood-estimator
type: topic
---

# Continuous-Time Maximum Likelihood Estimation

Continuous-time maximum likelihood estimation denotes a family of likelihood-based procedures in which the statistical model is specified in continuous time and the estimator is obtained by maximizing an exact, approximate, or quasi-likelihood derived from that model. In the cited literature, this includes exact pathwise likelihoods under continuous observation, likelihoods built from finite-time transition laws for discretely or irregularly sampled processes, hidden-Markov reformulations of continuous-time state-space models, and transformed likelihoods for non-Markovian or function-valued dynamics [1403.2954], [2010.14883], [2108.12649], [2104.04636]. The common objective is to infer structural parameters—drift, diffusion, generator, mean-function, or interaction parameters—from observations that are indexed by continuous time even when the data themselves are sampled discretely, irregularly, or through a latent layer.

## 1. Observation schemes and the scope of the estimator

The literature treats several observation regimes as instances of continuous-time likelihood inference. At one extreme is full path observation on a time interval \([0,T]\), as in continuously observed diffusion, jump-diffusion, threshold diffusion, Ornstein–Uhlenbeck, CIR, Volterra Ornstein–Uhlenbeck, and mixed fractional Vasicek models [1403.2954], [1711.02140], [1803.05408], [2404.05554], [2003.13351]. At another extreme are continuous-time models observed only at irregular or low-frequency time points, such as \(0=t_0<t_1<\cdots<t_T\) for general state-space models or \(Y_k^{(h)}=Y(kh)\) for cointegrated continuous-time state-space models [2010.14883], [1712.08665]. Between these cases lie partially observed diffusions, continuous-time models with latent states, and continuous-time processes approximated by continuous-time Markov chains or hidden Markov models [1611.00170], [2108.12649], [1504.01804].

This breadth matters because “continuous-time MLE” does not refer to a single formula. In some models the likelihood is a Radon–Nikodym derivative on path space. In others it is a product of transition probabilities \(p_\Delta(x_t\mid x_s)\), a matrix-product HMM likelihood, or a pseudo-Gaussian contrast built from pseudo-innovations. The cited papers therefore use the same inferential principle—maximize a likelihood-type criterion—under markedly different stochastic structures [2010.14883], [1712.08665], [2105.04390].

A recurring distinction is between models whose continuous-time transition law is available in closed form and those for which it is not. When \(p_\Delta(x_t\mid x_s)\) is explicit, irregular sampling can be handled directly through the elapsed-time gaps \(\Delta_\tau=t_\tau-t_{\tau-1}\), as in the general continuous-time state-space framework and in several diffusion examples [2010.14883]. When the transition density is unavailable, the literature replaces it by state-space discretisation, CTMC approximation, functional Fokker–Planck transition densities, or quasi-likelihood constructions [2108.12649], [2104.04636], [1712.08665].

## 2. Exact continuous-observation likelihoods

For continuously observed semimartingale models, the central construction is a likelihood ratio between path measures, typically derived by Girsanov-type arguments. In the Lévy-driven Ornstein–Uhlenbeck model
\[
dX_t=-aX_t\,dt+dL_t,
\]
the likelihood relative to \(P^0\) on \(\mathcal F_T\) is
\[
\frac{dP_T^a}{dP_T^0} = \exp\!\left( -\frac{a}{\sigma^2}\int_0^T X_s\,dX_s^c -\frac{a^2}{2\sigma^2}\int_0^T X_s^2\,ds \right),
\]
which yields the explicit continuous-time MLE
\[
\hat a_T = -\frac{\int_0^T X_s\,dX_s^c}{\int_0^T X_s^2\,ds}.
\]
The estimator is strongly consistent, asymptotically normal, and asymptotically efficient in the Hájek–Le Cam sense when the stated conditions hold [1403.2954].

A similar linear-quadratic likelihood appears in continuously observed CIR-type models. For the \(\alpha\)-stable CIR process
\[
dY_t=(a-bY_t)\,dt+\sigma\sqrt{Y_t}\,dW_t+\delta\,Y_{t-}^{1/\alpha}\,dL_t,
\]
the paper derives the explicit MLE
\[
\widehat b_T = \frac{Y_T-y_0-aT-\delta\int_0^T Y_{u-}^{1/\alpha}\,dL_u}{\int_0^T Y_u\,du},
\]
together with the identity
\[
\widehat b_T-b = -\frac{\sigma\int_0^T \sqrt{Y_u}\,dW_u}{\int_0^T Y_u\,du}.
\]
This identity becomes the basis for regime-dependent asymptotics in the subcritical, critical, and supercritical cases [1711.02140]. The jump-type CIR model driven by a subordinator has the analogous explicit estimator
\[
\widehat b_T = \frac{Y_T-y_0-aT-J_T}{\int_0^T Y_s\,ds},
\]
again with a martingale-ratio representation for the estimation error [1609.05865].

Threshold diffusions provide a distinct exact-likelihood example in which discontinuity at zero enters through occupation times and local time. For drifted Oscillating Brownian motion with piecewise constant drift \(b_\pm\) and volatility \(\sigma_\pm\), the continuous-time likelihood ratio under the driftless model is
\[
G_T(b_+,b_-) = \exp\!\left( \frac{b_+}{\sigma_+^2}R_T^+ +\frac{b_-}{\sigma_-^2}R_T^- -\frac{b_+^2}{\sigma_+^2}Q_T^+ -\frac{b_-^2}{\sigma_-^2}Q_T^- \right),
\]
and the MLEs are
\[
\beta_T^+ = \frac{R_T^+}{Q_T^+},\qquad \beta_T^- = \frac{R_T^-}{Q_T^-}.
\]
Using Itô–Tanaka, these can be written in terms of the terminal value and the symmetric local time \(L_T(\xi)\), which makes the threshold contribution explicit [1803.05408].

Exact continuous-time likelihoods also occur in higher-order linear systems. For the second-order Gaussian autoregression
\[
\ddot X(t)=\theta_1 \dot X(t)+\theta_2 X(t)+\dot W(t),
\]
the MLE is obtained from the continuous-time Gaussian likelihood and can be written as a matrix inverse involving
\[
S_{11}(T)=\int_0^T X(t)^2\,dt,\quad
S_{12}(T)=\int_0^T X(t)\dot X(t)\,dt,\quad
S_{22}(T)=\int_0^T \dot X(t)^2\,dt.
\]
The paper proves strong consistency for every \((\theta_1,\theta_2)\in\mathbb R^2\) and a full regime classification of asymptotic behavior based on the roots of \(\lambda^2-\theta_1\lambda-\theta_2=0\) [1206.1379].

## 3. Approximate and quasi-likelihood formulations for sampled data

When the underlying model is continuous time but the data are sampled discretely or irregularly, the likelihood is often unavailable in closed form. One response is numerical integration over the latent state path. In the general continuous-time state-space model with irregular observations \(Y_{t_1},\ldots,Y_{t_T}\), the exact likelihood is a multiple integral over latent states:
\[
\mathcal{L}_T = \int \cdots \int p(x_0)p(y_0\mid x_0) \prod_{\tau=1}^T p_{\Delta_\tau}(x_\tau\mid x_{\tau-1})\,p(y_\tau\mid x_\tau)\, dx_T\cdots dx_0.
\]
The paper approximates this by discretising the state space into bins \(B_i\), which turns the model into a continuous-time HMM with structured, time-inhomogeneous transitions and yields the matrix-product approximation
\[
\mathcal{L}_T \approx \boldsymbol{\delta}\,\mathbf{P}(y_0) \Bigl(\prod_{\tau=1}^T \boldsymbol{\Gamma}_{\Delta_\tau}\mathbf{P}(y_\tau)\Bigr)\mathbf{1}.
\]
The forward recursion reduces computation to order \(\mathcal{O}(T m^2)\) [2010.14883].

A different route is state-space discretisation without time discretisation. For a univariate diffusion
\[
dS_t=\mu(S_t,\theta)\,dt+\sigma(S_t,\theta)\,dW_t,
\]
the CTMC approximation constructs a finite-state generator \(\mathbf Q(\theta)\) on a spatial grid and uses
\[
\mathbf T(\Delta)=\exp(\mathbf Q(\theta)\Delta)
\]
as the exact transition matrix of the approximating CTMC. With transition counts \(\mathbf C(\Delta)\), the log-likelihood becomes
\[
L_{N,m}(\theta,\Delta)=\sum_{i,j}\mathbf C(\Delta)_{ij}\ln \mathbf T(\Delta)_{ij}.
\]
The paper emphasizes that this introduces no time-discretization error during parameter estimation and proves \(\widehat{\theta}_{N,m}\to \widehat{\theta}_N\) as \(m\to\infty\) under its smoothness assumptions [2108.12649].

For finite-state continuous-time Markov jump processes observed every lag time \(\tau\), the likelihood may be written directly in terms of the generator \(\mathbf K\) through
\[
\mathbf T(\tau)=\exp(\mathbf K\tau),\qquad
P(x \mid \mathbf{K}, x_0)=\prod_{i,j}\mathbf{T}(\tau)_{ij}^{\mathbf{C}(\tau)_{ij}}.
\]
The resulting MLE estimates \(\mathbf K\) itself rather than a lag-dependent discrete-time transition matrix. The same paper develops a reversible parameterization based on a symmetric matrix \(\mathbf S\) and stationary distribution \(\pi\), which enforces detailed balance exactly during optimization [1504.01804].

Quasi-likelihood constructions arise when exact innovations are unavailable. In cointegrated continuous-time linear state-space models observed at low frequencies, the estimator is based on pseudo-innovations from the Kalman filter and the pseudo-Gaussian criterion
\[
\mathcal{L}_n^{(h)}(\vartheta) = \frac{1}{n}\sum_{k=1}^n \Big[ d\log(2\pi) + \log\det V^{(h)} + \varepsilon_k^{(h)}(\vartheta)^\mathsf{T} \big(V^{(h)}\big)^{-1} \varepsilon_k^{(h)}(\vartheta) \Big].
\]
In locally stationary continuous-time models, the analogous object is a kernel-localized \(M\)-contrast, and in the state-space case it becomes a localized quasi maximum likelihood criterion [1712.08665], [2105.04390].

## 4. Functional states, hidden structure, and transformed likelihoods

Several papers extend continuous-time MLE beyond finite-dimensional Markov diffusions by enlarging the state or transforming the process. In the continuous-time higher-order Markov framework, the recent history over \([t-T,t]\) is encoded by a state function \(H_t\), and the observed scalar process satisfies
\[
dy(t)=v(H_t)\,dt+s(H_t)\,dW.
\]
The associated functional state evolves as
\[
dH_t(x)=p(H_t(x))\,dt+o(H_t(x))\,dB(x).
\]
After deriving a Chapman–Kolmogorov relation over curves and a functional Fokker–Planck equation,
\[
\frac{\partial P(H_t\mid F_0)}{\partial t} = -\frac{\partial}{\partial H_t}\big[V(H_t)P(H_t\mid F_0)\big] +\frac{1}{2}\frac{\partial^2}{\partial H_t^2}\big[D(H_t)P(H_t\mid F_0)\big],
\]
the likelihood of an observed trajectory is formed as a product of segment-to-segment transition densities \(p(H_{t_{i+1}}\mid H_{t_i})\) and maximized over a finite-dimensional parameterization \(\Omega_m\) of the drift and diffusion [2104.04636].

For partially observed diffusions,
\[
dX_t = f(X_t,\theta)\,dt + g(X_t,\theta)\,dW_t,\qquad
dY_t = h(X_t,\theta)\,dt + dV_t,
\]
the incomplete-data log-likelihood is expressed on the observation filtration through the innovation process
\[
I_t = Y_t - \int_0^t \hat h_s(\theta)\,ds,
\qquad
\hat h_s(\theta)=\mathbb E_\theta[h(X_s,\theta)\mid\mathcal F_s^Y],
\]
leading to
\[
\mathcal{L}_t(\theta)=\int_0^t \hat h_s(\theta)\cdot dY_s -\frac12\int_0^t \|\hat h_s(\theta)\|^2\,ds.
\]
The gradient involves the tangent filter, and the paper studies a continuous-time stochastic gradient ascent recursion for online maximum likelihood [1611.00170].

Non-Markovian models are handled by transformation to a semimartingale. In the ergodic Volterra Ornstein–Uhlenbeck process,
\[
X_t = x_0 + \int_0^t K(t-s)\bigl(b+\beta X_s\bigr)\,ds + \sigma \int_0^t K(t-s)\,dB_s,
\]
the transformed process
\[
Z_t(X)=\int_{[0,t]}(X_{t-s}-x_0)\,L(ds)
\]
satisfies
\[
Z_t(X)=\int_0^t (b+\beta X_s)\,ds+\sigma B_t,
\]
which restores a Girsanov likelihood and produces explicit estimators for \(b\) and \(\beta\) [2404.05554]. The mixed fractional Vasicek model uses an analogous canonical representation. After defining
\[
Z_t=\int_0^t g(s,t)\,dX_s,\qquad
Q_t=\frac{d}{d\langle M\rangle_t}\int_0^t g(s,t)X_s\,ds,
\]
the transformed dynamics become
\[
dZ_t=(a-\beta Q_t)\,d\langle M\rangle_t+dM_t,
\]
and the likelihood takes the Gaussian-shift form needed for explicit MLEs of \(a\) and \(\beta\) [2003.13351].

For Gaussian processes with continuous observations, the paper on mean-function estimation identifies the class
\[
h=K\mu,\qquad \mu\in M(T),
\]
for which the likelihood can be written explicitly. The contrast is
\[
\Phi_\varepsilon(\theta)=\langle \mu_\theta,X^\varepsilon\rangle-\frac12\langle \mu_\theta,K\mu_\theta\rangle,
\]
and the parametric MLE is \(\hat\theta_\varepsilon\in\arg\max_{\theta\in\Theta}\Phi_\varepsilon(\theta)\) [2507.05628].

## 5. Asymptotic theory, rates, and efficiency

The asymptotic behavior of continuous-time MLEs is highly model-dependent. In ergodic settings the standard pattern is strong consistency and a Gaussian limit with \(\sqrt{T}\) normalization. This occurs for the continuous-time Lévy-driven Ornstein–Uhlenbeck MLE, for the ergodic Volterra Ornstein–Uhlenbeck estimator, and in the subcritical regimes of several CIR-type models [1403.2954], [2404.05554], [1711.02140]. In the Gaussian-process small-noise setting, the analogue of increasing information is \(\varepsilon\to0\), the model is locally asymptotically normal with normalization \(\varphi(\varepsilon)=\varepsilon\Sigma^{-1/2}\), and the MLE is asymptotically efficient with asymptotic covariance \(\Sigma^{-1}\) [2507.05628].

Outside the ergodic setting, the asymptotics need not be normal and may depend on the dynamical regime. For the \(\alpha\)-stable CIR process, the subcritical case yields strong consistency and asymptotic normality, the supercritical case yields strong consistency and asymptotic mixed normality, and in the critical case “the description of the asymptotic behavior of the MLE in question remains open” [1711.02140]. For the jump-type CIR process driven by a subordinator, the subcritical case is asymptotically normal, the critical case has a non-standard limit expressed through a critical diffusion CIR process, and the supercritical case has a mixed normal limit involving the almost sure limit of an exponentially rescaled process [1609.05865].

Heavy tails and recurrence structure alter both rate and limit law. In the heavy-tailed Lévy-driven Ornstein–Uhlenbeck model, the experiment is locally asymptotically mixed normal, the correct rate is \(\sqrt{T\varphi_T}\), and the MLE limit is a Gaussian scale mixture
\[
\sqrt{T\varphi_T}\,(\widehat\theta_T-\theta_0)\Longrightarrow \frac{N}{\sqrt{\theta_0 S(\alpha/2)}},
\]
with random information driven by an \(\alpha/2\)-stable variable [1911.11202]. In threshold diffusion, the occupation times \(Q_T^\pm\) determine whether the estimator is consistent and whether the limit is Gaussian, mixed normal, or nonstandard; the paper treats ergodic, transient, and null recurrent cases separately [1803.05408]. For the second-order Gaussian autoregression, the MLE is always strongly consistent, but the paper identifies nine distinct asymptotic regimes. One of its main conclusions is that when \(p>0\) and \(q<p\), the MLE convergence rate is governed by the smaller root while the normalized likelihood ratio is governed by the larger root [1206.1379].

Efficiency results are likewise model-specific. The Lévy-driven Ornstein–Uhlenbeck drift estimator is asymptotically efficient in the Hájek–Le Cam sense [1403.2954]. The heavy-tailed Lévy-driven Ornstein–Uhlenbeck MLE is asymptotically efficient in the convolution-theorem sense under the LAMN structure [1911.11202]. The Gaussian-process mean-function MLE is asymptotically efficient under LAN, and the discrete-sample \(M\)-estimator matches the continuous-time efficiency when
\[
\varepsilon^{-1}\|K^n-K\|_{L^\infty}\to 0,
\]
with the concrete sufficient condition \(n\varepsilon\to\infty\) under the stated mesh and Lipschitz assumptions [2507.05628].

## 6. Computation, constraints, and methodological issues

A persistent computational theme is the replacement of intractable infinite-dimensional or continuous-state likelihoods by structured finite-dimensional surrogates. In the continuous-time higher-order Markov model, the drift and diffusion are assumed smooth or piecewise smooth and approximated by polynomials or splines, precisely to avoid the discrete higher-order parameter explosion \((m-1)m^n\) [2104.04636]. In the general irregularly sampled continuous-time state-space model, the approximation can be made arbitrarily accurate in principle by increasing the number of bins \(m\), but larger \(m\) increases computation time; the paper reports that \(m\ge 100\) was a conservative and often effective choice in its examples [2010.14883].

For finite-state CTMC likelihoods, computational efficiency is tied to gradient evaluation. One paper derives an \(O(n^3)\) gradient computation through eigendecomposition and a reusable matrix \(\mathbf Z\), contrasting this with earlier \(O(n^5)\) or \(O(n^6)\) approaches, and uses L-BFGS-B for optimization [1504.01804]. The diffusion-to-CTMC approximation similarly exploits matrix exponentials and count compression; its stated complexity is
\[
\mathcal O\big(N+N_c(m^3+B\,m)\big),
\]
where the raw sample is summarized once into a count matrix [2108.12649].

The literature is also explicit about methodological limits. Maximum approximate likelihood for general continuous-time state-space models requires that the transition density be available in closed form, which excludes many diffusions without tractable transitions [2010.14883]. The finite-state generator MLE is non-convex and may have local minima; embeddability problems also mean that directly taking a matrix logarithm of an empirical transition matrix can produce invalid generators [1504.01804]. Quasi-likelihood methods are used in low-frequency cointegrated state-space models precisely because the sampled process is not in standard innovation form and the true innovations are not directly available from finitely many observations [1712.08665].

A further methodological issue is correct likelihood specification. The highway-capacity paper argues that the Kaplan–Meier product-limit method is structurally mismatched to stochastic capacity estimation and that an earlier parametric likelihood was flawed because it used the capacity PDF where the CDF was required. The corrected likelihood is
\[
L(\theta)=\prod_{i=1}^{n}\left[F_c(I_i)^{\delta_i}\left(1-F_c(I_i)\right)^{1-\delta_i}\right],
\]
with \(P_B(I_i)=P(C<I_i)=F_c(I_i)\) [2507.00893]. This example does not concern diffusion inference, but it illustrates a broader point already visible across the continuous-time MLE literature: the asymptotic theory and computational machinery are secondary to specifying the correct likelihood object.

Taken together, these results show that continuous-time maximum likelihood estimation is not a single methodology but a stratified class of exact, approximate, and quasi-likelihood procedures. Exact pathwise likelihoods remain the canonical benchmark when continuous observation and measure equivalence are available. Approximate transition-density, CTMC, HMM, and quasi-innovation constructions extend the framework to sampled, hidden, irregular, or non-Markovian settings. Transformations, state augmentation, and stationary approximations are the principal devices that make these extensions possible, while consistency, LAN or LAMN structure, mixed normality, and efficiency are determined by the geometry of the underlying continuous-time model rather than by a universal asymptotic template.

Source: https://www.emergentmind.com/topics/continuous-time-maximum-likelihood-estimator