Empirical Semi-Markov Kernel
- Empirical Semi-Markov Kernel is a data-derived joint law capturing both state transitions and associated sojourn times, extending traditional Markov matrices.
- It is estimated via nonparametric plug-in methods or Bayesian formulations, using observed transitions and waiting-time distributions from high-frequency data.
- Applications span finance, graph dynamics, and thermodynamics, providing critical insights into temporal memory and system behaviors.
An empirical semi-Markov kernel is a data-derived specification of the joint law of the next state and the associated sojourn time in a semi-Markov process. In the cited literature, the object appears in several closely related senses: as a nonparametric plug-in estimate obtained from transition and waiting-time counts, as a predictive kernel in Bayesian nonparametrics, and as an empirical object reconstructed from empirical measures and flows along long trajectories. Across these formulations, the kernel extends a Markov transition matrix by coupling state transitions with non-geometric or non-exponential holding-time laws, thereby encoding temporal memory that is absent from ordinary Markov models (D'Amico et al., 2011, Arfè et al., 2018, Jia et al., 2022).
1. Formal definition and canonical factorizations
The standard finite-state formulation defines the semi-Markov kernel by
where is the embedded chain and the transition time. The corresponding embedded transition probabilities are
and the waiting-time cdf conditional on the next state is
A more general state-space version is
and, under absolute continuity,
so that inference may be posed directly for the kernel density (D'Amico et al., 2011, Barbu et al., 2019).
A recurrent special case is a factorization between jump direction and waiting time. In the graph-dynamics construction, if the Markov chain and the sojourn times are independent, then
and a state-dependent variant replaces by 0. In stochastic thermodynamics, the analogous condition is direction-time independence (DTI),
1
which states that the waiting time at 2 does not depend on the destination 3. This restriction is structurally important but not generic: later large-deviation results explicitly treat the non-DTI case (Raberto et al., 2011, Ertel et al., 2021, Maier et al., 18 Sep 2025).
The basic misconception to avoid is that a semi-Markov kernel is merely an enriched transition matrix. It is instead a joint transition-and-duration object. The empirical problem is therefore intrinsically bivariate: estimating where the process goes next and how long it remains in the current state before moving.
2. Empirical construction from observed trajectories
In direct nonparametric estimation, the empirical semi-Markov kernel is assembled from observed state transitions and waiting times. In the high-frequency return model, the embedded Markov matrix 4 is estimated by transition counts between states, 5 is estimated as the empirical distribution of waiting times before an 6 move, 7 is the empirical probability mass for waiting time 8, and the full kernel 9 is populated using empirical frequencies. This is a literal plug-in empirical kernel on a discretized return state space (D'Amico et al., 2011).
A more elaborate construction appears in the indexed semi-Markov model for price changes. There, the kernel is conditioned not only on the present return state 0 but also on a memory index 1 built from a moving average of past squared returns: 2 The empirical procedure counts, for each triple 3, the fraction of occasions on which the process is in state 4, the index is in regime 5, the next state is 6, and the sojourn time is 7. These counts yield estimates of 8. The paper also discretizes returns into 9 symmetric states and the volatility index into 5 levels for estimation and simulation purposes (D'Amico et al., 2011).
When the state process is not directly observed, the empirical kernel may be constructed through latent-state estimation. For a finite-state discrete-time semi-Markov chain observed with Gaussian noise, the transition matrix at step 0 depends on the occupation time 1 since the last jump: 2 The empirical kernel at step 3 is then 4, obtained by plugging in an estimated occupation time 5 recovered by filtering or smoothing from the noisy observations. In this setting, the empirical kernel is not a single static object but a sequence of occupation-time-conditioned transition matrices driven by hidden-state inference (Elliott, 2019).
These constructions share a common template: empirical frequencies are accumulated over observed or inferred visits to the current state, and the result is normalized to recover the joint law of next state and waiting time.
3. Predictive and Bayesian formulations
In Bayesian nonparametrics, the empirical semi-Markov kernel often appears as a predictive distribution rather than a pure frequency estimator. The semi-Markov beta-Stacy process places an independent Dirichlet process prior on each row of the transition matrix and an independent beta-Stacy process prior on each holding-time distribution. After observing a trajectory, the posterior remains in the same family, and the one-step-ahead predictive law is described in the source as the empirical semi-Markov kernel. Its transition probabilities are weighted mixtures of prior and empirical sample, determined by transition counts 6, sojourn counts 7, and the current tenure 8. The source emphasizes that this formula generalizes the empirical transition kernel familiar from Markov chains to the semi-Markov setting (Arfè et al., 2018).
A second Bayesian line studies inference directly for the kernel density 9. Given observed data
0
a prior 1 is placed on the space of semi-Markov kernel densities, and the posterior is updated by Bayes’ rule. The central asymptotic question is whether the posterior contracts around the true kernel 2. Under assumptions 3, the posterior mass outside an 4-ball around 5 vanishes: 6 with distance measured by a Hellinger-type semidistance. The same framework constructs robust tests between Hellinger balls around semi-Markov kernels and supplies sufficient prior conditions involving ergodicity, metric entropy, Kullback-Leibler support, and sieve control (Barbu et al., 2019).
This suggests a useful distinction. In frequentist usage, the empirical kernel is the count-based plug-in object. In Bayesian usage, a closely related role is played by the predictive kernel and by posterior concentration for the latent “true” kernel. The underlying estimand is the same joint transition-duration law.
4. Empirical measure, empirical flow, and large-deviation structure
In the large-deviation literature, the empirical semi-Markov kernel is reconstructed from long-run occupation and transition statistics. For Markov renewal processes on a countable state space, the empirical measure is
7
and the empirical flow is
8
When the balance condition holds, these determine the empirical transition kernel and empirical waiting-time law through
9
The pair 0 is the empirical semi-Markov kernel in implicit form, and the joint large deviation principle for 1 yields rate functions written in terms of relative entropies between empirical and nominal transition and waiting-time components (Jia et al., 2022).
For semi-Markov processes satisfying DTI, the empirical flow and empirical measure also underlie large deviation principles for empirical currents. The empirical current is
2
and its rate function is obtained by contracting the joint rate function for 3. This formulation makes the empirical kernel relevant not only for estimation but also for nonequilibrium fluctuation theory (Faggionato, 2017).
A more recent result gives a direct large-deviation rate function for the empirical semi-Markov kernel itself, without assuming DTI. If 4 is the empirical frequency of observing a transition 5 after waiting time 6, then
7
The source describes this as a Kullback-Leibler divergence between the empirical and true semi-Markov kernels, weighted by the empirical frequency 8. The same paper uses contraction to derive bounds for the rate function of empirical entropy production and a lower bound on the variance of the mean entropy production rate measured along a finite-time trajectory (Maier et al., 18 Sep 2025).
One implication is methodological: empirical kernels can be studied either through finite-sample counting procedures or through asymptotic probability laws for the entire empirical object.
5. Conditioning variables: memory, actions, and modes
Semi-Markov kernels are often empirically indexed by auxiliary variables that alter either the jump law or the holding-time law. In financial econometrics, the memory index 9 conditions the kernel on recent volatility. The empirical finding reported for this indexed model is that a plain semi-Markov specification leaves the autocorrelation of squared returns dying off quickly, whereas the indexed semi-Markov model can mimic uncorrelated returns together with persistent autocorrelation of squared returns, with an optimal memory length 0 (D'Amico et al., 2011).
In graph dynamics, the continuous-time network process is built by subordinating a Markov chain on graph states to a renewal counting process: 1 The semi-Markov kernel governs the timing of transitions, while the Markov kernel governs the next graph state. Under independence, the two factors separate; with state dependence, the survival function becomes 2. The paper emphasizes undirected graphs with a fixed number of nodes for simplicity, but also presents an interbank market example where directed or weighted graphs are meaningful (Raberto et al., 2011).
In stochastic games, the kernel is state-action dependent: 3 with 4 the state and 5 the players’ actions. Here the kernel simultaneously specifies the distribution of the sojourn time and the next state under a given action pair, and it enters the Shapley operator through integrals over 6 and the associated sojourn-time distribution 7 (Yu et al., 2021).
In control of continuous-time semi-Markov jump linear systems, the kernel is written as
8
where 9 is the embedded transition probability and 0 the cdf of the holding time before switching from mode 1 to 2. The stabilization criterion is then expressed through probabilities obtained by integrating the corresponding densities 3 over intervals such as 4, 5, and 6 (Wang et al., 2021).
These variants show that empirical semi-Markov kernels need not be homogeneous. They may be conditioned on volatility regimes, occupancy age, controls, or mode pairs, while retaining the same underlying role as the joint law of next event and waiting time.
6. Applications, diagnostics, and interpretive issues
High-frequency finance provides the most explicit empirical case studies. One model uses intraday returns as a homogeneous semi-Markov process and overnight returns as a Markov chain, all on the same finite discretized state space. The data are one-minute returns from the Italian stock market from first of January 2007 until end of December 2010. A nonparametric test of the geometric-sojourn-time hypothesis rejects the Markovian null for 15 out of 20 possible transition pairs in the 5-state model, and synthetic data generated from the estimated semi-Markov kernel reproduce first passage time and volatility autocorrelation more accurately than the Markov alternative (D'Amico et al., 2011). A related indexed model, using the same market and period, reports that there exists an optimal 7: too short does not capture regime, while too long blends regimes (D'Amico et al., 2011).
In thermodynamics, the semi-Markov kernel is treated as an operationally accessible object measurable from time series of jumps between discrete states. Under DTI, the ratio
8
must be independent of 9, and the source states that DTI is necessary for thermodynamic consistency in equilibrium. A violation of the Markov-based thermodynamic uncertainty relation can then be used as an inference tool to conclude that an observed semi-Markov process cannot result from coarse-graining an underlying Markov one (Ertel et al., 2021). The later large-deviation treatment without DTI strengthens this diagnostic perspective by linking empirical-kernel fluctuations directly to entropy-production bounds (Maier et al., 18 Sep 2025).
Graph dynamics provides a different empirical perspective. Monte Carlo simulations in the graph-renewal model draw sojourn times from a Mittag-Leffler distribution, advance the process according to the subordinate Markov kernel, and track events such as the appearance of triangles or graph connectivity. The reported outputs include boxplots, means, and medians for stopping times under different values of the tail parameter 0, illustrating how the waiting-time component of the kernel changes observed network event statistics (Raberto et al., 2011).
Two recurring interpretive issues follow from these examples. First, the empirical semi-Markov kernel is not exhausted by transition probabilities; waiting-time estimation is equally constitutive. Second, DTI is a useful simplifying assumption and, in some thermodynamic settings, a structural requirement, but it is not universal. The later large-deviation results for non-DTI kernels therefore expand the empirical scope of semi-Markov analysis rather than merely refining an already complete theory (Ertel et al., 2021, Maier et al., 18 Sep 2025).