Approximate Markovianity
- Approximate Markovianity is a framework that replaces complex non-Markovian dynamics with simplified Markov models preserving selected observables.
- It employs rigorous metrics—such as weak convergence, total variation, and Kullback–Leibler divergence—to quantify and control approximation errors.
- Its applications span diffusion processes, quantum dynamics, symbolic chains, and numerical model reduction, providing practical, error-controlled analysis methods.
Approximate Markovianity denotes a family of approximation principles in which dynamics that are non-Markovian, singular, non-stationary, continuous-state, or otherwise analytically intractable are replaced by Markov models with controlled error. In the literature, this control is expressed in several distinct but technically precise ways: weak convergence of path laws and semigroups, convergence of resolvents and Dirichlet forms, small Ornstein -distance to a finite-order Markov chain, small distance to a Markovian process tensor, exponential decay of conditional dependence with a finite Markov length, bounded total variation or Kullback–Leibler discrepancy, and exact matching of low-order moments on discrete grids (Xu, 2012, Gallo et al., 2011, Figueroa-Romero et al., 2018).
1. Scope and principal formalisms
The term is used across several research areas, but the common structure is stable: one starts from dynamics whose effective description is difficult, and replaces them by a simpler Markov object that preserves a specified class of observables, transition laws, or information-theoretic signatures. The Markov object may be a lattice jump chain, a finite-order symbolic chain, a process tensor with factorized Choi state, a reduced finite-state model, or a controlled continuous-time Markov chain.
| Domain | Object being approximated | Main approximation criterion |
|---|---|---|
| Jump and diffusion processes | Continuous-state Markov process or diffusion | Weak convergence, semigroup/resolvent convergence, pathwise coupling |
| Infinite-memory symbolic processes | Chain of infinite order | Small -distance to the canonical -step Markov approximation |
| Quantum multi-time dynamics | Process tensor or open-system evolution | Small trace/diamond distance to a Markovian process tensor; small QCMI |
| Noisy circuit outputs | Classical measurement distribution | Finite Markov length via exponential decay of conditional dependence |
| Model reduction and numerics | High-dimensional or continuous model | TV/KL control, Bellman approximation, kernel error, moment matching |
A useful consequence of this breadth is that “approximate Markovianity” is not tied to one metric. In some settings it is a statement about generators and Dirichlet forms; in others it is about memory depth, conditional independence, or numerical surrogate models. This suggests that the topic is best understood as a structural theme rather than a single definition (Mimica et al., 2016, Tzortzis et al., 2014, Zhang et al., 7 Oct 2025).
2. Approximation of continuous and singular stochastic dynamics by Markov chains
A central classical usage is the approximation of continuous-state Markov processes by discrete-state Markov chains. For singular stable-like processes supported on the union of coordinate axes, the process on is defined by a Dirichlet form
where the jump kernel has the axis-supported form
with and . The approximating chains live on lattices 0 with conductances 1 that mimic this singular geometry. The key approximation theorem states that if the discrete Lévy measures converge weakly on shells 2, then for each 3 and 4, the laws of 5 converge weakly in 6 to the law of 7. The proof passes through heat kernel bounds, exit time estimates, Hölder regularity, resolvent convergence 8, semigroup convergence, and tightness in Skorokhod space (Xu, 2012).
The same pattern appears in a much broader non-symmetric setting. For pure jump processes with kernel 9, the approximating chains 0 on 1 have generator
2
Approximate Markovianity is formulated as weak convergence 3 in 4, established either by non-symmetric Mosco convergence of semi-Dirichlet forms or by convergence of semimartingale characteristics. The discrete forms 5 approximate the continuous form 6, and strong semigroup convergence 7 yields convergence of finite-dimensional distributions; tightness then upgrades this to process-level convergence (Mimica et al., 2016).
A stronger formulation replaces weak approximation by pathwise coupling. Under bounded coefficients, local ellipticity or specified degenerate structure, and Hölder/Lipschitz regularity, one can construct a Markov chain and a limiting diffusion on a common probability space so that the probability of exact coincidence on a discrete grid is at least 8, and the linearly interpolated chain stays uniformly close to the diffusion with explicitly controlled probability on the whole time interval. This converts weak convergence into a strong approximation principle: the discrete Markov chain is not only distributionally close but can be realized as a near pathwise replica of the diffusion (Konakov et al., 9 Jun 2026).
The same objective motivates recombining lattice constructions. For SDEs
9
the paper on recombination on lattice trees constructs discrete-time, discrete-space Markov chains whose conditional increment moments locally match the diffusion’s first two moments. By iterating Carathéodory’s theorem, each row of the transition matrix has only a small number of nonzero entries, while all nodes lie on a “universal lattice” 0. The number of attainable states at time 1 grows at most polynomially, yielding sparse transition matrices and recombination rather than exponential branching (Cosentino et al., 2021).
Approximate Markovianity can also refer to time-inhomogeneous Markov models that are close to stationary ones. For slowly changing non-stationary chains with transition matrices 2, performance measures such as discounted reward, reward to hitting time, state reward at time 3, and cumulative reward over 4 admit asymptotic expansions whose leading term is the corresponding stationary quantity for a frozen matrix 5, with first- and second-order corrections computed by linear systems involving 6 or 7. In that setting, approximate Markovianity means that a time-inhomogeneous chain behaves, for the target functional, like a stationary Markov chain plus explicit perturbative corrections (Zheng et al., 2018).
3. Finite-memory approximation of chains of infinite order
In symbolic dynamics and ergodic theory, approximate Markovianity is formulated as finite-memory approximation of a chain whose transition kernel depends on the entire past. A kernel
8
defines a chain of infinite order unless it depends only on the last 9 symbols. The canonical 0-step Markov approximation is the stationary 1-step chain with transition kernel
2
Closeness is measured by the Ornstein 3-distance
4
which, for ergodic processes, also admits the interpretation of the minimal asymptotic proportion of sites that must be changed to transform one realization into the other. In this framework, a chain of infinite order is approximately Markov of order 5 when 6 is small, and asymptotically Markovian when 7 as 8 (Gallo et al., 2011).
The estimates are constructive. Using coupling from the past, common i.i.d. uniforms, range partitions, and coalescence times, the paper proves bounds such as
9
under summable continuity rate and very weak non-nullness, as well as
0
under localized continuity assumptions, and
1
under weak non-nullness with an appropriate CFTP construction. The theory covers non-necessarily continuous kernels and kernels with null transition probabilities, and it yields explicit rates, including exponential decay in renewal-type examples and polynomial–logarithmic decay in non-summable continuity regimes (Gallo et al., 2011).
A major significance of this formulation is that approximate Markovianity is strong enough to imply ergodic-theoretic structure. Because ergodic Markov chains are Bernoulli and the class of Bernoulli shifts is 2-closed, convergence of 3 to 4 in 5 implies that the limiting infinite-order chain is Bernoulli (Gallo et al., 2011).
4. Quantum formulations: process tensors, information flow, and Markov length
In quantum theory, approximate Markovianity is formulated at the level of multi-time processes rather than one-step marginals. A 6-step process is represented by a process tensor with Choi state 7. A process is Markovian exactly when its Choi state factorizes into a tensor product of single-step maps,
8
and approximate Markovianity is quantified by the distance to the closest Markovian process. One metric is
9
For Haar-random closed-system dynamics on 0, there is a concentration bound
1
and 2 as 3 for fixed 4. At fixed global dimension, however, long-time limits can be highly non-Markovian: for 5, the expected non-Markovianity need not remain small. The same large-deviation picture extends from Haar-random dynamics to approximate unitary designs, a phenomenon termed Markovianization (Figueroa-Romero et al., 2018, Figueroa-Romero et al., 2020).
A related information-theoretic formulation uses quantum conditional mutual information. If an approximately CPTP channel 6 satisfies
7
then the resulting conditional mutual information obeys
8
This gives an approximate quantum Markov condition in the Buscemi–Das–Wilde sense and links approximate Markovianity to quantum Darwinism: the same paper argues that Darwinistic plateau behavior forces information backflow and non-Markovianity to be small (Guo et al., 2022).
For noisy quantum circuits, the relevant object is the final classical output distribution 9. Approximate Markovianity is encoded by a finite Markov length 0, defined by
1
for all tripartitions 2. If 3, then local conditionals can be truncated to neighborhoods of radius 4, which yields a sequential classical sampler running in quasi-polynomial time. The paper proves this property for arbitrary circuits above a constant depolarizing-noise threshold and gives analytical and numerical evidence that random circuits satisfy it for any constant noise rate (Zhang et al., 7 Oct 2025).
Closed-system quantum histories provide yet another variant. There, one defines history probabilities through projectors 5, measures approximate consistency by
6
and approximate one-step Markovianity by
7
Numerical evidence for Heisenberg-type spin systems with 12–20 spins shows that both quantities are small for suitable coarse observables and that the approximation improves with system size (Schmidtke et al., 2016).
5. Markovianization as a numerical and modeling strategy
In several applied areas, approximate Markovianity is valuable primarily because it converts hard continuous or nonlocal problems into tractable Markov models.
In rough volatility, the rough Heston model is non-Markovian and, except for the classical Heston limit 8, not a semimartingale in the usual sense. The key approximation replaces the fractional kernel by a finite sum of exponentials,
9
which yields a finite-dimensional affine Markov process. The paper proves that for broad classes of European payoffs the weak pricing error is controlled by the 0-kernel error 1, not by the 2-error used in strong analyses. It further gives Gaussian-type constructions with super-polynomial convergence, including
3
valid for all 4, including the hyper-rough regime (Bayer et al., 2023).
For stationary and weak KAM Hamilton–Jacobi equations on the torus, the approximation proceeds in the opposite direction: a deterministic calculus-of-variations problem is replaced by a continuous-time Markov decision problem on a lattice 5. The controlled generator 6 induces a Bellman equation
7
whose solution is the value function of the CTMDP. The paper proves
8
for the discounted problem, and
9
for the effective Hamiltonian. Mather measures of the continuous problem arise as limits of invariant state-action measures of the approximating Markov chains (Averboukh, 2024).
Another constructive route is exact moment matching on nonuniform grids. For processes with
0
one can build a nearest-neighbor Markov chain 1 on a nonuniform grid 2 with transition probabilities 3 chosen so that
4
for all 5. This covers heat diffusion and log-GBM. In the GBM example, numerical experiments with a nonuniform grid yield consistently small empirical Wasserstein-1 distances at long time horizons and outperform a comparable uniform grid in the long-time regime (Kim et al., 25 Nov 2025).
Approximate Markovianity also appears in finite-state reduction. A finite-state Markov chain can be approximated by a lower-dimensional process by optimizing over a total-variation ball either at the level of transition probabilities or invariant distributions. The resulting solutions have a water-filling structure, and the final reduced Markov chain is obtained by minimizing a lifted Kullback–Leibler divergence between the original chain and the reduced one (Tzortzis et al., 2014).
6. Distinctions, limitations, and recurring misconceptions
A first distinction is between weak and strong approximation. Weak convergence of 6 to 7 in Skorokhod space, semigroup convergence, or weak pricing error control does not imply a pathwise coupling. This contrast is explicit between non-symmetric pure-jump chain approximations, which are process-level but weak (Mimica et al., 2016), and strong diffusion approximations, which maximize exact pathwise coincidence on grids and control uniform interpolation error on a common probability space (Konakov et al., 9 Jun 2026). A parallel distinction appears in rough Heston, where 8-kernel control is sufficient for weak option-pricing error, whereas earlier strong analyses relied on 9-kernel error (Bayer et al., 2023).
A second distinction concerns the strength of the Markov notion itself. In quantum theory, process-tensor Markovianity is stronger than CP-divisibility or trace-distance backflow criteria; factorization of the Choi state eliminates all higher-order temporal correlations, not only two-time memory effects (Figueroa-Romero et al., 2018). In noisy-circuit sampling, finite Markov length is a classical conditional-independence condition on output distributions, not a statement about open-system generators (Zhang et al., 7 Oct 2025). In symbolic dynamics, 00-approximation by finite-order chains is an ergodic-theoretic memory-depth statement, not a semigroup statement (Gallo et al., 2011).
A third distinction is temporal scale. Approximate Markovianity can hold on finite horizons and fail at longer ones. Haar-random quantum processes become close to Markovian when the environment is sufficiently large compared with the subsystem and the number of sampled times is fixed, but this can break down as the number of time steps 01 grows at fixed total dimension (Figueroa-Romero et al., 2018). Closed quantum histories show a related dependence on measurement spacing: approximate consistency and Markovianity improve when the time step is not too small relative to the relaxation time (Schmidtke et al., 2016). For slowly varying non-stationary Markov chains, the approximation is local in logarithmic time windows, not globally stationary (Zheng et al., 2018).
A final recurrent misconception is that approximate Markovianity always means state-space reduction. In fact, some of the most important formulations keep or even enlarge the state description: lattice approximations to diffusions, finite-factor rough Heston models, and process-tensor descriptions all preserve a Markov structure by changing representation rather than merely compressing it (Cosentino et al., 2021, Bayer et al., 2023, Figueroa-Romero et al., 2020).
Taken together, these developments show that approximate Markovianity is a unifying research program for replacing difficult dynamics by tractable Markov surrogates while preserving a chosen set of signatures—path laws, operators, memory depth, information flow, prices, or moments. The specific metric changes across fields, but the underlying objective remains the same: to isolate the scale or representation at which a memory-bearing system can be treated, with quantified error, as effectively Markovian.