Papers
Topics
Authors
Recent
Search
2000 character limit reached

Approximate Markovianity

Updated 14 July 2026
  • Approximate Markovianity is a framework that replaces complex non-Markovian dynamics with simplified Markov models preserving selected observables.
  • It employs rigorous metrics—such as weak convergence, total variation, and Kullback–Leibler divergence—to quantify and control approximation errors.
  • Its applications span diffusion processes, quantum dynamics, symbolic chains, and numerical model reduction, providing practical, error-controlled analysis methods.

Approximate Markovianity denotes a family of approximation principles in which dynamics that are non-Markovian, singular, non-stationary, continuous-state, or otherwise analytically intractable are replaced by Markov models with controlled error. In the literature, this control is expressed in several distinct but technically precise ways: weak convergence of path laws and semigroups, convergence of resolvents and Dirichlet forms, small Ornstein dˉ\bar d-distance to a finite-order Markov chain, small distance to a Markovian process tensor, exponential decay of conditional dependence with a finite Markov length, bounded total variation or Kullback–Leibler discrepancy, and exact matching of low-order moments on discrete grids (Xu, 2012, Gallo et al., 2011, Figueroa-Romero et al., 2018).

1. Scope and principal formalisms

The term is used across several research areas, but the common structure is stable: one starts from dynamics whose effective description is difficult, and replaces them by a simpler Markov object that preserves a specified class of observables, transition laws, or information-theoretic signatures. The Markov object may be a lattice jump chain, a finite-order symbolic chain, a process tensor with factorized Choi state, a reduced finite-state model, or a controlled continuous-time Markov chain.

Domain Object being approximated Main approximation criterion
Jump and diffusion processes Continuous-state Markov process or diffusion Weak convergence, semigroup/resolvent convergence, pathwise coupling
Infinite-memory symbolic processes Chain of infinite order Small dˉ\bar d-distance to the canonical kk-step Markov approximation
Quantum multi-time dynamics Process tensor or open-system evolution Small trace/diamond distance to a Markovian process tensor; small QCMI
Noisy circuit outputs Classical measurement distribution Finite Markov length via exponential decay of conditional dependence
Model reduction and numerics High-dimensional or continuous model TV/KL control, Bellman approximation, kernel error, moment matching

A useful consequence of this breadth is that “approximate Markovianity” is not tied to one metric. In some settings it is a statement about generators and Dirichlet forms; in others it is about memory depth, conditional independence, or numerical surrogate models. This suggests that the topic is best understood as a structural theme rather than a single definition (Mimica et al., 2016, Tzortzis et al., 2014, Zhang et al., 7 Oct 2025).

2. Approximation of continuous and singular stochastic dynamics by Markov chains

A central classical usage is the approximation of continuous-state Markov processes by discrete-state Markov chains. For singular stable-like processes supported on the union of coordinate axes, the process XX on Rd\mathbb{R}^d is defined by a Dirichlet form

E(f,f)=12RdRd(f(y)f(x))2J(x,y)m(dy)dx,\mathcal{E}(f,f) = \frac{1}{2}\int_{\mathbb{R}^d}\int_{\mathbb{R}^d}(f(y)-f(x))^2 J(x,y)\,m(dy)\,dx,

where the jump kernel has the axis-supported form

J(x,y)={c(x,y)xy1+α,yxi=1dRi{0}, 0,otherwise,J(x,y)= \begin{cases} \dfrac{c(x,y)}{|x-y|^{1+\alpha}}, & y-x \in \bigcup_{i=1}^d \mathbb{R}_i\setminus\{0\},\ 0, & \text{otherwise}, \end{cases}

with 0<K1c(x,y)K20<K_1\le c(x,y)\le K_2 and α(0,2)\alpha\in(0,2). The approximating chains YnY^n live on lattices dˉ\bar d0 with conductances dˉ\bar d1 that mimic this singular geometry. The key approximation theorem states that if the discrete Lévy measures converge weakly on shells dˉ\bar d2, then for each dˉ\bar d3 and dˉ\bar d4, the laws of dˉ\bar d5 converge weakly in dˉ\bar d6 to the law of dˉ\bar d7. The proof passes through heat kernel bounds, exit time estimates, Hölder regularity, resolvent convergence dˉ\bar d8, semigroup convergence, and tightness in Skorokhod space (Xu, 2012).

The same pattern appears in a much broader non-symmetric setting. For pure jump processes with kernel dˉ\bar d9, the approximating chains kk0 on kk1 have generator

kk2

Approximate Markovianity is formulated as weak convergence kk3 in kk4, established either by non-symmetric Mosco convergence of semi-Dirichlet forms or by convergence of semimartingale characteristics. The discrete forms kk5 approximate the continuous form kk6, and strong semigroup convergence kk7 yields convergence of finite-dimensional distributions; tightness then upgrades this to process-level convergence (Mimica et al., 2016).

A stronger formulation replaces weak approximation by pathwise coupling. Under bounded coefficients, local ellipticity or specified degenerate structure, and Hölder/Lipschitz regularity, one can construct a Markov chain and a limiting diffusion on a common probability space so that the probability of exact coincidence on a discrete grid is at least kk8, and the linearly interpolated chain stays uniformly close to the diffusion with explicitly controlled probability on the whole time interval. This converts weak convergence into a strong approximation principle: the discrete Markov chain is not only distributionally close but can be realized as a near pathwise replica of the diffusion (Konakov et al., 9 Jun 2026).

The same objective motivates recombining lattice constructions. For SDEs

kk9

the paper on recombination on lattice trees constructs discrete-time, discrete-space Markov chains whose conditional increment moments locally match the diffusion’s first two moments. By iterating Carathéodory’s theorem, each row of the transition matrix has only a small number of nonzero entries, while all nodes lie on a “universal lattice” XX0. The number of attainable states at time XX1 grows at most polynomially, yielding sparse transition matrices and recombination rather than exponential branching (Cosentino et al., 2021).

Approximate Markovianity can also refer to time-inhomogeneous Markov models that are close to stationary ones. For slowly changing non-stationary chains with transition matrices XX2, performance measures such as discounted reward, reward to hitting time, state reward at time XX3, and cumulative reward over XX4 admit asymptotic expansions whose leading term is the corresponding stationary quantity for a frozen matrix XX5, with first- and second-order corrections computed by linear systems involving XX6 or XX7. In that setting, approximate Markovianity means that a time-inhomogeneous chain behaves, for the target functional, like a stationary Markov chain plus explicit perturbative corrections (Zheng et al., 2018).

3. Finite-memory approximation of chains of infinite order

In symbolic dynamics and ergodic theory, approximate Markovianity is formulated as finite-memory approximation of a chain whose transition kernel depends on the entire past. A kernel

XX8

defines a chain of infinite order unless it depends only on the last XX9 symbols. The canonical Rd\mathbb{R}^d0-step Markov approximation is the stationary Rd\mathbb{R}^d1-step chain with transition kernel

Rd\mathbb{R}^d2

Closeness is measured by the Ornstein Rd\mathbb{R}^d3-distance

Rd\mathbb{R}^d4

which, for ergodic processes, also admits the interpretation of the minimal asymptotic proportion of sites that must be changed to transform one realization into the other. In this framework, a chain of infinite order is approximately Markov of order Rd\mathbb{R}^d5 when Rd\mathbb{R}^d6 is small, and asymptotically Markovian when Rd\mathbb{R}^d7 as Rd\mathbb{R}^d8 (Gallo et al., 2011).

The estimates are constructive. Using coupling from the past, common i.i.d. uniforms, range partitions, and coalescence times, the paper proves bounds such as

Rd\mathbb{R}^d9

under summable continuity rate and very weak non-nullness, as well as

E(f,f)=12RdRd(f(y)f(x))2J(x,y)m(dy)dx,\mathcal{E}(f,f) = \frac{1}{2}\int_{\mathbb{R}^d}\int_{\mathbb{R}^d}(f(y)-f(x))^2 J(x,y)\,m(dy)\,dx,0

under localized continuity assumptions, and

E(f,f)=12RdRd(f(y)f(x))2J(x,y)m(dy)dx,\mathcal{E}(f,f) = \frac{1}{2}\int_{\mathbb{R}^d}\int_{\mathbb{R}^d}(f(y)-f(x))^2 J(x,y)\,m(dy)\,dx,1

under weak non-nullness with an appropriate CFTP construction. The theory covers non-necessarily continuous kernels and kernels with null transition probabilities, and it yields explicit rates, including exponential decay in renewal-type examples and polynomial–logarithmic decay in non-summable continuity regimes (Gallo et al., 2011).

A major significance of this formulation is that approximate Markovianity is strong enough to imply ergodic-theoretic structure. Because ergodic Markov chains are Bernoulli and the class of Bernoulli shifts is E(f,f)=12RdRd(f(y)f(x))2J(x,y)m(dy)dx,\mathcal{E}(f,f) = \frac{1}{2}\int_{\mathbb{R}^d}\int_{\mathbb{R}^d}(f(y)-f(x))^2 J(x,y)\,m(dy)\,dx,2-closed, convergence of E(f,f)=12RdRd(f(y)f(x))2J(x,y)m(dy)dx,\mathcal{E}(f,f) = \frac{1}{2}\int_{\mathbb{R}^d}\int_{\mathbb{R}^d}(f(y)-f(x))^2 J(x,y)\,m(dy)\,dx,3 to E(f,f)=12RdRd(f(y)f(x))2J(x,y)m(dy)dx,\mathcal{E}(f,f) = \frac{1}{2}\int_{\mathbb{R}^d}\int_{\mathbb{R}^d}(f(y)-f(x))^2 J(x,y)\,m(dy)\,dx,4 in E(f,f)=12RdRd(f(y)f(x))2J(x,y)m(dy)dx,\mathcal{E}(f,f) = \frac{1}{2}\int_{\mathbb{R}^d}\int_{\mathbb{R}^d}(f(y)-f(x))^2 J(x,y)\,m(dy)\,dx,5 implies that the limiting infinite-order chain is Bernoulli (Gallo et al., 2011).

4. Quantum formulations: process tensors, information flow, and Markov length

In quantum theory, approximate Markovianity is formulated at the level of multi-time processes rather than one-step marginals. A E(f,f)=12RdRd(f(y)f(x))2J(x,y)m(dy)dx,\mathcal{E}(f,f) = \frac{1}{2}\int_{\mathbb{R}^d}\int_{\mathbb{R}^d}(f(y)-f(x))^2 J(x,y)\,m(dy)\,dx,6-step process is represented by a process tensor with Choi state E(f,f)=12RdRd(f(y)f(x))2J(x,y)m(dy)dx,\mathcal{E}(f,f) = \frac{1}{2}\int_{\mathbb{R}^d}\int_{\mathbb{R}^d}(f(y)-f(x))^2 J(x,y)\,m(dy)\,dx,7. A process is Markovian exactly when its Choi state factorizes into a tensor product of single-step maps,

E(f,f)=12RdRd(f(y)f(x))2J(x,y)m(dy)dx,\mathcal{E}(f,f) = \frac{1}{2}\int_{\mathbb{R}^d}\int_{\mathbb{R}^d}(f(y)-f(x))^2 J(x,y)\,m(dy)\,dx,8

and approximate Markovianity is quantified by the distance to the closest Markovian process. One metric is

E(f,f)=12RdRd(f(y)f(x))2J(x,y)m(dy)dx,\mathcal{E}(f,f) = \frac{1}{2}\int_{\mathbb{R}^d}\int_{\mathbb{R}^d}(f(y)-f(x))^2 J(x,y)\,m(dy)\,dx,9

For Haar-random closed-system dynamics on J(x,y)={c(x,y)xy1+α,yxi=1dRi{0}, 0,otherwise,J(x,y)= \begin{cases} \dfrac{c(x,y)}{|x-y|^{1+\alpha}}, & y-x \in \bigcup_{i=1}^d \mathbb{R}_i\setminus\{0\},\ 0, & \text{otherwise}, \end{cases}0, there is a concentration bound

J(x,y)={c(x,y)xy1+α,yxi=1dRi{0}, 0,otherwise,J(x,y)= \begin{cases} \dfrac{c(x,y)}{|x-y|^{1+\alpha}}, & y-x \in \bigcup_{i=1}^d \mathbb{R}_i\setminus\{0\},\ 0, & \text{otherwise}, \end{cases}1

and J(x,y)={c(x,y)xy1+α,yxi=1dRi{0}, 0,otherwise,J(x,y)= \begin{cases} \dfrac{c(x,y)}{|x-y|^{1+\alpha}}, & y-x \in \bigcup_{i=1}^d \mathbb{R}_i\setminus\{0\},\ 0, & \text{otherwise}, \end{cases}2 as J(x,y)={c(x,y)xy1+α,yxi=1dRi{0}, 0,otherwise,J(x,y)= \begin{cases} \dfrac{c(x,y)}{|x-y|^{1+\alpha}}, & y-x \in \bigcup_{i=1}^d \mathbb{R}_i\setminus\{0\},\ 0, & \text{otherwise}, \end{cases}3 for fixed J(x,y)={c(x,y)xy1+α,yxi=1dRi{0}, 0,otherwise,J(x,y)= \begin{cases} \dfrac{c(x,y)}{|x-y|^{1+\alpha}}, & y-x \in \bigcup_{i=1}^d \mathbb{R}_i\setminus\{0\},\ 0, & \text{otherwise}, \end{cases}4. At fixed global dimension, however, long-time limits can be highly non-Markovian: for J(x,y)={c(x,y)xy1+α,yxi=1dRi{0}, 0,otherwise,J(x,y)= \begin{cases} \dfrac{c(x,y)}{|x-y|^{1+\alpha}}, & y-x \in \bigcup_{i=1}^d \mathbb{R}_i\setminus\{0\},\ 0, & \text{otherwise}, \end{cases}5, the expected non-Markovianity need not remain small. The same large-deviation picture extends from Haar-random dynamics to approximate unitary designs, a phenomenon termed Markovianization (Figueroa-Romero et al., 2018, Figueroa-Romero et al., 2020).

A related information-theoretic formulation uses quantum conditional mutual information. If an approximately CPTP channel J(x,y)={c(x,y)xy1+α,yxi=1dRi{0}, 0,otherwise,J(x,y)= \begin{cases} \dfrac{c(x,y)}{|x-y|^{1+\alpha}}, & y-x \in \bigcup_{i=1}^d \mathbb{R}_i\setminus\{0\},\ 0, & \text{otherwise}, \end{cases}6 satisfies

J(x,y)={c(x,y)xy1+α,yxi=1dRi{0}, 0,otherwise,J(x,y)= \begin{cases} \dfrac{c(x,y)}{|x-y|^{1+\alpha}}, & y-x \in \bigcup_{i=1}^d \mathbb{R}_i\setminus\{0\},\ 0, & \text{otherwise}, \end{cases}7

then the resulting conditional mutual information obeys

J(x,y)={c(x,y)xy1+α,yxi=1dRi{0}, 0,otherwise,J(x,y)= \begin{cases} \dfrac{c(x,y)}{|x-y|^{1+\alpha}}, & y-x \in \bigcup_{i=1}^d \mathbb{R}_i\setminus\{0\},\ 0, & \text{otherwise}, \end{cases}8

This gives an approximate quantum Markov condition in the Buscemi–Das–Wilde sense and links approximate Markovianity to quantum Darwinism: the same paper argues that Darwinistic plateau behavior forces information backflow and non-Markovianity to be small (Guo et al., 2022).

For noisy quantum circuits, the relevant object is the final classical output distribution J(x,y)={c(x,y)xy1+α,yxi=1dRi{0}, 0,otherwise,J(x,y)= \begin{cases} \dfrac{c(x,y)}{|x-y|^{1+\alpha}}, & y-x \in \bigcup_{i=1}^d \mathbb{R}_i\setminus\{0\},\ 0, & \text{otherwise}, \end{cases}9. Approximate Markovianity is encoded by a finite Markov length 0<K1c(x,y)K20<K_1\le c(x,y)\le K_20, defined by

0<K1c(x,y)K20<K_1\le c(x,y)\le K_21

for all tripartitions 0<K1c(x,y)K20<K_1\le c(x,y)\le K_22. If 0<K1c(x,y)K20<K_1\le c(x,y)\le K_23, then local conditionals can be truncated to neighborhoods of radius 0<K1c(x,y)K20<K_1\le c(x,y)\le K_24, which yields a sequential classical sampler running in quasi-polynomial time. The paper proves this property for arbitrary circuits above a constant depolarizing-noise threshold and gives analytical and numerical evidence that random circuits satisfy it for any constant noise rate (Zhang et al., 7 Oct 2025).

Closed-system quantum histories provide yet another variant. There, one defines history probabilities through projectors 0<K1c(x,y)K20<K_1\le c(x,y)\le K_25, measures approximate consistency by

0<K1c(x,y)K20<K_1\le c(x,y)\le K_26

and approximate one-step Markovianity by

0<K1c(x,y)K20<K_1\le c(x,y)\le K_27

Numerical evidence for Heisenberg-type spin systems with 12–20 spins shows that both quantities are small for suitable coarse observables and that the approximation improves with system size (Schmidtke et al., 2016).

5. Markovianization as a numerical and modeling strategy

In several applied areas, approximate Markovianity is valuable primarily because it converts hard continuous or nonlocal problems into tractable Markov models.

In rough volatility, the rough Heston model is non-Markovian and, except for the classical Heston limit 0<K1c(x,y)K20<K_1\le c(x,y)\le K_28, not a semimartingale in the usual sense. The key approximation replaces the fractional kernel by a finite sum of exponentials,

0<K1c(x,y)K20<K_1\le c(x,y)\le K_29

which yields a finite-dimensional affine Markov process. The paper proves that for broad classes of European payoffs the weak pricing error is controlled by the α(0,2)\alpha\in(0,2)0-kernel error α(0,2)\alpha\in(0,2)1, not by the α(0,2)\alpha\in(0,2)2-error used in strong analyses. It further gives Gaussian-type constructions with super-polynomial convergence, including

α(0,2)\alpha\in(0,2)3

valid for all α(0,2)\alpha\in(0,2)4, including the hyper-rough regime (Bayer et al., 2023).

For stationary and weak KAM Hamilton–Jacobi equations on the torus, the approximation proceeds in the opposite direction: a deterministic calculus-of-variations problem is replaced by a continuous-time Markov decision problem on a lattice α(0,2)\alpha\in(0,2)5. The controlled generator α(0,2)\alpha\in(0,2)6 induces a Bellman equation

α(0,2)\alpha\in(0,2)7

whose solution is the value function of the CTMDP. The paper proves

α(0,2)\alpha\in(0,2)8

for the discounted problem, and

α(0,2)\alpha\in(0,2)9

for the effective Hamiltonian. Mather measures of the continuous problem arise as limits of invariant state-action measures of the approximating Markov chains (Averboukh, 2024).

Another constructive route is exact moment matching on nonuniform grids. For processes with

YnY^n0

one can build a nearest-neighbor Markov chain YnY^n1 on a nonuniform grid YnY^n2 with transition probabilities YnY^n3 chosen so that

YnY^n4

for all YnY^n5. This covers heat diffusion and log-GBM. In the GBM example, numerical experiments with a nonuniform grid yield consistently small empirical Wasserstein-1 distances at long time horizons and outperform a comparable uniform grid in the long-time regime (Kim et al., 25 Nov 2025).

Approximate Markovianity also appears in finite-state reduction. A finite-state Markov chain can be approximated by a lower-dimensional process by optimizing over a total-variation ball either at the level of transition probabilities or invariant distributions. The resulting solutions have a water-filling structure, and the final reduced Markov chain is obtained by minimizing a lifted Kullback–Leibler divergence between the original chain and the reduced one (Tzortzis et al., 2014).

6. Distinctions, limitations, and recurring misconceptions

A first distinction is between weak and strong approximation. Weak convergence of YnY^n6 to YnY^n7 in Skorokhod space, semigroup convergence, or weak pricing error control does not imply a pathwise coupling. This contrast is explicit between non-symmetric pure-jump chain approximations, which are process-level but weak (Mimica et al., 2016), and strong diffusion approximations, which maximize exact pathwise coincidence on grids and control uniform interpolation error on a common probability space (Konakov et al., 9 Jun 2026). A parallel distinction appears in rough Heston, where YnY^n8-kernel control is sufficient for weak option-pricing error, whereas earlier strong analyses relied on YnY^n9-kernel error (Bayer et al., 2023).

A second distinction concerns the strength of the Markov notion itself. In quantum theory, process-tensor Markovianity is stronger than CP-divisibility or trace-distance backflow criteria; factorization of the Choi state eliminates all higher-order temporal correlations, not only two-time memory effects (Figueroa-Romero et al., 2018). In noisy-circuit sampling, finite Markov length is a classical conditional-independence condition on output distributions, not a statement about open-system generators (Zhang et al., 7 Oct 2025). In symbolic dynamics, dˉ\bar d00-approximation by finite-order chains is an ergodic-theoretic memory-depth statement, not a semigroup statement (Gallo et al., 2011).

A third distinction is temporal scale. Approximate Markovianity can hold on finite horizons and fail at longer ones. Haar-random quantum processes become close to Markovian when the environment is sufficiently large compared with the subsystem and the number of sampled times is fixed, but this can break down as the number of time steps dˉ\bar d01 grows at fixed total dimension (Figueroa-Romero et al., 2018). Closed quantum histories show a related dependence on measurement spacing: approximate consistency and Markovianity improve when the time step is not too small relative to the relaxation time (Schmidtke et al., 2016). For slowly varying non-stationary Markov chains, the approximation is local in logarithmic time windows, not globally stationary (Zheng et al., 2018).

A final recurrent misconception is that approximate Markovianity always means state-space reduction. In fact, some of the most important formulations keep or even enlarge the state description: lattice approximations to diffusions, finite-factor rough Heston models, and process-tensor descriptions all preserve a Markov structure by changing representation rather than merely compressing it (Cosentino et al., 2021, Bayer et al., 2023, Figueroa-Romero et al., 2020).

Taken together, these developments show that approximate Markovianity is a unifying research program for replacing difficult dynamics by tractable Markov surrogates while preserving a chosen set of signatures—path laws, operators, memory depth, information flow, prices, or moments. The specific metric changes across fields, but the underlying objective remains the same: to isolate the scale or representation at which a memory-bearing system can be treated, with quantified error, as effectively Markovian.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Approximate Markovianity.