---
title: Tree-Indexed Markov Chains Overview
url: https://www.emergentmind.com/topics/tree-indexed-markov-chains
type: topic
---

# Tree-Indexed Markov Chains Overview

Searching arXiv for the cited tree-indexed Markov chain literature.
Tree-indexed Markov chains are stochastic processes whose indices form a rooted tree rather than a line. In the Ulam–Harris–Neveu formalism, a branching Markov process \(X=(X_u)_{u\in\mathbb T_\infty}\) on a Polish state space \((\mathcal X,\mathcal B)\) is specified by an initial law \(v\) and a transition kernel \(Q\) such that for every finite subtree \(T\subset\mathbb T_\infty\) containing the root,
\[
\Pr\bigl(X_u\in dx_u,\ u\in T\bigr)
=
v(dx_{\varnothing})\prod_{u\in T\setminus\{\varnothing\}}Q\bigl(x_{p(u)};dx_u\bigr)
\]
[2403.16505]. On full binary trees, bifurcating Markov chains replace the single-child kernel by a mother-to-two-daughters kernel \(P(x,dy,dz)\), while in finite-alphabet graphical models a “Markov chain on a tree” is characterized by graph-separation conditional independence [2012.04741, 2307.15844]. The subject therefore encompasses several closely related formalisms in which Markovian dependence is propagated along ancestral structure rather than along a one-dimensional time axis.

## 1. Indexing trees, notation, and canonical constructions

A standard index set is the Ulam–Harris–Neveu tree
\[
\mathbb T_\infty=\bigcup_{k=0}^\infty (\mathbb N^*)^k,
\qquad
(\mathbb N^*)^0=\{\varnothing\},
\]
with root \(\varnothing\), parent map \(p(u)\), height \(h(u)\), latest common ancestor \(u\wedge v\), and graph distance
\[
d(u,v)=h(u)+h(v)-2\,h(u\wedge v)
\]
[2403.16505]. For binary models one often writes
\[
G_k=\{0,1\}^k,\qquad T_k=\bigcup_{0\le r\le k}G_r,\qquad T=\bigcup_{r\ge0}G_r,
\]
so that \(G_k\) is the \(k\)-th generation and \(T_k\) is the tree up to generation \(k\) [2012.04741, 2106.07711]. For rooted \(d\)-ary trees,
\[
T=\bigcup_{n\ge0}\Sigma^n,\qquad \Sigma=\{1,2,\dots,d\},
\]
and \(|T_n|=(d^{n+1}-1)/(d-1)\) for the first \(n\) levels [2401.05320].

The line graph remains a limiting benchmark. In Weibel’s formulation, the finite line graph
\[
L_n=\{u_1,u_2,\dots,u_n\},
\]
with edges \(u_i\text{--}u_{i+1}\), is the special case in which the tree-indexed process reduces to an ordinary length-\(n\) Markov chain [2403.16505]. This comparison is structurally important because many asymptotic and variance formulas can be read as tree analogues of classical chain results.

In bifurcating models, the transition mechanism is not a single-child kernel but a joint law for siblings. If \(P(x,dy,dz)\) is the mother-to-two-daughter kernel, its one-dimensional marginals are
\[
P_0(x,A)=P(x,A\times S),\qquad P_1(x,A)=P(x,S\times A),
\]
and the averaged kernel
\[
Q=\frac{P_0+P_1}{2}
\]
defines the auxiliary lineage chain obtained by following a typical random branch [2012.04741, 2105.09619]. This auxiliary chain is central in many-to-one identities, ergodic theory, and fluctuation results.

## 2. Markov properties, factorization, and neighboring formalisms

For tree-indexed processes in the rooted sense, the branching-Markov property states that descendants evolve conditionally independently across different parents once the current generation is fixed. In the binary case, conditionally on the \(\sigma\)-field generated by the first \(k\) generations, the pairs \(\{(X_{i0},X_{i1}):i\in G_k\}\) are independent and each has conditional distribution \(P(X_i,\cdot)\) [2012.04741]. The associated many-to-one formulas identify expectations and second moments of generation sums in terms of iterates of the lineage kernel \(Q\) [2012.04741, 2105.09619].

A different but related formalism assigns variables to the vertices of a finite tree \(\mathcal G=(\mathcal M,\mathcal E)\) and defines a Markov chain on the tree by edgewise conditional independence. For every edge \((i,j)\in\mathcal E\),
\[
X_j\;\perp\!\!\!\perp\;X_{\mathcal B(i\leftarrow j)\setminus\{i\}}\mid X_i,
\]
and this is equivalent to a global Markov property based on graph separation:
\[
S\text{ separates }A,B \Longrightarrow X_A\;\perp\!\!\!\perp\;X_B\mid X_S
\]
[2307.15844]. In the strictly positive directed case this yields the familiar factorization
\[
P(x_1,\dots,x_m)=P(x_{\mathrm{root}})\prod_{i\ne \mathrm{root}}P(x_i\mid x_{\mathrm{pa}(i)}).
\]

The literature also contains a block version. An \(o\)-block Markov chain on a rooted tree satisfies
\[
\mu[\xi\text{ on }S_o(x)\mid \xi\text{ on }V\setminus T_o'(x)]
=
\mu[\xi\text{ on }S_o(x)\mid \xi(x)],
\]
so that the joint law of the immediate children depends only on the parent state and is independent of the exterior of the descendant subtree [2008.09978]. Measures that are block Markov chains for every choice of root are classical Markov chains on the tree, and they form a strict subclass of Markov random fields; in the one-dimensional case the class of block Markov chains coincides with the class of Markov chains [2008.09978].

The terminology is not uniform across the literature. A separate line studies Markov chains on rooted trees where the tree is the state space rather than the index set. For almost upper-directed kernels on infinite locally finite trees, every irreducible kernel has some invariant measures, and an \(h\)-invariant measure is given by a determinantal formula; recurrence and positive recurrence are then analyzed through explicit criteria and a leaf addition algorithm [2411.07158]. This is a distinct problem from labeling a tree by a Markovian field.

## 3. Ergodicity, empirical averages, and law-of-large-numbers phenomena

For arbitrary finite subsets \(A_n\subset\mathbb T_\infty\), empirical averages are defined by
\[
M_{A_n}(f)=\sum_{u\in A_n}f(X_u),\qquad
\overline M_{A_n}(f)=\frac{1}{|A_n|}M_{A_n}(f).
\]
Under a geometrical assumption requiring that two uniformly sampled vertices in \(A_n\) are typically far apart, together with either tightness of the height of their latest common ancestor or a strong-ergodicity condition on the kernel \(Q\), one has
\[
\overline M_{A_n}(f)\xrightarrow[n\to\infty]{\Pr} c_f
\]
[2403.16505]. In the ergodic case with unique invariant law \(p\) and bounded continuous \(f\), the limit is \(c_f=p(f)\) [2403.16505].

Several tree families satisfy the geometric hypotheses naturally. Any infinite rooted tree of degree \(\le D<\infty\) satisfies the required distance condition for every \((A_n)\) with \(|A_n|\to\infty\); spherically symmetric trees with generation sets \(A_n=G_n\) satisfy both the distance and ancestor conditions; and for a supercritical Galton–Watson tree conditioned on non-extinction, both \(A_n=G_n\) and \(A_n=T_n\) satisfy the hypotheses almost surely [2403.16505]. These examples show that the ergodic theorem is not restricted to regular binary trees.

A complementary law of large numbers arises from observing the tree-indexed chain along an independent random walk on the complete binary tree. If the random-walk observation process \(\widetilde R^n\) converges to a càdlàg Feller process \(R\), then the empirical measure process over generation \(\lfloor nt\rfloor\),
\[
\widetilde Z^n_t=\frac{1}{2^{\lfloor nt\rfloor}}\sum_{\sigma\in \mathcal T_{\lfloor nt\rfloor}}\delta_{X^n_\sigma},
\]
converges to
\[
Z_t=\delta_{\mathcal L(R_t)}
\]
in \(D_{\mathcal P(E)}[0,\infty)\) [1406.3768]. The mechanism is asymptotic independence of two uniformly chosen leaves in a deep generation, since their most recent common ancestor lies at depth \(o(n)\) [1406.3768].

The same framework yields an optimization statement relevant for Markov-chain Monte Carlo. When the underlying chain is stationary and reversible, and \(Q\) is compact self-adjoint on \(L^2(p)\), the variance of \(\overline M_A(f)\) for an eigenfunction \(f\) is proportional to the Hosoya–Wiener polynomial
\[
H_A(\lambda)=\sum_{u,v\in A}\lambda^{d(u,v)}.
\]
Among all subtrees of fixed size \(n\), the line graph \(L_n\) uniquely minimizes \(H_T(\lambda)\) for each nonzero eigenvalue \(\lambda\in(-1,1)\setminus\{0\}\), so the ordinary chain yields the smallest non-asymptotic variance for the empirical average in that regime [2403.16505].

## 4. Fluctuations, concentration, and large deviations

For bifurcating Markov chains, additive functionals over generations or whole trees exhibit a three-regime phase structure governed by the comparison between the branching factor \(2\) and the ergodicity rate \(\alpha\) of the lineage kernel. If
\[
N_{n,\varnothing}(f)=|G_n|^{-1/2}\sum_{i\in T_n}\bigl(f_{|i|}(X_i)-\langle\mu,f_{|i|}\rangle\bigr),
\]
then under geometric ergodicity one has the following trichotomy: in the subcritical case \(2\alpha^2<1\), \(N_{n,\varnothing}(f)\) converges to a centered Gaussian limit; in the critical case \(2\alpha^2=1\), \(n^{-1/2}N_{n,\varnothing}(f)\) converges to a centered Gaussian limit with a different variance; and in the supercritical case \(2\alpha^2>1\), \((2\alpha^2)^{-n/2}N_{n,\varnothing}(f)\) is asymptotically governed by martingale limits rather than a Gaussian law [2012.04741, 2106.07711]. The proofs rely on martingale decompositions, predictable quadratic variation, and spectral projection arguments.

Moderate deviations refine the central-limit regime for one-variable additive functionals of bifurcating Markov chains. Under uniform geometric ergodicity of the lineage kernel \(Q\), if \(b_n\to\infty\) with \(b_n/\sqrt{|G_n|}\to0\) in the subcritical case, then \(b_n^{-1}N_{n,\varnothing}(f^\cdot)\) satisfies a moderate deviation principle on \(\mathbb R\) with speed \(b_n^2\) and quadratic rate function \(I(x)=\frac12 x^2/\Sigma\); in the critical case \(2\alpha^2=1\), the normalization becomes \(b_n^{-1}n^{-1/2}N_{n,\varnothing}(f^\cdot)\) under the second spectral gap assumption [2105.09619].

Concentration inequalities can be obtained without full spectral analysis. Under transportation cost-information assumptions \(H_p(C)\) on the initial law and local branching kernel, the law of the bifurcating Markov chain on \(T_n\) belongs to \(T_p(C_N)\), which yields Gaussian-type tail bounds for Lipschitz empirical means [1501.06693]. In particular, under \(H_1(C)\), the generation average
\[
\overline M_{G_n}(f)=2^{-n}\sum_{i\in G_n}f(X_i)
\]
satisfies a sub-Gaussian Laplace bound with variance proxy \(\gamma_n\), and \(\gamma_n=O(1/|G_n|)\) when \(r<\sqrt2\), where \(r=r_0+r_1\) is the sum of the Wasserstein contraction constants of the two marginals [1501.06693]. Analogous subtree-level bounds hold for \(\overline M_{T_n}(f)\), with \(\tau_n=O(1/|T_n|)\) when \(r<\sqrt2\) [1501.06693].

These results show that branching geometry does not merely change constants. It produces threshold phenomena, new critical renormalizations, and non-Gaussian supercritical limits that have no analogue for ordinary one-dimensional chains with the same kernel.

## 5. Information-theoretic, statistical, and geometric extensions

In the information-theoretic formulation of a Markov chain on a finite tree, the shared information of the entire collection admits an explicit edgewise characterization:
\[
\mathrm{SI}(X_1,\dots,X_m)=\min_{(i,j)\in\mathcal E} I(X_i\wedge X_j).
\]
The minimizer is the weakest edge of the tree, and this converts estimation of shared information into a best-arm identification problem over edges [2307.15844]. When the joint law is unknown but the tree is known, the paper develops a uniform-sampling multiarmed bandit algorithm based on empirical mutual information and establishes error and sample-complexity bounds [2307.15844].

Statistical inference for hidden models has also been developed in the tree-indexed setting. For hidden Markov models indexed by the complete binary tree, with hidden branching Markov chain \(X\) and conditionally independent observations \(Y_u\sim G(X_u,\cdot)\), the maximum-likelihood estimator based on \(Y_{T_n}\) is strongly consistent and asymptotically normal under the stated dominance, Doeblin, regularity, and smoothness assumptions [2409.06295]. The proofs use ergodic theorems for Markov chains indexed by trees with neighborhood-dependent functions, exponential forgetting of conditional initial distributions, a score martingale-array central limit theorem, and a law of large numbers for the observed information matrix [2409.06295].

Finite-state tree-indexed chains on rooted \(d\)-trees admit a large-deviation and dimension-theoretic theory via the method of types. For a transition matrix \(M\) and a positive weight matrix \(A\), the empirical average
\[
Y_n=\frac1{|T_n|}\sum_{g\in T_n,\ g\ne\epsilon}\log A_{X_{\tilde g},X_g}
\]
satisfies a Cramér-type theorem along periodic subsequences, the empirical averages converge almost surely to the unique zero of the rate function in the irreducible case, and the Hausdorff dimension of the associated Markov hom tree-shift is given by a nonlinear Perron–Frobenius variational formula [2401.05320].

A continuum extension replaces the discrete rooted tree by a Lévy tree. In that setting one constructs a Markov process indexed by the Lévy tree through the snake property, defines local time at a regular and instantaneous point, proves a Poisson decomposition of excursions away from that point, and shows that the genealogy of excursions is itself encoded by a Lévy tree called the tree coded by the local time [2411.12717]. This recovers, in particular, the excursion theory of Abraham and Le Gall for Brownian motion indexed by the Brownian tree [2411.12717].

## 6. Poisson representability and phase transitions on finite and infinite trees

A recent development studies \(\{0,1\}\)-valued tree-indexed Markov chains through Poisson representations. For a rooted tree \(T=(V(T),E(T))\), parameters \(r,p\in[0,1]\), and a parent–child edge \((u,v)\), the transition matrix is
\[
P_{u,v}(x,y)=(1-p)\delta_{x,y}+p\,[r\,1_{y=1}+(1-r)\,1_{y=0}],
\]
equivalently
\[
P_{u,v}=
\begin{pmatrix}
1-pr & pr\\[6pt]
p(1-r) & 1-p(1-r)
\end{pmatrix},
\]
with \(X(o)\sim\mathrm{Bernoulli}(r)\); along each parent-child link the state flips to an independent fresh \(\mathrm{Bernoulli}(r)\) with probability \(p\) and otherwise retains its parent’s value [2501.14428].

Poisson-representability is formulated on a countable index set \(S\). A \(\{0,1\}\)-valued process \(X=(X_i:i\in S)\) belongs to the class \(\mathcal R\) if there exists a measure \(v\) on nonempty subsets of \(S\) such that, for a Poisson point process \(Y\) on \(\mathcal P(S)\setminus\{\varnothing\}\) with intensity \(v\),
\[
X_i=1_{\{\,i\in\bigcup_{B\in Y}B\,\}}.
\]
Equivalently, the unique signed intensity measure satisfies
\[
P\bigl(X(A)=0\bigr)=\exp\!\bigl[-v(\mathcal S_A)\bigr],
\qquad
\mathcal S_A=\{B\subset S:B\ne\varnothing,\ B\cap A\ne\varnothing\},
\]
and \(X\in\mathcal R\) precisely when this \(v\) is nonnegative on all finite subsets [2501.14428].

For finite trees that are not simple paths, representability is controlled by two thresholds depending only on the leaf boundary size \(k=|\partial T|\) and maximal degree \(m=\max\{\deg(v):v\in V(T)\}\). The paper defines \(r_0(k)\) from complementary Bell numbers and \(r_1(m)\) from the largest negative root of \(\mathrm{Li}_{1-m}\), and proves:
- if \(r>r_0(k)\) and \(p\to0^+\), then \(X\in\mathcal R\);
- if \(r<r_0(k)\) and \(p\to0^+\), then \(X\notin\mathcal R\);
- if \(r>r_1(m)\) and \(p\to1^-\), then \(X\in\mathcal R\);
- if \(r<r_1(m)\) and \(p\to1^-\), then \(X\notin\mathcal R\).

Thus every non-path finite tree exhibits a genuine phase transition in \((r,p)\) for Poisson-representability [2501.14428].

The infinite case is not degenerate. For the “octopus” tree \(T_m\), consisting of one central vertex of degree \(m\) with \(m\) infinite rays attached, there is again a phase transition solely in \(r\): if \(r<r_1(m)\) then \(X\notin\mathcal R\) for all \(p\in(0,1)\); there exists \(0<r_2(m)<1\), independent of \(p\), such that if \(r\ge r_2(m)\) then \(X\in\mathcal R\) for all \(p\); and for \(m=3\) the critical value is exactly \(r_2(3)=r_1(3)=1/2\), independent of \(p\) [2501.14428].

The proofs are combinatorial and rely on Möbius inversion for the signed intensity measure, an explicit inclusion–exclusion formula for \(v(S)\), Taylor expansion around \(p=0\) and \(p=1\), and restriction lemmas extended to signed measures [2501.14428]. The same paper also sharpens earlier results by showing that any signed measure \(v\) arising from a Markov field has support only on connected subsets, and by extending line-chain formulas to arbitrary tree-indexed chains [2501.14428].

Tree-indexed Markov chains are therefore not a single theorem but a broad research area in which branching geometry, ancestral overlap, and local kernel structure govern ergodic averages, fluctuations, concentration, inference, information measures, and representability. The common theme is that the tree replaces linear time by a partially ordered genealogy, and this change introduces new thresholds, new factorization phenomena, and new links to graphical models, large deviations, and random-tree geometry.

Source: https://www.emergentmind.com/topics/tree-indexed-markov-chains