Papers
Topics
Authors
Recent
Search
2000 character limit reached

Coalescent Point Process Overview

Updated 13 April 2026
  • Coalescent point process is a stochastic framework that defines the genealogy of extant individuals through i.i.d. node depths in branching models.
  • It enables explicit likelihood computations and fast simulations by associating coalescence times to pairs of individuals using statistical properties.
  • Extensions include marked, multi-type, and varying-environment CPPs, linking discrete models with continuum limits such as Feller diffusion and CSBP.

A coalescent point process (CPP) is a stochastic framework that encodes the genealogy of extant individuals in branching, birth–death, or population models by associating to each neighboring pair of individuals a coalescence time (“node depth,” “split time”)—the time back to their most recent common ancestor (MRCA). This process is fundamentally characterized by sequences of independent and identically distributed (i.i.d.) random variables in the simplest settings, admits multiple generalizations to multi-type, marked, and varying-environment populations, and under scaling limits connects to continuum branching mechanisms and critical diffusions. The CPP approach enables explicit likelihoods, fast simulations, and tractable inference for phylogenies and population genetic data, offering deep connections with continuous-state branching, Lévy processes, and modern coalescent theory.

1. Formal Definition and Basic Construction

A canonical CPP construction arises in reconstructed trees of birth–death models or splitting trees. Given stem age tt, one sequentially draws i.i.d. nonnegative random variables H1,H2,H_1, H_2, \dots with common distribution ff, terminating at the first index NN with HN>tH_N > t. The vector (H1,,HN1)(H_1, \dots, H_{N-1}) (all less than tt) determines the node depths of an ultrametric tree with n=Nn=N leaves at time tt (Lambert et al., 2013, Burden et al., 7 Jan 2026). The topology is always uniform on ranked oriented trees (URT). For each pair of individuals i,ji, j in a planar embedding, their coalescence time H1,H2,H_1, H_2, \dots0 is given by

H1,H2,H_1, H_2, \dots1

where H1,H2,H_1, H_2, \dots2 is the coalescence time between individuals H1,H2,H_1, H_2, \dots3 and H1,H2,H_1, H_2, \dots4 (Lambert et al., 2011, Blancas et al., 2022). The process H1,H2,H_1, H_2, \dots5—the coalescent point process—carries the full genealogy.

2. Distributional Properties and Model Classes

The law of node depths in a CPP arises from the underlying diversification (speciation/extinction) rates. For general time-dependent rates H1,H2,H_1, H_2, \dots6 (speciation at H1,H2,H_1, H_2, \dots7, extinction at H1,H2,H_1, H_2, \dots8 depending only on non-heritable traits H1,H2,H_1, H_2, \dots9), the density ff0 of node depths has an explicit Volterra representation, and all ff1 are i.i.d. (Lambert et al., 2013):

ff2

where ff3 is an explicit inverse-tail function. In the standard birth–death case (ff4 constant), ff5 and ff6 simplify considerably.

Only in models where diversification rates depend on non-heritable quantities (time, age) does the reconstructed tree correspond to a CPP; if rates depend on the current diversity ff7, only the ranked tree shape is URT, but the node depths are no longer independent (tree is not a CPP). For heritable trait-dependent rates, neither URT nor CPP structures hold (Lambert et al., 2013).

3. Extensions: Marked, Multi-type, and Varying Environment CPPs

Marked CPPs

In models with mutations (e.g., neutral mutations at birth), the CPP is enriched by marking each node depth where a mutation event occurred, encoding the history of mutations as a point measure along each branch (Delaporte, 2013). In large-population scaling limits (splitting tree with rare/independent mutations), the marked CPP converges to a Poisson point process on ff8, where ff9 denotes the set of finite point measures on node depths and mutation marks. The limit law is explicitly characterized using excursion and ladder-height theory of spectrally positive Lévy processes. In the Brownian (critical branching) case, the node depths are distributed as the excursion depths of Brownian motion, with Poissonian marks specifying mutation events.

Multi-type CPPs

For multi-type branching processes, the CPP encodes not only coalescence times but also the type along each ancestral lineage. The construction involves recording the ancestral index process and types in a planar embedding, and the statistical properties (e.g., MRCA time, coalescence between same-type individuals) can be explicitly calculated, notably in linear-fractional cases (Popovic et al., 2013).

Varying-Environment CPPs

When the offspring distribution varies across generations (“varying environment”), the sequence NN0 is not Markov in general, but there exists a finite-dimensional Markov process NN1 that fully reconstructs the genealogy. In the special linear-fractional case, NN2 becomes i.i.d., and thus Markov (Blancas et al., 2022).

4. Limit Theorems and Scaling Relations

Scaling limits of branching processes lead to continuum analogues of CPPs:

  • Feller diffusion limit: CPPs constructed from birth–death trees converge (after rescaling) to the genealogical structure of the Feller diffusion. Node height distributions admit explicit forms mirroring the Bernoulli-sampled BD process, and Poisson (NN3-)sampling schemes transfer to the diffusion setting (Burden et al., 7 Jan 2026).
  • CSBP limit: For critical or subcritical Bienaymé–Galton–Watson (BGW) processes, scaling the population and time yields continuous-state branching processes coded by spectrally positive Lévy processes. The limit provides a Markov process NN4 on point measures, recovering genealogies and coalescence times in the continuum tree (Lambert et al., 2011).
  • Kingman/Brownian connection: Under suitable sampling and scaling, the genealogy of the Fleming–Viot process converges to the Brownian CPP, with coalescence times of NN5 sampled individuals becoming i.i.d. NN6 in the small-time limit (Lambert et al., 2016).

5. Likelihoods, Simulation, and Efficient Inference

The core advantage of the CPP framework is analytic and computational tractability:

  • The likelihood of a reconstructed tree is given by a product of node depth densities, admitting closed-form or semi-analytic expressions for broad model classes (Lambert et al., 2013).
  • Simulation of reconstructed trees proceeds via sampling NN7 i.i.d. node depths and performing NN8 grafting steps.
  • The URT property means that shape and branch length information can be separated, simplifying hypothesis testing and model fitting.
  • Extensions (e.g., incomplete sampling, mass extinction, multi-type populations) are naturally handled within the CPP framework (Lambert et al., 2013, Popovic et al., 2013).
  • Fast simulation for large trees with high extinction or under complex macroevolutionary scenarios is feasible—orders-of-magnitude faster than full forward-time process simulation.

6. Pfaffian and Spatial Generalizations

In spatial settings, such as coalescing random walks or particles on a lattice or Brownian motion in NN9, the set of surviving particles at time HN>tH_N > t0 forms a coalescent point process whose HN>tH_N > t1-point correlation functions can be represented as Pfaffians with explicit antisymmetric kernels. The checkerboard duality yields a powerful framework valid for general (possibly inhomogeneous, non-symmetric) models, encompassing both discrete and continuum cases (Śniady, 26 Feb 2026).

7. Biological, Probabilistic, and Methodological Implications

CPPs provide a natural interface between forward-time macroevolutionary models and reconstructed genealogies, yielding a unified language for modeling, statistical inference, and simulation. The marked process approach directly links to population genetics quantities (e.g., site-frequency spectra, allelic partitions) and extends Ewens-sampling-like results to complex, non-constant population scenarios (Delaporte, 2013). The general framework, resting on branching and Lévy process theory, underpins connections among classical coalescent, critical branching, Feller diffusion, and continuum genealogies, facilitating analysis across discrete and continuum timescales.


Model/System CPP Structure Key Properties
Birth–Death i.i.d. node depths, explicit HN>tH_N > t2 URT, closed-form likelihoods, efficient simulation
Multi-type GW Functional of Markov chain, types tracked Explicit coalescence/statistics in linear-fractional case
Marked/mutations Poisson-marked, ladder-height via Lévy Inhomogeneous regenerative set of mutations
Varying-env GW Not Markov unless linear-fractional Finite-dimensional Markov process HN>tH_N > t3 for genealogy
Feller/CSBP limit Continuum genealogical tree Explicit scaling, Poisson/mixture sampling, CPP limits

The CPP framework unifies discrete and continuum genealogy, supports rapid inference and simulation, and enables tractable analysis of models beyond constant-size or exchangeable-population assumptions, fostering broad applicability in phylogenetics, population genetics, and stochastic process theory (Lambert et al., 2013, Lambert et al., 2011, Delaporte, 2013, Burden et al., 7 Jan 2026, Blancas et al., 2022, Śniady, 26 Feb 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Coalescent Point Process.