Coalescent Point Process Overview
- Coalescent point process is a stochastic framework that defines the genealogy of extant individuals through i.i.d. node depths in branching models.
- It enables explicit likelihood computations and fast simulations by associating coalescence times to pairs of individuals using statistical properties.
- Extensions include marked, multi-type, and varying-environment CPPs, linking discrete models with continuum limits such as Feller diffusion and CSBP.
A coalescent point process (CPP) is a stochastic framework that encodes the genealogy of extant individuals in branching, birth–death, or population models by associating to each neighboring pair of individuals a coalescence time (“node depth,” “split time”)—the time back to their most recent common ancestor (MRCA). This process is fundamentally characterized by sequences of independent and identically distributed (i.i.d.) random variables in the simplest settings, admits multiple generalizations to multi-type, marked, and varying-environment populations, and under scaling limits connects to continuum branching mechanisms and critical diffusions. The CPP approach enables explicit likelihoods, fast simulations, and tractable inference for phylogenies and population genetic data, offering deep connections with continuous-state branching, Lévy processes, and modern coalescent theory.
1. Formal Definition and Basic Construction
A canonical CPP construction arises in reconstructed trees of birth–death models or splitting trees. Given stem age , one sequentially draws i.i.d. nonnegative random variables with common distribution , terminating at the first index with . The vector (all less than ) determines the node depths of an ultrametric tree with leaves at time (Lambert et al., 2013, Burden et al., 7 Jan 2026). The topology is always uniform on ranked oriented trees (URT). For each pair of individuals in a planar embedding, their coalescence time 0 is given by
1
where 2 is the coalescence time between individuals 3 and 4 (Lambert et al., 2011, Blancas et al., 2022). The process 5—the coalescent point process—carries the full genealogy.
2. Distributional Properties and Model Classes
The law of node depths in a CPP arises from the underlying diversification (speciation/extinction) rates. For general time-dependent rates 6 (speciation at 7, extinction at 8 depending only on non-heritable traits 9), the density 0 of node depths has an explicit Volterra representation, and all 1 are i.i.d. (Lambert et al., 2013):
2
where 3 is an explicit inverse-tail function. In the standard birth–death case (4 constant), 5 and 6 simplify considerably.
Only in models where diversification rates depend on non-heritable quantities (time, age) does the reconstructed tree correspond to a CPP; if rates depend on the current diversity 7, only the ranked tree shape is URT, but the node depths are no longer independent (tree is not a CPP). For heritable trait-dependent rates, neither URT nor CPP structures hold (Lambert et al., 2013).
3. Extensions: Marked, Multi-type, and Varying Environment CPPs
Marked CPPs
In models with mutations (e.g., neutral mutations at birth), the CPP is enriched by marking each node depth where a mutation event occurred, encoding the history of mutations as a point measure along each branch (Delaporte, 2013). In large-population scaling limits (splitting tree with rare/independent mutations), the marked CPP converges to a Poisson point process on 8, where 9 denotes the set of finite point measures on node depths and mutation marks. The limit law is explicitly characterized using excursion and ladder-height theory of spectrally positive Lévy processes. In the Brownian (critical branching) case, the node depths are distributed as the excursion depths of Brownian motion, with Poissonian marks specifying mutation events.
Multi-type CPPs
For multi-type branching processes, the CPP encodes not only coalescence times but also the type along each ancestral lineage. The construction involves recording the ancestral index process and types in a planar embedding, and the statistical properties (e.g., MRCA time, coalescence between same-type individuals) can be explicitly calculated, notably in linear-fractional cases (Popovic et al., 2013).
Varying-Environment CPPs
When the offspring distribution varies across generations (“varying environment”), the sequence 0 is not Markov in general, but there exists a finite-dimensional Markov process 1 that fully reconstructs the genealogy. In the special linear-fractional case, 2 becomes i.i.d., and thus Markov (Blancas et al., 2022).
4. Limit Theorems and Scaling Relations
Scaling limits of branching processes lead to continuum analogues of CPPs:
- Feller diffusion limit: CPPs constructed from birth–death trees converge (after rescaling) to the genealogical structure of the Feller diffusion. Node height distributions admit explicit forms mirroring the Bernoulli-sampled BD process, and Poisson (3-)sampling schemes transfer to the diffusion setting (Burden et al., 7 Jan 2026).
- CSBP limit: For critical or subcritical Bienaymé–Galton–Watson (BGW) processes, scaling the population and time yields continuous-state branching processes coded by spectrally positive Lévy processes. The limit provides a Markov process 4 on point measures, recovering genealogies and coalescence times in the continuum tree (Lambert et al., 2011).
- Kingman/Brownian connection: Under suitable sampling and scaling, the genealogy of the Fleming–Viot process converges to the Brownian CPP, with coalescence times of 5 sampled individuals becoming i.i.d. 6 in the small-time limit (Lambert et al., 2016).
5. Likelihoods, Simulation, and Efficient Inference
The core advantage of the CPP framework is analytic and computational tractability:
- The likelihood of a reconstructed tree is given by a product of node depth densities, admitting closed-form or semi-analytic expressions for broad model classes (Lambert et al., 2013).
- Simulation of reconstructed trees proceeds via sampling 7 i.i.d. node depths and performing 8 grafting steps.
- The URT property means that shape and branch length information can be separated, simplifying hypothesis testing and model fitting.
- Extensions (e.g., incomplete sampling, mass extinction, multi-type populations) are naturally handled within the CPP framework (Lambert et al., 2013, Popovic et al., 2013).
- Fast simulation for large trees with high extinction or under complex macroevolutionary scenarios is feasible—orders-of-magnitude faster than full forward-time process simulation.
6. Pfaffian and Spatial Generalizations
In spatial settings, such as coalescing random walks or particles on a lattice or Brownian motion in 9, the set of surviving particles at time 0 forms a coalescent point process whose 1-point correlation functions can be represented as Pfaffians with explicit antisymmetric kernels. The checkerboard duality yields a powerful framework valid for general (possibly inhomogeneous, non-symmetric) models, encompassing both discrete and continuum cases (Śniady, 26 Feb 2026).
7. Biological, Probabilistic, and Methodological Implications
CPPs provide a natural interface between forward-time macroevolutionary models and reconstructed genealogies, yielding a unified language for modeling, statistical inference, and simulation. The marked process approach directly links to population genetics quantities (e.g., site-frequency spectra, allelic partitions) and extends Ewens-sampling-like results to complex, non-constant population scenarios (Delaporte, 2013). The general framework, resting on branching and Lévy process theory, underpins connections among classical coalescent, critical branching, Feller diffusion, and continuum genealogies, facilitating analysis across discrete and continuum timescales.
| Model/System | CPP Structure | Key Properties |
|---|---|---|
| Birth–Death | i.i.d. node depths, explicit 2 | URT, closed-form likelihoods, efficient simulation |
| Multi-type GW | Functional of Markov chain, types tracked | Explicit coalescence/statistics in linear-fractional case |
| Marked/mutations | Poisson-marked, ladder-height via Lévy | Inhomogeneous regenerative set of mutations |
| Varying-env GW | Not Markov unless linear-fractional | Finite-dimensional Markov process 3 for genealogy |
| Feller/CSBP limit | Continuum genealogical tree | Explicit scaling, Poisson/mixture sampling, CPP limits |
The CPP framework unifies discrete and continuum genealogy, supports rapid inference and simulation, and enables tractable analysis of models beyond constant-size or exchangeable-population assumptions, fostering broad applicability in phylogenetics, population genetics, and stochastic process theory (Lambert et al., 2013, Lambert et al., 2011, Delaporte, 2013, Burden et al., 7 Jan 2026, Blancas et al., 2022, Śniady, 26 Feb 2026).