Papers
Topics
Authors
Recent
Search
2000 character limit reached

Coalescent Point Processes

Updated 9 July 2026
  • Coalescent point processes are representations of genealogies that encode consecutive coalescence times to fully capture tree structures.
  • They enable explicit likelihood calculations and simulation schemes across discrete and continuous branching models.
  • Advanced formulations include enriched Markov chain descriptions and diffusion limits, which are pivotal for analyzing reconstructed phylogenies.

Searching arXiv for recent and foundational papers on coalescent point processes. Coalescent point processes (CPPs) are point-process representations of ultrametric genealogies in which the basic observables are coalescence times of consecutive extant individuals in a planar ordering, or equivalently node depths added sequentially in a reconstructed tree. In a planar Bienaymé–Galton–Watson genealogy or a Galton–Watson process in varying environment, if ai(n)\mathfrak a_i(n) denotes the ancestor of individual (0,i)(0,i) in generation n-n, the pairwise coalescence time is

Ci,j:=min{n1:ai(n)=aj(n)},Ai:=Ci,i+1,C_{i,j}:=\min\{\,n\ge1:\mathfrak a_i(n)=\mathfrak a_j(n)\}, \qquad A_i:=C_{i,i+1},

and the sequence A=(Ai,i1)\mathbf A=(A_i,i\ge1) is the coalescent point process. Its fundamental property is

Ci,j=max{Ai,Ai+1,,Aj1},C_{i,j}=\max\{A_i,A_{i+1},\dots,A_{j-1}\},

so A\mathbf A encodes the entire genealogy of the present-day population (Lambert et al., 2011, Blancas et al., 2022). In reconstructed phylogenies, the same notion appears through i.i.d. node depths H1,H2,H_1,H_2,\dots stopped at a stem age TT, yielding product-form likelihoods, uniform ranked tree shapes, and direct simulation schemes (Lambert et al., 2013).

1. Formal structure of the process

The discrete genealogical definition of a CPP begins with a planar embedding of a branching tree, with individuals arranged left-to-right within each generation. Consecutive extant individuals (0,i)(0,i) and (0,i)(0,i)0 determine the elementary coalescence times (0,i)(0,i)1, and all pairwise coalescence times are recovered as maxima over consecutive segments. This makes the CPP a complete encoding of the reduced ancestral tree of the standing population (Blancas et al., 2022).

A second, widely used formulation is the reconstructed-phylogeny construction. A CPP of stem age (0,i)(0,i)2 is built by sampling i.i.d. nonnegative speciation-time variables (0,i)(0,i)3 from a common density (0,i)(0,i)4 on (0,i)(0,i)5 until the first (0,i)(0,i)6. If the first exceedance occurs at the (0,i)(0,i)7-th draw, the resulting tree has (0,i)(0,i)8 tips, and conditional on (0,i)(0,i)9, the n-n0 node depths are i.i.d. copies of n-n1 truncated to n-n2. The ranked topology is then uniform ranked tree (URT) (Lambert et al., 2013).

This phylogenetic formulation is conveniently parameterized by the inverse-tail function

n-n3

with density

n-n4

In the birth–death CPP and its diffusion limits, this inverse-tail description is also the natural entry point for likelihoods, sampling transformations, and scaling arguments (Lambert et al., 2013, Burden et al., 7 Jan 2026).

2. Branching-tree genealogies and the non-Markov issue

For doubly infinite planar Bienaymé–Galton–Watson genealogies, Lambert and Popović defined the CPP and derived the basic law of the first branch length: n-n5 where n-n6 and n-n7 is the Galton–Watson process with generating function n-n8 (Lambert et al., 2011). This identifies the tail of the consecutive coalescence time with a one-particle conditional survival event.

In a Galton–Watson process in varying environment, the offspring law may change from generation to generation. Blancas–Palau considered an arbitrary large present-day population descending from an unspecified arbitrary large time in the past, with reproduction driven by a Galton–Watson process in a varying environment n-n9. The genealogy is still determined by Ci,j:=min{n1:ai(n)=aj(n)},Ai:=Ci,i+1,C_{i,j}:=\min\{\,n\ge1:\mathfrak a_i(n)=\mathfrak a_j(n)\}, \qquad A_i:=C_{i,i+1},0, but in general the process is not Markov: the law of Ci,j:=min{n1:ai(n)=aj(n)},Ai:=Ci,i+1,C_{i,j}:=\min\{\,n\ge1:\mathfrak a_i(n)=\mathfrak a_j(n)\}, \qquad A_i:=C_{i,i+1},1 given Ci,j:=min{n1:ai(n)=aj(n)},Ai:=Ci,i+1,C_{i,j}:=\min\{\,n\ge1:\mathfrak a_i(n)=\mathfrak a_j(n)\}, \qquad A_i:=C_{i,i+1},2 depends on deeper information than Ci,j:=min{n1:ai(n)=aj(n)},Ai:=Ci,i+1,C_{i,j}:=\min\{\,n\ge1:\mathfrak a_i(n)=\mathfrak a_j(n)\}, \qquad A_i:=C_{i,i+1},3 alone, because the survival-and-splitting structure of the relevant subtree depends on the whole sequence Ci,j:=min{n1:ai(n)=aj(n)},Ai:=Ci,i+1,C_{i,j}:=\min\{\,n\ge1:\mathfrak a_i(n)=\mathfrak a_j(n)\}, \qquad A_i:=C_{i,i+1},4 (Blancas et al., 2022).

A central misconception addressed in this setting is that the consecutive-coalescence sequence should inherit a simple one-step Markov structure from the constant-environment case. Blancas–Palau gave a counterexample showing that the point-measure process proposed by Lambert and Popović for constant environment does not have the Markov property in varying environment; concretely, two histories Ci,j:=min{n1:ai(n)=aj(n)},Ai:=Ci,i+1,C_{i,j}:=\min\{\,n\ge1:\mathfrak a_i(n)=\mathfrak a_j(n)\}, \qquad A_i:=C_{i,i+1},5 ending in the same Ci,j:=min{n1:ai(n)=aj(n)},Ai:=Ci,i+1,C_{i,j}:=\min\{\,n\ge1:\mathfrak a_i(n)=\mathfrak a_j(n)\}, \qquad A_i:=C_{i,i+1},6 can induce different laws for Ci,j:=min{n1:ai(n)=aj(n)},Ai:=Ci,i+1,C_{i,j}:=\min\{\,n\ge1:\mathfrak a_i(n)=\mathfrak a_j(n)\}, \qquad A_i:=C_{i,i+1},7 (Blancas et al., 2022). The same paper also isolates a broader misconception in reconstructed phylogenies: URT topology does not by itself imply a CPP representation, because node-depth correlations can persist even when ranked shapes remain uniform (Lambert et al., 2013).

3. Markovian reconstructions beyond the raw coalescence sequence

Because Ci,j:=min{n1:ai(n)=aj(n)},Ai:=Ci,i+1,C_{i,j}:=\min\{\,n\ge1:\mathfrak a_i(n)=\mathfrak a_j(n)\}, \qquad A_i:=C_{i,i+1},8 does not generally form a Markov chain, a standard strategy is to enrich the state space with additional information on surviving side branches. In the single-type Bienaymé–Galton–Watson setting, Lambert and Popović introduced counts Ci,j:=min{n1:ai(n)=aj(n)},Ai:=Ci,i+1,C_{i,j}:=\min\{\,n\ge1:\mathfrak a_i(n)=\mathfrak a_j(n)\}, \qquad A_i:=C_{i,i+1},9 of younger offshoots at level A=(Ai,i1)\mathbf A=(A_i,i\ge1)0 that still matter for future coalescences, and defined the finite point measure

A=(Ai,i1)\mathbf A=(A_i,i\ge1)1

The smallest atom location A=(Ai,i1)\mathbf A=(A_i,i\ge1)2 satisfies A=(Ai,i1)\mathbf A=(A_i,i\ge1)3, so the CPP is recovered as the first point mass of the enriched process (Lambert et al., 2011).

Blancas–Palau adapted this program to varying environment with a vector-valued chain. The state space is

A=(Ai,i1)\mathbf A=(A_i,i\ge1)4

and for each individual A=(Ai,i1)\mathbf A=(A_i,i\ge1)5,

A=(Ai,i1)\mathbf A=(A_i,i\ge1)6

Here A=(Ai,i1)\mathbf A=(A_i,i\ge1)7 counts “excess” daughters at level A=(Ai,i1)\mathbf A=(A_i,i\ge1)8 to the right of the A=(Ai,i1)\mathbf A=(A_i,i\ge1)9-th spine that survive to the present. If Ci,j=max{Ai,Ai+1,,Aj1},C_{i,j}=\max\{A_i,A_{i+1},\dots,A_{j-1}\},0 is the first nonzero coordinate of Ci,j=max{Ai,Ai+1,,Aj1},C_{i,j}=\max\{A_i,A_{i+1},\dots,A_{j-1}\},1, then

Ci,j=max{Ai,Ai+1,,Aj1},C_{i,j}=\max\{A_i,A_{i+1},\dots,A_{j-1}\},2

Theorem 3.2 states that Ci,j=max{Ai,Ai+1,,Aj1},C_{i,j}=\max\{A_i,A_{i+1},\dots,A_{j-1}\},3 is a time-homogeneous Markov chain on Ci,j=max{Ai,Ai+1,,Aj1},C_{i,j}=\max\{A_i,A_{i+1},\dots,A_{j-1}\},4, started at Ci,j=max{Ai,Ai+1,,Aj1},C_{i,j}=\max\{A_i,A_{i+1},\dots,A_{j-1}\},5, and that each state carries only finitely many integers yet suffices to update the future (Blancas et al., 2022).

The transition kernel is explicit. Conditionally on Ci,j=max{Ai,Ai+1,,Aj1},C_{i,j}=\max\{A_i,A_{i+1},\dots,A_{j-1}\},6, the next state is formed by decrementing the coordinate at Ci,j=max{Ai,Ai+1,,Aj1},C_{i,j}=\max\{A_i,A_{i+1},\dots,A_{j-1}\},7, preserving coordinates above Ci,j=max{Ai,Ai+1,,Aj1},C_{i,j}=\max\{A_i,A_{i+1},\dots,A_{j-1}\},8, and filling lower or newly created levels with independent variables Ci,j=max{Ai,Ai+1,,Aj1},C_{i,j}=\max\{A_i,A_{i+1},\dots,A_{j-1}\},9 having law

A\mathbf A0

The proof uses a stopping-line strong branching property: after conditioning on the information stored in A\mathbf A1, the remaining future subtree splits into independent Galton–Watson subtrees in varying environment, and the explicit “number of surviving sisters” laws provide the fresh randomness needed for A\mathbf A2 (Blancas et al., 2022).

4. Explicit distributions and solvable subclasses

The most tractable discrete subclass is the linear-fractional environment. If each offspring law A\mathbf A3 is linear-fractional with parameters A\mathbf A4, Blancas–Palau showed that each A\mathbf A5 is geometric with success probability

A\mathbf A6

By induction, the entries A\mathbf A7 are independent geometricA\mathbf A8, and therefore

A\mathbf A9

In particular, the variables H1,H2,H_1,H_2,\dots0 are i.i.d. and Markov (Blancas et al., 2022).

A basic constant-environment example already exhibits heavy tails. If H1,H2,H_1,H_2,\dots1 and H1,H2,H_1,H_2,\dots2, then

H1,H2,H_1,H_2,\dots3

so H1,H2,H_1,H_2,\dots4 has a discrete Pareto-type tail (Blancas et al., 2022). In a two-generation varying example with H1,H2,H_1,H_2,\dots5 geometricH1,H2,H_1,H_2,\dots6, H1,H2,H_1,H_2,\dots7 geometricH1,H2,H_1,H_2,\dots8, and thereafter deterministic one-offs,

H1,H2,H_1,H_2,\dots9

and TT0 for TT1 (Blancas et al., 2022).

Multi-type branching processes admit an analogous but richer CPP. Popović and Rivas defined a multi-type coalescent point process that records both consecutive coalescence depths and type sequences along ancestral paths, and encoded it through a Markov chain TT2 of finite type-vectors (Popovic et al., 2013). In the multi-type linear-fractional case, the formulas simplify sharply: TT3 where TT4 is the same-type coalescence time. Here the TT5 are again i.i.d., while same-type statistics depend on the type-specific parameters TT6 (Popovic et al., 2013). In a two-type illustration, symmetric and asymmetric offspring laws can produce the same law for the overall coalescence times TT7 but different laws for same-type coalescences, showing that type asymmetry modifies ancestral-type structure without necessarily changing the unlabeled CPP (Popovic et al., 2013).

5. Reconstructed phylogenies, model classes, and likelihood theory

In macro-evolutionary applications, CPPs describe reconstructed phylogenies obtained after pruning extinct lineages. Lambert and Stadler characterized a large class of forward-time diversification models for which the reconstructed tree is exactly a CPP (Lambert et al., 2013).

Forward-time dependence Reconstructed topology CPP status
TT8, TT9 with (0,i)(0,i)0 non-heritable URT exactly CPP
Dependence on species number or heritable traits URT in general not CPP

The critical distinction is between URT topology and the stronger CPP property. If speciation rates depend only on time and extinction rates depend only on time and on a non-heritable trait such as age, then the reconstructed tree is a CPP. If rates depend on the number of coexisting species or on heritable traits, the reconstructed tree may still have URT topology, but node depths generally develop correlations and the iid CPP description fails (Lambert et al., 2013).

For the generic node depth (0,i)(0,i)1, the inverse-tail function in the asymmetric-trait model is

(0,i)(0,i)2

where (0,i)(0,i)3 is the probability that a lineage born at time (0,i)(0,i)4 leaves no extant descendant at time (0,i)(0,i)5. In the time-dependent birth–death case,

(0,i)(0,i)6

and in the constant-rate case,

(0,i)(0,i)7

These formulas feed directly into the coalescent density (0,i)(0,i)8 and hence into closed-form tree likelihoods (Lambert et al., 2013).

If an oriented reconstructed tree (0,i)(0,i)9 of stem age (0,i)(0,i)00 has (0,i)(0,i)01 tips and node depths (0,i)(0,i)02, then

(0,i)(0,i)03

where (0,i)(0,i)04 is (0,i)(0,i)05 if orientation is retained and (0,i)(0,i)06 otherwise, with (0,i)(0,i)07 the number of cherries. Conditional on exactly (0,i)(0,i)08 tips,

(0,i)(0,i)09

Bernoulli (0,i)(0,i)10-sampling preserves the CPP form through

(0,i)(0,i)11

whereas uniform (0,i)(0,i)12-sampling destroys the iid node-depth property even though the ranked topology remains URT. Age-dependent extinction and mass-extinction pulses also remain tractable in the CPP formalism, with the latter implemented by an explicit thinning transformation (0,i)(0,i)13 (Lambert et al., 2013).

6. Continuous-state, Brownian, marked, and diffusion limits

The discrete CPP has several continuous analogues and scaling limits. Lambert and Popović introduced the discrete “great-aunt” measure (0,i)(0,i)14 associated with the survival-conditioned spine of a Bienaymé–Galton–Watson tree and proved its convergence, under appropriate scaling, to a continuous-state great-aunt measure (0,i)(0,i)15. The limiting CSB genealogy then carries a discretized coalescent point process with multiplicities, and the first block has tail

(0,i)(0,i)16

The corresponding enriched process (0,i)(0,i)17 is again a Markov chain on finite point measures, and a full invariance principle identifies rescaled discrete coalescent processes with their CSB limits (Lambert et al., 2011).

For general continuous-state branching processes, Foucart, Ma, and Mallein obtained a related but dynamically formulated genealogy through a flow of nested subordinators and its inverse flow. Sampling independent Poisson arrival times along the continuous population yields non-exchangeable Markovian coalescent processes called consecutive coalescents, because only consecutive blocks can merge. Their merger rates are

(0,i)(0,i)18

In the quadratic case this reproduces Popović’s Brownian coalescent point process, and in the stable case it recovers the Beta-coalescent point process of Lambert–Popović (Foucart et al., 2018).

The Brownian CPP itself appears as a Poisson point measure of comb “teeth” on (0,i)(0,i)19 with intensity

(0,i)(0,i)20

Lambert and Schertzer showed that it arises as a local limit of the Kingman coalescent under conditional sampling: if (0,i)(0,i)21 sampled individuals are conditioned to have most recent common ancestor shorter than (0,i)(0,i)22, the rescaled genealogy converges to a Brownian CPP killed at an independent Gamma(0,i)(0,i)23 random variable. In this limit, the consecutive coalescence times of the (0,i)(0,i)24 sampled individuals, normalized by the killing level, are i.i.d. Uniform(0,i)(0,i)25 (Lambert et al., 2016).

CPPs also support explicit diffusion limits of branching trees. The Feller diffusion was studied as the limit of a CPP whose node-height density is skewed toward zero. Under near-critical scaling of birth and death rates,

(0,i)(0,i)26

the rescaled population converges to the Feller diffusion with generator

(0,i)(0,i)27

Poisson sampling of the diffusion limit preserves the CPP form: the reconstructed tree of a Poisson-sampled Feller diffusion has node-height distribution

(0,i)(0,i)28

which has the same algebraic form as the Bernoulli-sampled birth–death case (Burden et al., 7 Jan 2026).

Finally, neutral mutations at birth lead to marked coalescent point processes. In splitting trees with clonally inherited types and mutation marks at birth, the marked CPP records for each extant lineage both its coalescence depth and the mutation levels along that lineage. Under large-population rescaling, and under either a rare-mutation regime or a more general lifetime-dependent regime, the marked CPP converges to a Poisson point measure on (0,i)(0,i)29. In the rare-mutation case, the limit is described through the marked ladder-height process of a spectrally positive Lévy process, and in the critical branching process with exponential lifetimes the limiting object is the Poisson point process of Brownian excursion depths with Poissonian mutations on lineages (Delaporte, 2013).

A plausible implication of these constructions is that CPPs are not a single model but a family of equivalent genealogical encodings whose utility depends on the ambient category: planar Galton–Watson trees, reconstructed phylogenies, Lévy splitting trees, CSBPs, and diffusion limits each preserve the consecutive-coalescence representation while modifying the state enrichment, independence structure, and sampling transformations.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Coalescent Point Processes.