Papers
Topics
Authors
Recent
Search
2000 character limit reached

Colless Index in Phylogenetics

Updated 10 July 2026
  • The Colless index is a measure of asymmetry in rooted bifurcating trees, defined as the sum of absolute differences between descendant counts of child subtrees at each internal node.
  • It reaches a maximum with caterpillar trees and a minimum with maximally balanced trees, with normalized variants enabling comparisons across different tree sizes.
  • Widely used in empirical phylogenetics, the Colless index underpins null model tests such as the Yule model and has spurred generalizations and alternative formulations for broader applications.

The Colless index is a classical measure of imbalance for rooted bifurcating phylogenetic trees. In its standard form, it assigns to each internal node the absolute difference between the numbers of descendant leaves in its two child subtrees, and then sums these local imbalances over the tree. It is one of the oldest and most widely used balance indices in phylogenetics, where it serves both as a descriptive statistic of tree shape and as a test statistic for null models such as the Yule model of speciation (Cardona et al., 2012).

1. Definition and normalizations

For a rooted bifurcating tree TT, the Colless index is commonly written as

C(T)=∑v∈V˚(T)∣κT(v1)−κT(v2)∣\mathcal{C}(T)=\sum_{v\in \mathring{V}(T)} |\kappa_T(v_1)-\kappa_T(v_2)|

where V˚(T)\mathring{V}(T) is the set of internal nodes, v1v_1 and v2v_2 are the children of vv, and κT(w)\kappa_T(w) is the number of descendant leaves of ww (Coronado et al., 2019). Equivalent formulations use Vint(T)V_{\mathrm{int}}(T) for the internal-node set and balT(v)=∣κT(v1)−κT(v2)∣bal_T(v)=|\kappa_T(v_1)-\kappa_T(v_2)|, so that C(T)=∑v∈V˚(T)∣κT(v1)−κT(v2)∣\mathcal{C}(T)=\sum_{v\in \mathring{V}(T)} |\kappa_T(v_1)-\kappa_T(v_2)|0 (Mir et al., 2012).

The index measures asymmetry at each bifurcation. A larger value indicates a more lopsided tree, whereas a smaller value indicates a more even pattern of subdivision. In the standard theoretical literature, the unnormalized sum is the default object of study.

A normalized variant is also used in empirical comparisons. For a rooted strict binary tree with C(T)=∑v∈V˚(T)∣κT(v1)−κT(v2)∣\mathcal{C}(T)=\sum_{v\in \mathring{V}(T)} |\kappa_T(v_1)-\kappa_T(v_2)|1 leaves, if C(T)=∑v∈V˚(T)∣κT(v1)−κT(v2)∣\mathcal{C}(T)=\sum_{v\in \mathring{V}(T)} |\kappa_T(v_1)-\kappa_T(v_2)|2 and C(T)=∑v∈V˚(T)∣κT(v1)−κT(v2)∣\mathcal{C}(T)=\sum_{v\in \mathring{V}(T)} |\kappa_T(v_1)-\kappa_T(v_2)|3 denote the numbers of leaves in the left and right subtrees of inner node C(T)=∑v∈V˚(T)∣κT(v1)−κT(v2)∣\mathcal{C}(T)=\sum_{v\in \mathring{V}(T)} |\kappa_T(v_1)-\kappa_T(v_2)|4, then

C(T)=∑v∈V˚(T)∣κT(v1)−κT(v2)∣\mathcal{C}(T)=\sum_{v\in \mathring{V}(T)} |\kappa_T(v_1)-\kappa_T(v_2)|5

Under this normalization, C(T)=∑v∈V˚(T)∣κT(v1)−κT(v2)∣\mathcal{C}(T)=\sum_{v\in \mathring{V}(T)} |\kappa_T(v_1)-\kappa_T(v_2)|6 for a perfectly balanced complete binary tree and C(T)=∑v∈V˚(T)∣κT(v1)−κT(v2)∣\mathcal{C}(T)=\sum_{v\in \mathring{V}(T)} |\kappa_T(v_1)-\kappa_T(v_2)|7 for a completely imbalanced comb tree (Keller-Schmidt et al., 2011). This normalization is useful when comparing trees of different sizes, but the main combinatorial and probabilistic results are usually stated for the unnormalized index.

Two scope conditions recur throughout the literature. First, the classical index is defined for rooted binary or rooted bifurcating trees. Second, its dependence is purely topological: it is invariant under isomorphisms and leaf relabelings (Mir et al., 2012).

2. Extremal values and minimizers

The maximal Colless index on rooted binary trees with C(T)=∑v∈V˚(T)∣κT(v1)−κT(v2)∣\mathcal{C}(T)=\sum_{v\in \mathring{V}(T)} |\kappa_T(v_1)-\kappa_T(v_2)|8 leaves is uniquely achieved by the caterpillar tree, with value

C(T)=∑v∈V˚(T)∣κT(v1)−κT(v2)∣\mathcal{C}(T)=\sum_{v\in \mathring{V}(T)} |\kappa_T(v_1)-\kappa_T(v_2)|9

This gives the classical upper extreme of binary-tree imbalance (Fischer et al., 2019).

The minimal value is substantially more subtle. Let

VËš(T)\mathring{V}(T)0

for rooted bifurcating trees with VËš(T)\mathring{V}(T)1 leaves. The minimum satisfies the recurrence

VËš(T)\mathring{V}(T)2

equivalently

VËš(T)\mathring{V}(T)3

A closed expression is available in terms of the binary expansion VËš(T)\mathring{V}(T)4, VËš(T)\mathring{V}(T)5: VËš(T)\mathring{V}(T)6 The same minimum also admits a formulation through the Takagi, or Blancmange, curve, revealing a fractal structure in the sequence of minimal values (Coronado et al., 2019).

Maximally balanced trees always attain this minimum, but they are generally not the only minimizers. The literature identifies two distinguished extremal classes among Colless-minimal trees: the maximally balanced trees and the greedy-from-the-bottom trees. All trees with minimum Colless index also have minimum Sackin index, but the converse is false, so the Colless criterion is strictly finer in this sense (Coronado et al., 2019).

These results matter for normalization. Once both the minimum VËš(T)\mathring{V}(T)7 and the caterpillar maximum VËš(T)\mathring{V}(T)8 are known, one can define

VËš(T)\mathring{V}(T)9

which rescales the index across leaf numbers while respecting the actual minimum for each v1v_10 (Coronado et al., 2019).

3. Behavior under stochastic tree models

A central use of the Colless index is as a statistic under random tree models. Under the Yule model, the expected value is known exactly. If v1v_11 denotes the Colless index of a binary phylogenetic tree with v1v_12 leaves generated according to the Yule model, then

v1v_13

where v1v_14 if v1v_15 is odd and v1v_16 otherwise (Mir et al., 2012).

The variance under the Yule model is also available in closed form. An explicit exact formula, valid for every v1v_17, was derived in terms of harmonic numbers, and the same analysis gives an asymptotic expansion. The principal qualitative point is that the variance of the Colless index under the Yule model grows as v1v_18 (Cardona et al., 2012).

Under the uniform model, the asymptotic scale differs with the sampling space. For tree shapes, meaning full binary rooted trees with unlabeled leaves sampled uniformly, the expected Colless index satisfies

v1v_19

where v2v_20 is determined by the singularity of the relevant generating function. In the same regime, the expected Sackin and Colless indices share the same leading asymptotic behavior, and their difference is v2v_21 (Goh et al., 2022). For labeled phylogenetic trees under the uniform model, the comparison literature cited in the cophenetic-index work states that the expected Colless and Sackin indices are both asymptotic to v2v_22 (Mir et al., 2012).

Setting Result for Colless index Source
Yule model Exact v2v_23 formula (Mir et al., 2012)
Yule model Exact variance; growth v2v_24 (Cardona et al., 2012)
Uniform tree shapes v2v_25 (Goh et al., 2022)

These formulas are not merely descriptive. They provide the baseline needed for significance testing, standardization, and model comparison across different tree sizes.

4. Role in empirical phylogenetics

Balance indices are used to test null models of evolutionary processes, and the Colless index is one of the canonical examples. Closed-form expressions for moments under the Yule model permit the calculation of standardized statistics and support hypothesis tests of whether an observed tree is consistent with Yule evolution. The availability of exact variance formulas for all v2v_26 is particularly important for tests involving trees of arbitrary size, rather than only asymptotic approximations (Cardona et al., 2012).

The index also functions as an empirical discriminator among generative models. In a branching-process model based on rare innovations and exhaustive combinations of features, the mean values and standard deviations of the Colless index for simulated trees were reported to be compatible with empirical phylogenies from TreeBASE and PANDIT. In the same study, both the Yule/ERM model and Aldous’ branching model showed larger discrepancies with the empirical data. Figure-based comparisons in that work further indicated that, after rescaling by v2v_27, only the innovation model remained close to the real data across the full range of v2v_28 considered (Keller-Schmidt et al., 2011).

A recurrent methodological implication is that the Colless index is most informative when used against an explicit null model and with size-aware normalization or standardization. This suggests that raw values alone can be misleading across heterogeneous datasets, whereas model-based calibration turns the index into a comparably interpretable test statistic.

5. Generalizations, alternatives, and modifications

A persistent limitation of the classical Colless index is that it is tailored to bifurcating trees. The multifurcating case motivated the introduction of Colless-like indices v2v_29, defined by first assigning an vv0-size

vv1

to each subtree, then evaluating a dissimilarity vv2 on the vector of child-subtree vv3-sizes at each internal node, and finally summing those local balance values: vv4 For suitable choices such as vv5 or vv6, these indices are sound in the sense that they vanish exactly on fully symmetric trees. On bifurcating trees, some of these Colless-like indices reduce to the classical Colless index up to a constant scale factor (Mir et al., 2018).

A different line of modification replaces absolute differences by squared differences. The quadratic Colless index is

vv7

This variant was proposed to overcome two drawbacks of the original index: the fact that the minimum is usually attained by non-maximally-balanced trees as well, and the analytical difficulty introduced by the absolute value. The squared version preserves the basic local-to-global construction, while making expectations and variances under the Yule and uniform models tractable and reducing ties between non-isomorphic trees (Coronado et al., 2020).

Comparisons with other balance statistics sharpen the profile of the Colless index. The total cophenetic index has a larger range and greater resolution power than Colless and makes sense for arbitrary trees, whereas the vv8 index extends naturally to phylogenetic networks, unlike Colless, whose standard definition is tree-specific [(Mir et al., 2012); (Bienvenu et al., 2020)]. A common misconception is therefore that the Colless index is a universal balance functional; the literature instead presents it as a particularly effective binary-tree statistic with well-defined but limited scope.

6. Structural decompositions and relations to other indices

Recent work has recast the Colless index as a compound statistic built from more elementary components. For each internal node vv9, let κT(w)\kappa_T(w)0 be the sizes of the larger and smaller child pending subtrees, and define

κT(w)\kappa_T(w)1

Then

κT(w)\kappa_T(w)2

so the Sackin index is the sum of these two building blocks whereas the Colless index is their difference (Knüver et al., 5 Sep 2025).

This decomposition yields several consequences. It implies κT(w)\kappa_T(w)3 for all binary trees. It also isolates the role of the smaller pending subtrees: κT(w)\kappa_T(w)4 In the cited study, κT(w)\kappa_T(w)5 is itself treated as an imbalance index, while κT(w)\kappa_T(w)6 is treated as a balance index. The same work relates κT(w)\kappa_T(w)7 to the stairs2 index and uses these components to analyze disagreement between Sackin and Colless rankings (Knüver et al., 5 Sep 2025).

This structural reinterpretation suggests a broader point. The classical Colless index is not merely a monolithic asymmetry score: it can be decomposed into interpretable contributions associated with large and small child subtrees. In that sense, the modern literature treats the index both as a historical standard and as a gateway to a more granular theory of tree balance.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Colless index.