Papers
Topics
Authors
Recent
Search
2000 character limit reached

Colless minus Sackin Index

Updated 10 July 2026
  • Colless minus Sackin is a tree statistic defined as the difference between the Colless and Sackin indices, isolating the contribution of smaller pending subtrees.
  • It is derived from an algebraic decomposition into N_a and N_b components, providing a clear interpretation of tree balance versus imbalance.
  • The index informs extremal tree-shape analysis, linking minimal differences in balanced trees with maximal values in caterpillar trees.

Searching arXiv for the primary and related papers to ground the article in current literature. Colless minus Sackin is the tree statistic defined by the difference between the Colless index and the Sackin index on rooted binary trees. In recent work on the “building blocks” of classical balance indices, this difference is identified as a particularly simple derived quantity: for a rooted binary tree TT, C(T)S(T)C(T)-S(T) equals negative twice the sum of the smaller pending-subtree sizes over all internal nodes, and thus isolates an elementary component latent in both classical indices (Knüver et al., 5 Sep 2025). This representation places the statistic at the intersection of imbalance theory, extremal tree-shape analysis, and the comparative study of phylogenetic balance indices. It also clarifies why the difference is not merely an ad hoc linear combination, but a structurally meaningful index in its own right (Knüver et al., 5 Sep 2025).

1. Formal definition and algebraic decomposition

For a rooted binary tree TT, let V˚(T)\mathring V(T) denote the inner vertices and V1(T)V^1(T) the leaves. For each inner vertex vv, let nvanvbn_{v_a}\ge n_{v_b} be the sizes of its two maximal pending subtrees. The Sackin index is defined by

S(T)=vV(T)nv=xV1(T)δx,S(T)=\sum_{v\in V(T)} n_v=\sum_{x\in V^1(T)} \delta_x,

where nvn_v is the number of leaves in the pending subtree rooted at vv, and C(T)S(T)C(T)-S(T)0 is the depth of leaf C(T)S(T)C(T)-S(T)1 (Knüver et al., 5 Sep 2025). Using C(T)S(T)C(T)-S(T)2, one obtains

C(T)S(T)C(T)-S(T)3

with

C(T)S(T)C(T)-S(T)4

The Colless index is

C(T)S(T)C(T)-S(T)5

and under the convention C(T)S(T)C(T)-S(T)6,

C(T)S(T)C(T)-S(T)7

(Knüver et al., 5 Sep 2025).

The decomposition

C(T)S(T)C(T)-S(T)8

is the central algebraic fact. It shows that the Sackin and Colless indices are themselves compound expressions built from two more elementary tree statistics, C(T)S(T)C(T)-S(T)9 and TT0 (Knüver et al., 5 Sep 2025). Within this framework, Colless minus Sackin is not a secondary comparison between two unrelated indices; it is the exact extraction of one underlying component.

2. The identity TT1

The difference

TT2

admits the explicit simplification

TT3

Equivalently,

TT4

This is given as Lemma 1 in “Revealing the building blocks of tree balance: fundamental units of the Sackin and Colless Indices” (Knüver et al., 5 Sep 2025).

This identity yields the most direct interpretation of Colless minus Sackin: it is negative twice the total size of the smaller child-subtree over all internal nodes. In the terminology of that work, the statistic is controlled entirely by the elementary “building block” TT5 (Knüver et al., 5 Sep 2025). The converse difference is

TT6

An earlier asymptotic study on uniform random tree shapes introduced the same difference in equivalent form,

TT7

which is nonnegative (Goh et al., 2022). This is the same quantity as TT8, expressed with different notation. Accordingly, Colless minus Sackin is exactly the negative of this nonnegative accumulation of local smaller-subtree sizes.

3. Index-theoretic status: imbalance and balance

A central theorem establishes that, for every TT9,

V˚(T)\mathring V(T)0

is an imbalance index, whereas

V˚(T)\mathring V(T)1

are balance indices on V˚(T)\mathring V(T)2 (Knüver et al., 5 Sep 2025). Since a balance index is the negative of an imbalance index, this is consistent with V˚(T)\mathring V(T)3.

The extremal behavior explains this classification. To be an imbalance index, a statistic must be maximized by the caterpillar V˚(T)\mathring V(T)4 and minimized by the fully balanced tree V˚(T)\mathring V(T)5 when V˚(T)\mathring V(T)6. For Colless minus Sackin, the upper bound

V˚(T)\mathring V(T)7

is attained only by the caterpillar V˚(T)\mathring V(T)8. Equivalently,

V˚(T)\mathring V(T)9

and the caterpillar uniquely minimizes V1(T)V^1(T)0 (Knüver et al., 5 Sep 2025).

For powers of two, if V1(T)V^1(T)1, the fully balanced tree V1(T)V^1(T)2 is the unique minimizer of V1(T)V^1(T)3, with

V1(T)V^1(T)4

and correspondingly

V1(T)V^1(T)5

Thus V1(T)V^1(T)6 is maximized there (Knüver et al., 5 Sep 2025).

This result is distinct from the older literature in which Sackin and Colless were primarily treated as separate classical indices. For example, work on the minimum Colless index established its own extremal structure and recursive characterization, but did not study V1(T)V^1(T)7 as an independent quantity (Coronado et al., 2019). The newer analysis therefore reframes part of the older extremal theory through a derived index that is algebraically simpler.

4. Minimizers and the connection with Colless-minimal trees

A major theorem gives a lower bound for V1(T)V^1(T)8 on V1(T)V^1(T)9: vv0 and the bound is tight (Knüver et al., 5 Sep 2025). The principal structural consequence is that every Colless-minimal tree also minimizes vv1. Hence every Colless-minimal tree also maximizes vv2 and vv3 (Knüver et al., 5 Sep 2025).

This relationship is nontrivial because both vv4 and vv5 are individually minimized on the Colless-minimal class, yet their difference is also minimized there (Knüver et al., 5 Sep 2025). A related but earlier result had already shown that every rooted bifurcating tree with minimum Colless index also has minimum Sackin index (Coronado et al., 2019). That implication was later emphasized again in an extremal treatment of the Colless index, where Colless-minimal trees were shown to lie inside the Sackin-minimal class, while the converse fails (Fischer et al., 2019). The Colless-minus-Sackin theory sharpens this relation: not only are Colless-minimal trees Sackin-minimal, they also minimize the specific contrast vv6 (Knüver et al., 5 Sep 2025).

However, they are not the only minimizers. The characterization theorem states that a rooted binary tree vv7 minimizes vv8 and maximizes vv9 and nvanvbn_{v_a}\ge n_{v_b}0 if and only if every internal node nvanvbn_{v_a}\ge n_{v_b}1 has maximal pending subtree sizes nvanvbn_{v_a}\ge n_{v_b}2 that form either a Colless partition or an echelon partition (Knüver et al., 5 Sep 2025). This recursive local criterion yields a broader minimizer class than the class of Colless-minimal trees alone.

5. Relation to asymptotic comparison of Sackin and Colless

Before the 2025 decomposition result, the difference between Sackin and Colless had already appeared in asymptotic analysis of random tree shapes. For uniformly random rooted full binary tree shapes, the expected Sackin and Colless indices satisfy

nvanvbn_{v_a}\ge n_{v_b}3

so they have the same leading-order asymptotics (Goh et al., 2022). In that setting, the difference

nvanvbn_{v_a}\ge n_{v_b}4

was used to measure the gap between them, with

nvanvbn_{v_a}\ge n_{v_b}5

and

nvanvbn_{v_a}\ge n_{v_b}6

which implies

nvanvbn_{v_a}\ge n_{v_b}7

(Goh et al., 2022).

This shows that, in expectation under the uniform tree-shape model, the difference between Sackin and Colless is lower order relative to the common nvanvbn_{v_a}\ge n_{v_b}8 leading term (Goh et al., 2022). The later identity

nvanvbn_{v_a}\ge n_{v_b}9

provides an exact local-subtree explanation for that earlier gap variable (Knüver et al., 5 Sep 2025). A plausible implication is that what appeared asymptotically as a small correction term is, structurally, the cumulative contribution of the smaller branch at every internal split.

The same contrast is not directly developed in other expectation-focused papers. For instance, exact Yule-model formulas are available for S(T)=vV(T)nv=xV1(T)δx,S(T)=\sum_{v\in V(T)} n_v=\sum_{x\in V^1(T)} \delta_x,0 and S(T)=vV(T)nv=xV1(T)δx,S(T)=\sum_{v\in V(T)} n_v=\sum_{x\in V^1(T)} \delta_x,1, and one can formally write

S(T)=vV(T)nv=xV1(T)δx,S(T)=\sum_{v\in V(T)} n_v=\sum_{x\in V^1(T)} \delta_x,2

but those works did not derive a structural theory of S(T)=vV(T)nv=xV1(T)δx,S(T)=\sum_{v\in V(T)} n_v=\sum_{x\in V^1(T)} \delta_x,3 itself (Cardona et al., 2012). Likewise, exact expected-value papers on Colless under Yule and Sackin under the uniform model treat the two indices comparatively rather than through their difference as a named object (Mir et al., 2012).

6. Conceptual significance and relation to other balance indices

The broader significance of Colless minus Sackin lies in the re-interpretation of the classical indices as non-atomic. The 2025 study explicitly argues that Sackin and Colless are compound in nature and can be decomposed into more elementary components that independently satisfy the defining properties of tree balance or imbalance indices (Knüver et al., 5 Sep 2025). In this view, S(T)=vV(T)nv=xV1(T)δx,S(T)=\sum_{v\in V(T)} n_v=\sum_{x\in V^1(T)} \delta_x,4 is an imbalance index, S(T)=vV(T)nv=xV1(T)δx,S(T)=\sum_{v\in V(T)} n_v=\sum_{x\in V^1(T)} \delta_x,5 is a balance index, and

S(T)=vV(T)nv=xV1(T)δx,S(T)=\sum_{v\in V(T)} n_v=\sum_{x\in V^1(T)} \delta_x,6

Colless minus Sackin therefore isolates the “balanced” contribution embedded inside both classical statistics (Knüver et al., 5 Sep 2025).

This decomposition also offers a new angle on a longstanding empirical observation: Colless and Sackin are often strongly correlated. Comparative work on the S(T)=vV(T)nv=xV1(T)δx,S(T)=\sum_{v\in V(T)} n_v=\sum_{x\in V^1(T)} \delta_x,7 index describes them as the two most prominent classical indices and emphasizes that they are “extremely correlated” and often redundant as a pair (Bienvenu et al., 2020). The building-block perspective suggests a structural reason for that correlation: both are assembled from the same two components, differing only by the sign with which S(T)=vV(T)nv=xV1(T)δx,S(T)=\sum_{v\in V(T)} n_v=\sum_{x\in V^1(T)} \delta_x,8 enters. This is an interpretation rather than an explicit theorem of the comparative paper.

The statistic also clarifies an asymmetry between local and global notions of balance. Colless is defined from local split asymmetries, whereas Sackin sums leaf depths or, equivalently, descendant counts over internal nodes (Coronado et al., 2019). Their difference removes the S(T)=vV(T)nv=xV1(T)δx,S(T)=\sum_{v\in V(T)} n_v=\sum_{x\in V^1(T)} \delta_x,9 contribution entirely and retains only the accumulated smaller-subtree mass. This makes Colless minus Sackin more directly interpretable than either parent index in isolation when the analytic goal is to quantify how much “small-side” structure is present across the entire tree.

A recurrent point in the literature is that older papers on Colless or Sackin generally do not study the combined quantity nvn_v0. The paper “The minimum value of the Colless index” explicitly does not contain a theorem about a “Colless minus Sackin” quantity, nor any direct inequality relating nvn_v1 and nvn_v2 beyond general comparative remarks (Coronado et al., 2019). Likewise, “Revisiting Shao and Sokal’s nvn_v3 index of phylogenetic balance” provides extensive comparative context for Colless and Sackin but does not define a formal difference statistic nvn_v4 (Bienvenu et al., 2020). Similar remarks apply to work on the total cophenetic index, where the key subtraction identity involves Sackin and nvn_v5, not Colless and Sackin (Mir et al., 2012).

A common misconception would therefore be to treat Colless minus Sackin as a long-established classical index. The evidence instead indicates that its explicit identification as

nvn_v6

and its formal classification as an imbalance index are recent developments (Knüver et al., 5 Sep 2025). Earlier literature contained ingredients that are algebraically compatible with this viewpoint, including the identity

nvn_v7

for tree shapes (Goh et al., 2022), but did not elevate the quantity to the status of a distinct object with its own extremal theory.

Another misconception would be to assume that minimizers of nvn_v8 are exactly the Colless-minimal trees. The 2025 characterization shows that every Colless-minimal tree minimizes nvn_v9, but not every minimizer of vv0 is Colless-minimal; the full class is generated recursively from Colless partitions and echelon partitions (Knüver et al., 5 Sep 2025). This broader class is essential to the statistic’s independent identity.

In current phylogenetic balance theory, Colless minus Sackin thus occupies a specific niche: it is a derived but structurally fundamental imbalance index, canonically equal to negative twice the total smaller-subtree size over internal nodes, closely tied to Colless-minimal and Sackin-minimal phenomena, and useful for analyzing where classical indices agree or diverge in their assessment of tree shape (Knüver et al., 5 Sep 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Colless minus Sackin.