Papers
Topics
Authors
Recent
Search
2000 character limit reached

Sackin minus Colless: Balance Index in Phylogenetics

Updated 10 July 2026
  • The paper shows that Sackin minus Colless is a balance index that aggregates twice the size of the smaller subtree at each split, enabling efficient linear-time computation.
  • Explicit formulas and extremal characterizations reveal that maximally balanced trees maximize this difference, while caterpillar trees uniquely minimize it.
  • Under diverse stochastic models, the difference isolates lower-order balance signals, offering nuanced insights into the structural properties of phylogenetic trees.

Sackin minus Colless is the difference between two classical rooted-binary-tree shape functionals, the Sackin index and the Colless index. For a rooted binary phylogenetic tree TT with nn leaves, this difference is

D(T)=S(T)C(T),D(T)=S(T)-C(T),

and recent work identifies it as a balance index rather than an imbalance index: larger values correspond to more balanced recursive subdivisions, while smaller values correspond to more imbalanced, caterpillar-like shapes (Knüver et al., 5 Sep 2025). The quantity is structurally simple but mathematically rich. It admits an exact per-node decomposition, sharp extremal characterizations, explicit formulas for its maximal value in terms of the minimum Sackin and Colless indices, and asymptotic interpretations under several stochastic tree models (Goh et al., 2022, Coronado et al., 2019).

1. Definition and basic identities

Let TT be a rooted binary tree with root ρ\rho. For each leaf \ell, let depthT()\mathrm{depth}_T(\ell) be the number of edges from ρ\rho to \ell. For each internal node vv, let nn0 be the pending subtree rooted at nn1, with nn2. If nn3 has children nn4, adopt the convention nn5 (Knüver et al., 5 Sep 2025).

Under these conventions, the two classical indices are

nn6

and

nn7

Equivalent formulations of the Sackin index as a sum of leaf depths and as a sum over internal descendant-leaf counts are standard in the cited literature (Mir et al., 2012, Coronado et al., 2019).

A central identity is that

nn8

where nn9 is the number of leaves below D(T)=S(T)C(T),D(T)=S(T)-C(T),0 and D(T)=S(T)C(T),D(T)=S(T)-C(T),1 are its children (Goh et al., 2022). In the notation of Knüver and Fischer, if

D(T)=S(T)C(T),D(T)=S(T)-C(T),2

then

D(T)=S(T)C(T),D(T)=S(T)-C(T),3

This makes the interpretation immediate: each internal node contributes twice the size of its smaller child-subtree, so D(T)=S(T)C(T),D(T)=S(T)-C(T),4 aggregates the “smaller split sizes” throughout the tree (Knüver et al., 5 Sep 2025).

The same decomposition yields elementary inequalities valid for every rooted full binary tree with D(T)=S(T)C(T),D(T)=S(T)-C(T),5 leaves: D(T)=S(T)C(T),D(T)=S(T)-C(T),6 The lower bound reflects that every internal node contributes at least D(T)=S(T)C(T),D(T)=S(T)-C(T),7, since the smaller side of a binary split has at least one leaf (Goh et al., 2022).

2. Structural interpretation as a balance index

The most direct conceptual result is that D(T)=S(T)C(T),D(T)=S(T)-C(T),8 is a balance index. In the terminology of Knüver and Fischer, D(T)=S(T)C(T),D(T)=S(T)-C(T),9 is itself a balance index, while TT0 is an imbalance index, and

TT1

Thus the classical indices are compound quantities, and their difference isolates the contribution of the smaller side of each split (Knüver et al., 5 Sep 2025).

This decomposition clarifies why TT2 behaves differently from either constituent index. The Sackin index is sensitive to accumulated root-to-leaf depths, and the Colless index is sensitive to absolute split asymmetries. Their difference suppresses the large-side contribution and retains only the total mass placed on the smaller side of each internal bifurcation. A tree with many near-TT3 splits has large TT4, hence large TT5; a tree with repeated TT6 splits has small TT7, hence small TT8 (Goh et al., 2022, Knüver et al., 5 Sep 2025).

A useful recursive formulation follows from a decomposition TT9 with ρ\rho0: ρ\rho1 where ρ\rho2. Equivalently,

ρ\rho3

This locality implies linear-time computability by a single post-order traversal, because it suffices to compute child-subtree sizes and accumulate the smaller one at each internal node (Knüver et al., 5 Sep 2025).

A common misconception is to regard ρ\rho4 as merely a derived algebraic curiosity with no independent shape content. The decomposition above shows the opposite: ρ\rho5 is itself an interpretable index with its own extremal set, recurrence, and optimization theory (Knüver et al., 5 Sep 2025).

3. Extremal trees and exact bounds

The minimum of ρ\rho6 over rooted binary trees with ρ\rho7 leaves is attained uniquely by the caterpillar ρ\rho8, the unique rooted binary tree with exactly one cherry and every internal vertex adjacent to a leaf. The exact bound is

ρ\rho9

with equality if and only if \ell0 (Knüver et al., 5 Sep 2025). This is equivalent to \ell1, with equality exactly for the caterpillar.

For powers of two, the maximal case is especially simple. If \ell2, then the fully balanced tree \ell3 is the unique maximizer of \ell4, and

\ell5

Hence for perfect balance, Sackin minus Colless coincides with the Sackin index itself (Knüver et al., 5 Sep 2025).

For general \ell6, let \ell7, let

\ell8

and let the minimum Colless value be

\ell9

where depthT()\mathrm{depth}_T(\ell)0. Then the maximal value of depthT()\mathrm{depth}_T(\ell)1 is

depthT()\mathrm{depth}_T(\ell)2

This maximum is achieved by every Colless-minimal tree, including the maximally balanced tree depthT()\mathrm{depth}_T(\ell)3 and the greedy-from-the-bottom tree depthT()\mathrm{depth}_T(\ell)4; the echelon tree depthT()\mathrm{depth}_T(\ell)5 also achieves the maximum (Knüver et al., 5 Sep 2025).

The same phenomenon can be expressed through minimality results for depthT()\mathrm{depth}_T(\ell)6 and depthT()\mathrm{depth}_T(\ell)7. Coronado et al. proved that every tree with minimum Colless index also has minimum Sackin index, though the converse is false (Coronado et al., 2019). Therefore every Colless-minimal tree satisfies

depthT()\mathrm{depth}_T(\ell)8

The non-converse is concretely exhibited at depthT()\mathrm{depth}_T(\ell)9: there exist two 12-leaf trees ρ\rho0 with

ρ\rho1

both minimal, but

ρ\rho2

so only the first is Colless-minimal (Coronado et al., 2019). This establishes that maximizing ρ\rho3 is stricter than minimizing Sackin alone.

Small exact values illustrate the range. For the caterpillar on ρ\rho4 leaves,

ρ\rho5

For a Colless-minimal 5-leaf tree such as ρ\rho6,

ρ\rho7

For ρ\rho8, the fully balanced tree has

ρ\rho9

These values are explicitly given in the cited sources (Knüver et al., 5 Sep 2025, Keller-Schmidt et al., 2011).

4. Closed forms at the optimum

The minimum Colless value \ell0 admits several exact representations. It satisfies the recursion

\ell1

equivalently,

\ell2

It also has a closed form in terms of the binary expansion of \ell3, and another in terms of the Takagi–Blancmange curve (Coronado et al., 2019).

Combining this with the exact minimum Sackin value

\ell4

one obtains the exact optimum of Sackin minus Colless: \ell5 This identity holds for every Colless-minimal tree, including maximally balanced trees and greedy-from-the-bottom trees (Coronado et al., 2019).

Several immediate consequences follow. If \ell6 is a power of two, then \ell7, so

\ell8

If \ell9, then vv0, and

vv1

More generally, since vv2, one obtains explicit lower bounds on the maximal possible value of vv3, while vv4 yields the trivial upper bound

vv5

Asymptotically,

vv6

so the optimum grows on the order of vv7 (Coronado et al., 2019).

A plausible implication is that the optimization problem for vv8 is governed by the same fine binary-arithmetic structure that governs minimum Colless trees. The cited papers make this precise through the role of Colless partitions and, in the 2025 analysis, through the admissibility of either Colless partitions or echelon partitions at each internal node (Knüver et al., 5 Sep 2025).

5. Expected behavior under stochastic models

For uniformly random tree shapes—rooted full binary unlabeled trees sampled uniformly from the Otter/Polya class—the expected Sackin and Colless indices have the same leading asymptotic: vv9 with nn00, so nn01 (Goh et al., 2022). Because the leading terms coincide, they cancel in the difference, and

nn02

Thus nn03 isolates a lower-order but still substantial balance signal that is invisible at the dominant nn04 scale (Goh et al., 2022).

For uniformly random labeled phylogenetic trees under the PDA model, the exact expected Sackin value is

nn05

with asymptotic form

nn06

The same paper notes that nn07, but does not provide a closed form for the expectation of the Colless index in that model (Goh et al., 2022).

A different pair of results concerns mixed-model expectations: under the Yule model, the expected Colless index is

nn08

with asymptotic expansion

nn09

where nn10 if nn11 is odd and nn12 otherwise (Mir et al., 2012). The same paper gives an exact hypergeometric expression for nn13, but it explicitly cautions that subtracting these two quantities would mix distinct null models and therefore would not define nn14 under any single probability law (Mir et al., 2012).

This caution is methodologically important. The difference nn15 is well defined treewise, but its expectation is model-dependent. Statements about nn16 require both expectations to be taken under the same generative model.

6. Phylogenetic interpretation and innovation-based models

Within macroevolutionary modeling, the innovation-triggered branching model of Keller-Schmidt and Klemm provides an instructive comparison. In that model, the normalized Sackin index—reported as average leaf depth,

nn17

scales as

nn18

so the unnormalized Sackin index satisfies

nn19

A deterministic-growth surrogate gives the more explicit approximation

nn20

hence

nn21

The paper reports that the actual innovation model has the same asymptotic scaling but with a smaller prefactor than nn22 (Keller-Schmidt et al., 2011).

For the Colless index, the same work does not provide an explicit asymptotic formula under the innovation model, but numerical evidence shows that the mean and standard deviation of the normalized Colless statistic match empirical phylogenies closely across the observed size range, and that the fit is better than Yule/ERM and better than Aldous’ branching with respect to both depth and imbalance (Keller-Schmidt et al., 2011). Because the paper normalizes the two indices differently,

nn23

it emphasizes that nn24 in unnormalized form is not directly discussed there (Keller-Schmidt et al., 2011).

Still, the paper’s mechanism provides useful context. Rare innovations create deep backbones, while loss-driven diversification within innovation-induced subtrees yields comparatively balanced within-level structure. The reported implication is that the model can simultaneously produce high Sackin values and realistic Colless values, unlike ERM trees, which are too shallow, or AB trees, which, though deep, deviate more strongly in imbalance (Keller-Schmidt et al., 2011). This suggests that nn25 is particularly sensitive to models that generate deep yet locally balanced hierarchical organization.

Several tree families recur in the analysis of nn26. The maximally balanced tree nn27 satisfies nn28 at every internal node and is Colless-minimal. The greedy-from-the-bottom tree nn29, built by repeatedly merging the two smallest current subtrees, is also Colless-minimal. The echelon tree nn30, defined recursively by

nn31

likewise maximizes nn32 for all nn33 (Knüver et al., 5 Sep 2025).

Knüver and Fischer prove a full characterization of the maximizers of nn34: a tree is nn35-maximal if and only if at every internal node the split is either a Colless partition of the corresponding subtree size or the echelon partition nn36 for the appropriate nn37 (Knüver et al., 5 Sep 2025). This is stronger than merely identifying a few extremal examples; it describes the complete local grammar of globally optimal shapes.

The comparison with the nn38 index is also informative. The cited work states that nn39, and therefore nn40, is typically less resolved than Sackin, Colless, or nn41 in ranking trees, but it is useful for diagnosing disagreements among indices. In particular, because nn42, it directly tracks how much leaf mass accumulates on the smaller side of splits, which is precisely the component suppressed when one looks only at nn43 or only at nn44 (Knüver et al., 5 Sep 2025).

At the level of qualitative interpretation, balanced trees have nn45 on the order of nn46, combs have nn47 on the order of nn48, and random shape-uniform trees have expected nn49 despite nn50 and nn51 individually being of order nn52 (Goh et al., 2022). The difference therefore functions as a second-order balance statistic: it is much smaller than either raw index under some models, yet it captures a distinct structural aspect of recursive subdivision that neither constituent index isolates on its own.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Sackin minus Colless.