Sackin minus Colless: Balance Index in Phylogenetics
- The paper shows that Sackin minus Colless is a balance index that aggregates twice the size of the smaller subtree at each split, enabling efficient linear-time computation.
- Explicit formulas and extremal characterizations reveal that maximally balanced trees maximize this difference, while caterpillar trees uniquely minimize it.
- Under diverse stochastic models, the difference isolates lower-order balance signals, offering nuanced insights into the structural properties of phylogenetic trees.
Sackin minus Colless is the difference between two classical rooted-binary-tree shape functionals, the Sackin index and the Colless index. For a rooted binary phylogenetic tree with leaves, this difference is
and recent work identifies it as a balance index rather than an imbalance index: larger values correspond to more balanced recursive subdivisions, while smaller values correspond to more imbalanced, caterpillar-like shapes (Knüver et al., 5 Sep 2025). The quantity is structurally simple but mathematically rich. It admits an exact per-node decomposition, sharp extremal characterizations, explicit formulas for its maximal value in terms of the minimum Sackin and Colless indices, and asymptotic interpretations under several stochastic tree models (Goh et al., 2022, Coronado et al., 2019).
1. Definition and basic identities
Let be a rooted binary tree with root . For each leaf , let be the number of edges from to . For each internal node , let 0 be the pending subtree rooted at 1, with 2. If 3 has children 4, adopt the convention 5 (Knüver et al., 5 Sep 2025).
Under these conventions, the two classical indices are
6
and
7
Equivalent formulations of the Sackin index as a sum of leaf depths and as a sum over internal descendant-leaf counts are standard in the cited literature (Mir et al., 2012, Coronado et al., 2019).
A central identity is that
8
where 9 is the number of leaves below 0 and 1 are its children (Goh et al., 2022). In the notation of Knüver and Fischer, if
2
then
3
This makes the interpretation immediate: each internal node contributes twice the size of its smaller child-subtree, so 4 aggregates the “smaller split sizes” throughout the tree (Knüver et al., 5 Sep 2025).
The same decomposition yields elementary inequalities valid for every rooted full binary tree with 5 leaves: 6 The lower bound reflects that every internal node contributes at least 7, since the smaller side of a binary split has at least one leaf (Goh et al., 2022).
2. Structural interpretation as a balance index
The most direct conceptual result is that 8 is a balance index. In the terminology of Knüver and Fischer, 9 is itself a balance index, while 0 is an imbalance index, and
1
Thus the classical indices are compound quantities, and their difference isolates the contribution of the smaller side of each split (Knüver et al., 5 Sep 2025).
This decomposition clarifies why 2 behaves differently from either constituent index. The Sackin index is sensitive to accumulated root-to-leaf depths, and the Colless index is sensitive to absolute split asymmetries. Their difference suppresses the large-side contribution and retains only the total mass placed on the smaller side of each internal bifurcation. A tree with many near-3 splits has large 4, hence large 5; a tree with repeated 6 splits has small 7, hence small 8 (Goh et al., 2022, Knüver et al., 5 Sep 2025).
A useful recursive formulation follows from a decomposition 9 with 0: 1 where 2. Equivalently,
3
This locality implies linear-time computability by a single post-order traversal, because it suffices to compute child-subtree sizes and accumulate the smaller one at each internal node (Knüver et al., 5 Sep 2025).
A common misconception is to regard 4 as merely a derived algebraic curiosity with no independent shape content. The decomposition above shows the opposite: 5 is itself an interpretable index with its own extremal set, recurrence, and optimization theory (Knüver et al., 5 Sep 2025).
3. Extremal trees and exact bounds
The minimum of 6 over rooted binary trees with 7 leaves is attained uniquely by the caterpillar 8, the unique rooted binary tree with exactly one cherry and every internal vertex adjacent to a leaf. The exact bound is
9
with equality if and only if 0 (Knüver et al., 5 Sep 2025). This is equivalent to 1, with equality exactly for the caterpillar.
For powers of two, the maximal case is especially simple. If 2, then the fully balanced tree 3 is the unique maximizer of 4, and
5
Hence for perfect balance, Sackin minus Colless coincides with the Sackin index itself (Knüver et al., 5 Sep 2025).
For general 6, let 7, let
8
and let the minimum Colless value be
9
where 0. Then the maximal value of 1 is
2
This maximum is achieved by every Colless-minimal tree, including the maximally balanced tree 3 and the greedy-from-the-bottom tree 4; the echelon tree 5 also achieves the maximum (Knüver et al., 5 Sep 2025).
The same phenomenon can be expressed through minimality results for 6 and 7. Coronado et al. proved that every tree with minimum Colless index also has minimum Sackin index, though the converse is false (Coronado et al., 2019). Therefore every Colless-minimal tree satisfies
8
The non-converse is concretely exhibited at 9: there exist two 12-leaf trees 0 with
1
both minimal, but
2
so only the first is Colless-minimal (Coronado et al., 2019). This establishes that maximizing 3 is stricter than minimizing Sackin alone.
Small exact values illustrate the range. For the caterpillar on 4 leaves,
5
For a Colless-minimal 5-leaf tree such as 6,
7
For 8, the fully balanced tree has
9
These values are explicitly given in the cited sources (Knüver et al., 5 Sep 2025, Keller-Schmidt et al., 2011).
4. Closed forms at the optimum
The minimum Colless value 0 admits several exact representations. It satisfies the recursion
1
equivalently,
2
It also has a closed form in terms of the binary expansion of 3, and another in terms of the Takagi–Blancmange curve (Coronado et al., 2019).
Combining this with the exact minimum Sackin value
4
one obtains the exact optimum of Sackin minus Colless: 5 This identity holds for every Colless-minimal tree, including maximally balanced trees and greedy-from-the-bottom trees (Coronado et al., 2019).
Several immediate consequences follow. If 6 is a power of two, then 7, so
8
If 9, then 0, and
1
More generally, since 2, one obtains explicit lower bounds on the maximal possible value of 3, while 4 yields the trivial upper bound
5
Asymptotically,
6
so the optimum grows on the order of 7 (Coronado et al., 2019).
A plausible implication is that the optimization problem for 8 is governed by the same fine binary-arithmetic structure that governs minimum Colless trees. The cited papers make this precise through the role of Colless partitions and, in the 2025 analysis, through the admissibility of either Colless partitions or echelon partitions at each internal node (Knüver et al., 5 Sep 2025).
5. Expected behavior under stochastic models
For uniformly random tree shapes—rooted full binary unlabeled trees sampled uniformly from the Otter/Polya class—the expected Sackin and Colless indices have the same leading asymptotic: 9 with 00, so 01 (Goh et al., 2022). Because the leading terms coincide, they cancel in the difference, and
02
Thus 03 isolates a lower-order but still substantial balance signal that is invisible at the dominant 04 scale (Goh et al., 2022).
For uniformly random labeled phylogenetic trees under the PDA model, the exact expected Sackin value is
05
with asymptotic form
06
The same paper notes that 07, but does not provide a closed form for the expectation of the Colless index in that model (Goh et al., 2022).
A different pair of results concerns mixed-model expectations: under the Yule model, the expected Colless index is
08
with asymptotic expansion
09
where 10 if 11 is odd and 12 otherwise (Mir et al., 2012). The same paper gives an exact hypergeometric expression for 13, but it explicitly cautions that subtracting these two quantities would mix distinct null models and therefore would not define 14 under any single probability law (Mir et al., 2012).
This caution is methodologically important. The difference 15 is well defined treewise, but its expectation is model-dependent. Statements about 16 require both expectations to be taken under the same generative model.
6. Phylogenetic interpretation and innovation-based models
Within macroevolutionary modeling, the innovation-triggered branching model of Keller-Schmidt and Klemm provides an instructive comparison. In that model, the normalized Sackin index—reported as average leaf depth,
17
scales as
18
so the unnormalized Sackin index satisfies
19
A deterministic-growth surrogate gives the more explicit approximation
20
hence
21
The paper reports that the actual innovation model has the same asymptotic scaling but with a smaller prefactor than 22 (Keller-Schmidt et al., 2011).
For the Colless index, the same work does not provide an explicit asymptotic formula under the innovation model, but numerical evidence shows that the mean and standard deviation of the normalized Colless statistic match empirical phylogenies closely across the observed size range, and that the fit is better than Yule/ERM and better than Aldous’ branching with respect to both depth and imbalance (Keller-Schmidt et al., 2011). Because the paper normalizes the two indices differently,
23
it emphasizes that 24 in unnormalized form is not directly discussed there (Keller-Schmidt et al., 2011).
Still, the paper’s mechanism provides useful context. Rare innovations create deep backbones, while loss-driven diversification within innovation-induced subtrees yields comparatively balanced within-level structure. The reported implication is that the model can simultaneously produce high Sackin values and realistic Colless values, unlike ERM trees, which are too shallow, or AB trees, which, though deep, deviate more strongly in imbalance (Keller-Schmidt et al., 2011). This suggests that 25 is particularly sensitive to models that generate deep yet locally balanced hierarchical organization.
7. Related extremal families and comparative use
Several tree families recur in the analysis of 26. The maximally balanced tree 27 satisfies 28 at every internal node and is Colless-minimal. The greedy-from-the-bottom tree 29, built by repeatedly merging the two smallest current subtrees, is also Colless-minimal. The echelon tree 30, defined recursively by
31
likewise maximizes 32 for all 33 (Knüver et al., 5 Sep 2025).
Knüver and Fischer prove a full characterization of the maximizers of 34: a tree is 35-maximal if and only if at every internal node the split is either a Colless partition of the corresponding subtree size or the echelon partition 36 for the appropriate 37 (Knüver et al., 5 Sep 2025). This is stronger than merely identifying a few extremal examples; it describes the complete local grammar of globally optimal shapes.
The comparison with the 38 index is also informative. The cited work states that 39, and therefore 40, is typically less resolved than Sackin, Colless, or 41 in ranking trees, but it is useful for diagnosing disagreements among indices. In particular, because 42, it directly tracks how much leaf mass accumulates on the smaller side of splits, which is precisely the component suppressed when one looks only at 43 or only at 44 (Knüver et al., 5 Sep 2025).
At the level of qualitative interpretation, balanced trees have 45 on the order of 46, combs have 47 on the order of 48, and random shape-uniform trees have expected 49 despite 50 and 51 individually being of order 52 (Goh et al., 2022). The difference therefore functions as a second-order balance statistic: it is much smaller than either raw index under some models, yet it captures a distinct structural aspect of recursive subdivision that neither constituent index isolates on its own.