Papers
Topics
Authors
Recent
Search
2000 character limit reached

Separable Weighted Leaf-Collision Proximities

Updated 13 January 2026
  • The paper introduces SWLCP as a novel approach that leverages tree ensemble leaf-collision structures to define supervised similarity measures.
  • It employs separable, sample-local weighting and sparse matrix factorization to reduce the computational cost from quadratic to near-linear scaling.
  • Empirical results demonstrate that SWLCPs achieve significant improvements in runtime and memory usage on large-scale datasets.

Separable Weighted Leaf-Collision Proximities (SWLCPs) constitute a mathematically rigorous family of supervised similarity measures defined via the leaf co-occurrence structure in tree ensembles such as Random Forests and Gradient Boosted Trees. SWLCPs generalize the notion that tree ensembles induce proximities based on the extent to which sample pairs collide, i.e., are assigned to the same leaf, with collisions modulated through separable, sample-local weighting schemes. The SWLCP structure enables exact, scalable computation by leveraging sparse matrix factorization, circumventing the quadratic time or memory complexities inherent to traditional explicit pairwise proximity formulations (Aumon et al., 6 Jan 2026).

1. Formal Framework and Definition

Let NN denote the number of samples and T={1,…,T}\mathcal{T} = \{1, \ldots, T\} the set of TT trees in the ensemble. For a sample xix_i and tree tt, denote by ℓi(t)\ell_i(t) the index of the unique leaf of tree tt containing xix_i. The Weighted Leaf-Collision Proximity (WLCP) matrix P∈RN×NP \in \mathbb{R}^{N \times N} is defined as

Pij=∑t=1Twijt⋅I(ℓi(t)=ℓj(t)),P_{ij} = \sum_{t=1}^T w_{ijt} \cdot I(\ell_i(t) = \ell_j(t)),

where T={1,…,T}\mathcal{T} = \{1, \ldots, T\}0 is a collision weight and T={1,…,T}\mathcal{T} = \{1, \ldots, T\}1 is the indicator function. A WLCP is termed separable if

T={1,…,T}\mathcal{T} = \{1, \ldots, T\}2

for nonnegative, sample-local vectors T={1,…,T}\mathcal{T} = \{1, \ldots, T\}3, where T={1,…,T}\mathcal{T} = \{1, \ldots, T\}4 and T={1,…,T}\mathcal{T} = \{1, \ldots, T\}5 are specific to individual samples and trees but independent of the paired sample index.

Common proximities, including those underlying classical Random Forests, are instances of the separable form. For example, setting T={1,…,T}\mathcal{T} = \{1, \ldots, T\}6 and T={1,…,T}\mathcal{T} = \{1, \ldots, T\}7 recovers the original Random Forest (RF) proximity.

2. Sparse Matrix Factorization

The defining property of SWLCPs is that they admit an exact sparse matrix factorization, which enables scalable computation. Denote T={1,…,T}\mathcal{T} = \{1, \ldots, T\}8 as the set of all leaf nodes across all trees, with T={1,…,T}\mathcal{T} = \{1, \ldots, T\}9. Construct sparse matrices TT0 as follows:

  • For each leaf TT1, associate a unique column.
  • For sample TT2 and leaf TT3, set TT4 and TT5, where TT6 identifies the tree containing leaf TT7.

With these definitions, the SWLCP matrix factorizes as

TT8

with each row of TT9 and xix_i0 containing at most one nonzero per tree (and thus at most xix_i1 nonzeros per row). This formulation restricts computation to leaf-level collisions and avoids the xix_i2 cost of explicit pairwise comparisons.

Proximity variant xix_i3 xix_i4
RF (original) xix_i5 xix_i6
RF-GAP xix_i7 xix_i8
GBT xix_i9 tt0

3. Computational Workflow and Pseudocode

The scalable computation of SWLCPs proceeds through the following steps:

  1. Leaf-to-Column Mapping: Enumerate all unique leaves across all trees and map each to a unique column index, yielding tt1 columns.
  2. Sparse Matrix Construction: Build index and data arrays for tt2 and tt3 with nonzero entries determined by the sample-leaf assignments and the respective tt4, tt5.
  3. Sparse Matrix Multiplication: Form tt6 and tt7 as CSR/CSC matrices and compute tt8 using optimized sparse linear algebra routines.

A high-level pseudocode summary is as follows:

Pij=∑t=1Twijt⋅I(ℓi(t)=ℓj(t)),P_{ij} = \sum_{t=1}^T w_{ijt} \cdot I(\ell_i(t) = \ell_j(t)),9

By construction, tt9.

4. Computational Complexity Analysis

The classical explicit pairwise approach evaluating all ℓi(t)\ell_i(t)0 sample pairs per tree incurs ℓi(t)\ell_i(t)1 computational cost and ℓi(t)\ell_i(t)2 memory. The sparse factorization method introduced for SWLCPs reduces this to near-linear in ℓi(t)\ell_i(t)3:

  • Tree traversal and leaf encoding: ℓi(t)\ell_i(t)4, where ℓi(t)\ell_i(t)5 is average tree height.
  • Sparse matrix assembly: ℓi(t)\ell_i(t)6.
  • Sparse multiplication: Each ℓi(t)\ell_i(t)7 has ℓi(t)\ell_i(t)8 nonzeros; each ℓi(t)\ell_i(t)9 contains at most tt0 nonzeros (for average leaf size tt1), leading to tt2.
  • Total time complexity: tt3.

Memory usage is dominated by the storage of tt4, tt5, and sparse tt6, with tt7 and tt8 nonzeros, respectively, as opposed to tt9 for dense approaches.

5. Illustrative Construction: Toy Example

Consider xix_i0 samples and xix_i1 trees with leaves xix_i2 and xix_i3. Assigning samples as xix_i4, xix_i5, and xix_i6, and using the RF proximity (xix_i7, xix_i8), the xix_i9 columns of P∈RN×NP \in \mathbb{R}^{N \times N}0 and P∈RN×NP \in \mathbb{R}^{N \times N}1 encode sample-leaf memberships. The resulting P∈RN×NP \in \mathbb{R}^{N \times N}2 is:

P∈RN×NP \in \mathbb{R}^{N \times N}3

This corresponds to the fraction of trees for which sample pairs share a leaf.

6. Practical Implementation Using Python

Efficient construction of SWLCPs leverages NumPy and SciPy's sparse matrix routines. The primary operations consist of assigning sample-leaf-weight triples to sparse index/value arrays, constructing P∈RN×NP \in \mathbb{R}^{N \times N}4 and P∈RN×NP \in \mathbb{R}^{N \times N}5 as CSR matrices, and computing their product. The following fragment demonstrates core code elements:

T={1,…,T}\mathcal{T} = \{1, \ldots, T\}00

This organization ensures that only a linear number of nonzeros (P∈RN×NP \in \mathbb{R}^{N \times N}6) are handled in sparse arrays and products, and dense P∈RN×NP \in \mathbb{R}^{N \times N}7 structures are never explicitly formed unless required.

7. Empirical Performance and Scalability

Empirical evaluation on the Fashion-MNIST dataset (P∈RN×NP \in \mathbb{R}^{N \times N}8) demonstrates the practical impact of the SWLCP factorization strategy (Aumon et al., 6 Jan 2026):

  • Runtime: Traditional pairwise computation scales quadratically, exceeding 20 minutes for P∈RN×NP \in \mathbb{R}^{N \times N}9. The SWLCP method achieves near-linear scaling (empirical exponent Pij=∑t=1Twijt⋅I(ℓi(t)=ℓj(t)),P_{ij} = \sum_{t=1}^T w_{ijt} \cdot I(\ell_i(t) = \ell_j(t)),0), requiring Pij=∑t=1Twijt⋅I(ℓi(t)=ℓj(t)),P_{ij} = \sum_{t=1}^T w_{ijt} \cdot I(\ell_i(t) = \ell_j(t)),1 seconds for the same Pij=∑t=1Twijt⋅I(ℓi(t)=ℓj(t)),P_{ij} = \sum_{t=1}^T w_{ijt} \cdot I(\ell_i(t) = \ell_j(t)),2.
  • Memory: The explicit approach exhausts memory at Pij=∑t=1Twijt⋅I(ℓi(t)=ℓj(t)),P_{ij} = \sum_{t=1}^T w_{ijt} \cdot I(\ell_i(t) = \ell_j(t)),3–Pij=∑t=1Twijt⋅I(ℓi(t)=ℓj(t)),P_{ij} = \sum_{t=1}^T w_{ijt} \cdot I(\ell_i(t) = \ell_j(t)),4; the SWLCP approach remains below 4 GB for Pij=∑t=1Twijt⋅I(ℓi(t)=ℓj(t)),P_{ij} = \sum_{t=1}^T w_{ijt} \cdot I(\ell_i(t) = \ell_j(t)),5, scaling linearly due to Pij=∑t=1Twijt⋅I(ℓi(t)=ℓj(t)),P_{ij} = \sum_{t=1}^T w_{ijt} \cdot I(\ell_i(t) = \ell_j(t)),6 storage.

This confirms that restricting to leaf-level collisions and utilizing separable weighting reduces real-world computational cost from intractable Pij=∑t=1Twijt⋅I(ℓi(t)=ℓj(t)),P_{ij} = \sum_{t=1}^T w_{ijt} \cdot I(\ell_i(t) = \ell_j(t)),7 to essentially linear in Pij=∑t=1Twijt⋅I(ℓi(t)=ℓj(t)),P_{ij} = \sum_{t=1}^T w_{ijt} \cdot I(\ell_i(t) = \ell_j(t)),8.


For detailed derivations, algorithmic improvements, and further results, see "Scalable Tree Ensemble Proximities in Python" (Aumon et al., 6 Jan 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Separable Weighted Leaf-Collision Proximities.