Separable Weighted Leaf-Collision Proximities
- The paper introduces SWLCP as a novel approach that leverages tree ensemble leaf-collision structures to define supervised similarity measures.
- It employs separable, sample-local weighting and sparse matrix factorization to reduce the computational cost from quadratic to near-linear scaling.
- Empirical results demonstrate that SWLCPs achieve significant improvements in runtime and memory usage on large-scale datasets.
Separable Weighted Leaf-Collision Proximities (SWLCPs) constitute a mathematically rigorous family of supervised similarity measures defined via the leaf co-occurrence structure in tree ensembles such as Random Forests and Gradient Boosted Trees. SWLCPs generalize the notion that tree ensembles induce proximities based on the extent to which sample pairs collide, i.e., are assigned to the same leaf, with collisions modulated through separable, sample-local weighting schemes. The SWLCP structure enables exact, scalable computation by leveraging sparse matrix factorization, circumventing the quadratic time or memory complexities inherent to traditional explicit pairwise proximity formulations (Aumon et al., 6 Jan 2026).
1. Formal Framework and Definition
Let denote the number of samples and the set of trees in the ensemble. For a sample and tree , denote by the index of the unique leaf of tree containing . The Weighted Leaf-Collision Proximity (WLCP) matrix is defined as
where 0 is a collision weight and 1 is the indicator function. A WLCP is termed separable if
2
for nonnegative, sample-local vectors 3, where 4 and 5 are specific to individual samples and trees but independent of the paired sample index.
Common proximities, including those underlying classical Random Forests, are instances of the separable form. For example, setting 6 and 7 recovers the original Random Forest (RF) proximity.
2. Sparse Matrix Factorization
The defining property of SWLCPs is that they admit an exact sparse matrix factorization, which enables scalable computation. Denote 8 as the set of all leaf nodes across all trees, with 9. Construct sparse matrices 0 as follows:
- For each leaf 1, associate a unique column.
- For sample 2 and leaf 3, set 4 and 5, where 6 identifies the tree containing leaf 7.
With these definitions, the SWLCP matrix factorizes as
8
with each row of 9 and 0 containing at most one nonzero per tree (and thus at most 1 nonzeros per row). This formulation restricts computation to leaf-level collisions and avoids the 2 cost of explicit pairwise comparisons.
| Proximity variant | 3 | 4 |
|---|---|---|
| RF (original) | 5 | 6 |
| RF-GAP | 7 | 8 |
| GBT | 9 | 0 |
3. Computational Workflow and Pseudocode
The scalable computation of SWLCPs proceeds through the following steps:
- Leaf-to-Column Mapping: Enumerate all unique leaves across all trees and map each to a unique column index, yielding 1 columns.
- Sparse Matrix Construction: Build index and data arrays for 2 and 3 with nonzero entries determined by the sample-leaf assignments and the respective 4, 5.
- Sparse Matrix Multiplication: Form 6 and 7 as CSR/CSC matrices and compute 8 using optimized sparse linear algebra routines.
A high-level pseudocode summary is as follows:
9
By construction, 9.
4. Computational Complexity Analysis
The classical explicit pairwise approach evaluating all 0 sample pairs per tree incurs 1 computational cost and 2 memory. The sparse factorization method introduced for SWLCPs reduces this to near-linear in 3:
- Tree traversal and leaf encoding: 4, where 5 is average tree height.
- Sparse matrix assembly: 6.
- Sparse multiplication: Each 7 has 8 nonzeros; each 9 contains at most 0 nonzeros (for average leaf size 1), leading to 2.
- Total time complexity: 3.
Memory usage is dominated by the storage of 4, 5, and sparse 6, with 7 and 8 nonzeros, respectively, as opposed to 9 for dense approaches.
5. Illustrative Construction: Toy Example
Consider 0 samples and 1 trees with leaves 2 and 3. Assigning samples as 4, 5, and 6, and using the RF proximity (7, 8), the 9 columns of 0 and 1 encode sample-leaf memberships. The resulting 2 is:
3
This corresponds to the fraction of trees for which sample pairs share a leaf.
6. Practical Implementation Using Python
Efficient construction of SWLCPs leverages NumPy and SciPy's sparse matrix routines. The primary operations consist of assigning sample-leaf-weight triples to sparse index/value arrays, constructing 4 and 5 as CSR matrices, and computing their product. The following fragment demonstrates core code elements:
00
This organization ensures that only a linear number of nonzeros (6) are handled in sparse arrays and products, and dense 7 structures are never explicitly formed unless required.
7. Empirical Performance and Scalability
Empirical evaluation on the Fashion-MNIST dataset (8) demonstrates the practical impact of the SWLCP factorization strategy (Aumon et al., 6 Jan 2026):
- Runtime: Traditional pairwise computation scales quadratically, exceeding 20 minutes for 9. The SWLCP method achieves near-linear scaling (empirical exponent 0), requiring 1 seconds for the same 2.
- Memory: The explicit approach exhausts memory at 3–4; the SWLCP approach remains below 4 GB for 5, scaling linearly due to 6 storage.
This confirms that restricting to leaf-level collisions and utilizing separable weighting reduces real-world computational cost from intractable 7 to essentially linear in 8.
For detailed derivations, algorithmic improvements, and further results, see "Scalable Tree Ensemble Proximities in Python" (Aumon et al., 6 Jan 2026).