Papers
Topics
Authors
Recent
Search
2000 character limit reached

Branchless SIMD in B^S-tree

Updated 24 January 2026
  • The paper introduces a branchless SIMD approach in B^S-trees, eliminating conditional branches to streamline node search operations.
  • It presents a novel methodology leveraging SIMD instructions to traverse tree nodes in parallel, improving throughput and accuracy.
  • Empirical results reveal marked reductions in latency and enhanced scalability compared to conventional B-tree implementations.

Parallel Voting Decision Tree (PV-Tree) is a distributed algorithm designed to efficiently construct decision trees in large-scale, data-parallel environments. The core objective is to minimize communication costs, especially in contexts such as gradient boosting decision trees (GBDT) and random forests, where repeated computation of attribute histograms can otherwise overwhelm distributed systems. PV-Tree introduces a two-stage voting mechanism—local and global voting—to identify promising attributes for splitting, enabling near-optimal accuracy while keeping communication independent of the total number of attributes or dataset size. The theoretical foundation establishes probabilistic guarantees for identifying the globally best split, and empirical evaluation demonstrates favorable trade-offs between accuracy and efficiency relative to standard data-parallel and attribute-parallel baselines (Meng et al., 2016).

1. Problem Formulation and Setting

Consider a dataset D={(xi,yi)}i=1ND = \{(x_i, y_i)\}_{i=1}^N where xi∈Rdx_i \in \mathbb{R}^d, yi∈Yy_i \in \mathcal{Y}, and the goal is to construct a decision tree in parallel across MM machines. The dataset is partitioned horizontally into D=⋃m=1MDmD = \bigcup_{m=1}^{M} D_m, with ∣Dm∣=n=N/M|D_m| = n = N/M. At each node OO of the current tree, the challenge is to identify the attribute j∗∈[d]j^* \in [d] and threshold w∗∈Wj∗w^* \in W_{j^*} that optimize an informativeness criterion Δj(w;O)\Delta_{j}(w; O). For classification, this is typically information gain: xi∈Rdx_i \in \mathbb{R}^d0 and for regression, variance gain. xi∈Rdx_i \in \mathbb{R}^d1 and xi∈Rdx_i \in \mathbb{R}^d2 are split probabilities, and xi∈Rdx_i \in \mathbb{R}^d3 is conditional entropy (Meng et al., 2016).

2. Algorithmic Workflow: PV-Tree in Detail

PV-Tree operates by iteratively computing splits at each active tree node using three main phases:

2.1 Local Voting Phase

Each machine constructs histograms xi∈Rdx_i \in \mathbb{R}^d4 for all attributes xi∈Rdx_i \in \mathbb{R}^d5 over its local subset xi∈Rdx_i \in \mathbb{R}^d6, optionally pre-binning continuous features into xi∈Rdx_i \in \mathbb{R}^d7 bins to allow fast evaluation. The best local gain per attribute is computed: xi∈Rdx_i \in \mathbb{R}^d8 The local top-xi∈Rdx_i \in \mathbb{R}^d9 attributes are selected based on yi∈Yy_i \in \mathcal{Y}0, yielding sets yi∈Yy_i \in \mathcal{Y}1, yi∈Yy_i \in \mathcal{Y}2 for each machine.

2.2 Global Voting Phase

Each machine broadcasts its local candidate set yi∈Yy_i \in \mathcal{Y}3 (e.g., via MPI_AllGather). For each attribute yi∈Yy_i \in \mathcal{Y}4, aggregate the vote count yi∈Yy_i \in \mathcal{Y}5. The yi∈Yy_i \in \mathcal{Y}6 attributes with highest yi∈Yy_i \in \mathcal{Y}7 are selected to form the global candidate set yi∈Yy_i \in \mathcal{Y}8.

2.3 Final Split Selection Phase

For each yi∈Yy_i \in \mathcal{Y}9, all machines send their local histograms MM0 (of MM1 bins) to the master (or combine via all-reduce). The master aggregates these to form MM2 for each attribute in MM3. A scan over the binned histograms identifies the globally best split MM4.

The key pseudocode is as follows: Δj(w;O)\Delta_{j}(w; O)9 (Meng et al., 2016)

3. Communication Complexity Analysis

The PV-Tree communication protocol is designed so that its total per-split communication is MM5, where MM6 is the number of machines, MM7 is the number of local candidates per machine, and MM8 is the histogram bin count. This is achieved as follows:

  • Local voting: Each of the MM9 machines transmits D=⋃m=1MDmD = \bigcup_{m=1}^{M} D_m0 attribute indices (D=⋃m=1MDmD = \bigcup_{m=1}^{M} D_m1 words).
  • Histogram gathering: Each D=⋃m=1MDmD = \bigcup_{m=1}^{M} D_m2 sends D=⋃m=1MDmD = \bigcup_{m=1}^{M} D_m3 histograms with D=⋃m=1MDmD = \bigcup_{m=1}^{M} D_m4 bins (D=⋃m=1MDmD = \bigcup_{m=1}^{M} D_m5 words).

Thus, the total cost is D=⋃m=1MDmD = \bigcup_{m=1}^{M} D_m6 per split—independent of D=⋃m=1MDmD = \bigcup_{m=1}^{M} D_m7 (the feature dimension) and dataset size D=⋃m=1MDmD = \bigcup_{m=1}^{M} D_m8. In contrast, baseline data-parallel requires D=⋃m=1MDmD = \bigcup_{m=1}^{M} D_m9, and attribute-parallel approaches may require ∣Dm∣=n=N/M|D_m| = n = N/M0. Empirical measurements in large datasets (e.g., ∣Dm∣=n=N/M|D_m| = n = N/M1) show that PV-Tree with ∣Dm∣=n=N/M|D_m| = n = N/M2 reduces communication to ∣Dm∣=n=N/M|D_m| = n = N/M3 MB per full tree (depth=6), far less than the ∣Dm∣=n=N/M|D_m| = n = N/M4 MB (attribute-parallel) or ∣Dm∣=n=N/M|D_m| = n = N/M5 MB (data-parallel) required by alternatives (Meng et al., 2016).

4. Theoretical Accuracy Guarantees

PV-Tree incorporates probabilistic guarantees of optimal split selection. Let ∣Dm∣=n=N/M|D_m| = n = N/M6 denote the true global information gains. Define ∣Dm∣=n=N/M|D_m| = n = N/M7 for ∣Dm∣=n=N/M|D_m| = n = N/M8, and

∣Dm∣=n=N/M|D_m| = n = N/M9

where OO0 as OO1, and OO2 constants. The probability that PV-Tree with OO3 machines, local sample size OO4, OO5 local votes, and OO6 global votes selects the best attribute is at least

OO7

which approaches OO8 as OO9. The guarantee follows from uniform convergence of local gain estimates (VC-type analysis) and the fact that majority voting with set size j∗∈[d]j^* \in [d]0 preserves the top-scoring attribute with high probability (Meng et al., 2016).

5. Empirical Evaluation and Comparative Performance

PV-Tree was benchmarked inside GBDT on real-world learning-to-rank (LTR) and click-through rate (CTR) tasks:

Task #Training #Test d Machines PV-Tree Time Data-Parallel Attr-Parallel Sequential
LTR 11M 1M 1200 8 5,825 s 32,260 s 14,660 s 28,690 s
CTR 235M 31M 800 32 5,349 s 9,209 s 26,928 s 154,112 s

PV-Tree achieves significant speed-ups (e.g., j∗∈[d]j^* \in [d]1 and j∗∈[d]j^* \in [d]2 over sequential for LTR and CTR, respectively). Communication costs, measured for a depth-6 tree on j∗∈[d]j^* \in [d]3, j∗∈[d]j^* \in [d]4, are j∗∈[d]j^* \in [d]5 MB for PV-Tree, compared to j∗∈[d]j^* \in [d]6 MB (attribute-parallel) and j∗∈[d]j^* \in [d]7 MB (data-parallel). In terms of accuracy, PV-Tree matches or exceeds these baselines for practical values of j∗∈[d]j^* \in [d]8 (j∗∈[d]j^* \in [d]9) and sufficiently large w∗∈Wj∗w^* \in W_{j^*}0 (Meng et al., 2016).

6. Design Trade-offs and Scalability Considerations

PV-Tree’s efficiency depends critically on the choice of w∗∈Wj∗w^* \in W_{j^*}1 and w∗∈Wj∗w^* \in W_{j^*}2. Increasing the number of machines (w∗∈Wj∗w^* \in W_{j^*}3) accelerates local computation but reduces per-machine sample size (w∗∈Wj∗w^* \in W_{j^*}4), which can eventually impair local top-w∗∈Wj∗w^* \in W_{j^*}5 selection accuracy. Selecting w∗∈Wj∗w^* \in W_{j^*}6 too small can degrade global split quality, while too large a w∗∈Wj∗w^* \in W_{j^*}7 increases communication costs unnecessarily. Empirical evidence suggests w∗∈Wj∗w^* \in W_{j^*}8 to w∗∈Wj∗w^* \in W_{j^*}9 is effective when Δj(w;O)\Delta_{j}(w; O)0 is large. For a fixed Δj(w;O)\Delta_{j}(w; O)1, there is an optimal Δj(w;O)\Delta_{j}(w; O)2 balancing computational and statistical efficiency (Meng et al., 2016).

7. Context and Relation to Other Parallel Decision Tree Methods

PV-Tree contrasts with attribute-parallel (vertical partitioning) approaches, which partition features across machines but require reshuffling all samples from the best attribute’s split at each split, incurring a communication cost proportional to Δj(w;O)\Delta_{j}(w; O)3. Data-parallel approaches, which aggregate all histograms across attributes, incur Δj(w;O)\Delta_{j}(w; O)4 communication, scaling linearly with Δj(w;O)\Delta_{j}(w; O)5. PV-Tree uniquely achieves communication Δj(w;O)\Delta_{j}(w; O)6, decoupling network cost from Δj(w;O)\Delta_{j}(w; O)7 or Δj(w;O)\Delta_{j}(w; O)8. The theoretical and empirical findings position PV-Tree as an effective solution for distributed tree induction in both high-dimensional and large-scale regimes (Meng et al., 2016).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Branchless SIMD in B$^S$-tree.