---
title: Parallel Voting Decision Tree (PV-Tree)
url: https://www.emergentmind.com/topics/parallel-decision-tree-learning-pv-tree
type: topic
---

# Parallel Voting Decision Tree (PV-Tree)

Parallel Voting Decision Tree (PV-Tree) is a communication-efficient distributed algorithm for learning decision trees and their ensemble variants, such as Gradient Boosted Decision Trees (GBDT) and Random Forest, in large-scale settings. The core innovation of PV-Tree lies in its two-tiered voting protocol for attribute selection, which enables substantial reduction in communication overhead while preserving the statistical fidelity of split selection. PV-Tree is designed for horizontally partitioned data distributed across multiple compute nodes, with formal theoretical guarantees and empirical validation on industry-scale datasets [1611.01276].

## 1. Problem Setting and Objective

PV-Tree operates on a training dataset $D = \{ (x_i, y_i) \}_{i=1}^N$, partitioned horizontally across $M$ machines, such that each machine $m$ holds $D_m$ with $|D_m| = n = N/M$. For each split at a tree node $O$, the goal is to identify the attribute $j^*$ and split point $w^*$ maximizing a node-informativeness criterion, such as information gain (for classification) or variance reduction (for regression):

- Information Gain (IG): 
  $$
  IG_j(w;O) = H(Y|O) - [ P_L H(Y|X_j < w) + P_R H(Y|X_j \geq w) ]
  $$
- Variance Gain (VG):
  $$
  VG_j(w;O) = \operatorname{Var}(Y|O) - [ P_L \operatorname{Var}(Y|X_j < w) + P_R \operatorname{Var}(Y|X_j \geq w) ]
  $$
with $P_L = P(X_j < w | O)$, $P_R = 1 - P_L$, and $H(\cdot|\cdot), \operatorname{Var}(\cdot|\cdot)$ denoting conditional entropy and variance, respectively.

## 2. Algorithmic Framework

The PV-Tree algorithm finds the best split at each node via three main phases: local voting, global voting, and final split selection:

1. **Local Voting**: Each machine computes the local gain $\Delta_{m,j}^* = \max_{w} \Delta_{m,j}(w; D_m)$ for every attribute $j$, selecting the top-$k$ locally best attributes $S_m$, using binning (typically $B$ bins) for continuous features.
2. **Global Voting**: All machines exchange their local top-$k$ sets $S_m$ (e.g., using MPI\_AllGather). Global vote counts $f_j = |\{ m : j \in S_m \}|$ are computed for each attribute. The top-2$k$ globally most-voted attributes are selected as $G = \{ j_{(1)}, \ldots, j_{(2k)} \}$.
3. **Final Split Selection**: Each machine sends full-binned histograms $H_{m,j}$ for $j \in G$ to a master (or all-reduce operation). The master aggregates global histograms $H_j = \sum_{m=1}^M H_{m,j}$ and scans for the split $(j^*, w^*) = \arg\max_{j\in G, w} \Delta_j(w; H_j)$.

The three phases are formalized by the following pseudocode:

```plaintext
Algorithm PV-Tree_FindBestSplit(D_m on m=1…M, k)
  1. Each machine m:
      - Build H_{m,j} for j=1 to d.
      - For each j, compute Δ_{m,j}^* = max_w Δ_j(w; H_{m,j}).
      - Select S_m <- top-k attributes by Δ_{m,j}^*.
  2. AllGather({S_m}_{m=1}^M) => {f_j}.
  3. G <- 2k attributes with highest f_j.
  4. Each m sends {H_{m,j}: j∈G} to master.
     Master forms H_j = Σ_m H_{m,j}.
  5. Return (j*,w*) = argmax_{j∈G, w} Δ_j(w; H_j).
```
[1611.01276]

## 3. Communication Complexity and Scalability

The communication efficiency of PV-Tree is achieved by restricting the exchange of full-grained histograms to a subset of attributes:

- **Local Voting**: Each machine sends $k$ attribute indices ($O(Mk)$ words).
- **Global Voting and Aggregation**: Each machine sends $2k$ histograms of $B$ bins each ($O(M \cdot 2k \cdot B)$ words).
- **Total**: $O(MkB)$ words per split iteration, independent of the total number of attributes $d$ or training instances $N$.

Comparative communication costs for one split iteration (per node):

| Method            | Communication per split | Dependent on $d$? | Dependent on $N$? |
|-------------------|------------------------|-------------------|-------------------|
| PV-Tree           | $O(MkB)$               | No                | No                |
| Data-parallel     | $O(MdB)$               | Yes               | No                |
| Attribute-parallel| $O(N)$                 | No                | Yes               |
[1611.01276]

This configuration supports scalable and efficient distributed training, with empirical evidence of significant reductions in communication volume (e.g., 10MB for PV-Tree vs. 424MB for data-parallel when $N=1$e9, $d=1200$, $k=15$).

## 4. Theoretical Analysis and Accuracy Guarantees

PV-Tree provides formal probabilistic guarantees that its two-stage voting protocol will, with high probability, select the statistically optimal split attribute:

Given true information gain rankings $IG_{(1)} \geq IG_{(2)} \geq \dots \geq IG_{(d)}$, the probability that PV-Tree selects the best attribute satisfies

$$
P(\text{best attribute selected}) \geq \sum_{m=\lfloor M/2 \rfloor + 1}^M {M \choose m} [1 - \sum_{j=k+1}^d \delta_{(j)}]^m [\sum_{j=k+1}^d \delta_{(j)}]^{M-m}
$$

where for $j > k$, $\delta_{(j)}(n, k) = \alpha_{(j)}(n) + 4 \exp(-c_{(j)} n l_{(j)}(k)^2)$ with $l_{(j)}(k) = |IG_{(1)} - IG_{(j)}|/2$, $c_{(j)} > 0$, $\alpha_{(j)}(n) \rightarrow 0$ as $n \rightarrow \infty$. As $n \rightarrow \infty$, this probability approaches 1 for fixed $M$, $k$, and $d$.

This result arises from a combination of concentration bounds comparing empirical and true information gain, and a combinatorial Majoritarian (binomial) argument for the global voting phase [1611.01276].

## 5. Empirical Performance and Trade-offs

Extensive experiments demonstrate the superior performance of PV-Tree in terms of wall-clock convergence and communication efficiency across industrial-scale tasks, specifically for GBDT learning:

Example summary (training on Gradient Boosted Decision Trees):

| Task | # Training | # Test | $d$ | Machines | Sequential | Data-parallel | Attr.-parallel | PV-Tree    |
|------|------------|--------|-----|----------|------------|---------------|----------------|------------|
| LTR  | 11M        | 1M     | 1200|     8    | 28,690 s   | 32,260 s      | 14,660 s       | 5,825 s    |
| CTR  | 235M       | 31M    |  800|    32    |154,112 s   | 9,209 s       | 26,928 s       | 5,349 s    |
[1611.01276]

PV-Tree achieves up to $4.9 \times$ speedup over sequential training on LTR data using eight machines and $28.8 \times$ speedup on CTR using 32 machines. Communication cost analyses indicate an order-of-magnitude lower transfer volume relative to alternatives at equivalent accuracy.

A recognized trade-off is that increasing $M$ reduces per-node data but increases parallelism; optimal $M$ depends on fixed $N$. Excessively small $k$ may degrade accuracy, but $k=5 \ldots 20$ is sufficient with large $n$. The design ensures statistical correctness is not compromised as communication is reduced.

## 6. Comparison to Other Parallel Decision Tree Frameworks

PV-Tree contrasts with both data-parallel and attribute-parallel (vertical) approaches:

- Data-parallel approaches exchange histogram data for all attributes, incurring $O(MdB)$ cost.
- Attribute-parallel methods require global reshuffling of example-indexed values for selected attributes, entailing $O(N)$ communication per split.
- Vertical Hoeffding Tree (VHT) [1607.08325] achieves parallelization by distributing features across workers with a Model Aggregator architecture but is tailored for streaming scenarios and uses a Hoeffding bound-based split protocol.

The two-stage voting of PV-Tree leverages horizontal partitioning and candidate restriction, thus scaling efficiently to scenarios with high $d$, large $N$, and many compute nodes.

## 7. Application Domains and Limitations

PV-Tree is directly applicable to GBDT and Random Forest learning on large-scale tabular data. By sharply lowering communication costs, it enables distributed model induction in environments with limited network bandwidth or where high feature dimensionality would otherwise bottleneck attribute selection. A plausible implication is enhanced applicability on real-world tasks in advertising (CTR), ranking (LTR), and other domains with massive datapoints and features.

Known limitations include reduced per-node sample size as $M$ increases, which *may* affect statistical power if not offset by larger $k$ or $n$. Careful tuning of $k$ is requisite for balancing accuracy and communication efficiency. These characteristics position PV-Tree as a method with broad practical scalability, strong theoretical underpinnings, and empirical validation as a state-of-the-art approach in scalable parallel decision tree induction [1611.01276].

Source: https://www.emergentmind.com/topics/parallel-decision-tree-learning-pv-tree