Papers
Topics
Authors
Recent
Search
2000 character limit reached

Concurrent Splay-Based Tree

Published 27 Jun 2026 in cs.DC and cs.DS | (2606.28889v1)

Abstract: Most work on efficient concurrent ordered indices, such as concurrent binary search trees, B-trees, skip lists, etc., has focused on data structures that provide good \emph{worst-case} guarantees. In real workloads, objects are often accessed at different rates, since access distributions may be non-uniform. Many efficient distribution-adaptive data structures exist in the sequential case; however, they are often complicated to make efficient in the concurrent case. The most prominent distribution-adaptive data structure is Splay Tree. Its most important advantage is that it does not store any balancing information and provides a reasonable performance improvement on extremely skewed workloads, such as Zipfian workloads. This paper proposes a splay-like rotation design for concurrent binary search trees. Instead of moving an accessed node to the root, rotations use two depth thresholds that are based on the static-optimality complexity computed from the number of accesses to the node: a node is rotated only when it is substantially deeper than the upper threshold, and rotations of the node stop before reaching the lower threshold. This design aims to preserve the main practical benefit of splaying on skewed workloads while reducing contention near the root. We present two variants of the rotation design: one using an exact 64-bit access counter per node and one using a 6-bit approximate counter. We prove static optimality for the corresponding sequential read-only tree and evaluate both rotation designs by implementing them on top of the concurrent AVL tree of Bronson et al. Our experiments show that the approach can improve throughput on several skewed workloads.

Summary

  • The paper introduces a novel rotating scheme for concurrent BSTs that employs probabilistic, depth-threshold splay-like rotations to adapt dynamically to non-uniform access patterns.
  • It details both exact and approximate counter variants, showing how reduced per-node state can maintain static optimality while mitigating contention.
  • Experimental evaluations demonstrate improved throughput under skewed workloads, underscoring the design’s scalability and practical advantages for key-value systems.

Summary of "Concurrent Splay-Based Tree" (2606.28889)

This essay presents a comprehensive analysis of the "Concurrent Splay-Based Tree" paper. The work introduces novel techniques for constructing concurrent binary search trees (BSTs) that are adaptive to access distributions, specifically extending splay-tree-style adaptivity to concurrent data structures without resorting to expensive root-centric operations. By modifying classic splay rotations and introducing probabilistic and depth-threshold-based control mechanisms, the authors achieve practical, scalable concurrency while maintaining theoretical guarantees such as static optimality. Experiments demonstrate improved throughput under skewed workloads compared to previous state-of-the-art concurrent BSTs.

Motivation and Background

Traditional concurrent ordered indices (e.g., skip lists, AVL trees, B-trees) favor data structure invariants that deliver robust worst-case guarantees, but neglect benefits from adaptivity to non-uniform access patterns—an attribute that sequential splay trees exploit efficiently. In heavily skewed workloads typical of industrial scenarios (e.g., following Zipfian distributions), adaptive structures can drastically reduce average access costs by relocating hot keys near the root. However, naive attempts to parallelize splay trees create substantial contention bottlenecks near the root due to all accesses converging there for splaying.

Alternative approaches (CBTree, Splay-List) attempted to augment concurrent BSTs with adaptivity but introduced significant memory overhead due to per-node access tracking and struggled with tuning and scaling. The absence of a concurrent splay tree that efficiently captures adaptivity without root contention persists as a gap in practical concurrent BST design.

Algorithmic Innovations

The proposed "splay-like" concurrent BST design deviates from classical splay trees by constraining the upward movement of accessed nodes. Instead of always rotating accessed nodes to the root, nodes are only rotated if their current depth substantially exceeds a dynamically computed, workload-sensitive threshold, and even then, only until they reach a shallower, complexity-informed target depth. This threshold is parameterized with constants AA and BB, and is proportional to log⁡(m/ac(x))\log(m/ac(x)), where ac(x)ac(x) denotes accesses to key xx and mm is the global count of accesses. The procedure includes a randomized trigger for splaying on each access, mitigating unnecessary contention.

Two variants are analyzed:

  • Exact Counter Variant: Each node maintains a 64-bit integer access counter. All rotations and threshold computations use these exact counts.
  • Approximate Counter Variant: Compact Morris-style probabilistic counters, requiring as little as 6 bits per node, are used to estimate the number of accesses. This reduces memory traffic and contention on global counters, albeit with some loss in static-optimality guarantees.

Pseudocode for the 'get' operation in the exact counter case calculates on-the-fly the target and upper-bound depths, conditionally triggers randomized rotation, and moves the node upwards only as necessary.

Theoretical Foundations

The central theoretical contribution is a formal proof that the sequential version of the rotation scheme delivers static optimality: the amortized access cost per key xx is O(log⁡(m/ac(x)))O(\log(m/ac(x))), matching classical splay tree guarantees. The proof uses a potential-function argument similar to Sleator and Tarjan's amortized analysis, extended to the randomized, depth-thresholded context. In essence, rotations are undertaken only if a node substantially violates static-optimality-implied depth, with the expected amortized cost bounded in terms of access frequencies.

The approximate-counter variant draws on probabilistic counting literature, specifically the Morris counter and its theoretical guarantees, albeit with a caveat: strictly optimal amortized bounds do not directly carry over due to variance in approximate counts, but implementations offer strong empirical results.

Experimental Evaluation

The authors implemented both splay-like variants atop the Bronson et al. concurrent AVL tree baseline and performed extensive throughput benchmarking. The evaluation covers a variety of realistic workload distributions (uniform, Zipfian, 90/10, 95/5, 99/1), and both read-only and update-heavy workloads. Key parameter choices include scaling the probability of splaying inversely in proportion to the number of threads, which empirically balances adaptivity and contention.

Strong experimental results include:

  • Read-only workloads: Original AVL outperforms on strictly uniform/random workloads due to caching benefits but falls behind on heavily skewed access patterns (95/5, 99/1). Splay-like and approximate-splay-like trees match or exceed baseline throughput under these regimes. Approximate counters deliver nearly identical throughput to exact counters.
  • Update workloads: Under mixed-insertion/removal scenarios, splay-like trees consistently outperform both the AVL and CBTree baselines except on uniform distributions.

Figure 1

Figure 1

Figure 1

Figure 1

Figure 1

Figure 1: Throughput of the concurrent splay-like, approximate-splay-like, AVL, and CBTree data structures under uniform read-only workloads with 10610^6 keys.

Figure 2

Figure 2

Figure 2

Figure 2

Figure 2

Figure 2: Throughput on update-heavy workloads (mixed read/insert/delete) under uniform access distributions, highlighting performance parity of approximate and exact counter variants.

Implications and Future Directions

The presented splay-like rotation protocol for concurrent BSTs achieves the practical adaptivity of splay trees on non-uniform workloads, with substantially lower synchronization bottlenecking than existing adaptive concurrent trees. The mechanism leverages minimal per-node state, scales robustly under high thread counts, and generalizes to both exact and approximate access frequency tracking.

Practically, the approach affords high-throughput, low-latency access in key-value storage engines or transactional memory systems subject to heavy-tailed hot-spotting—a ubiquitous scenario. The extension to approximate counters is promising for memory-constrained environments and paves the way for further hardware-friendly designs.

On the theoretical side, the methodology suggests that randomized, local, and frequency-aware reorganizations can achieve adaptivity in concurrent environments with provable bounds, and points to open questions about more refined probabilistic counting and tighter amortized-cost guarantees under concurrency.

Conclusion

The paper delivers a principled concurrent splay-based tree, introducing randomized, frequency-driven, depth-thresholded rotations to reconcile static optimality with high-performance concurrency. The experimental results substantiate the claim that the design outperforms prior adaptive structures, especially under skewed workloads, using either exact or approximate access counters. The findings invite further exploration of adaptive concurrent data structures and their integration into large-scale, workload-aware systems.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.