- The paper introduces ElasticDivide, a novel SSP scheduler that overlaps synchronization with computation to enhance parallel SpTrSV efficiency.
- The paper achieves significant speedups—7–30% on ARM and 19–60% on x86—with up to 4.5× acceleration in high-NUMA and high-synchronization regimes.
- The paper reports linear scheduling overhead and improved load balance through dynamic dependency management in irregular sparse computations.
Elasticity in Parallel Sparse Triangular Solve: A Technical Overview
Introduction and Motivation
Sparse triangular solve (SpTrSV) is a critical kernel in scientific computing, AI, and engineering workflows, typically invoked either directly in direct solvers (e.g., LU/Cholesky/QR factorization, Gauss-Seidel) or as a preconditioning step in iterative solvers such as conjugate gradients. Parallelizing SpTrSV effectively on modern multicore architectures poses significant challenges due to the fine-grained, irregular dependencies introduced by sparsity. Existing scheduling approaches either cluster dependencies to enable coarse-grained wavefront parallelism (e.g., via nested dissection or supernodes) or employ sophisticated dynamic scheduling schemes. The scaling limits of state-of-the-art SpTrSV schedulers arise largely from coordination (synchronization) overheads and the difficulty of maintaining load balance with good locality.
This work introduces the application of stale synchronous parallel (SSP) execution models to SpTrSV for the first time, developing a new DAG-based scheduler, ElasticDivide, which explicitly overlaps synchronization with computation. The practical outcome is shown to be significant geometric-mean speedups (7–30% on ARM, 19–60% on x86) over prior methods at 48-core scale, and up to 2×–4.5× acceleration in high-NUMA, high-synchronization regimes.
Stale Synchronous Parallel Scheduling Model
SSP is a generalization of bulk synchronous parallel (BSP) computation. In BSP, computation proceeds in well-defined supersteps with global barriers between them, enforcing strict synchronicity. SSP---formalized in prior distributed systems and ML work---permits compute units (cores/threads) to speculatively advance computation up to a staleness parameter s, i.e., a core can be ahead by up to s supersteps relative to another, but not more. Dependencies that cross core boundaries are thus "elongated," and synchronization can be amortized/overlapped with useful computation.
Algorithmically, given a DAG G=(V,E) with vertex weights (modeling task durations), an SSP schedule of staleness s consists of an assignment of vertices to cores and supersteps (π,σ) such that for every (v,w)∈E, σ(v)+s⋅δπ(v),π(w)≤σ(w). For s=1, this reduces to a classic BSP schedule.
The hypothesis and key insight is that the dependency structure of SpTrSV computational graphs is "elastic" enough that a substantial degree of overlapping is achievable without overflow in barrier (superstep) count or significant schedule degradation.
The ElasticDivide Scheduling Algorithm
ElasticDivide extends GrowLocal (a high-performance BSP SpTrSV scheduler) to staleness 4.5×0. The main innovations are:
- Delayed cross-core dependency resolution: Tasks whose dependencies cross core boundaries are assigned so that their readiness can be tracked at a 4.5×1 superstep lag, enabling maximal safe overlap between progress on local and remote subgraphs.
- Joint assignment of current and next superstep: Within each superstep, ElasticDivide balances between assigning immediately ready (current superstep) and imminently ready (next superstep) vertices, prioritizing both core exclusivity and task order.
Algorithmically, assignments are made by incrementally "filling" cores up to a guessed work quota per superstep, then using a parallelism score (inversely proportional to the simulated critical path in the absence of communication) to decide whether to grow or advance the superstep. The practical result is that as many vertices as possible are assigned to each superstep/core without violating staleness-constrained dependencies.

Figure 1: Geometric-mean speed-ups over Serial for various data sets on ARM Kunpeng with 10th to 90th percentile range shown, demonstrating efficiency scaling of ElasticDivide versus GrowLocal and BSP schedulers.
Implementation Highlights
- Barrier Mechanism: A "weak" (fuzzy) barrier is implemented by separating arrive and wait operations. Each core writes only its own progress flag (superstep counter); others read this, utilizing per-core caches for reduced contention. This design exploits the SSP model: barriers only enforce relative staleness, and busy-waiting is localized.
- SpTrSV SSP Kernel: The core solver consumes the computed schedule 4.5×2. Each thread executes its assigned rows per superstep, invoking wait only when needed.
Experimental Evaluation
Benchmarking was performed on three architectures: ARM Kunpeng 920 (48 cores/socket), AMD EPYC 7763 (64 cores/socket), Intel Xeon Gold 6238T (22 cores/socket), over three matrix sets: raw SuiteSparse, SuiteSparse with iChol-AMD reordering, and SuiteSparse with METIS nested dissection.

Figure 2: Performance profiles comparing SpTrSV scheduling algorithms across data sets and architectures, showing the proportion of runs within a speedup threshold of the best.
ElasticDivide consistently outperforms contemporary schedulers (GrowLocal, SpMP, HDagg), with marked advantages in nonuniform memory (NUMA) architectures and higher synchronization regimes.
Scalability Analysis

Figure 3: Scaling plots show geometric mean speed-up vs. core count, highlighting that ElasticDivide continues to scale positively where other schedulers plateau or regress as parallel resources increase.
Critical observations:
- ElasticDivide achieves 204.5×3 speed-up at 64 cores on ARM Metis, 4.5×4 better than GrowLocal.
- The advantage increases with core count and number of NUMA domains, as SSP is more effective at hiding costly synchronizations.
Synchronization Barrier Reduction
ElasticDivide matches GrowLocal in barrier count (when normalized for staleness), showing 4.5×5 to 4.5×6 the synchronizations, despite the additional flexibility. HDagg achieves larger reductions only in highly structured (e.g., nested dissection) cases but otherwise struggles.
Synchronization-Compute Overlap
The most dramatic effect of SSP scheduling emerges in speed-ups vs. classical BSP execution: geometric-mean ratios of up to 4.5×7 (ARM, large core count, difficult workloads), and improvements surpassing 4.5×8 in high-synchronization and high-NUMA settings.
Scheduling Overhead and Amortization

Figure 4: Scheduling time versus matrix nonzero count, fitted to power-law curves showing near-linear scaling of ElasticDivide and GrowLocal, with ElasticDivide exhibiting superior implementation efficiency.
ElasticDivide's serial scheduling overhead scales linearly in nonzeros and is approximately 4.5×9 lower than GrowLocal, yielding amortization break-even points as low as 12–25 solves in typical settings (s0 SpMP in specialized x86-optimized cases).
Implications and Future Directions
This work demonstrates that SpTrSV computational DAGs can in practice support SSP execution with minimal penalty, leveraging the natural elasticity of sparse dependency patterns. For end-users and system-level developers, deploying ElasticDivide can produce immediate, low-effort speedups in solver routines, especially on NUMA-heavy or highly parallel hardware.
From a theoretical standpoint, this suggests that staleness-allowing execution regimes deserve more attention in the design of task-parallel scheduling for irregular dependencies.
Obvious avenues for future research include:
- Extension to GPUs and accelerators: The proposed SSP scheduling can mask cross-device/data transfer latencies, suggesting promise for sparse linear algebra and deep learning on heterogeneous clusters.
- Adaptive staleness tuning: Dynamic adjustment of s1 based on observed system latency/load could further optimize performance.
- Integration with higher-level solvers and preconditioners: Since SpTrSV appears as a subroutine in broader sparse solver pipelines, propagating SSP scheduling benefits system-wide is an open engineering opportunity.
Conclusion
By bridging the gap between classical synchronous and fully asynchronous execution through stale-synchronous-parallel scheduling, ElasticDivide efficiently reduces synchronization-induced stalls in parallel SpTrSV. Its strong numerical results—particularly under adverse memory decompositions—and low amortization make it a compelling new standard for parallel triangular solvers. The approach is broadly applicable and likely to inspire new elastic scheduling methodologies in future AI and HPC systems.