- The paper introduces Cholesky-Bench, a reproducible framework comparing naive and collapsed fork-join with synchronous and asynchronous tasking across OpenMP and HPX on a 128-core AMD system.
- The results show that loop collapsing nearly closes the performance gap with tasking, while asynchronous execution adds 7% speedup for OpenMP and 14% for HPX over synchronous tasking.
- HPX outperforms OpenMP by up to 26% at optimal tile sizes and achieves approximately 2 microseconds per task versus 7.6 microseconds for GCC OpenMP, although fork-join can be faster for small problems.
Motivation and scope
Fork-join parallelism, as standardized by OpenMP, remains the dominant shared-memory programming model, but its implicit barriers at the end of parallel regions penalize algorithms with inhomogeneous work distributions. Asynchronous many-task (AMT) runtimes such as HPX express computation as a dependency graph of fine-grained tasks and can eliminate redundant synchronization entirely. Although this advantage is widely assumed, the actual benefit of asynchronous tasking over a carefully written fork-join baseline is rarely quantified side by side in a controlled setting.
This paper addresses that gap with Cholesky-Bench, an open-source framework that compares four parallelization variants of the tiled right-looking Cholesky decomposition across two runtimes: the OpenMP implementations shipped with GCC and LLVM, and HPX 1.11.0. The four variants are naive fork-join, collapsed fork-join using collapse(2), synchronous tasking (OpenMP 3.0-style taskwait), and asynchronous tasking with explicit data dependencies (depend clauses for OpenMP 4.0, future-based hpx::dataflow for HPX). All eight runtime–variant combinations are benchmarked on a dual-socket node with two AMD EPYC 7742 processors (128 physical cores), sweeping tile counts from 4 to 1024 per dimension and problem sizes from 28 to 216, with medians over 20 runs of FP64 factorizations using sequential OpenBLAS inside tasks so that the runtime is the sole source of parallelism.
The choice of tiled Cholesky is deliberate: its POTRF/TRSM/SYRK/GEMM dependency structure exposes abundant fine-grained parallelism with irregular workloads, making it a canonical stress test for task schedulers. The paper also includes PLASMA and multi-threaded LAPACKE/OpenBLAS as reference lines.
Parallelization variants
The naive fork-join variant parallelizes only the outer loops of the right-looking algorithm. Its weakness is structural: near the end of the factorization the outer loop has too few iterations to feed 128 threads, and the inner trailing-submatrix-update loop is never exposed to the scheduler. The collapsed fork-join variant fuses the SYRK and GEMM update loops into a single perfectly nested loop under collapse(2), dispatching between SYRK and GEMM via an index check; this substantially increases exposed parallelism per phase while retaining the same implicit barriers.
The synchronous tasking variant uses #pragma omp task with explicit taskwait barriers, exposing exactly the same parallelism as the collapsed variant — any performance difference therefore isolates task-creation and scheduling overhead relative to fork-join. The asynchronous tasking variant expresses BLAS-level data dependencies directly, removing all global barriers. Notably, annotating tasks with priorities (OpenMP 4.5) produced no measurable improvement and was dropped.
On the HPX side, dependencies are expressed through chained futures. The authors distinguish two styles: lightweight hpx::future<void> handles for dependency tracking only (low overhead but race-prone if dependencies are mis-specified) versus wrapping the tile data itself in futures (safer, at the cost of an indirection). Tiles shared by readers use hpx::shared_future<T>. Each tile is stored as a std::vector<T> whose .data() pointer satisfies the cBLAS interface without manual buffer management. The four HPX variants mirror the OpenMP set, with fork-join implemented via hpx::experimental::for_loop.
Tile-size scaling results
At a problem size of 216 on 128 threads, all variants exhibit a well-defined sweet spot balancing sufficient parallel work against per-task overhead amortization. Three findings stand out:
- Collapsing closes most of the gap to tasking. At the optimal tile size, collapsing the inner loop yields almost 30% speedup over naive OpenMP fork-join and brings it on par with synchronous tasking.
- Asynchrony adds a further gain: +7% for OpenMP and +14% for HPX when moving from synchronous to asynchronous tasking.
- HPX dominates across the board. At their respective best tile sizes, the HPX fork-join, collapsed fork-join, synchronous-tasking, and asynchronous-tasking variants are 30%, 15%, 21%, and 26% faster than their OpenMP counterparts.
Both stacks outperform the PLASMA reference at their best tile sizes; at PLASMA's default tile size the OpenMP variants merely match it, indicating the gap stems from tile-size configuration rather than implementation superiority.
The authors appropriately caution that the cross-runtime comparison holds algorithm, BLAS layer, compiler, pinning, and problem generator fixed, but OpenMP and HPX differ structurally in scheduling strategy — particularly for fork-join — so the cross-runtime numbers should be read as a comparison of the two stacks as typically deployed, not as a pure measurement of intrinsic runtime overhead.
Problem-size scaling and task overhead
Fixing tile counts at 16–128 per dimension and varying problem size isolates task-management costs. A no-op curve, in which all BLAS calls are replaced by stubs, approximates pure task creation, dependency tracking, and scheduling cost. Two results are notable:
- For OpenMP, classical fork-join outperforms asynchronous tasking up to a certain problem size, because OpenMP's task overhead is substantial relative to small amounts of BLAS work. The optimal tile count is problem-size dependent: finely splitting small problems to occupy all cores is counterproductive.
- For HPX, asynchronous tasking dominates fork-join across the entire problem-size range for 32, 64, and 128 tiles per dimension, with no crossover; only at 16 tiles do the variants become indistinguishable.
Dividing the measured no-op runtime by the analytic task count (n POTRF, n(n−1)/2 TRSM and SYRK each, n(n−1)(n−2)/6 GEMM tasks) yields an effective per-task overhead that is nearly independent of tile count, confirming linear overhead growth: approximately 2μs per task for HPX versus 7.6μs for GCC OpenMP — roughly 3.8× smaller for HPX. This is consistent with HPX being designed around asynchronous tasking from the outset rather than having tasks added to a fork-join-centric standard, though the authors note explicitly that establishing this advantage across architectures requires a broader study.
Compiler comparison
Recompiling identical source with LLVM/Clang 22.1.2 (OpenMP 5.1) instead of GCC 14.2.0 (OpenMP 4.5) reveals compiler-specific behavior. Tasking and naive fork-join performance are essentially identical between compilers at optimal tile sizes, and LLVM exhibits lower task-creation overhead for dependency-free tasks. However, GCC is 44% faster on the collapsed fork-join variant. The cause is a divergence in standard conformance: because the fused trailing-update loop is non-rectangular, the OpenMP specification forbids attaching an explicit schedule clause to a collapse(2) construct. GCC enforces this strictly (dynamic scheduling fails to compile); LLVM accepts it as a non-standard extension, and with dynamic scheduling enabled the gap largely closes. The reported numbers reflect the standard-conforming code path for both compilers — a practically important caveat for anyone tuning collapsed loops portably.
Limitations and open questions
The paper is candid about its boundaries. All measurements come from a single dual-socket AMD Zen 2 node, so the reported speedups are platform-specific; validation on Intel, ARM, and RISC-V microarchitectures is left open. Only the right-looking traversal is studied, leaving the question of whether left-looking or top-looking variants would change the conclusions. The framework covers only OpenMP and HPX; extending to TTG or StarPU under the same software stack remains future work. Finally, the study is confined to shared memory, so the findings say nothing about distributed AMT behavior or applications combining irregular computation with communication.
Conclusion
Cholesky-Bench provides a unified, reproducible basis for comparing fork-join and task-based parallelization of the tiled Cholesky decomposition. Its central quantitative findings are that collapsed fork-join recovers most of the load-imbalance penalty of naive fork-join and matches synchronous tasking; that removing redundant barriers via asynchronous tasking yields an additional 7% (OpenMP) to 14% (HPX); and that HPX outperforms OpenMP across all four variants at optimal tile size, with asynchronous HPX tasks up to 26% faster and exhibiting roughly 3.8× smaller per-task overhead than GCC OpenMP. At the same time, the results temper the narrative that fork-join is obsolete: for small problems, fork-join can beat asynchronous tasking outright, and the collapsed variant remains a strong low-overhead alternative. The GCC-versus-LLVM discrepancy on collapsed loops further underscores that measured performance depends on compiler-specific interpretations of the OpenMP standard, not only on the programming model chosen.