Papers
Topics
Authors
Recent
Search
2000 character limit reached

From Fork-Join to Asynchronous Tasks: Parallelizing Tiled Cholesky Decomposition with OpenMP and HPX

Published 10 Jun 2026 in cs.DC and cs.PF | (2606.11937v1)

Abstract: Fork-join parallelism, popularized by OpenMP, remains the dominant model for shared-memory parallel programming, but its implicit synchronization barriers can penalize algorithms with inhomogeneous workloads. Asynchronous many-task (AMT) runtimes sidestep these barriers by expressing work as a dependency graph of fine-grained tasks. Yet, the actual performance benefit over a carefully written fork-join baseline is rarely quantified. In this work, we introduce Cholesky-Bench and use it to revisit the tiled Cholesky decomposition, a canonical irregular kernel, comparing four parallelization variants of the right-looking algorithm across two runtimes: the OpenMP implementations shipped with GCC and LLVM, and the HPX AMT runtime. The variants span classical fork-join, a collapsed fork-join that exposes additional inner-loop parallelism, synchronous tasking, and asynchronous tasking with explicit data dependencies. We benchmark all eight combinations on a dual-socket 128-core AMD Zen 2 node across multiple tile sizes and problem sizes. Our results show that across all variants, HPX outperforms OpenMP at the optimal tile size by 15%-30%. Specifically, asynchronous HPX tasks are up to 26% faster than their OpenMP counterparts, and exhibit roughly 3.8x smaller task overhead. Furthermore, the collapsed fork-join variants close most of the gap to synchronous tasking. Removing redundant synchronization barriers yields an additional improvement of 7% (OpenMP) to 14% (HPX). A GCC-versus-LLVM comparison further reveals compiler-specific differences in fork-join scheduling and task-creation overheads.

Summary

  • The paper introduces Cholesky-Bench, a reproducible framework comparing naive and collapsed fork-join with synchronous and asynchronous tasking across OpenMP and HPX on a 128-core AMD system.
  • The results show that loop collapsing nearly closes the performance gap with tasking, while asynchronous execution adds 7% speedup for OpenMP and 14% for HPX over synchronous tasking.
  • HPX outperforms OpenMP by up to 26% at optimal tile sizes and achieves approximately 2 microseconds per task versus 7.6 microseconds for GCC OpenMP, although fork-join can be faster for small problems.

Motivation and scope

Fork-join parallelism, as standardized by OpenMP, remains the dominant shared-memory programming model, but its implicit barriers at the end of parallel regions penalize algorithms with inhomogeneous work distributions. Asynchronous many-task (AMT) runtimes such as HPX express computation as a dependency graph of fine-grained tasks and can eliminate redundant synchronization entirely. Although this advantage is widely assumed, the actual benefit of asynchronous tasking over a carefully written fork-join baseline is rarely quantified side by side in a controlled setting.

This paper addresses that gap with Cholesky-Bench, an open-source framework that compares four parallelization variants of the tiled right-looking Cholesky decomposition across two runtimes: the OpenMP implementations shipped with GCC and LLVM, and HPX 1.11.0. The four variants are naive fork-join, collapsed fork-join using collapse(2), synchronous tasking (OpenMP 3.0-style taskwait), and asynchronous tasking with explicit data dependencies (depend clauses for OpenMP 4.0, future-based hpx::dataflow for HPX). All eight runtime–variant combinations are benchmarked on a dual-socket node with two AMD EPYC 7742 processors (128 physical cores), sweeping tile counts from 4 to 1024 per dimension and problem sizes from 282^8 to 2162^{16}, with medians over 20 runs of FP64 factorizations using sequential OpenBLAS inside tasks so that the runtime is the sole source of parallelism.

The choice of tiled Cholesky is deliberate: its POTRF/TRSM/SYRK/GEMM dependency structure exposes abundant fine-grained parallelism with irregular workloads, making it a canonical stress test for task schedulers. The paper also includes PLASMA and multi-threaded LAPACKE/OpenBLAS as reference lines.

Parallelization variants

The naive fork-join variant parallelizes only the outer loops of the right-looking algorithm. Its weakness is structural: near the end of the factorization the outer loop has too few iterations to feed 128 threads, and the inner trailing-submatrix-update loop is never exposed to the scheduler. The collapsed fork-join variant fuses the SYRK and GEMM update loops into a single perfectly nested loop under collapse(2), dispatching between SYRK and GEMM via an index check; this substantially increases exposed parallelism per phase while retaining the same implicit barriers.

The synchronous tasking variant uses #pragma omp task with explicit taskwait barriers, exposing exactly the same parallelism as the collapsed variant — any performance difference therefore isolates task-creation and scheduling overhead relative to fork-join. The asynchronous tasking variant expresses BLAS-level data dependencies directly, removing all global barriers. Notably, annotating tasks with priorities (OpenMP 4.5) produced no measurable improvement and was dropped.

On the HPX side, dependencies are expressed through chained futures. The authors distinguish two styles: lightweight hpx::future<void> handles for dependency tracking only (low overhead but race-prone if dependencies are mis-specified) versus wrapping the tile data itself in futures (safer, at the cost of an indirection). Tiles shared by readers use hpx::shared_future<T>. Each tile is stored as a std::vector<T> whose .data() pointer satisfies the cBLAS interface without manual buffer management. The four HPX variants mirror the OpenMP set, with fork-join implemented via hpx::experimental::for_loop.

Tile-size scaling results

At a problem size of 2162^{16} on 128 threads, all variants exhibit a well-defined sweet spot balancing sufficient parallel work against per-task overhead amortization. Three findings stand out:

  1. Collapsing closes most of the gap to tasking. At the optimal tile size, collapsing the inner loop yields almost 30% speedup over naive OpenMP fork-join and brings it on par with synchronous tasking.
  2. Asynchrony adds a further gain: +7% for OpenMP and +14% for HPX when moving from synchronous to asynchronous tasking.
  3. HPX dominates across the board. At their respective best tile sizes, the HPX fork-join, collapsed fork-join, synchronous-tasking, and asynchronous-tasking variants are 30%, 15%, 21%, and 26% faster than their OpenMP counterparts.

Both stacks outperform the PLASMA reference at their best tile sizes; at PLASMA's default tile size the OpenMP variants merely match it, indicating the gap stems from tile-size configuration rather than implementation superiority.

The authors appropriately caution that the cross-runtime comparison holds algorithm, BLAS layer, compiler, pinning, and problem generator fixed, but OpenMP and HPX differ structurally in scheduling strategy — particularly for fork-join — so the cross-runtime numbers should be read as a comparison of the two stacks as typically deployed, not as a pure measurement of intrinsic runtime overhead.

Problem-size scaling and task overhead

Fixing tile counts at 16–128 per dimension and varying problem size isolates task-management costs. A no-op curve, in which all BLAS calls are replaced by stubs, approximates pure task creation, dependency tracking, and scheduling cost. Two results are notable:

  • For OpenMP, classical fork-join outperforms asynchronous tasking up to a certain problem size, because OpenMP's task overhead is substantial relative to small amounts of BLAS work. The optimal tile count is problem-size dependent: finely splitting small problems to occupy all cores is counterproductive.
  • For HPX, asynchronous tasking dominates fork-join across the entire problem-size range for 32, 64, and 128 tiles per dimension, with no crossover; only at 16 tiles do the variants become indistinguishable.

Dividing the measured no-op runtime by the analytic task count (nn POTRF, n(n−1)/2n(n-1)/2 TRSM and SYRK each, n(n−1)(n−2)/6n(n-1)(n-2)/6 GEMM tasks) yields an effective per-task overhead that is nearly independent of tile count, confirming linear overhead growth: approximately 2 μs2\,\mu\mathrm{s} per task for HPX versus 7.6 μs7.6\,\mu\mathrm{s} for GCC OpenMP — roughly 3.8×3.8\times smaller for HPX. This is consistent with HPX being designed around asynchronous tasking from the outset rather than having tasks added to a fork-join-centric standard, though the authors note explicitly that establishing this advantage across architectures requires a broader study.

Compiler comparison

Recompiling identical source with LLVM/Clang 22.1.2 (OpenMP 5.1) instead of GCC 14.2.0 (OpenMP 4.5) reveals compiler-specific behavior. Tasking and naive fork-join performance are essentially identical between compilers at optimal tile sizes, and LLVM exhibits lower task-creation overhead for dependency-free tasks. However, GCC is 44% faster on the collapsed fork-join variant. The cause is a divergence in standard conformance: because the fused trailing-update loop is non-rectangular, the OpenMP specification forbids attaching an explicit schedule clause to a collapse(2) construct. GCC enforces this strictly (dynamic scheduling fails to compile); LLVM accepts it as a non-standard extension, and with dynamic scheduling enabled the gap largely closes. The reported numbers reflect the standard-conforming code path for both compilers — a practically important caveat for anyone tuning collapsed loops portably.

Limitations and open questions

The paper is candid about its boundaries. All measurements come from a single dual-socket AMD Zen 2 node, so the reported speedups are platform-specific; validation on Intel, ARM, and RISC-V microarchitectures is left open. Only the right-looking traversal is studied, leaving the question of whether left-looking or top-looking variants would change the conclusions. The framework covers only OpenMP and HPX; extending to TTG or StarPU under the same software stack remains future work. Finally, the study is confined to shared memory, so the findings say nothing about distributed AMT behavior or applications combining irregular computation with communication.

Conclusion

Cholesky-Bench provides a unified, reproducible basis for comparing fork-join and task-based parallelization of the tiled Cholesky decomposition. Its central quantitative findings are that collapsed fork-join recovers most of the load-imbalance penalty of naive fork-join and matches synchronous tasking; that removing redundant barriers via asynchronous tasking yields an additional 7% (OpenMP) to 14% (HPX); and that HPX outperforms OpenMP across all four variants at optimal tile size, with asynchronous HPX tasks up to 26% faster and exhibiting roughly 3.8×3.8\times smaller per-task overhead than GCC OpenMP. At the same time, the results temper the narrative that fork-join is obsolete: for small problems, fork-join can beat asynchronous tasking outright, and the collapsed variant remains a strong low-overhead alternative. The GCC-versus-LLVM discrepancy on collapsed loops further underscores that measured performance depends on compiler-specific interpretations of the OpenMP standard, not only on the programming model chosen.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.