Papers
Topics
Authors
Recent
Search
2000 character limit reached

Exploiting Task-Based Parallelism for the Red-Black Gauss-Seidel Method on 2D Grids

Published 2 Jul 2026 in cs.DC and math.NA | (2607.01735v1)

Abstract: Gauss-Seidel is a well-established iterative method for the solution of linear systems, and multicoloring has been widely used to increase parallelism in iterative solution techniques. Implementing multi-color Gauss-Seidel with conventional divide-and-conquer parallelization strategies, however, may be inefficient due to global synchronization requirements and load imbalances. Task-based programming models can mitigate these issues by enabling fine-grained parallelism, removing global barriers and allowing updates of different colors to partially overlap in time. In this work, we implement the red-black Gauss-Seidel method using two task-based programming models and compare them with a classical divide-and-conquer parallel implementation to evaluate the impact of fine-grained parallelism on execution efficiency. The red-black scheme serves as a representative example, as task-based approaches naturally extend to more general multi-color schemes arising from unstructured grids and wider stencils. Using the solve of the 2D Poisson equation as benchmark, our results show that task-based implementations can achieve performance comparable to conventional divide-and-conquer parallelization while providing greater resilience to hardware-level asynchronicity.

Summary

  • The paper demonstrates that task-based scheduling, particularly with OmpSs-2, can outperform traditional fork-join models for solving the 2D Poisson equation using RBGS.
  • It employs rigorous benchmarking on many-core NUMA systems with a regular stencil computation, revealing enhanced throughput, stability, and NUMA-aware efficiency.
  • The study underscores the potential of dynamic scheduling for future scientific applications, exascale computing, and improved handling of hardware asynchronicity.

Task-Based Parallelism for Red-Black Gauss-Seidel on 2D Grids: Analysis and Performance

Introduction and Context

This work investigates task-based parallelism applied to the red-black Gauss-Seidel (RBGS) iterative method for solving the 2D Poisson equation, with a detailed comparison between OpenMP parallel for (fork-join), OpenMP task, and OmpSs-2 (task-based) models. The primary objective is to assess the efficiency, resilience to hardware asynchronicity, and programmability of these paradigms on contemporary many-core, NUMA architectures. By focusing on a regular stencil computation—a canonical regular workload in high-performance computing—the study provides quantitative and architectural insight into whether and when task-based scheduling yields tangible advantages over conventional fork-join paradigms.

Algorithmic Design and Parallelization Strategies

The RBGS approach partitions the 2D grid into red and black sites following a checkerboard pattern, such that all points of the same color can be updated independently. This decoupling is classically exploited for parallelization via either coarse- or fine-grained decomposition. The fork-join paradigm utilizes static work distribution with bulk-synchronous barriers after every color’s update phase, while task-based schemes dynamically schedule blocks of red/black sites to threads according to data dependencies, allowing for overlap and finer control over granularity.

Figure 1

Figure 1: Data decomposition of the 2D grid into horizontal stripes—each assigned as a task—enabling flexible granularity in task-based models.

The task-based implementation divides the grid into nbnb horizontal stripes, resulting in 2⋅nb2 \cdot nb tasks per iteration. Tasks are issued using the appropriate pragma constructs (OpenMP or OmpSs-2) with explicit data dependencies to guarantee correctness in the presence of overlapping task execution.

Experimental Setup and Benchmarking Methodology

Experiments were carried out on two hardware platforms: JUWELS (dual-socket Intel Xeon, two NUMA domains, mesh interconnect) and HAICGU (dual-socket Kunpeng 920, four NUMA domains, ring topology). Both platforms are representative of state-of-the-art many-core systems with distinct memory topologies and bandwidth characteristics.

Key benchmarking protocols include:

  • Execution of 100 iterations of RBGS for various grid sizes to capture scaling and resource utilization trends.
  • Use of a 1:1 core-to-thread mapping, strong memory affinity, and careful control of initialization to avoid NUMA biases.
  • Attainable performance estimation via the roofline model, confirming the problem regime is strictly memory-bound given the low arithmetic intensity of the RBGS kernel.

Performance Results

The empirical results demonstrate that task-based paradigms, and specifically OmpSs-2, can match or exceed the performance of hand-tuned fork-join parallel for implementations, contrary to prior studies suggesting minimal benefit for regular workloads. OmpSs-2’s NUMA-aware scheduling and memory allocation capabilities allow it to sustain higher, more stable throughput, particularly under increased hardware asynchronicity and resource contention.

Figure 2

Figure 2: Median performance of RBGS (in GFLOP/s) across matrix sizes for OpenMP parallel for, OpenMP task, OmpSs-2, and predicted attainable limits.

OmpSs-2 achieves performance at or near the peak memory bandwidth limit indicated by the roofline model. When using OmpSs-2’s NUMA-aware allocation, performance degradation due to NUMA misses is effectively mitigated, as evidenced by the approximately 40–43% drop observed if such features are not used. This underscores the impact of memory placement and runtime scheduling on stencil computation efficiency in many-core systems.

In addition to median throughput, the analysis exposes a sharp contrast in run-to-run performance variability. HAICGU, with its larger number of NUMA domains, exhibits significant execution-time variability under fork-join scheduling, resulting in increased idle time and greater performance variance.

Figure 3

Figure 3: Box-and-whisker plots of RBGS performance highlighting reduced variance for OmpSs-2 compared to OpenMP parallel for, especially on HAICGU.

Runtime Behavior and Execution Traces

Instrumentation and trace analysis using Extrae, Ovni, and Paraver enable a fine-grained view of runtime thread behavior. For OpenMP parallel for, the strict synchronization barriers result in well-demarcated execution phases, but frequent stalling occurs as slower threads create idle gaps.

Figure 4

Figure 4

Figure 4: Mapping OpenMP execution to CPU cores: blue segments show idle threads during synchronization.

In contrast, OmpSs-2 demonstrates high thread occupancy throughout execution, with partial overlaps between red and black site updates across iterations. This dynamic scheduling capability allows the runtime to absorb transient variability due to OS noise, hardware heterogeneity, or resource contention. The result is not only improved absolute throughput, but also increased robustness to non-deterministic runtime conditions.

Figure 5

Figure 5

Figure 5: OmpSs-2 mapping shows extensive overlapping of computation across threads—demonstrating resilience to asynchronicity and superior load balance.

Notably, cache reuse was investigated via hardware counters. For the grid sizes and task granularities considered, neither paradigm evidenced a meaningful difference in L2/L3 miss rates due to the working set size exceeding the cache capacity—minimizing cache-locality-induced variance between scheduling models.

Implications and Broader Impact

These findings contradict earlier assumptions that task-based models are inferior for regular stencil computations. As modern architectures push toward greater core counts, deeper memory hierarchies, and increased hardware asynchronicity, the practical advantages of fully dynamic, dependency-aware scheduling increase even for traditionally regular workloads. The results are pronounced for architectures with more NUMA domains and noisier resource sharing, but even on “simpler” platforms, the performance gap with fork-join models closes.

The programming effort required to transition from parallel for to task-based OpenMP or OmpSs-2 remains moderate, especially when compared to fully distributed-memory paradigms (e.g., MPI). The added flexibility and robustness may become increasingly valuable as future architectures increase hardware-induced variability and non-uniform resource costs.

For AI and scientific computing, these insights extend to a large class of stencil, PDE, and iterative methods beyond the Poisson equation, with implications for scalable solver development on exascale and heterogeneous systems. Task-based runtimes with improved NUMA-awareness and scheduling intelligence can facilitate high, stable throughput independent of computational regularity.

Conclusion

This study establishes that task-based execution, particularly with NUMA-aware runtimes such as OmpSs-2, is not just competitive but can outperform fork-join parallel for scheduling in RBGS for the 2D Poisson equation, especially as core count and architectural complexity increase. Fine-grained dynamic scheduling efficiently mitigates variability induced by the operating system and hardware topology, yielding both higher and lower-variance throughput. These findings suggest a reevaluation of task-based paradigms for other regular scientific workloads and support their adoption in next-generation scientific and AI software targeting emergent many-core platforms (2607.01735).

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.