Performance portability across GPU vendors for irregular kernels

Establish performance portability across GPU vendors for irregular kernels, achieving parity performance from a single source across NVIDIA CUDA and AMD ROCm/HIP targets.

Background

Awkward Array supports nested and jagged data structures whose index manipulation and irregular memory-access patterns create substantial performance challenges on GPUs. The paper explains that CUDA kernels for these workloads are comparatively mature, whereas direct CUDA-to-HIP ports can compile and run while suffering severe performance degradation on AMD hardware.

The authors identify the unresolved portability gap as more than a toolchain issue: differences between NVIDIA warps and AMD wavefronts require target-specific optimization. Although the paper demonstrates optimization patterns that recover parity for the studied kernels, it presents vendor-independent performance portability across the broader class of irregular kernels as an unresolved problem.

References

Performance portability across vendors---one source, parity performance on each target---remains substantially unsolved.

— Bridging the Vendor Gap: Enabling AMD GPU Support for Awkward Array via ROCm/HIP for the HL-LHC Era  (2609.24628 - Osborne et al., 21 Sep 2026) in Section 2, “Background and motivation,” subsection “Why GPUs, and why portability”