Quantitative prediction of progress divergence from memory pressure

Determine the degree of progress divergence among parallel GPU workers as a quantitative function of memory pressure, accounting for the fact that kernels operating at the same pressure can exhibit different divergence levels.

Background

The paper studies parallel scan kernels in which multiple GPU workers asynchronously scan the same sequence of blocks. Progress divergence causes workers’ phases to separate, reducing shared-cache reuse and increasing off-chip traffic. The authors experimentally vary occupancy, data path, software-pipeline depth, and compute work, and observe that memory pressure is related to divergence but does not uniquely determine it.

The unresolved issue is to derive a quantitative predictor or law for divergence from memory pressure and other relevant kernel characteristics. Such a result would improve understanding of when shared-cache traffic deteriorates and would support more principled prediction beyond the paper’s empirical calibration pipeline.

References

It is clear that progress divergence is highly relevant to memory subsystem contention; there is no progress divergence when the memory pressure is low. However, we cannot quantitatively calculate the divergence according to the memory pressure because kernels at the same pressure can present different degrees of divergence.

— PASCAL: A Phase-Aware Shared-Cache Model for Parallel Scans  (2609.10515 - Zhou et al., 9 Sep 2026) in Section 5.2, “Why dynamic prediction is necessary”