Bitwise portability across accelerator vendors

Validate whether the rf fused-upcast GEMM approach provides bitwise reproducibility across different accelerator vendors whose software stacks and hardware implementations differ, thereby establishing portability beyond NVIDIA GPU architectures.

Background

The paper presents rf, a reproducibility method for LLM inference that uses fixed-configuration fused-upcast GEMM kernels, IEEE-754 fused multiply-add arithmetic, and shape-dependent reduction orders. The authors validate bitwise-identical linear-layer outputs across NVIDIA Ampere, Ada, and Hopper GPUs, all within a single hardware-vendor ecosystem.

The paper explicitly identifies cross-vendor validation as unresolved because portability across different vendors involves differences in both hardware and software, rather than only differences among NVIDIA GPU architectures. Determining whether the method's bitwise guarantees survive those differences is therefore left as an open problem.

References

Validating bitwise portability across different vendors (where software, not just hardware, differ) is an open problem.

— Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures  (2609.25624 - Cooper et al., 22 Sep 2026) in Conclusion