Bitwise portability across accelerator vendors
Validate whether the rf fused-upcast GEMM approach provides bitwise reproducibility across different accelerator vendors whose software stacks and hardware implementations differ, thereby establishing portability beyond NVIDIA GPU architectures.
References
Validating bitwise portability across different vendors (where software, not just hardware, differ) is an open problem.
— Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures
(2609.25624 - Cooper et al., 22 Sep 2026) in Conclusion