Validate the deconstruction-limited interpretation of Rubin’s emulated-DGEMM performance
Determine whether NVIDIA’s approximately 200-TFLOPS Rubin Emulated-DGEMM specification lies on the deconstruction-limited branch of the reduced Ozaki Scheme II FP8 model, rather than reflecting a size-independent efficiency limitation.
References
Hypothesis: the published $\sim$200-TFLOPS Rubin Emulated-DGEMM figure lies on the deconstruction-limited branch of Eq.~eq:pdense---a deconstruction-bound, not tensor-bound, operating point at a model-inferred size.
— Ozaki 2.5: Engineering the Deconstruction Path of fp64-Emulated Dense Matrix Multiplication on FP8 Tensor Cores
(2609.09095 - Matsuoka, 8 Sep 2026) in Section 3, paragraph “Where the published figure falls”; Prediction 1, Section 12
To be measured: the L1 reuse schedule, $\eta_{\text{red}$, the concurrency parameter $\theta$, storage-mode boundaries, end-to-end knees, per-set numerical behaviour, and application replay.
— Ozaki 2.5: Engineering the Deconstruction Path of fp64-Emulated Dense Matrix Multiplication on FP8 Tensor Cores
(2609.09095 - Matsuoka, 8 Sep 2026) in Section 12, “Validation Plan and Falsifiable Predictions” and “Pass/fail criteria”