Validate the deconstruction-limited interpretation of Rubin’s emulated-DGEMM performance

Determine whether NVIDIA’s approximately 200-TFLOPS Rubin Emulated-DGEMM specification lies on the deconstruction-limited branch of the reduced Ozaki Scheme II FP8 model, rather than reflecting a size-independent efficiency limitation.

Background

The paper models FP64-emulated DGEMM on Rubin as being limited, below a size crossover, by the cost of converting FP64 operands into CRT residue planes before FP8 tensor-core multiplication. The authors infer that NVIDIA’s approximately 200-TFLOPS specification could correspond to a deconstruction-bound operating point at an unreported problem size of roughly 500–900 in the harmonic output dimension.

This interpretation is explicitly conditional because NVIDIA has not published the benchmark dimensions, algorithmic mode, or implementation details behind the figure. A DGEMM size sweep comparing CRT/FP8 and slicing-based INT8 controls is proposed to distinguish the deconstruction hypothesis from a flat library-efficiency explanation.

References

Hypothesis: the published $\sim$200-TFLOPS Rubin Emulated-DGEMM figure lies on the deconstruction-limited branch of Eq.~eq:pdense---a deconstruction-bound, not tensor-bound, operating point at a model-inferred size.

Ozaki 2.5: Engineering the Deconstruction Path of fp64-Emulated Dense Matrix Multiplication on FP8 Tensor Cores  (2609.09095 - Matsuoka, 8 Sep 2026) in Section 3, paragraph “Where the published figure falls”; Prediction 1, Section 12

To be measured: the L1 reuse schedule, $\eta_{\text{red}$, the concurrency parameter $\theta$, storage-mode boundaries, end-to-end knees, per-set numerical behaviour, and application replay.

Ozaki 2.5: Engineering the Deconstruction Path of fp64-Emulated Dense Matrix Multiplication on FP8 Tensor Cores  (2609.09095 - Matsuoka, 8 Sep 2026) in Section 12, “Validation Plan and Falsifiable Predictions” and “Pass/fail criteria”