Verify concurrent FP8 and integer-tensor service capacity

Determine whether the integer-tensor reduction GEMMs can execute concurrently with near-peak FP8 tensor-core multiplication at the independent-service rates assumed by the Ozaki 2.5 performance model.

Background

The tensor-migrated routes assume that the integer-tensor reduction workload can be serviced concurrently with the FP8 MMA stream. If the two operations instead contend for a shared tensor service, the predicted throughput must be replaced by a harmonic composition of the FP8-side and INT8-side rates.

The authors provide an uncertainty bracket between complete independence and complete serialization, but state that the actual concurrency behavior is not established. Matched-occupancy measurements are proposed to determine the interpolation parameter θ\theta.

References

The vendor int8 figure is a dense-tile rate, not a proven concurrent capacity beside near-peak fp#1{8} work (the two-endpoint concurrency bracket of \S~\ref{sec:proj} quantifies both extremes of that uncertainty).

Ozaki 2.5: Engineering the Deconstruction Path of fp64-Emulated Dense Matrix Multiplication on FP8 Tensor Cores  (2609.09095 - Matsuoka, 8 Sep 2026) in Section 5, paragraph “O2: the byte-plane reduction as tensor work—the exact two-limb encoding”; Section 9, paragraph “Tensor-service concurrency”