- The paper derives a closed-form collapse threshold, showing that at sink strength Δ≈7, forward iteration with S=1 can zero 82% of non-sink probabilities, while S=256 prevents underflow through Δ=7.
- The paper finds that reverse block iteration and S=256 scaling independently remove most P-cast damage, producing comparable results and reducing MSE by roughly 3–10× versus forward iteration with S=1.
- The paper characterizes S=256 as the unique bit-exact power-of-two scale that minimizes worst-case quantization error while maximizing E4M3 coverage, although S=448 offers deeper tail coverage at higher MSE.
Overview
This paper analyzes two implementation-level choices in FP8 (E4M3) attention kernels—the KV block iteration order and the static scaling factor applied to the softmax probability matrix P before casting—and shows that both are first-order determinants of output precision under the Attention Sink phenomenon. The work is explicitly positioned as a quantitative explanation rather than a proposal: reverse iteration and S=256 scaling are already deployed in FlashAttention-3/4 on engineering grounds, and the contribution is a closed-form account of why these choices work, plus a diagnostic threshold practitioners can use to predict P-cast failure. The analysis motivated updating Tencent's hpc-ops kernel from S=1 to S=256, and suggests analogous changes for FlashInfer and TensorRT-LLM XQA, which currently use S=448.
Background: E4M3 P-casting and the sink interaction
The E4M3 format provides 126 positive representable values, with a maximum of 448, minimum normal 2−6, minimum subnormal 2−9, and round-to-zero below 2−10. Within any normal binade there are exactly eight uniformly spaced values, giving a relative precision of roughly 12.5%—about 16× worse than BF16. FlashAttention-style FP8 kernels cast P to E4M3 before the S=2560 matmul while keeping the running sum S=2561 in FP32 from pre-cast probabilities; normalization is therefore exact, and all precision loss concentrates in the numerator.
Attention Sink [(2606.06521) cites Xiao et al. and follow-ups] places initial tokens at logit scores S=2562 above the mean, with reported S=2563 at multi-thousand-token contexts. Under forward block iteration, the sink inflates the running softmax maximum to S=2564, where S=2565 is the expected maximum of S=2566 standard Gaussians (S=2567; the paper notes the asymptotic S=2568 badly overestimates this at small S=2569). All subsequent P values are then S=10, and those below S=11 silently underflow to zero.
Quantifying P-collapse
The central result is a closed-form underflow fraction. A value S=12 casts to zero iff S=13, so for unit-normal scores the collapsed fraction is
S=14
with the caveat that treating per-row S=15 as fixed makes this a mean-shift estimate that slightly understates realized collapse due to convexity of S=16 in the lower tail. Simulations at S=17, S=18 show the threshold behavior starkly: at S=19, the zeroed fraction rises from 22.3% at S=2560 to 82% at S=2561 and essentially 100% by S=2562, whereas S=2563 yields 0% through S=2564. The effective information loss (non-sink mass times zeroed fraction) peaks near 42% at S=2565—the regime where positions carrying half or more of the probability mass have most of their P values destroyed. This yields the practical diagnostic S=2566: the sink strength at which the median non-sink P value crosses the round-to-zero boundary.
A companion MSE bound, derived under an idealized pairwise-uncorrelated-S=2567 assumption (the paper concedes real S=2568 are correlated, shifting the absolute constant by an S=2569 factor), gives S=4480. A key structural observation follows: because non-sink mass vanishes at large S=4481, the MSE damage is confined to the transition region S=4482—at extreme sink strength even total collapse barely affects the output.
Reverse iteration as a sufficiency fix
Reversing the block order defers the sink block to the last iteration. Before it arrives, the running maximum is bounded by extreme-value statistics at S=4483, keeping pre-sink P values well above the cast boundary. Formally, with S=4484 a value survives whenever S=4485; for S=4486 this gives an underflow probability below S=4487. The final S=4488-correction when the sink block is processed (S=4489 at 2−60, 2−61) acts on the FP32 accumulator and costs negligible precision. Notably, the paper shows via paired 2−62-tests that reverse+2−63 and forward+2−64 are statistically indistinguishable wherever P-collapse dominates (2−65); at larger 2−66 reverse is marginally better but by only 2−67 against MSE floors of 2−68–2−69. Either fix alone therefore saturates to the same precision floor, and the choice between them can be made on engineering grounds.
Characterizing 2−90
The scale-factor analysis introduces 2−91, the normalized worst-case quantization step over the mapped range 2−92, and proves it forms a sawtooth over the E4M3 number line with power-of-two values on its lower envelope at 2−93. Specifically, every 2−94 with 2−95 achieves 2−96; every non-power-of-two 2−97 strictly exceeds it; and 2−98 incurs saturation error that also exceeds it. Combined with bit-exactness of IEEE 754 scaling by powers of two (exponent-only operations, no rounding) and monotone improvement of normal-range coverage with 2−99, this yields a constructive characterization: restricted to 2−100, 2−101 is the unique scale satisfying all three conditions simultaneously.
The paper is careful about what this optimality does not claim. Condition (C3) alone—"largest 2−102"—already uniquely identifies 256 among powers of two, so (C2) is explanatory rather than discriminating. More importantly, dropping bit-exactness admits 2−103, which attains strictly better deep-tail coverage (2−104 vs. 2−105) but pays a 14% larger worst-case step, predicting ~30% higher MSE in the worst case. The measured gap is 10–15%, consistent with the bound being minimax rather than average-case. The 256-vs-448 choice is thus presented honestly as a trade-off resolved empirically, not a clean domination.
Experimental validation
Experiments use a kernel-faithful simulation matching production semantics (FP32 2−106 from pre-cast P, FP32 accumulator from post-cast P·V, round-to-nearest E4M3 cast, epilogue division by 2−107), with Q, K, V held in FP32 to isolate the P-cast effect—an explicit scoping decision, since production QKV quantization adds an orthogonal noise floor. Sweeping 2−108 at 2−109:
| Configuration |
MSE (16×0), Δ=7 |
N=8192 |
N=16384 |
| Forward, S=1 |
5.65 |
4.40 |
2.94 |
| Reverse, S=1 |
1.70 |
0.83 |
0.32 |
| Forward, S=448 |
1.81 |
0.90 |
0.32 |
| Forward, S=256 |
1.64 |
0.80 |
0.28 |
| Reverse, S=256 |
1.64 |
0.81 |
0.28 |
Forward+S=1 is 16×1 worse than optimized configurations at 16×2, growing to 16×3 at 16×4 as non-sink probability mass increases. At 16×5 all configurations converge since non-sink mass falls below 2%. Against production implementations, FlashAttention-3/4 and the updated hpc-ops satisfy all optimality conditions; FlashInfer and TensorRT-LLM XQA (both 16×6) could gain 10–15% MSE reduction from switching to 256 at negligible cost—two bit-exact FMAs against a tensor-core matmul.
Limitations and open questions
The paper states its boundaries plainly. The 16×7 analysis is minimax-optimal assuming 16×8 spans 16×9 uniformly, whereas actual post-softmax distributions are highly skewed; a dynamic per-block P0 scale could improve average-case precision, and whether it does remains unexamined. All results are isolated output MSE against an FP32 reference—no perplexity or task-accuracy measurements are reported, so the paper explicitly frames the P1 switch as zero-cost and never-worse on P-cast MSE rather than a guaranteed end-to-end quality win; where FP8 QKV noise dominates, the downstream benefit may be small. The scope is limited to E4M3 P-cast with one static scale: MXFP4's E8M0 block exponents remove the scale freedom the analysis addresses, and outlier-heavy activations require per-channel smoothing first. Finally, the model assumes a single sink cluster at position 0; distributed multi-sink patterns are handled only by the informal substitution of a per-block maximum gap, without formal treatment.
Conclusion
This paper converts two folklore engineering choices in FP8 attention into quantitative statements: P-collapse under Attention Sink is a threshold effect governed by P2, eliminated independently by either reverse iteration or P3 scaling, both of which saturate to the same precision floor; and P4 is constructively characterized within the bit-exact family by the P5 sawtooth's lower envelope plus maximal normal coverage. Kernel-faithful experiments confirm P6–P7 MSE improvements in the critical transition region, and the resulting recommendation—switching any P8 or P9 kernel to S=25600—has been applied in practice. The main unresolved question the paper leaves is empirical: whether removing P-cast error translates into measurable end-to-end model quality gains once the FP8 QKV noise floor is accounted for.