Papers
Topics
Authors
Recent
Search
2000 character limit reached

P-Cast Precision in FP8 Attention: Sink-Induced Collapse and the Optimality of S=2^8

Published 2 Jun 2026 in cs.AR, cs.AI, cs.DC, cs.LG, and cs.PF | (2606.06521v1)

Abstract: FP8 (E4M3) acceleration for attention computation offers significant throughput gains, but the 3-bit mantissa introduces precision challenges when the softmax probability matrix P is cast to FP8 before the P*V matrix multiplication. We analyze two implementation choices that affect output precision under the Attention Sink phenomenon: (1) the KV block iteration order, and (2) the static scaling factor applied to P before casting. We show that forward KV iteration causes "P-collapse" -- to leading order, a fraction Phi(Delta + delta_k - 6.93 - ln S) of non-sink P values underflow to zero, where the small shift delta_k ~ 1 (for k_sink = 4) is the expected within-sink-block score maximum -- and that reverse iteration removes it, with a zero-underflow guarantee when reverse is combined with S = 256. We further give a constructive characterization of S = 256 = 28 as the static scale that simultaneously satisfies (i) bit-exact IEEE 754 scaling, (ii) the lower envelope of a sawtooth function dp(S) over the E4M3 number line (dp = 2-4, the minimum worst-case quantization step), and (iii) the maximum normal-range coverage among bit-exact (2k) scales (a non-bit-exact scale such as 448 attains slightly higher coverage). Both optimizations are already deployed in FlashAttention-3/4 on engineering grounds; our contribution is a quantitative account of why these choices are good and a closed-form threshold Delta_c = 6.93 + ln S - delta_k for predicting kernel-level precision loss. Kernel-faithful experiments (Q, K, V in FP32 to isolate the P-cast effect) show 3-10x MSE improvement at moderate sink strengths, and paired tests confirm both fixes saturate to the same precision floor when combined.

Authors (1)

Summary

  • The paper derives a closed-form collapse threshold, showing that at sink strength Δ≈7, forward iteration with S=1 can zero 82% of non-sink probabilities, while S=256 prevents underflow through Δ=7.
  • The paper finds that reverse block iteration and S=256 scaling independently remove most P-cast damage, producing comparable results and reducing MSE by roughly 3–10× versus forward iteration with S=1.
  • The paper characterizes S=256 as the unique bit-exact power-of-two scale that minimizes worst-case quantization error while maximizing E4M3 coverage, although S=448 offers deeper tail coverage at higher MSE.

Overview

This paper analyzes two implementation-level choices in FP8 (E4M3) attention kernels—the KV block iteration order and the static scaling factor applied to the softmax probability matrix PP before casting—and shows that both are first-order determinants of output precision under the Attention Sink phenomenon. The work is explicitly positioned as a quantitative explanation rather than a proposal: reverse iteration and S=256S = 256 scaling are already deployed in FlashAttention-3/4 on engineering grounds, and the contribution is a closed-form account of why these choices work, plus a diagnostic threshold practitioners can use to predict P-cast failure. The analysis motivated updating Tencent's hpc-ops kernel from S=1S = 1 to S=256S = 256, and suggests analogous changes for FlashInfer and TensorRT-LLM XQA, which currently use S=448S = 448.

Background: E4M3 P-casting and the sink interaction

The E4M3 format provides 126 positive representable values, with a maximum of 448, minimum normal 2−62^{-6}, minimum subnormal 2−92^{-9}, and round-to-zero below 2−102^{-10}. Within any normal binade there are exactly eight uniformly spaced values, giving a relative precision of roughly 12.5%—about 16×16\times worse than BF16. FlashAttention-style FP8 kernels cast PP to E4M3 before the S=256S = 2560 matmul while keeping the running sum S=256S = 2561 in FP32 from pre-cast probabilities; normalization is therefore exact, and all precision loss concentrates in the numerator.

Attention Sink [(2606.06521) cites Xiao et al. and follow-ups] places initial tokens at logit scores S=256S = 2562 above the mean, with reported S=256S = 2563 at multi-thousand-token contexts. Under forward block iteration, the sink inflates the running softmax maximum to S=256S = 2564, where S=256S = 2565 is the expected maximum of S=256S = 2566 standard Gaussians (S=256S = 2567; the paper notes the asymptotic S=256S = 2568 badly overestimates this at small S=256S = 2569). All subsequent P values are then S=1S = 10, and those below S=1S = 11 silently underflow to zero.

Quantifying P-collapse

The central result is a closed-form underflow fraction. A value S=1S = 12 casts to zero iff S=1S = 13, so for unit-normal scores the collapsed fraction is

S=1S = 14

with the caveat that treating per-row S=1S = 15 as fixed makes this a mean-shift estimate that slightly understates realized collapse due to convexity of S=1S = 16 in the lower tail. Simulations at S=1S = 17, S=1S = 18 show the threshold behavior starkly: at S=1S = 19, the zeroed fraction rises from 22.3% at S=256S = 2560 to 82% at S=256S = 2561 and essentially 100% by S=256S = 2562, whereas S=256S = 2563 yields 0% through S=256S = 2564. The effective information loss (non-sink mass times zeroed fraction) peaks near 42% at S=256S = 2565—the regime where positions carrying half or more of the probability mass have most of their P values destroyed. This yields the practical diagnostic S=256S = 2566: the sink strength at which the median non-sink P value crosses the round-to-zero boundary.

A companion MSE bound, derived under an idealized pairwise-uncorrelated-S=256S = 2567 assumption (the paper concedes real S=256S = 2568 are correlated, shifting the absolute constant by an S=256S = 2569 factor), gives S=448S = 4480. A key structural observation follows: because non-sink mass vanishes at large S=448S = 4481, the MSE damage is confined to the transition region S=448S = 4482—at extreme sink strength even total collapse barely affects the output.

Reverse iteration as a sufficiency fix

Reversing the block order defers the sink block to the last iteration. Before it arrives, the running maximum is bounded by extreme-value statistics at S=448S = 4483, keeping pre-sink P values well above the cast boundary. Formally, with S=448S = 4484 a value survives whenever S=448S = 4485; for S=448S = 4486 this gives an underflow probability below S=448S = 4487. The final S=448S = 4488-correction when the sink block is processed (S=448S = 4489 at 2−62^{-6}0, 2−62^{-6}1) acts on the FP32 accumulator and costs negligible precision. Notably, the paper shows via paired 2−62^{-6}2-tests that reverse+2−62^{-6}3 and forward+2−62^{-6}4 are statistically indistinguishable wherever P-collapse dominates (2−62^{-6}5); at larger 2−62^{-6}6 reverse is marginally better but by only 2−62^{-6}7 against MSE floors of 2−62^{-6}8–2−62^{-6}9. Either fix alone therefore saturates to the same precision floor, and the choice between them can be made on engineering grounds.

Characterizing 2−92^{-9}0

The scale-factor analysis introduces 2−92^{-9}1, the normalized worst-case quantization step over the mapped range 2−92^{-9}2, and proves it forms a sawtooth over the E4M3 number line with power-of-two values on its lower envelope at 2−92^{-9}3. Specifically, every 2−92^{-9}4 with 2−92^{-9}5 achieves 2−92^{-9}6; every non-power-of-two 2−92^{-9}7 strictly exceeds it; and 2−92^{-9}8 incurs saturation error that also exceeds it. Combined with bit-exactness of IEEE 754 scaling by powers of two (exponent-only operations, no rounding) and monotone improvement of normal-range coverage with 2−92^{-9}9, this yields a constructive characterization: restricted to 2−102^{-10}0, 2−102^{-10}1 is the unique scale satisfying all three conditions simultaneously.

The paper is careful about what this optimality does not claim. Condition (C3) alone—"largest 2−102^{-10}2"—already uniquely identifies 256 among powers of two, so (C2) is explanatory rather than discriminating. More importantly, dropping bit-exactness admits 2−102^{-10}3, which attains strictly better deep-tail coverage (2−102^{-10}4 vs. 2−102^{-10}5) but pays a 14% larger worst-case step, predicting ~30% higher MSE in the worst case. The measured gap is 10–15%, consistent with the bound being minimax rather than average-case. The 256-vs-448 choice is thus presented honestly as a trade-off resolved empirically, not a clean domination.

Experimental validation

Experiments use a kernel-faithful simulation matching production semantics (FP32 2−102^{-10}6 from pre-cast P, FP32 accumulator from post-cast P·V, round-to-nearest E4M3 cast, epilogue division by 2−102^{-10}7), with Q, K, V held in FP32 to isolate the P-cast effect—an explicit scoping decision, since production QKV quantization adds an orthogonal noise floor. Sweeping 2−102^{-10}8 at 2−102^{-10}9:

Configuration MSE (16×16\times0), Δ=7 N=8192 N=16384
Forward, S=1 5.65 4.40 2.94
Reverse, S=1 1.70 0.83 0.32
Forward, S=448 1.81 0.90 0.32
Forward, S=256 1.64 0.80 0.28
Reverse, S=256 1.64 0.81 0.28

Forward+S=1 is 16×16\times1 worse than optimized configurations at 16×16\times2, growing to 16×16\times3 at 16×16\times4 as non-sink probability mass increases. At 16×16\times5 all configurations converge since non-sink mass falls below 2%. Against production implementations, FlashAttention-3/4 and the updated hpc-ops satisfy all optimality conditions; FlashInfer and TensorRT-LLM XQA (both 16×16\times6) could gain 10–15% MSE reduction from switching to 256 at negligible cost—two bit-exact FMAs against a tensor-core matmul.

Limitations and open questions

The paper states its boundaries plainly. The 16×16\times7 analysis is minimax-optimal assuming 16×16\times8 spans 16×16\times9 uniformly, whereas actual post-softmax distributions are highly skewed; a dynamic per-block PP0 scale could improve average-case precision, and whether it does remains unexamined. All results are isolated output MSE against an FP32 reference—no perplexity or task-accuracy measurements are reported, so the paper explicitly frames the PP1 switch as zero-cost and never-worse on P-cast MSE rather than a guaranteed end-to-end quality win; where FP8 QKV noise dominates, the downstream benefit may be small. The scope is limited to E4M3 P-cast with one static scale: MXFP4's E8M0 block exponents remove the scale freedom the analysis addresses, and outlier-heavy activations require per-channel smoothing first. Finally, the model assumes a single sink cluster at position 0; distributed multi-sink patterns are handled only by the informal substitution of a per-block maximum gap, without formal treatment.

Conclusion

This paper converts two folklore engineering choices in FP8 attention into quantitative statements: P-collapse under Attention Sink is a threshold effect governed by PP2, eliminated independently by either reverse iteration or PP3 scaling, both of which saturate to the same precision floor; and PP4 is constructively characterized within the bit-exact family by the PP5 sawtooth's lower envelope plus maximal normal coverage. Kernel-faithful experiments confirm PP6–PP7 MSE improvements in the critical transition region, and the resulting recommendation—switching any PP8 or PP9 kernel to S=256S = 25600—has been applied in practice. The main unresolved question the paper leaves is empirical: whether removing P-cast error translates into measurable end-to-end model quality gains once the FP8 QKV noise floor is accounted for.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.