Hardware validation of low-bit KV-compression latency overhead
Validate the cost-normalized comparison between tensor parallelism and KV-cache compression on physical hardware by measuring the decode-kernel dequantization overhead of low-bit KV-cache compression and determining whether that overhead reverses the reported cost ordering.
References
Published low-bit KV kernels report overheads well below the smallest of these, so the ordering is robust to the one assumption we could not measure, though confirming it on hardware remains the most valuable experiment this study leaves undone.
— More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving
(2608.23962 - Tumkur et al., 25 Aug 2026) in Section 4, subsection “RQ1: which is cheaper?”; Section 6, “Limitations”; Conclusion