Verification of Roofline Predictions for Layer-Wise Heterogeneous Quantization
Verify whether the roofline-based speedup prediction remains accurate for layer-wise heterogeneous quantization of Gemma 3, particularly when different layers use Q4 and Q8 precision and dequantization overheads may differ between the two formats.
References
However, it is possible that under layer-wise heterogeneous quantization the overhead of dequantizing weights when loading them into compute cores may differ for Q4 and Q8, potentially introducing an additional prediction error. Practical verification of this result requires tools with fine-grained control over the computation graph, such as ExLlamaV3 or TensorRT-LLM, which were not available for the Gemma 3 architecture within the scope of this study.
— A Layer Importance Metric for Quantization Accounting for the Speed-Quality Trade-off in Autoregressive Models
(2608.26926 - Safronov, 27 Aug 2026) in Section “Formalizing the Latency Parameters of FFN and Embedding (Speed Component),” subsection “FFN”