Cause of the Llama lossless-speculation deficit

Determine the cause of the −0.14 judged-quality deficit observed for the strict, lossless speculative-decoding arm on Llama-3.1-8B-Instruct, which was localized to long multi-step responses but remained unexplained within the study.

Background

The study used strict speculative decoding as an implementation null because, in exact arithmetic, its rejection rule should reproduce the target model’s output distribution. For Qwen2.5-7B-Instruct, this arm produced a judged-quality difference of +0.05 relative to the dual reference, whereas for Llama-3.1-8B-Instruct it produced −0.14, a value inside the prespecified ±0.3 equivalence bound but significantly below zero.

Additional controls localized the Llama deficit to long hard-verifiable responses and showed that it was not attributable solely to the NF4 drafter. The authors therefore identified and bounded a model- and implementation-specific effect but did not resolve its underlying cause.

References

On Llama the lossless speculative arm read −0.14, below zero but within the bound; we localized the deficit to long multi-step responses but did not determine its cause.

A Calibrated Instrument for Measuring How Inference Optimizations Affect Output Quality  (2609.18005 - Kaplan, 16 Sep 2026) in Section 7, Limitations, page 15