Speculative Decoding for Packed NPU Serving
Build and evaluate a composition of K-token speculative decoding with the static, mask-packed NPU serving mode, including correctness, throughput, cache management, and context-window effects.
References
The 0.2--0.3~ms marginal row cost would also make a $K$-token speculative verify nearly free on this path, a composition we have not yet built.
— Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets
(2608.19147 - Berenbaum et al., 19 Aug 2026) in Section 6.6, “Continuous Batching on CPU, GPU, and NPU”