Long-Generation Stability of the Full Distributed Stack

Re-measure sustained long-generation throughput and stability for the full Llama 3.1 8B distributed stack using v5_beam shards, speculative decoding, and micro-batching, rather than relying on measurements from an earlier export configuration.

Background

The paper reports flat or improving throughput over long generations for the monolithic model and reports earlier sustained-generation measurements for a two-stage distributed pipeline without speculative decoding using pre-v5_beam shards.

Those earlier distributed measurements do not cover the complete current configuration. Consequently, the effect of long contexts, cache growth, speculative decoding, and concurrent streams on sustained performance of the full stack remains unresolved.

References

On the 2-stage distributed pipeline (single-stream, no spec), earlier measurements on an extended-prompt workload showed no per-token degradation across $200$ tokens ($15.95$~tok/s), $1000$ tokens ($15.55$~tok/s), and 10 consecutive prompts ($14.46$~tok/s aggregate, $0.72$~QPS, with $36$~ms KV-cache reset between prompts). Those numbers predate $v_5$_beam; we have not re-measured sustained long-generation on the full stack.

Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets  (2608.19147 - Berenbaum et al., 19 Aug 2026) in Section 5.7, “Long Generation and Stability”