Long-Generation Stability of the Full Distributed Stack
Re-measure sustained long-generation throughput and stability for the full Llama 3.1 8B distributed stack using v5_beam shards, speculative decoding, and micro-batching, rather than relying on measurements from an earlier export configuration.
References
On the 2-stage distributed pipeline (single-stream, no spec), earlier measurements on an extended-prompt workload showed no per-token degradation across $200$ tokens ($15.95$~tok/s), $1000$ tokens ($15.55$~tok/s), and 10 consecutive prompts ($14.46$~tok/s aggregate, $0.72$~QPS, with $36$~ms KV-cache reset between prompts). Those numbers predate $v_5$_beam; we have not re-measured sustained long-generation on the full stack.
— Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets
(2608.19147 - Berenbaum et al., 19 Aug 2026) in Section 5.7, “Long Generation and Stability”