Determine the causal mechanism behind the cross-runtime TTFT leadership flip

Determine whether differences in scheduler behavior, rather than prefix-cache effectiveness, causally explain the reversal in time-to-first-token performance between vLLM and TensorRT-LLM under bursty and fixed-rate arrival patterns.

Background

The benchmark finds that vLLM has lower time-to-first-token latency at low concurrency, whereas TensorRT-LLM performs better on median latency under high-concurrency burst arrivals. Under fixed-rate realistic RAG traffic, the ranking reverses again, with vLLM exhibiting substantially lower tail latency. Cache-hit measurements do not explain these changes, suggesting that scheduling behavior above the cache is responsible.

The causal explanation remains unresolved because the study directly measures queueing and prefill decomposition for vLLM but lacks comparable scheduler counters for TensorRT-LLM. Consequently, the proposed scheduling account is supported indirectly by cache-hit elimination and latency-distribution patterns rather than by instrumentation that establishes causality.

References

TensorRT-LLM exposed no comparable counters in the tested version, so the cross-runtime flip remains an inference from cache-effect elimination plus the batching plateaus of Figure~\ref{fig:cdf}; we present the scheduling account as the best-supported hypothesis, not a fully instrumented causal claim (Section~\ref{sec:limits}).

— PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving  (2609.19657 - Shewale et al., 17 Sep 2026) in Section 3.2, “Burst concurrency: when does reuse survive load?”, and Section 7, Limitation (v)