Determine the causal mechanism behind the cross-runtime TTFT leadership flip
Determine whether differences in scheduler behavior, rather than prefix-cache effectiveness, causally explain the reversal in time-to-first-token performance between vLLM and TensorRT-LLM under bursty and fixed-rate arrival patterns.
References
TensorRT-LLM exposed no comparable counters in the tested version, so the cross-runtime flip remains an inference from cache-effect elimination plus the batching plateaus of Figure~\ref{fig:cdf}; we present the scheduling account as the best-supported hypothesis, not a fully instrumented causal claim (Section~\ref{sec:limits}).
— PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving
(2609.19657 - Shewale et al., 17 Sep 2026) in Section 3.2, “Burst concurrency: when does reuse survive load?”, and Section 7, Limitation (v)