Determine whether the three cost mechanisms generalize to other attention implementations and accelerators
Determine whether the conditions producing discrete batch alignment, KV-cache survival effects, and prefill interference in compile-time-static LLM serving also hold on other attention implementations and inference accelerators.
References
Future work should extend validation to repeated tool calls, longer contexts, and loads in which the top bucket of enlarged configurations is used, confirm whether the conditions that produce each mechanism hold on other attention implementations and accelerators, and automate configuration selection that accounts for tool waiting and re-arrival.
— Tool Waiting and Re-arrival in Compile-Time-Static LLM Serving: Cost Mechanisms and Configuration Selection
(2609.34663 - Jang et al., 28 Sep 2026) in Section 6.2, Conclusion