Validate repeated tool calls, longer contexts, and enlarged-bucket workloads

Extend validation of compile-time-static LLM serving configurations to repeated tool calls, longer contexts, and workloads in which the top bucket of enlarged configurations is selected, in order to determine whether the reported cost mechanisms and configuration-selection effects persist under those conditions.

Background

The experiments restrict each session to one tool call and therefore do not test how repeated re-arrivals, accumulating context, and recurring KV-cache reuse or eviction affect execution cost and interference. The validation also does not cover workloads that continuously select the largest bucket of an enlarged configuration. The authors explicitly identify these settings as requiring further validation rather than treating the reported results as established for them.

References

Future work should extend validation to repeated tool calls, longer contexts, and loads in which the top bucket of enlarged configurations is used, confirm whether the conditions that produce each mechanism hold on other attention implementations and accelerators, and automate configuration selection that accounts for tool waiting and re-arrival.

— Tool Waiting and Re-arrival in Compile-Time-Static LLM Serving: Cost Mechanisms and Configuration Selection  (2609.34663 - Jang et al., 28 Sep 2026) in Section 6.2, Conclusion