Determine whether the three cost mechanisms generalize to other attention implementations and accelerators

Determine whether the conditions producing discrete batch alignment, KV-cache survival effects, and prefill interference in compile-time-static LLM serving also hold on other attention implementations and inference accelerators.

Background

The reported KV allocation and eviction behavior, predefined bucket selection, and exclusive prefill/decode execution were observed on a specific Rebellions CA25 NPU stack using the eager attention path. Other attention implementations may use different KV allocation and eviction units, and other accelerators may implement different scheduling and execution structures. Consequently, the generality of the three mechanisms beyond the evaluated stack remains unresolved.

References

Future work should extend validation to repeated tool calls, longer contexts, and loads in which the top bucket of enlarged configurations is used, confirm whether the conditions that produce each mechanism hold on other attention implementations and accelerators, and automate configuration selection that accounts for tool waiting and re-arrival.

— Tool Waiting and Re-arrival in Compile-Time-Static LLM Serving: Cost Mechanisms and Configuration Selection  (2609.34663 - Jang et al., 28 Sep 2026) in Section 6.2, Conclusion