Integrating Prefix Reuse with Predictive Multi-Tier KV Caching

Integrate prefix reuse into Predictive Multi-Tier Cache Optimization (PMCO) and quantify its prefill benefits and memory costs, including the compatibility requirements for reused compressed prefixes.

Background

Prefix reuse allows multiple large-language-model requests that share system prompts or few-shot examples to reuse cached prompt segments, thereby reducing redundant prefill computation and time-to-first-token. The paper explains that supporting this capability would require extending PMCO’s per-request cache allocation to account for retained cache state. One proposed strategy is to keep shared prefixes at full rank while optimizing request-specific suffixes within the remaining memory budget; reusing compressed prefixes would additionally require compatible rank profiles and projection matrices.

The paper leaves unresolved both the system integration problem and the empirical characterization of its effects: it does not establish how prefix reuse should be incorporated into PMCO or quantify the resulting prefill improvements and memory overheads.

References

Integrating prefix reuse and quantifying its prefill benefits and memory costs remain open problems.

— TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching  (2609.21172 - Shu et al., 18 Sep 2026) in Section Discussion and Future Work, subsection “Prefix Caching”