Determine whether reused-KV inference preserves quality equivalence

Determine whether inference-time reuse of a Qwen3-1.7B backbone’s prefill KV cache across already-trained standard LoRA specialists can achieve quality equivalent to native specialist-prefill inference, particularly for the GSM8K math and HotpotQA extractive-question-answering workloads evaluated at different context lengths.

Background

The paper evaluates direct reuse of a backbone-computed prefix KV cache with adapters disabled, followed by specialist processing of a prompt suffix and generation. On the tested workloads, reuse reduces warm-cache time-to-first-token but can reduce task quality, with the observed penalty varying by generation budget, adapter initialization seed, task, and context length.

The authors explicitly distinguish the measured quality–latency tradeoff from equivalence. The GSM8K penalty is statistically significant in one 160-token configuration but not in the larger-budget or second-seed configurations, while the QA penalty is context-dependent. Establishing or refuting quality equivalence therefore requires controlled evaluations that resolve generation-budget, seed, backbone, and context-length effects.

References

We do not establish quality equivalence (the interval permits a loss up to ${\approx}5$pp) nor a general boundary-selection rule: \S\ref{sec-quality} does not support always recompute more'' ornever split representations.''

— Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs  (2609.17109 - Rajput, 15 Sep 2026) in Section 4, subsection “Interpretation”; reiterated in Section 6, “Limitations and future work,” and the Conclusion

We do not establish quality equivalence (the interval permits a loss up to ${\approx}5$pp) nor a general boundary-selection rule: \S\ref{sec-quality} does not support always recompute more'' ornever split representations.''

— Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs  (2609.17109 - Rajput, 15 Sep 2026) in Section 4, subsection “Interpretation”

The consistent direction (the mid-context boundary is the worst cell in every run we did) is a tendency we cannot yet attribute to a mechanism.

— Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs  (2609.17109 - Rajput, 15 Sep 2026) in Section 4, subsection “Interpretation”

The present comparisons do not establish specialist dependence; we do not inflate the sample to seek significance.

— Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs  (2609.17109 - Rajput, 15 Sep 2026) in Section 5, subsection “Specialist dependence is not established” and Appendix A, “Specialist-dependence contrasts”

The higher cap-hit is evidence of changed generation behavior; its causal contribution to the accuracy gap is not established.

— Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs  (2609.17109 - Rajput, 15 Sep 2026) in Section 4, subsection “Generation budget”; Section 6, “Limitations and future work”