Determine whether reused-KV inference preserves quality equivalence
Determine whether inference-time reuse of a Qwen3-1.7B backbone’s prefill KV cache across already-trained standard LoRA specialists can achieve quality equivalent to native specialist-prefill inference, particularly for the GSM8K math and HotpotQA extractive-question-answering workloads evaluated at different context lengths.
References
We do not establish quality equivalence (the interval permits a loss up to ${\approx}5$pp) nor a general boundary-selection rule: \S\ref{sec-quality} does not support always recompute more'' ornever split representations.''
We do not establish quality equivalence (the interval permits a loss up to ${\approx}5$pp) nor a general boundary-selection rule: \S\ref{sec-quality} does not support always recompute more'' ornever split representations.''
The consistent direction (the mid-context boundary is the worst cell in every run we did) is a tendency we cannot yet attribute to a mechanism.
The present comparisons do not establish specialist dependence; we do not inflate the sample to seek significance.
The higher cap-hit is evidence of changed generation behavior; its causal contribution to the accuracy gap is not established.