Establish end-to-end cached-versus-uncached exactness for Activated LoRA

Establish end-to-end cached-versus-uncached generation parity for the Activated LoRA adapter under correctly functioning invocation-sequence activation gating, and conduct a matched quality comparison between Activated LoRA and standard LoRA adapters.

Background

Activated LoRA is designed so that a base-prefix KV cache can be reused exactly: adapter weights activate only after a specified invocation sequence. The paper includes a preliminary QA comparison, but the activation gating was not verified, making the observed native-quality difference uninterpretable.

A plain forward pass does not exercise the generation-time gating behavior required for exact reuse. The paper therefore leaves unresolved whether the implementation achieves end-to-end parity and how its quality compares with standard LoRA under matched training and evaluation conditions.

References

Our structural cached-vs-uncached parity check for the aLoRA adapter was inconclusive: a plain forward pass does not exercise the generation-time activation gating that makes aLoRA's reuse exact, so we could not confirm end-to-end exactness in our harness.

— Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs  (2609.17109 - Rajput, 15 Sep 2026) in Section 5, subsection “Standard LoRA vs Activated LoRA”