Combining HHR with KV-cache quantization

Investigate the effect of combining Hierarchical Hash Retrieval with KV-cache quantization to determine whether the two compression techniques provide complementary benefits for efficient long-context inference.

Background

Hierarchical Hash Retrieval (HHR) reduces attention cost through hierarchical candidate routing and learned hash projection, whereas KV quantization compresses the stored key-value cache. The paper reports that HHR is orthogonal to KV quantization, meaning that the techniques could potentially be applied together, but does not evaluate their combined behavior. Establishing the effect of this combination is therefore left unresolved and is explicitly identified as future work.

References

Moreover, HHR is orthogonal to other compression techniques such as KV quantization. The effect of combining HHR with them remains an interesting question but is beyond the scope of this paper, and we leave it for future work.

— HHR: Hierarchical Hash Retrieval for Efficient LLM Generation  (2610.01230 - Liu et al., 1 Oct 2026) in Section 6, “Limitations and Future Work”