Transfer of DASC to purely linear-attention and state-space models

Determine whether Decay-Aware State Compression (DASC) transfers effectively from hybrid linear-attention architectures to purely linear-attention architectures and state-space models.

Background

The paper develops and evaluates DASC for hybrid architectures that combine full-attention layers with recurrent linear-attention layers, specifically Kimi-KDA and Qwen-GDN. The method exploits decay heterogeneity in recurrent states to compress persistent state checkpoints during prefix caching.

The authors explicitly delimit the scope of their evaluation to hybrid architectures and identify transfer to purely linear-attention or state-space models as unresolved. Establishing such transfer would determine whether the decay-aware state-selection and checkpoint-compression approach generalizes beyond the architectural setting studied in the paper.

References

We evaluate hybrid architectures only; transfer to purely linear-attention or state-space models remains open.

— DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving  (2608.30386 - Yu et al., 31 Aug 2026) in Section 6, “Limitations”