Obtain missing LongBench evaluations for larger and hybrid models

Obtain LongBench v1 On-Demand Attention results for Qwen3-8B, Qwen3.5-2B, and Gemma-4-12B-it to determine whether the reported selective-recall quality–access trade-off extends to these model scales, families, and hybrid architectures.

Background

The main experiments report LongBench v1 results for Qwen3-1.7B, while evaluations on Qwen3-8B, Qwen3.5-2B, and Gemma-4-12B-it provide only partial benchmark coverage. This prevents direct comparison of ODA with Full and Local attention on LongBench across the larger and architecturally distinct models.

The missing evaluations are relevant because the paper presents ODA as applicable across model scales, model families, and hybrid-attention backbones. Completing these measurements would test whether that claimed generality holds on the same long-context benchmark.

References

LongBench ODA results for these two models and Qwen3.5-2B are not yet available.

On-Demand Attention: Language Models Know When to Recall  (2609.20734 - Feng et al., 17 Sep 2026) in Appendix B.1, “Generation protocol”