Complete the training-resource accounting for ODA

Determine the complete training cost of the On-Demand Attention recall head, including total GPU-hours, peak memory, the exact number of unique consumed training examples, and the number of eligible training targets, in order to assess whether the method provides a measured total-training-cost advantage over alternative methods.

Background

Although only the recall head is trainable, its supervision requires computing the frozen Full trunk and both Local/Full counterfactual features and labels. Thus, the trainable parameter count alone does not characterize the method’s total training expense.

The paper explicitly reports that several resource quantities remain unavailable and therefore declines to claim a measured training-cost advantage. A complete accounting would enable comparison with methods that jointly train attention policies or require additional supervision.

References

Complete GPU-hours, peak memory, and the exact number of unique consumed examples and eligible targets are not yet available. We therefore do not claim a measured total-training-cost advantage over other methods.

On-Demand Attention: Language Models Know When to Recall  (2609.20734 - Feng et al., 17 Sep 2026) in Appendix B.9, “Training resources”