Generalizing ResAdapt beyond video resizing

Establish whether extending the ResAdapt training mixture to jointly include image and video data and implementing alternative pre-encoding visual budget operators, particularly hard frame selection, can generalize the learned input-side allocation policy beyond continuous resizing and yield consistent efficiency-preserving performance on image-centric benchmarks.

Background

ResAdapt is instantiated and trained primarily for video tasks using continuous per-frame resizing as the pre-encoding operator. The authors observe that transfer beyond this regime is uneven: while the policy sometimes increases fidelity for specific static images (e.g., charts), it fails to deliver uniformly efficiency-preserving gains on image-centric benchmarks.

To broaden validation and address the uneven transfer, the authors explicitly identify two open directions: (1) extending the training mixture to include both image and video data, and (2) exploring alternative pre-encoding operators beyond resizing, such as hard frame selection. These directions aim to test and potentially improve the generality of input-side allocation across modalities and operators.

References

Extending the training mixture to image–video data and exploring alternative operators, such as hard frame selection, remain open problems.

— ResAdapt: Adaptive Resolution for Efficient Multimodal Reasoning  (2603.28610 - Liao et al., 30 Mar 2026) in Limitations and Future Work, item (iii)

These systems demonstrate that joint allocation works, but a joint gain has at least three possible sources: better timestamps, better resolution policy, or simply more temporal coverage. None of these papers separates the three, so the mechanism behind their improvements remains open.

— Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs  (2609.03820 - Khatri, 3 Sep 2026) in Related Work, Section 2, paragraph “Selection with spatial adaptation”

This is evidence against a large OMP-specific interaction rather than proof the two are identical; with 203 discordant pairs the test excludes only large effects. The defensible claim is that temporal reinvestment works with OMP, points the same way under uniform sampling, and that its selector-specificity remains open at roughly the one-point scale.

— Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs  (2609.03820 - Khatri, 3 Sep 2026) in Section 5.4, “Spending the savings on more frames”