Pruning Unnecessary Zoom-In Tool Calls

Develop a reliable method for pruning unnecessary zoom-in tool calls while preserving beneficial tool use by multimodal language models.

Background

The M{content}Chihuahua experiments reveal that reinforcement-learning-trained zoom-in models may develop a reflex to invoke the tool even when the entire image already provides sufficient information. In the reported 1×1-grid setting, both evaluated models almost always call the tool despite the task not requiring it, and a simple constant negative tool-use reward causes the model to abandon tool use altogether. The authors therefore identify selective suppression of unnecessary calls as an unresolved problem.

References

Pruning unnecessary tool calls remains a hard problem (cf. Section 4.5. in \citet{bai_qwen3-vl_2025}).

Learning to Zoom Efficiently with a Contrastive Curriculum  (2609.03206 - Helm et al., 2 Sep 2026) in Appendix, subsection “Excessive Tool Use”