Validate FoCUS Scalability to Larger Vision-Language Models

Validate the scalability of Fine-grained Captioning Control Using Scene Rewards (FoCUS) to significantly larger vision-language models beyond the two evaluated backbones.

Background

FoCUS is evaluated on only two relatively small vision-LLM backbones: Qwen2.5-VL-3B-Instruct and InternVL3-2B. Although the method improves controllability and fine-grained caption quality for these models, the paper does not establish whether its scene-graph parsing, LLM-based reward evaluation, and GRPO optimization remain effective or computationally feasible when applied to substantially larger models. The authors therefore identify scalability to significantly larger models as an unresolved validation problem.

References

Second, while we demonstrate effectiveness on two VLM backbones, the scalability of our approach to significantly larger models remains to be validated.

Controllable Image Captioning with Prompt-Conditioned Scene Rewards  (2609.00709 - Hyun et al., 1 Sep 2026) in Limitations section