Validate FoCUS Scalability to Larger Vision-Language Models
Validate the scalability of Fine-grained Captioning Control Using Scene Rewards (FoCUS) to significantly larger vision-language models beyond the two evaluated backbones.
References
Second, while we demonstrate effectiveness on two VLM backbones, the scalability of our approach to significantly larger models remains to be validated.
— Controllable Image Captioning with Prompt-Conditioned Scene Rewards
(2609.00709 - Hyun et al., 1 Sep 2026) in Limitations section