Determine whether targeted fine-tuning resolves the temporal capability gap
Conduct a targeted fine-tuning probe on Vision-Language Models using physical-AI and densely timestamped video supervision to determine whether weak temporal localization and dense video captioning reflect a capability gap or an elicitation gap.
References
A targeted fine-tuning probe would distinguish a capability gap from an elicitation gap, and we leave this test to future work.
— VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models
(2609.09396 - Bhat et al., 8 Sep 2026) in Section 4.3, “Returns to Scale and the Temporal Lag”