Determine whether targeted fine-tuning resolves the temporal capability gap

Conduct a targeted fine-tuning probe on Vision-Language Models using physical-AI and densely timestamped video supervision to determine whether weak temporal localization and dense video captioning reflect a capability gap or an elicitation gap.

Background

The paper compares model families with progressively greater amounts of physical-AI training data and finds highly uneven transfer. Tracking, object localization, and event verification improve substantially, whereas Temporal Localization and Dense Video Captioning improve only slightly.

The authors hypothesize that bounding-box supervision is abundant while densely timestamped video supervision is scarce. They explicitly leave unresolved whether the poor temporal results arise from a fundamental capability limitation or from an elicitation problem that could be corrected through targeted fine-tuning.

References

A targeted fine-tuning probe would distinguish a capability gap from an elicitation gap, and we leave this test to future work.

VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models  (2609.09396 - Bhat et al., 8 Sep 2026) in Section 4.3, “Returns to Scale and the Temporal Lag”