Identify the cause of SALI’s performance drop on short MSR-VTT clips

Determine whether the 1.4-point overall R@1 drop of SALI with greedy-max shot-level matching on MSR-VTT is caused by the short duration of the clips or by the small frame budget of 12 frames per video.

Background

SALI improves retrieval for relational queries by representing videos with shot-level embeddings and matching query components against individual shots. However, on the short MSR-VTT clips, which use only 12 frames per video, SALI with greedy-max matching reduces overall R@1 by 1.4 points relative to CLIP4Clip-meanP. The comparable ActivityNet experiments do not show this degradation: ActivityNet videos have approximately the same median shot count but are about two minutes long and use 64 frames. The authors therefore leave unresolved whether clip duration or frame-budget sparsity explains the MSR-VTT performance loss.

References

Our experiments do not tell whether the short clips or the small frame budget (12 frames) cause the drop on MSR-VTT.

— SALI: Shot-Aware Late Interaction for Cross-Shot Relation Matching in Text-to-Video Retrieval using Film-Grammar Knowledge  (2609.29721 - Oyama et al., 24 Sep 2026) in Limitation section, immediately before Section Conclusion