In-the-wild and egocentric generalization of MotionBlind findings

Determine whether the motion-understanding failures observed on MotionBlind generalize to in-the-wild and egocentric video settings beyond the benchmark’s single-actor indoor recordings.

Background

MotionBlind contains 60 self-recorded instances designed to isolate speed, magnitude, and direction while controlling appearance and scene context. All clips feature a single actor indoors, which supports controlled diagnosis but limits conclusions about deployment conditions.

The authors explicitly identify generalization to in-the-wild and egocentric video as unresolved. Establishing performance in those settings would test whether the observed motion-perception failures reflect a broader limitation of Video-LLMs rather than an artifact of the benchmark’s recording conditions.

References

Our study has limitations. MotionBlind's 60 self-recorded instances trade scale for controlled isolation of physical variables, so results carry small-sample variance; all clips feature a single actor indoors, leaving in-the-wild and egocentric generalization open; and constrained yes/no parsing, while removing free-form ambiguity, may understate reasoning a model cannot verbalize in binary form.

MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs  (2609.09528 - Bhatia et al., 8 Sep 2026) in Section 6.1, “Limitations”