Separating architectural limitations from scale limitations

Determine whether the performance gap on MotionBlind is attributable primarily to Video-LLM architectural limitations or to the restricted parameter scale of the evaluated open-weight models.

Background

The study evaluates open-weight Video-LLMs with at most 12 billion parameters and reports that increasing model size within the tested range does not close the MotionBlind gap. However, the available experiments do not establish whether the failure arises from model architecture, insufficient scale, or both.

The authors explicitly state that they cannot fully disentangle these explanations. Larger-scale and architecturally varied evaluations would be needed to identify the source of the observed motion-understanding limitation.

References

Our open-weight models span $\le$12B parameters, so we cannot fully separate architectural from scale limits, though a larger model not closing the gap is suggestive.

MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs  (2609.09528 - Bhatia et al., 8 Sep 2026) in Section 6.1, “Limitations”