Learning and reasoning over long-horizon vision–language sequences
Establish effective methods to learn from and reason over long-horizon interleaved vision–language sequences that capture extended temporal dynamics and semantic coherence.
References
Recent advances in short-clip video generation have demonstrated the ability to capture short-term dynamics, but learning from and reasoning over long-horizon vision-language sequences remains a central open challenge.
— Emu3.5: Native Multimodal Models are World Learners
(2510.26583 - Cui et al., 30 Oct 2025) in Section 1 (Introduction)
The open question is no longer whether models can be made to reason over time, but whether certificate-long horizons can be reached under the compute budget of a wearable rather than a server.
— Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI
(2608.18671 - Zamani et al., 19 Aug 2026) in Section 11.1, “Better Temporal Reasoning” (Sec. future-temporal)