Scope of modeling future uncertainty in video understanding
Determine how far video understanding models should extend beyond deterministic prediction to represent and reason about future uncertainty, including modeling multiple plausible futures and integrating uncertainty into planning and decision-making processes.
References
Beyond deterministic objectives such as video labeling, an open question is how far video understanding should extend toward understanding future uncertainty.
Despite its high prediction fidelity, \algname is still prone to hallucinations, like other video world models. Mitigating these hallucinations is critical for trustworthy integration in diverse robotics applications, such as planning, policy evaluation, and policy finetuning, presenting an exciting direction for future work on hallucination detection and mitigation, e.g., via uncertainty quantification~\citep{mei2026worldmodelsknowdont}.
The existing version also exhibits limitations. First, the hierarchical slow--fast architecture introduces additional computational and memory overhead, motivating future research on model compression and asynchronous inference. Second, the current model does not explicitly capture multimodal futures or predictive uncertainty, which may limit its performance in ambiguous and rare driving scenarios. We leave solving them as future works.