Simultaneously achieving real-time latency and long-term geometric consistency in interactive world models
Determine whether and how to construct an interactive world modeling system for autoregressive streaming video generation that simultaneously achieves low latency sufficient for real-time user interaction and high long-term geometric consistency such that scenes remain coherent upon revisiting previously observed locations.
References
As summarized in Table~\ref{tab:compare_related_works}, the simultaneous achievement of both low latency and high consistency remains an open problem.
Several directions remain open. First, the current geometric world state primarily preserves coarse scene structure, while fine-grained consistency of object identity, appearance, and local details remains limited. Richer object-level or semantic world representations may provide stronger long-term identity consistency. Second, a persistent world should model not only static geometry but also dynamic state, including object motion, state transitions, and their long-term evolution. Developing explicit representations that can continuously update such dynamic world state is an important next step. Finally, further inference acceleration remains necessary for truly real-time interaction, including higher-compression video VAEs, more efficient few-step generators, and lower-cost geometric conditioning.