Long-horizon consistency in interactive generative world models
Establish mechanisms for sustaining long-horizon consistency in interactive generative systems, including action-conditioned video and scene-based world simulators, so that object identity, context, and coherence are preserved over extended real-time interaction.
References
Despite the leap to real-time interaction, sustaining long-horizon consistency remains unsolved.
Several directions remain open. First, the current geometric world state primarily preserves coarse scene structure, while fine-grained consistency of object identity, appearance, and local details remains limited. Richer object-level or semantic world representations may provide stronger long-term identity consistency.
Several directions remain open. First, the current geometric world state primarily preserves coarse scene structure, while fine-grained consistency of object identity, appearance, and local details remains limited. Richer object-level or semantic world representations may provide stronger long-term identity consistency. Second, a persistent world should model not only static geometry but also dynamic state, including object motion, state transitions, and their long-term evolution. Developing explicit representations that can continuously update such dynamic world state is an important next step.