Dynamic, real-time interactive worlds in unified models
Develop unified multimodal generative models that support continuous, real-time closed-loop interaction to create truly dynamic and interactive worlds, overcoming the current limitation of visual-prior unified models for text-to-image and text-to-video that are restricted to single-shot synthesis or stepwise editing.
References
Thus, while Stage II unified architectures, the creation of truly dynamic and interactive worlds remains an open challenge and motivates Stage III.
Fully real-time generated text remains an open challenge.
Together with I/O-level engineering of the estimation loop, the system runs $3.4\times$ faster at lower memory (\cref{tab:runtime_speed}). The remaining budget is dominated by the VAE round trips, which efficient or VAE-free video generators are designed to remove; combined with the trend towards real-time online estimators, real-time rates appear within reach and are left as future work.
While recent work explores joint audio-visual generation in robotics~\citep{nvidia2026cosmos3}, enabling continuous, multimodal participation across general and game scenes remains an open challenge.