Dynamic, real-time interactive worlds in unified models

Develop unified multimodal generative models that support continuous, real-time closed-loop interaction to create truly dynamic and interactive worlds, overcoming the current limitation of visual-prior unified models for text-to-image and text-to-video that are restricted to single-shot synthesis or stepwise editing.

Background

Stage II consolidates processing and generation for multiple modalities into a single backbone and paradigm, reducing fragmentation and enabling cross-modal transfer. However, visual-prior unified models remain constrained to one-shot synthesis or incremental editing and lack continuous, real-time interaction capabilities.

The authors explicitly identify achieving truly dynamic and interactive worlds within the unified modeling paradigm as an open challenge, motivating the transition to Stage III focused on interactive generative models.

References

Thus, while Stage II unified architectures, the creation of truly dynamic and interactive worlds remains an open challenge and motivates Stage III.

From Masks to Worlds: A Hitchhiker's Guide to World Models  (2510.20668 - Bai et al., 23 Oct 2025) in Section 4.2 (Benefits and Gaps)

Fully real-time generated text remains an open challenge.

Solaris: Towards Interfaces That Are Generated, Not Coded  (2609.00776 - Alaluf et al., 1 Sep 2026) in Section 4, “Limitations”

Together with I/O-level engineering of the estimation loop, the system runs $3.4\times$ faster at lower memory (\cref{tab:runtime_speed}). The remaining budget is dominated by the VAE round trips, which efficient or VAE-free video generators are designed to remove; combined with the trend towards real-time online estimators, real-time rates appear within reach and are left as future work.

PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control  (2609.17521 - Chen et al., 15 Sep 2026) in Section 4, Subsection 4.4, “Runtime Analysis”

While recent work explores joint audio-visual generation in robotics~\citep{nvidia2026cosmos3}, enabling continuous, multimodal participation across general and game scenes remains an open challenge.

EchoWM: Open and Enterable Omnimodal World Models  (2608.23189 - Zhang et al., 24 Aug 2026) in Section 1, Introduction