Scalable Training of Large-Scale Text-to-Video Foundation Models
Develop scalable training methodologies for large-scale text-to-video foundation models that effectively handle the complexities introduced by modeling motion, in order to synthesize realistic, temporally coherent videos under stringent memory, compute, and data scale constraints.
References
However, training large-scale text-to-video (T2V) foundation models remains an open challenge due to the added complexities that motion introduces.
— Lumiere: A Space-Time Diffusion Model for Video Generation
(2401.12945 - Bar-Tal et al., 2024) in Section 1, Introduction
While these settings allow us to isolate the effect of the latent representation, it remains unclear whether the advantages of V-RAE extend to large-scale open-domain text-to-video generation with larger datasets, higher resolutions, longer videos, and substantially larger DiT backbones.
— V-RAE: Rethinking Video Latent Spaces for Generation
(2608.13556 - Guo et al., 13 Aug 2026) in Section Limitations