- The paper demonstrates a novel action-conditioned drifting-based world model that replaces iterative diffusion with a single-step rollout for video generation.
- It employs an action-conditioned U-Net with frame-wise FiLM conditioning and a contrastive drifting loss to ensure high visual fidelity and precise action alignment.
- It achieves a 17× speedup and superior visual metrics, enabling efficient robotic policy planning and reliable offline simulation.
DriftWorld: Fast World Modeling through Drifting
Introduction and Motivation
DriftWorld introduces a novel action-conditioned world model optimized for high-speed and high-fidelity video generation in robotic policy planning and offline simulation. Existing world models, especially those based on diffusion or autoregressive transformers, achieve fine-grained visual accuracy but are fundamentally constrained by multi-step, iterative generation, resulting in untenable inference times for closed-loop control and large-scale action search. DriftWorld leverages drifting generative models to achieve single-step rollout of future frames, eliminating the multi-step denoising of diffusion-based counterparts (2607.15065). This directly addresses the computational bottleneck present in policy improvement and policy rank evaluation scenarios.

Figure 1: DriftWorld overview—enabling single-pass, action-conditioned rollout for efficient high-quality generation, planning, and offline simulation.
Model Architecture and Learning Framework
The core architectural innovation lies in the adaptation of the drifting paradigm—originally proposed for class-conditional image generation—to conditional video generation. DriftWorld employs an action-conditioned U-Net, directly mapping Gaussian noise, observation history, and specified future action sequences to the corresponding rollout frames. Crucially, conditioning is realized in a frame-wise manner using FiLM, ensuring precise action-outcome alignment at each predicted step.

Figure 2: The U-Net model is conditioned on observation history and future actions, with a contrastive drifting loss driving generated samples toward ground truth.
Training utilizes a contrastive drifting loss. For a given conditioning context, the model generates negative rollouts which are pushed towards the unique ground-truth trajectory (positive sample) and repelled from negative rollouts. The loss is computed either in pixel space (for simple domains) or the latent feature space derived from DINOv2/v3 encoders (for visually complex scenes), providing robust semantic similarity gradients and better visual sharpness.
Motion weighting is applied to further penalize inaction in dynamic regions, avoiding degenerate solutions where the model ignores action conditioning by replicating backgrounds. Accentuation parameters modulate the strength of action-following, and self-forcing techniques are incorporated during training for improved autoregressive rollout performance.
Numerical Results and Empirical Analysis
DriftWorld demonstrates 17× average speedup over state-of-the-art diffusion-based world models while either matching or exceeding them in all critical visual metrics (SSIM, PSNR, FID, FVD, LPIPS). On Push-T and Robomimic, it consistently delivers superior rollout accuracy (e.g., Push-T SSIM of 0.9925, PSNR of 33.78) at a per-frame generation cost of 0.0037 s on a single H100, compared with 0.0104 s (GPC) and up to 1.77 s (Ctrl-World). Multi-view and two-view rollouts on Robomimic further confirm its multi-modal precision.

Figure 3: In Push-T, DriftWorld maintains object integrity and precise interaction in 140-frame rollouts, surpassing diffusion (GPC) and MSE baselines which fail in long-horizon consistency.

Figure 4: In Bridge-V2, DriftWorld accurately simulates fine-grained physical contacts and manipulations, with competing models producing motion artifacts and implausible interactions.

Figure 5: RT-1 task visualization—DriftWorld faithfully reproduces complex grasping and drawer-opening actions.
Quantitative policy evaluation shows Pearson correlation coefficients up to 0.99 between DriftWorld-simulated scores and ground-truth performance across Push-T and Robomimic Lift/Can, outperforming both GPC diffusion and Ctrl-World models in absolute and relative rankings. In inference-time GPC-RANK policy improvement, DriftWorld boosts top-1 IoU scores (e.g., from 0.635 to 0.781 for Push-T) at a fraction of the runtime, enabling exhaustive action proposal rollouts for closed-loop selection.

Figure 6: Policy success rate and IoU predictions correlate with ground-truth performance, validating DriftWorld as an offline policy evaluator for practical robotics.
Ablation studies reveal the necessity of feature space drift and motion weighting for high-resolution, visually complex domains. Omission immediately induces blurring and stagnation in robot manipulators. Varying the accentuation strength at inference allows dynamic modulation of action adherence without retraining.

Figure 7: Varying the accentuation scale parameter (α) dramatically alters the extent to which the model follows action input, allowing online modulation of adherence.

Figure 8: Applied motion weighting prevents the model from ignoring small but critical action-induced displacements.
Theoretical and Practical Implications
DriftWorld defines a new operating point in world modeling: it maintains the high visual fidelity of diffusion/transformer approaches with an order-of-magnitude reduction in inference cost. The contrastive drifting loss, especially when implemented in a robust semantic feature space, provides high-level video realism and temporal consistency while remaining directly conditional on complex action sequences and even language instructions. The approach is explicitly generalizable to any robotic scenario involving high-frequency, action-conditioned video prediction, including multi-agent, real-robot, and multimodal setups.
However, the model's reliance on high-capacity feature encoders (DINOv2/v3) for complex scenes, and the increased memory footprint during training due to the necessity of generating multiple negative rollouts, present real hardware and scaling challenges. These may be mitigated by integrating temporal compression or hierarchical VAEs and by leveraging sparsified context representations.
Future Perspectives
Extending DriftWorld may involve scaling to longer horizons via sparse temporal conditioning, incorporating multimodal (text or tactile) inputs, and integrating policy learning directly into the drifting loss framework. The principle of contrastive generation in semantically rich spaces is naturally extensible to model-predictive control, reinforcement learning from pixels, and robust evaluation of out-of-distribution generalization.
Conclusion
DriftWorld provides formal evidence that single-step, drifting-based generative models can supplant multi-step diffusion as the basis for high-fidelity, high-throughput world modeling in robot learning. Its architectural adaptations—contrastive, action-accentuated drifting; feature-driven loss computation; and robust U-Net conditioning—collectively enable both efficient online control and reliable offline evaluation, with immediate implications for model-based reinforcement learning and policy deployment in real-world robotics.