Papers
Topics
Authors
Recent
Search
2000 character limit reached

Learning Transferable Dynamics Priors from Action to World Modeling

Published 28 Jun 2026 in cs.RO | (2606.29501v1)

Abstract: We study action-conditioned world modeling as a scalable way to learn transferable dynamics priors for robot learning. By pretraining a model to predict how actions drive visual scene evolution, the resulting world model captures reusable interaction dynamics beyond appearance-level video generation. Concretely, we pretrain a multi-view interactive base diffusion world model, A2World, on large-scale robot manipulation data with real action annotations. We validate the learned dynamics priors from two complementary perspectives. First, we adapt A2World into a task- or scene-specialized real-world simulator, A2World-sim, whose long-horizon rollouts support simulator-based policy evaluation and scalable what-if analysis by replacing real-robot rollouts with world model rollouts. Second, starting from the same pretrained weights, we adapt A2World into a video-action joint prediction model, A2World-policy, that predicts actions under visual and instruction conditioning. Experiments across simulation benchmarks and real-robot settings demonstrate that action-conditioned world model pretraining yields transferable dynamics priors that benefit both simulator-centric and policy-centric robot learning.

Summary

  • The paper introduces A2World, a 2.5B-parameter multi-view latent diffusion model pretrained on 2.156 million robot trajectories, achieving strong action-faithful video prediction across embodiments and camera setups.
  • The paper develops A2World-sim with history-aware autoregressive rollouts and pose-guided frame selection, reaching a simulator-to-reality evaluation correlation of Spearman ρ=0.916 and Pearson r=0.965 across eight policy-task pairs.
  • The paper transfers the learned dynamics prior to A2World-policy, which achieves 98.6% LIBERO success and 88.5% on the OOD LIBERO-Plus Spatial benchmark, showing that action-conditioned pretraining supports both simulation and control.

Motivation and core claim

The paper investigates action-conditioned world modeling as a mechanism for learning transferable dynamics priors for robot learning. The central argument is that actions constitute a natural causal supervision signal in manipulation: while objects, scenes, and viewpoints vary widely across datasets, state transitions are governed by how actions induce contacts, grasps, pushes, and releases. Pretraining a model to predict visual scene evolution conditioned on actions therefore encodes controllable interaction dynamics rather than appearance-level video statistics. The authors instantiate this with A2World, a multi-view latent diffusion world model pretrained on 2,156k robot manipulation trajectories spanning more than 20 embodiments, and demonstrate that the resulting prior transfers to two distinct downstream uses: a history-aware autoregressive simulator (A2World-sim) for policy evaluation, and a video-action joint prediction policy (A2World-policy). A notable methodological choice is that pretraining uses real action annotations directly, without auxiliary latent-action models producing pseudo-labels, which the authors argue avoids the embodiment and camera-setup constraints of prior work.

The base world model

A2World follows a latent video diffusion pipeline built on the WAN2.1 tokenizer, with a DiT backbone initialized from the Cosmos-Predict2-2B-Video2World checkpoint, yielding a 2.5B-parameter model trained with the EDM denoising score matching objective. Given an initial frame and a 20-step action chunk, it forecasts the next 20 frames. Action conditioning is injected by adding an MLP-encoded action embedding to the diffusion timestep embedding consumed by adaptive layer normalization in every DiT block; the cross-attention conditioning is deliberately zeroed so that the model relies solely on action-conditioned temporal modulation.

Two architectural components address multi-view consistency. Learnable view embeddings are concatenated to latent tokens before patch embedding to provide explicit camera identity, and cross-view attention modules are inserted into each DiT block so tokens from one view attend to tokens from other views. Multi-view latents from all VV cameras are packed along the temporal dimension, so multi-view generation reduces to a single temporally concatenated video diffusion pass. Pretraining mixes seven datasets (AgiBot, DROID, Open X-Embodiment, InternData-A1, InternData-M1, RoboCOIN, Galaxea), with all actions unified into a shared 14-dimensional dual-arm end-effector format (zero-padded for single-arm robots) and dataset-consistent batching to avoid camera-convention conflicts.

Qualitatively, the pretrained model exhibits counterfactual controllability: from the same initial frame it can be steered to grasp different objects or reproduce failed grasps under the corresponding actions, and it simulates full-DoF arm control on RoboCOIN that never appears in pretraining data. Rollouts on unseen scenes and camera setups (RoboMind, VIOLA) remain coherent.

A2World-sim: long-horizon simulation and policy evaluation

To convert the short-horizon predictor into a simulator, A2World-sim conditions on a history buffer and rolls out chunks autoregressively. History is injected through two paths: history tokens replace the (zeroed) cross-attention context, and the same tokens are projected into key/value memories concatenated with the self-attention stream. A pose-guided history sampling algorithm selects a fixed budget of frames uniformly along a weighted arc-length computed from relative end-effector translations and rotations, preserving motion-informative states such as turning points under a fixed token budget. Training adopts a Self-forcing-style strategy that periodically conditions the model on its own generated frames, exploiting the fact that under action conditioning and an initial frame, the future trajectory is largely determined by dynamics—no separate teacher model is required.

On rollout quality, A2World-sim achieves the best results on both LIBERO and the custom Flexiv real-robot data against Cosmos-Predict2, Ctrl-World, Prophet, and a text-conditioned variant pretrained on the same robot data (T-pre). On LIBERO it reaches PSNR 26.64, SSIM 0.8957, and optical-flow action-faithfulness metrics (EPE 0.3498, cosine 0.6045) that surpass all baselines; the text-pretrained variant is competitive in-distribution but degrades on real-robot data (PSNR 24.64 vs. 25.95), supporting the claim that action-conditioned pretraining yields a stronger dynamics prior than text-conditioned pretraining on identical data. On RoboNet, A2World-sim attains an FVD of 146.1 versus 175.3 for SAMPO, with the lowest LPIPS (8.9), though its PSNR (24.1) is slightly below SAMPO (25.3)—a limitation the paper reports but does not analyze.

Most consequential for the simulator use case, closed-loop policy evaluation inside A2World-sim correlates strongly with real-world outcomes across eight policy–task pairs: Spearman ρ=0.916\rho=0.916, Pearson r=0.965r=0.965, R2=0.930R^2=0.930, based on roughly 25 real and 64 simulator rollouts per policy with manual outcome verification. Under distribution shift (LIBERO fine-tuning, LIBERO-Plus Spatial evaluation), A2World-sim improves action-faithfulness over DreamDojo (EPE 0.1301 vs. 0.2738) and better preserves novel scene appearance qualitatively, though its temporal SSIM (0.7401) is lower than DreamDojo's (0.7778). The ablation on history sampling shows pose-guided selection consistently outperforms sliding-window and no-history baselines on all five metrics, indicating that which frames enter the history buffer materially affects long-horizon stability.

A2World-policy: transferring the prior to control

A2World-policy is a 3.0B-parameter MoE-like joint video-action diffusion model: video and action tokens share a single self-attention module per block, while each modality retains its own AdaLN and MLP denoising branches. The action branch is initialized by copying video-branch parameters from A2World, cross-attention is re-initialized from Cosmos-Predict2 (since pretraining zeroed it), and a T5 encoder provides instruction conditioning. Training uses a shared base noise level scaled per modality, and inference applies modality-wise classifier-free guidance with separately tunable guidance strengths for video and action.

On LIBERO, A2World-policy achieves 98.6% average success across the four suites, exceeding π0.5\pi_{0.5} (96.9%), OpenVLA-OFT (97.1%), CogVLA (97.4%), and Cosmos Policy (98.5%), though it is not uniformly best per suite. The OOD evaluation is the more informative comparison: fine-tuning on LIBERO and testing on LIBERO-Plus Spatial, the action-video-pretrained variant (A-pre) reaches 88.5% average success versus 80.2% for direct Cosmos initialization (C-init) and 85.8% for text-conditioned robot-data pretraining (T-pre), and is essentially tied with policy-targeted pretraining (P-pre, 88.6%) that uses a downstream-matched objective. The implication is that action-to-video pretraining already supplies the dynamics prior needed for policy transfer, with policy-targeted pretraining contributing only a marginal in-domain gain (98.8% vs. 98.6% on LIBERO) while sacrificing simulator reusability. On the custom five-task real-robot suite—covering contact-rich insertion, articulated switches, and deformable chains—A2World-policy outperforms π0.5\pi_{0.5} and LingBot-VA in both task progress and success rate, with the largest gains on long-horizon contact-rich tasks where baselines stall early; all evaluations were performed by third-party operators under a standardized protocol.

Video-action coupling and ablations

A consistent positive coupling between video prediction quality and action quality is observed across training checkpoints: checkpoints with better optical-flow video consistency also score better on a normalized composite action error (translation, geodesic rotation, gripper F1), and fully joint training reaches a stronger frontier than freezing the video branch (the frozen variant drops to 86.2% on LIBERO). This supports the paper's framing that shared forecasting and control representations reinforce each other, rather than competing for capacity.

Limitations and open questions

The paper concedes several points implicitly or explicitly. The policy evaluation protocol on real robots relies on manually assigned milestone-based progress scores and manually verified simulator outcomes, introducing subjective judgment into the headline correlations and success rates. The simulator-evaluation correlation is computed over only N=8N=8 policy–task pairs, so the strength of the ρ=0.916\rho=0.916 estimate carries substantial sampling uncertainty. The model is restricted to 20-frame chunks at 10 fps with a 20-frame history, and the multi-view design assumes dataset-consistent camera configurations within each batch, which constrains how heterogeneous datasets can be mixed. The RoboNet PSNR deficit relative to SAMPO suggests appearance fidelity is not uniformly superior. Open questions left by the paper include whether the action-to-video prior scales to higher-frequency, closed-loop control beyond the 10 fps rollout regime, whether the dual-arm 14-dimensional action unification loses fidelity for high-DoF whole-body or dexterous-hand embodiments, and whether simulator-based post-training (reward models plus RL) built on A2World-sim inherits the same fidelity as evaluation-only use.

Conclusion

The paper establishes that large-scale action-to-video pretraining on real robot data produces a dynamics prior that is demonstrably stronger than text-conditioned pretraining on identical data, and that this single prior serves two downstream roles—long-horizon simulator for policy evaluation and initialization for a competitive joint video-action policy—with only lightweight adaptation. The strongest quantitative evidence is the near-tie between action-video pretraining and downstream-matched policy pretraining under OOD transfer, combined with high simulator-to-reality evaluation correlation, which together argue that action-conditioned world modeling is an efficient, reusable foundation for both simulator-centric and policy-centric robot learning.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.