Capturing action-conditioned physical dynamics for robotic manipulation

Develop action-conditioned video-based world models that accurately capture contact-rich interactions, robot kinematics, and fine-grained physical dynamics necessary to predict object motion under specified 7-DoF end-effector actions and to support closed-loop planning and control for robotic manipulation tasks in RLBench.

Background

The paper evaluates visual world models across four embodied tasks and finds that while these models often aid perception and navigation, their benefits for robotic manipulation are modest. The authors attribute this to the difficulty of accurately modeling contact-rich interactions, robot kinematics, and fine-grained dynamics that are essential for manipulation.

In RLBench-based manipulation experiments, even post-trained video generators show only small improvements over strong baselines, indicating that current models struggle with precise action-conditioned predictions in physically complex settings. This motivates the explicit identification of robust physical dynamics modeling as an unresolved challenge for embodied world models.

References

This gap suggests that while current visual world models can effectively guide perception and navigation, capturing fine-grained physical dynamics and action-conditioned object motion remains an open challenge.

World-in-World: World Models in a Closed-Loop World  (2510.18135 - Zhang et al., 20 Oct 2025) in Section 4.1 (Benchmark Results), Robotic Manipulations

We cannot currently certify that contribution as statistically distinguishable from noise, and a reader should treat the B5-vs.-B4 comparison as suggestive rather than established until either the episode budget is increased or a paired-episode test (matched seed and initial state) is run in place of the unpaired comparison used here.

Calibrated Predictive Safety for Heterogeneous Robots: An Action-Conditioned JEPA Framework with Model-Based Safety Shields  (2608.17496 - Zhong et al., 18 Aug 2026) in Section 5, subsection “Closed-Loop Performance (Level 4, executed in simulation)”

One iteration recovers a large part of the exploitation gap, and we found empirically that the "grasp-and-miss" hallucinations were significantly reduced in LeWM. However, despite mitigating exploitation, characterizing when it recurs remains open.

Reinforced Planning with Latent World Models  (2608.18669 - Sommer et al., 19 Aug 2026) in Section 7.4, “World-Model Hallucination and Dyna Finetuning” (p. 8)

However, whether these models learn the causal and physical structure required for reliable control, rather than only statistical regularities in observed trajectories, remains a central open question.

Toward Unified Robot Learning: Bridging Representation, Vision-Language-Action, and World Models  (2609.03927 - Mehta et al., 3 Sep 2026) in Section 6.1, Need for World Models