Capturing action-conditioned physical dynamics for robotic manipulation
Develop action-conditioned video-based world models that accurately capture contact-rich interactions, robot kinematics, and fine-grained physical dynamics necessary to predict object motion under specified 7-DoF end-effector actions and to support closed-loop planning and control for robotic manipulation tasks in RLBench.
References
This gap suggests that while current visual world models can effectively guide perception and navigation, capturing fine-grained physical dynamics and action-conditioned object motion remains an open challenge.
We cannot currently certify that contribution as statistically distinguishable from noise, and a reader should treat the B5-vs.-B4 comparison as suggestive rather than established until either the episode budget is increased or a paired-episode test (matched seed and initial state) is run in place of the unpaired comparison used here.
One iteration recovers a large part of the exploitation gap, and we found empirically that the "grasp-and-miss" hallucinations were significantly reduced in LeWM. However, despite mitigating exploitation, characterizing when it recurs remains open.
However, whether these models learn the causal and physical structure required for reliable control, rather than only statistical regularities in observed trajectories, remains a central open question.