4DCodeBench: Teaching Agents to See Physics in Motion

This presentation examines 4DCodeBench, a new benchmark that challenges multimodal AI agents to reconstruct dynamic physical scenes by writing executable graphics code. Unlike static 3D reconstruction, agents must infer object geometry, material behavior, contact relationships, and temporal evolution from a single RGB video, then generate code that simulates and renders the scene. The benchmark reveals a striking capability gap: even the strongest models reconstruct static appearance and geometry far better than they capture deformation, flow, fracture, and persistent motion, exposing fundamental limitations in how current systems represent physical dynamics.
Script
Most vision AI can tell you what moved, but can it write the code that makes it move? 4DCodeBench asks agents to watch a video and reconstruct not just geometry, but the physics, material behavior, and temporal evolution behind it.
The benchmark does something methodologically important: it does not prescribe how agents should represent motion. They can use analytic trajectories, physics engines, or custom simulators like MPM or SPH. That choice becomes part of what we measure, not an implementation detail hidden inside a fixed pipeline.
Here is the central finding. The strongest model, GPT 6 Astra Max, scores 0.91 on static families like appearance and geometry, but only 0.67 on dynamic families that measure actual motion and deformation. Strong visual reconstruction does not imply physical reconstruction.
When you look at the outputs, the pattern is clear. Stronger models recover both shape and temporal evolution accurately. Weaker models approximate scenes with simple primitives, static arrangements, or motion fields that do not match the underlying physics. The 4D export prevents them from hiding behind a single plausible frame.
Perhaps most surprising: 67 percent of agent solutions use analytic motion, not physical simulation. Even when the task is explicitly about dynamics, most agents fit closed-form trajectories rather than infer and implement a simulator. That choice is not inherently wrong, but it reveals that current models often favor task-specific fitting over transferable physical representations.
The benchmark measures reconstruction fidelity, but it cannot yet distinguish a physically valid simulator from a carefully fitted kinematic program. Future work will need intervention tests, longer horizons, and moving cameras to determine whether agents truly understand dynamics or just memorize motion. Visit EmergentMind.com to explore this research further and create your own videos on the latest AI breakthroughs.