Papers
Topics
Authors
Recent
Search
2000 character limit reached

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

Published 2 Oct 2026 in cs.CV, cs.AI, and cs.GR | (2610.03715v1)

Abstract: We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstractions such as physical simulations to reproduce complex behavior. To evaluate this capability, we curate a set of real-world videos and construct synthetic scenes spanning diverse physical phenomena, including deformation, fluid flow, and fracture. We perform extensive benchmarking of frontier models, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics. 4DCodeBench provides a testbed for tracking progress toward agents that can interpret the dynamics of the world through code. Our benchmark is available at https://github.com/4DCodeBench/4DCodeBench

Summary

  • The paper introduces 4DCodeBench, a benchmark that evaluates agents' ability to reconstruct dynamic 3D scenes from videos, focusing on both visual fidelity and underlying physical dynamics, the code produces a results in a deterministic `build.sh` program that reconstructs the world without reading the reference video during execution.
  • The benchmark includes 200 scenes, consisting of real-world and synthetic videos, to test agents on various dynamic and material scenarios, and 90.1% are fully executable; proprietary models show 96.5%
  • The evaluation framework measures reconstruction quality across five principal families: Perceptual, 2D Dynamics, 2.5D Geometry, 3D Geometry, and 3D Dynamics with strong correlation of VLM Elo and human Elo at Spearman $\rho = 0.980$

Problem formulation and benchmark design

“4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes” (2610.03715) formulates dynamic-scene reconstruction as executable program synthesis. Given only an RGB video, an agent must generate graphics code that constructs a 3D scene, specifies its evolution through time, renders the result in Blender, and exports an explicit 4D representation. The benchmark therefore evaluates more than image-to-image imitation: the output must expose geometry, temporal state, and persistent correspondences for moving matter.

This formulation transfers model specification into the inference problem. Agents must decide whether to use analytic trajectories, keyframed motion, Blender’s physics systems, or custom simulators such as MPM, PBD, SPH, or FEM. The benchmark does not prescribe a physical representation or simulator. This is methodologically important because the choice of abstraction becomes an observable component of performance rather than an implementation detail hidden by a fixed reconstruction pipeline.

The task is constrained operationally but broad representationally. Each run receives a reference video and a fixed task specification, while scene-specific metadata, semantic descriptions, object lists, and ground-truth worlds are withheld. The agent operates in an isolated container with Blender, numerical libraries, Taichi, Warp, and related tooling. It must produce a deterministic build.sh program that reconstructs the world without reading the reference video during execution. The exported world includes camera parameters, per-frame meshes, dynamic-matter trajectories, and the rendered video. Consequently, a submission can be evaluated both through its rendered appearance and through its underlying 3D state.

The benchmark contains 200 scenes: 100 real-world videos and 100 synthetic scenes with full 4D ground truth. The real subset captures uncontrolled appearance and physical behavior but lacks complete world-state supervision. The synthetic subset enables direct comparison of geometry and motion in 3D, although it necessarily inherits the assumptions and biases of its simulators. The scene ontology covers rigid and articulated bodies, deformable solids, co-dimensional structures such as cloth and rods, and flowing materials including fluids, grains, and viscoplastic substances. The dataset is interaction-heavy: 83% of scenes contain multiple dynamic objects and 66% contain multiple material types.

Figure 1

Figure 1: Dataset examples and distributions across matter families, object counts, material counts, and passive versus driven motion.

This composition makes the benchmark substantially more demanding than static object reconstruction. The agent must infer not only object shape and camera configuration but also contact relationships, material behavior, external driving forces, topology changes, and temporal persistence. The authors deliberately exclude most humans and animals and favor limited occlusion and stationary-camera footage, narrowing the benchmark’s scope to physical scene dynamics rather than full articulated human-scene understanding.

Evaluation framework

The evaluation separates reconstruction quality into five principal families: Perceptual, 2D Dynamics, 2.5D Geometry, 3D Geometry, and 3D Dynamics. The Overall score is the mean of these families. This decomposition avoids treating pixel-level or semantic appearance as a sufficient proxy for physical reconstruction.

Perceptual similarity uses DINOv3 features over rendered and reference frames. Dynamic IoU measures the spatial extent of moving objects. For real videos, the benchmark additionally estimates depth, optical flow, and point trajectories from the reference and compares them with quantities analytically rendered from the reconstructed 4D world. Synthetic scenes permit stronger supervision: frame-zero geometry is evaluated with Chamfer distance, while motion is evaluated using both trajectory-level DTW and frame-to-frame displacement distributions via sliced Wasserstein distance.

The distinction between Lagrangian and Eulerian motion metrics is particularly useful. Trajectory DTW tests whether persistent matter follows the correct 3D path, requiring correspondence across time. EMD step instead compares distributions of instantaneous displacements without requiring persistent identity. The former detects incorrect material assignment and long-horizon trajectory errors; the latter captures aggregate velocity and direction while remaining applicable to fluids and topology-changing matter.

The benchmark supplements these metrics with two VLM-based evaluations. Scene-specific VQA contains 1,251 binary questions covering initial state, final state, key events, and contact. A separate pairwise judge compares two reconstructions against the reference and produces model-level Elo ratings. The authors also conduct a human preference study using 3,587 judgments from 76 participants.

These validation results support the automated evaluation, but also identify its operating range. Human and VLM Elo correlate at Spearman ρ=0.980\rho=0.980, while human Elo correlates with the Overall score at ρ=0.96\rho=0.96. On 916 comparisons assessed by both humans and the VLM, agreement is 89.3% with κ=0.761\kappa=0.761, compared with 92.1% agreement among repeated human judgments. However, VLM comparisons become unreliable for closely matched models: agreement is near chance within a 50-point Elo difference and rises to 90% only at a 343-point separation. Thus, VLM ranking is well suited to broad model comparison but should not be interpreted as a precise evaluator of small performance differences.

Overall model performance

The study evaluates 18 multimodal coding models, including proprietary and open-weight systems, with one run per model per scene. GPT-6 Astra [Max] ranks first on the Overall metric, followed by Claude Opus 5.5 [High], GPT-6 Astra [High], Claude Fable 5.1 [High], and GPT-6 Astra [Low]. Proprietary models generally outperform open-weight models, although the open-weight leaderboard itself contains substantial variation.

Figure 2

Figure 2: Overall leaderboard across 18 models, including human and VLM Elo, VQA accuracy, metric-family scores, and the aggregate Overall score.

Execution reliability is part of the evaluation rather than an independently reported engineering statistic. Across all 3,600 model-scene runs, 90.1% are fully executable. Every proprietary model reaches at least 96.5% executability, whereas open-weight models range from 29.0% to 95.5%. Only 5.4% of runs that pass the video gate fail subsequent 4D-world validation, indicating that the main bottleneck is not output formatting. The benchmark therefore supports the authors’ interpretation that many failures reflect deficiencies in 4D reasoning, scene abstraction, and simulation rather than inability to satisfy the file specification.

VQA provides a more granular view of reconstruction reliability. GPT-6 Astra [Max] achieves 87.6% mean per-scene accuracy, compared with 78.8% for Claude Fable 5.1 [High] and 29.5% for GLM 5.3 Flash [Max]. Scene-level aggregation reveals a highly nonuniform error pattern: Astra [Max] answers every question correctly on 59% of scenes, whereas Mistral answers none correctly on 94% of scenes.

Figure 3

Figure 3: VQA evaluates initial state, final state, key events, and contact using questions applied to the reference and edited reconstruction videos.

The results contradict a simple compute-scaling interpretation of agentic reconstruction. Across different models, token use and agent-step count have little relationship to Overall quality; raw step and token correlations are reported as ∣ρ∣≤0.10|\rho|\leq 0.10. Qwen3.8 Flash [XHigh], for example, uses an average of 752 agent steps and 1.5 million output tokens but ranks eleventh. This indicates that additional interaction is not intrinsically productive when the model’s scene abstractions or debugging strategies are inadequate.

The static–dynamic capability gap

The most consistent empirical finding is a separation between appearance and static geometry on one hand, and motion reconstruction on the other. Even GPT-6 Astra [Max], the strongest evaluated system, obtains approximately 0.91 on the static families—Perceptual and 3D Geometry—but only 0.67 on the dynamic families—2D and 3D Dynamics. The gap persists across models and across both image-space and 3D evaluations.

This result is central to the paper’s claim that strong visual reconstruction does not imply physical reconstruction. A model can produce a recognizable object with plausible initial geometry while failing to reproduce deformation, fracture, flow, contact, or the correct final state. The moderate within-scene concordance between metrics reinforces this distinction: appearance and layout metrics correlate relatively strongly, with DINOv3 and Dynamic IoU at ρ=0.76\rho=0.76, whereas trajectory DTW is more independent, correlating with other metrics at only approximately $0.39$–$0.50$.

The independence of trajectory DTW is diagnostically valuable. Per-frame visual similarity and displacement statistics may remain high when the wrong matter follows the wrong path or when temporal correspondences are inconsistent. In this sense, explicit 4D state export prevents a model from receiving full credit for a visually plausible but physically misassigned sequence.

Qualitative results illustrate the same pattern. Stronger models recover both object geometry and temporal evolution more effectively, while weaker models often approximate the scene with simple primitives, static arrangements, or incorrect motion fields.

Figure 4

Figure 4: Qualitative comparison showing that higher-ranked models more accurately recover both scene geometry and dynamics.

The benchmark also reveals that mesh validity is not equivalent to reconstruction fidelity. Weaker systems often produce simple watertight primitives that score well on manifoldness and related checks, while stronger systems attempt fractured solids, thin shells, and fluid surfaces that are structurally more difficult to mesh. Muse Glimmer [High] leads or ties on all four reported mesh-integrity checks, whereas GPT-6 Astra [Max] reaches only 0.830 watertightness. Interpenetration is relatively high across all models, ranging from 0.851 to 0.959 when expressed as the reported no-interpenetration score. The authors therefore correctly keep structural mesh diagnostics outside the Overall metric.

Scene-type effects and real-world generalization

Performance varies systematically with scene composition. Real scenes produce lower VQA and Perceptual scores than synthetic scenes, with standardized effects of −0.73-0.73 and −0.88-0.88, respectively. Scenes containing multiple matter types likewise reduce VQA and Perceptual performance by −0.57-0.57 and ρ=0.96\rho=0.960. These effects are not isolated to a small number of models; they recur across most of the evaluated systems.

Figure 5

Figure 5

Figure 5: Standardized category effects across matter families, scene sources, object counts, and motion regimes.

Co-dimensional structures are particularly difficult for 2D Dynamics and 2.5D Geometry, with effects of ρ=0.96\rho=0.961 and ρ=0.96\rho=0.962. This is consistent with the sensitivity of cloth, rods, and shells to thin geometry, self-occlusion, topology, and long-range deformation. Driven scenes have lower VQA but higher 3D Geometry, with effects of ρ=0.96\rho=0.963 and ρ=0.96\rho=0.964. External driving can make object motion easier to localize while still producing complex event semantics. Flowing scenes show lower 3D Dynamics by ρ=0.96\rho=0.965, although this effect does not satisfy the authors’ consistency threshold across models.

The real–synthetic comparison exposes a second form of generalization failure. Most models score lower on real videos while using 17% fewer tokens on them. Claude Opus 5.5 [High] nearly matches GPT-6 Astra [Max] on synthetic VQA, with scores of 0.88 and 0.90, but falls substantially behind on real footage, where Astra reaches 0.85 and Opus 5.5 reaches 0.78. GPT-6 Astra [Max] exhibits the strongest transfer from synthetic to real scenes among the tested systems. The result suggests that synthetic-scene competence is not a reliable proxy for robustness to real appearance, imperfect visibility, and unconstrained material behavior.

The dataset’s paired real and synthetic structure is therefore analytically useful, but its interpretation remains bounded by the synthetic construction process. Synthetic scenes are generated by multiple simulators, including MPM, SPH, IPC, ABD, and PPF-based systems, and are reviewed for numerical artifacts. Nevertheless, they encode simulator-specific priors that may favor models capable of reproducing common procedural patterns rather than models with generally valid physical representations.

How agents represent motion

The executable outputs permit analysis of the representations chosen by the models. Across all runs, 67% use analytic motion, 19% custom simulation, 10% Blender physics, and 3% keyframing. Analytic strategies include spline interpolation, prescribed vertex displacements, ballistic formulas, pose tables, and manually parameterized evolving shapes.

This distribution is a consequential result: most agents do not attempt to infer or implement a physical simulator, even though the task is explicitly framed around dynamic scenes. The benchmark permits this behavior because reconstruction fidelity, not mechanistic purity, is the primary objective. A closed-form trajectory can be preferable when the observation contains a single short event and the model can fit the visible motion more reliably than it can construct a stable solver.

Model-level strategy distributions differ substantially. Claude Opus 5.5 [High] uses custom simulation in 61% of solutions, whereas GPT-6 Astra [Low] uses analytic motion in 85% and GPT-5.6 Terra [High] uses it in 100%. These differences demonstrate that model rankings cannot be explained solely by a shared reconstruction algorithm. The agents are selecting distinct computational representations under the same environment and prompt.

Figure 6

Figure 6: Agent loop in which the model writes 4D code, renders the scene, compares the result with the input, and iteratively edits the program.

The paper gives concrete examples of representation choice. For dam-break scenes, nearly every model uses Blender’s Mantaflow FLIP solver. For rod-like pasta dynamics, most models implement PBD-style simulations. A bread-tearing scene produces more divergent approaches: one model uses two-field MLS-MPM with a fracture threshold, another uses XPBD ligament snapping, and GPT-6 Astra [Max] uses an analytic cohesive-fracture front without a solver. These cases show that simulation sophistication and reconstruction quality are not equivalent. A physically structured solver may generalize better under perturbation, but a task-specific analytic abstraction can match the observed clip more closely.

This distinction motivates one of the paper’s principal unresolved questions: whether simulation improves reconstruction quality when the objective is merely to reproduce a fixed observation, and whether requiring physics-based simulation would change the relative model ranking.

Inference-time reasoning and computational cost

Within the GPT-6 Astra family, additional reasoning effort improves performance monotonically. Increasing effort from Low to High to Max raises Overall from 0.73 to 0.77 to 0.79. The corresponding 2D Dynamics score increases from 0.48 to 0.55 to 0.60, while 3D Dynamics rises from 0.63 to 0.69 to 0.73. VQA increases from 78.2% to 85.0% to 87.6%.

The effect is accompanied by substantial additional computation: output tokens increase from approximately 13,000 to 29,000 to 67,000, and cost increases from about $\rho=0.96$611.30 per task. Thus, within a fixed model family, test-time reasoning is beneficial, particularly for dynamics. The result does not contradict the weak cross-model relationship between token volume and quality; it indicates that model capability and reasoning policy interact, so aggregate token counts are not a sufficient causal explanation of performance.</p> <p>The full benchmark requires considerable resources: 3,600 runs consume approximately 5,983 hours and $16,495 in reported compute. Open-weight agents use more steps and tokens on average—247 steps and 39.6 million tokens per task versus 101 steps and 11.8 million tokens for proprietary models—but cost less because of lower token pricing. Approximately 97% of tokens are cache reads. These measurements make the benchmark useful not only for quality comparison but also for studying the cost–quality frontier.

The results nevertheless argue against treating long agent trajectories as a proxy for effective reasoning. Qwen3.8 Flash [XHigh] has extreme interaction volume without commensurate quality, while additional reasoning for Astra yields consistent gains. The relevant variable is therefore the productivity of render–compare–edit cycles and the quality of the hypotheses they generate, not merely the number of cycles.

Limitations and open questions

The benchmark evaluates reconstruction fidelity, not the correctness of the inferred physical mechanism. An agent can prescribe a trajectory that matches the observed video while encoding no transferable model of material properties, contact laws, or external forces. The explicit 4D output improves observability but does not resolve this identifiability problem. In particular, the current metrics cannot distinguish a physically valid simulator from a carefully fitted kinematic program when both reproduce the recorded sequence.

The real-video metrics also depend on estimated depth, optical flow, and point tracks from pretrained vision systems. These estimates introduce model-dependent noise and may be especially unreliable for transparent, thin, fast-moving, or textureless objects. Six real scenes are excluded from several rasterized metrics because transparent objects defeat the surface-rasterization procedure. Synthetic evaluation provides stronger geometric supervision, but its physical diversity remains constrained by simulator assumptions and author-designed scenarios.

The one-run-per-scene protocol limits analysis of stochasticity and recovery from failure. Although agents operate in an iterative render–compare loop, the benchmark primarily scores final submissions. It does not measure how quickly a model improves, which intermediate hypotheses it rejects, when gains saturate, or whether models use visual feedback effectively. The authors explicitly identify intervention tests—such as changing initial conditions or external forces—and longer-horizon prediction as necessary to test whether a reconstruction encodes transferable dynamics rather than a clip-specific fit.

Finally, the benchmark focuses on mostly stationary-camera videos with limited occlusion and excludes most humans and animals. The results therefore address executable inverse graphics for object-centric dynamic scenes, not general video world modeling. Whether the static–dynamic gap persists under moving cameras, severe occlusion, multi-view inputs, or longer temporal horizons remains open.

Conclusion

4DCodeBench establishes executable dynamic-scene reconstruction as a benchmark for multimodal coding agents. Its combination of real videos, simulator-generated scenes, explicit 4D outputs, image-space and world-space metrics, VQA, VLM ranking, and human preference evaluation provides a comparatively detailed diagnosis of current capabilities.

The main result is a robust static–dynamic asymmetry: models can recover recognizable appearance and geometry substantially better than they can infer and reproduce material interactions, deformation, flow, fracture, and persistent motion. The strongest model reaches approximately 0.91 on static families but only 0.67 on dynamic families. Human and automated rankings align strongly, while scene-level and category-level analyses show that real footage, multi-material interactions, co-dimensional structures, and flowing matter remain especially difficult. The benchmark consequently isolates 4D abstraction and dynamic representation—not merely rendering quality—as the central unresolved capability tested by the paper.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

The paper introduces 4DCodeBench, a test for AI systems that can look at a video and write computer code to recreate what happens in it.

The word 4D means:

  • 3D space: where objects are and what they look like
  • Time: how those objects move and change

For example, an AI might watch a video of a ball hitting a soft object, water flowing, cloth moving, or something breaking. It must then write a program that rebuilds the scene and shows the same events happening over time.

This is called inverse graphics. Normally, computer graphics code creates an image or video from a description. In inverse graphics, the process is reversed: the computer watches an image or video and tries to figure out the hidden 3D objects, materials, and movements that created it.

2. What questions are the researchers asking?

The researchers mainly want to know:

  1. Can AI coding agents understand moving scenes from video?
  2. Can they recreate not only what objects look like, but also how they move and change?
  3. Are AI systems better at copying appearance than understanding motion and physics?
  4. Which kinds of scenes are hardest for AI?
  5. Do AI systems use real physics simulations, simple mathematical movements, or other tricks?
  6. Do automatic computer scores agree with what human viewers think looks most accurate?

The important challenge is that the AI is not given a list of objects or a written explanation. It only receives the video and must work out what is happening by itself.

3. How did the researchers conduct the study?

Building the benchmark

The researchers created a collection of 200 scenes:

  • 100 real-world videos, taken from physics videos, robot experiments, and online footage
  • 100 computer-generated scenes, made using physics simulators

The synthetic scenes are useful because the researchers know the exact 3D shape and movement of every object. This is like testing a student with an answer key: they can compare the AI’s reconstruction with the true answer.

The scenes include many types of materials and actions:

  • Hard objects moving or colliding
  • Robots moving objects
  • Cloth and ropes bending
  • Soft materials stretching or squashing
  • Fluids flowing
  • Sand and grains moving
  • Objects breaking or cracking

Asking AI models to write graphics programs

The researchers tested 18 multimodal coding models. These are AI systems that can understand images or videos and also write code.

Each model received:

  • One reference video
  • A task description
  • A computer environment containing tools such as Blender

The model had to produce an executable program. When the program ran, it needed to:

  1. Build the objects in 3D
  2. Describe how the objects change over time
  3. Create a new video
  4. Provide the objects’ 3D positions and shapes at different moments

The program was not allowed to look at the original video while running. This prevented the model from simply copying individual frames.

Different ways of creating motion

The AI models could choose how to reproduce movement. For example, they could use:

  • Analytic motion: using formulas, such as moving an object along a curved path
  • Keyframing: manually specifying where an object should be at important times
  • Physics simulation: using rules about forces, collisions, liquids, and materials
  • Custom simulation code: writing their own program to imitate physical behavior

A useful analogy is drawing a bouncing ball. One method is to tell the ball exactly where to be in every frame. Another is to use physics rules so that gravity and collisions make the ball bounce naturally.

Measuring the results

The researchers used several kinds of measurements:

  • Appearance: Does the recreated video look like the original?
  • 2D motion: Do moving objects travel across the screen in the same way?
  • 2.5D geometry: Does the estimated depth, or distance from the camera, look correct?
  • 3D geometry: Are the reconstructed objects shaped correctly?
  • 3D motion: Do the objects move through 3D space correctly?

They also used:

  • A vision-LLM as an automatic judge
  • Questions about the scene, such as what happened at the beginning or end
  • Human comparisons in which people chose which of two reconstructions was better

In total, 76 people made 3,587 comparisons.

4. What were the main findings?

AI is better at copying appearance than motion

The clearest result is that AI models can often recreate how a scene looks, but they have much more trouble recreating how it changes.

The best-performing model, called Astra [Max] in the paper, achieved:

  • About 0.91 on static aspects, such as appearance and geometry
  • About 0.67 on dynamic aspects, such as motion

This means the model might correctly build something that looks like the original object, but its movement, bending, flowing, or breaking may be wrong.

For example, an AI might create a good-looking piece of cloth but fail to make it fold in the correct way. It might recreate a cup and a table but make the cup fall too slowly or bounce incorrectly.

The strongest model performed best overall

The overall ranking placed Astra [Max] first. Other strong models included Claude Opus 5.5, Astra [High], Fable, and Astra [Low].

Proprietary models generally performed better than the open-weight models tested. However, the open models showed large differences between one another.

About 90.1% of all attempts successfully produced executable results. Proprietary models had especially high success rates, while some open models often produced programs that did not run correctly or did not meet the required format.

More reasoning helped within the same model

When the researchers gave Astra more time and effort to reason, its performance improved:

Reasoning level Overall score 2D motion score 3D motion score
Low 0.73 0.48 0.63
High 0.77 0.55 0.69
Max 0.79 0.60 0.73

Its ability to answer questions about the reconstructed scenes also improved, from 78.2% at the Low setting to 87.6% at the Max setting.

However, simply using more computer tokens did not always make one model better than another. This suggests that how an AI reasons may matter more than how much it writes.

AI models used different strategies

Most solutions used simple mathematical descriptions of movement:

  • 67% used analytic motion
  • 19% used custom physics simulations
  • 10% used Blender’s physics tools
  • 3% used keyframing

The models differed greatly in their choices. Some mostly guessed paths directly, while others tried to build physical simulations.

This is important because it shows that AI systems may reach similar-looking results using very different ideas about what is happening.

Real and complicated scenes were harder

Real-world videos were generally harder than synthetic scenes. Real videos contain details and physical effects that are difficult to measure exactly.

Scenes with several different kinds of materials were also harder. For example, a scene involving a hard object, cloth, and water is more difficult than a scene involving only one solid object.

Other difficult cases included:

  • Cloth and rope
  • Flowing materials
  • Multiple objects touching each other
  • Scenes with complicated interactions
  • Events involving deformation, breaking, or changing shape

Automatic scores agreed well with human opinions

The automatic vision-language judge ranked the models very similarly to human viewers.

The researchers found a very strong agreement between human and automatic rankings. On shared comparisons, humans and the AI judge agreed about 89.3% of the time.

This suggests that the benchmark’s automatic tests are useful for comparing models without requiring people to watch every single reconstruction.

5. Why are these results important?

This research shows that understanding a moving world is much harder than recognizing objects or copying a single image.

To recreate a dynamic scene, an AI must work out things such as:

  • Which objects exist
  • What materials they are made from
  • Which objects are touching
  • What forces are acting
  • How an object will bend, stretch, flow, or break
  • What caused the movement
  • How the scene will look in the future

A video does not directly reveal all of this information. For example, a ball moving across a screen could be rolling, sliding, being pushed, or following a preplanned path. The AI has to choose a reasonable explanation.

Conclusion: What could this research lead to?

4DCodeBench gives researchers a standard way to measure whether AI systems truly understand dynamic scenes rather than merely producing attractive pictures.

In the future, better systems could help with:

  • Robots learning how to handle objects
  • Computer animation
  • Video editing and special effects
  • Virtual reality and game worlds
  • Scientific simulations
  • Predicting how materials and objects will behave

However, the paper also shows that current AI systems still struggle with true physical understanding. They can often make something that looks right, but they may not understand why it moves that way.

The researchers suggest that future tests should ask AI systems to change the starting conditions or apply new forces. For example, instead of only recreating a ball’s original bounce, the AI could be asked what would happen if the ball were dropped from a different height. Such tests would show whether the AI learned the underlying physics or simply memorized one particular movement.

Overall, the paper presents 4DCodeBench as a useful measuring tool for progress toward AI systems that can understand, simulate, and interact with the changing physical world.

Knowledge Gaps

Knowledge Gaps, Limitations, and Open Questions

  • The benchmark evaluates reconstruction fidelity but does not determine whether an agent has recovered the underlying physical mechanism; prescribed trajectories can obtain high scores without learning transferable dynamics.
  • It remains unknown whether physics-based simulation produces better generalization than analytic motion, keyframing, or other kinematic approximations when scenes, initial conditions, or viewpoints change.
  • The study does not test interventions such as modified initial conditions, external forces, object properties, or contact configurations, so causal physical understanding is not measured.
  • Longer-horizon prediction beyond the observed video is not evaluated, leaving open whether reconstructed programs can forecast future dynamics rather than merely reproduce the input sequence.
  • The real-world subset lacks complete 3D and 4D ground truth, limiting the ability to distinguish correct reconstruction from visually plausible but geometrically or physically incorrect explanations.
  • The synthetic subset depends on simulator assumptions and parameterizations, which may not reflect real material behavior, contact dynamics, topology changes, sensor noise, or rendering artifacts.
  • The extent to which models overfit to the particular simulators, object meshes, material categories, and procedural distributions used to create the synthetic scenes is not tested.
  • The dataset contains only 200 scenes, and the paper does not establish how performance scales with substantially larger, more varied, or more systematically balanced datasets.
  • Scene selection favors stationary cameras, limited occlusion, continuous footage, and the absence of humans and animals; consequently, performance under camera motion, severe occlusion, cuts, clutter, articulated humans, and animal motion remains unresolved.
  • The benchmark largely excludes challenging real-world factors such as lighting changes, reflections, transparency, motion blur, sensor noise, imperfect segmentation, and dynamic backgrounds.
  • The impact of monocular ambiguity is not isolated from other sources of difficulty; the study does not compare monocular input with stereo, multi-view, depth, optical-flow, or other additional observations.
  • Camera and illumination parameters are left for agents to reconstruct, but the paper does not separately quantify errors in camera estimation, lighting estimation, and material appearance versus errors in object geometry and dynamics.
  • The benchmark does not evaluate whether reconstructed programs remain valid under novel viewpoints, camera trajectories, lighting conditions, or rendering engines.
  • The evaluation is primarily based on a single observed trajectory, so it cannot establish whether the inferred scene representation supports counterfactual rendering or manipulation.
  • The overall score averages heterogeneous perceptual, geometric, and dynamic metrics, but the paper does not justify the equal weighting or analyze how alternative weightings change model rankings.
  • Several evaluation signals rely on off-the-shelf vision models for depth, optical flow, point tracking, and DINO features; their biases and failure modes may systematically affect scores, especially for fluids, transparent objects, deformable materials, and large motion.
  • The use of analytically exposed geometric states for some metrics may favor submissions that provide convenient internal representations, even when those representations do not correspond to visually observable or physically meaningful matter.
  • The alignment procedure used before synthetic 3D evaluation may remove meaningful reconstruction errors, and its sensitivity to alignment choices is not investigated.
  • The trajectory and displacement metrics may not adequately evaluate topology-changing phenomena such as fracture, splitting, merging, cutting, or fluid surface evolution; their suitability across all matter classes remains uncertain.
  • Mesh-validity diagnostics identify structural defects but are excluded from the main score, leaving unresolved how geometric validity should be balanced against visual fidelity and physical accuracy.
  • The benchmark does not assess conservation laws, stability, contact forces, material parameters, energy behavior, or other physical properties that could distinguish physically plausible simulations from visually matching animations.
  • The analysis reports broad solution categories such as analytic motion and custom simulation, but does not measure the quality, correctness, or transferability of the underlying simulators used within those categories.
  • It is unclear whether agents choose simulation or analytic motion because of genuine task-appropriate reasoning, limited tool knowledge, computational constraints, or the scoring incentives of the benchmark.
  • Each model is run once per scene, so the study does not quantify stochastic variability, best-of-NN performance, or the reliability of repeated attempts.
  • The comparison across proprietary and open-weight models is potentially confounded by differences in agent harnesses, tool interfaces, context windows, inference settings, and execution infrastructure.
  • The benchmark does not provide controlled ablations of input resolution, video length, frame rate, prompt wording, available libraries, GPU resources, or token budgets, making it difficult to identify the source of performance differences.
  • The reported relationship between token use and quality does not establish causal effects of computation; models may differ in tokenization, hidden reasoning, tool-call overhead, or the efficiency of their generated code.
  • The reasoning-effort ablation is performed only for one model family, so the benefit of additional inference-time reasoning may not generalize to other proprietary or open-weight models.
  • The evaluation focuses on final submissions and does not analyze intermediate code, renders, failed attempts, or render-and-compare iterations, leaving the effectiveness of visual feedback and iterative debugging unknown.
  • The paper does not measure how often agents identify and correct specific errors, when their improvements saturate, or whether longer interaction leads to better solutions per unit of cost.
  • Human evaluation covers 17 of 18 models and only a subset of comparisons is judged by multiple participants, leaving uncertainty about the robustness of human rankings, especially for closely matched systems.
  • Human and VLM judgments measure overall perceived similarity but do not independently validate each metric family or determine whether judges can reliably assess 3D geometry and physical mechanisms from monocular rendered videos.
  • The VQA questions are scene-specific and gated by the presence of selected objects or events; their coverage, difficulty, and sensitivity to partial or shortcut-based reconstructions are not fully established.
  • The benchmark does not test whether agents can explain their inferred dynamics, identify uncertainty, distinguish observable facts from assumptions, or recognize when the video is insufficient to determine a unique reconstruction.
  • Ambiguity and non-identifiability are not explicitly modeled: multiple geometries, materials, cameras, and physical programs may explain the same video, but the evaluation generally compares against one reference reconstruction.
  • The paper does not investigate whether agents can represent uncertainty or produce multiple plausible hypotheses when the visual evidence underdetermines scene structure or physical parameters.
  • The real and synthetic splits differ in appearance, ground-truth availability, and scene construction, so the reported performance gap cannot be cleanly attributed to realism, domain shift, or the absence of 3D supervision alone.
  • The effects of individual scene properties—such as object count, material count, occlusion, deformation magnitude, topology change, and interaction type—are not fully disentangled because many are correlated within the curated dataset.
  • The benchmark does not assess computational efficiency of the generated programs, including simulation time, memory usage, rendering cost, numerical stability, or scalability to longer videos and higher spatial resolution.
  • The security and reliability implications of executing model-generated graphics and simulation code are not examined, including unsafe resource consumption, nondeterministic behavior, dependency failures, or malicious code generation.
  • The extent to which benchmark-specific conventions, output formats, Blender APIs, and available documentation influence performance is not evaluated, limiting conclusions about general 4D inverse-graphics ability beyond this environment.

Practical Applications

Immediate Applications

  • Multimodal-agent evaluation and model selection — AI/software industry, academia. Use 4DCodeBench as a standardized regression suite for multimodal coding agents that generate Blender, Taichi, Warp, or other simulation code from video. Organizations can compare models not only on rendered appearance, but also on executability, static geometry, 2D motion, 3D motion, and human-aligned VLM scores. Potential workflow: run candidate models on the 200 scenes, track Overall and dynamics-specific scores, inspect failure categories, and select models for downstream graphics or robotics pipelines. Dependencies: the benchmark’s 200 scenes may not represent every target domain; scores should not be treated as evidence of safe deployment in unobserved environments.
  • Regression testing for graphics-code agents — software engineering and digital content creation. The executable-output requirement can be incorporated into continuous integration for systems that generate Blender or simulation scripts. A submission can be automatically checked for valid video output, correct frame rate and resolution, valid per-frame geometry, non-manifold meshes, self-intersections, and object interpenetration. Potential product: a “4D code compiler” or CI service that runs generated scene programs in isolated containers and returns structural and visual diagnostics. Dependencies: secure sandboxing is essential because generated code may be unsafe, computationally expensive, or capable of accessing unintended files.
  • Automated evaluation of dynamic-scene reconstruction tools — VFX, animation, and game development. Studios can use the benchmark’s perceptual, trajectory, depth, and geometry metrics to test video-to-3D and video-to-animation systems. The separate static-versus-dynamic scores are particularly useful for distinguishing systems that reproduce a convincing first frame from systems that reproduce deformation, collisions, flow, or fracture. Potential workflow: use a model for initial reconstruction, render the scene, compare it with source footage, and route low-scoring cases to an artist or simulation specialist. Dependencies: monocular real-world video is inherently ambiguous; a visually similar reconstruction may have incorrect physical parameters or hidden geometry.
  • Human-aligned quality assurance for generated reconstructions — media production and research. The reported correlation between VLM rankings and human preferences supports using VLM pairwise comparison as a first-pass triage tool. For example, a studio could compare several generated reconstructions and send only close or high-value cases to human reviewers. Dependencies: VLM agreement is weaker for closely matched models, and the reported agreement was established on this task distribution. Human review remains necessary for production-critical decisions.
  • Training and benchmarking curricula for physical-world reasoning — academia and education. Researchers can use the dataset’s taxonomy—rigid and articulated bodies, deformable solids, cloth and rods, fluids, grains, fracture, and multi-material interactions—to construct staged curricula for multimodal agents. Models can first learn static geometry, then externally driven motion, and finally passive multi-material dynamics. Potential research tool: a benchmark dashboard reporting separate appearance, geometry, contact, trajectory, and flow performance rather than a single aggregate score. Dependencies: the synthetic portion is generated under simulator assumptions, while the real portion lacks complete 4D ground truth; both are needed for balanced evaluation.
  • Diagnosis of failure modes in physical reasoning — AI research. The consistent gap between static reconstruction and dynamic reconstruction can be used to target model improvements. Researchers can evaluate whether a new method improves material-property inference, contact handling, topology change, or long-horizon motion rather than merely improving image similarity. Dependencies: current submissions frequently use analytic or prescribed motion—67% of analyzed solutions use analytic motion—so a high visual score does not necessarily imply transferable physical understanding.
  • Interactive video-to-scene prototyping for artists and engineers — design, visualization, and engineering. Today’s agents can already generate executable approximations for relatively simple rigid-body or driven-motion scenes. A user could provide a short video, obtain a Blender scene and animation script, and manually correct object geometry, camera, materials, or trajectories. Potential products: storyboard reconstruction, rapid animation blocking, educational demonstrations, and initial CAD or simulation setup. Dependencies: substantial human correction is likely for real scenes, especially those involving occlusion, transparent objects, fluids, cloth, fracture, or multiple interacting materials.
  • Robotics dataset inspection and simulation bootstrapping — robotics. Robot-manipulation videos in the real subset can be used to test whether an agent recovers object geometry, contacts, and externally driven motion. Approximate reconstructions may help researchers label demonstrations, visualize failed manipulation attempts, or initialize a simulator before manual refinement. Dependencies: the paper shows that driven scenes can be easier for 3D geometry but harder for some event-level judgments; reconstructed scenes should not directly control a robot without independent state estimation and safety validation.
  • Teaching and public demonstrations of physics — education and daily life. The executable representation enables learners to inspect how a scene is constructed, modify parameters, and rerun the animation. A classroom tool could compare an observed video with a student-written model of bouncing, deformation, fluid flow, or fracture. Dependencies: generated code may use kinematic shortcuts rather than correct physical laws, so educational interfaces should expose assumptions and distinguish “visual match” from “physically valid model.”
  • Benchmark-informed procurement and policy evaluation of AI systems — public-sector technology governance. Government or institutional buyers can require reporting of executability, dynamic-scene performance, open-weight versus proprietary model behavior, inference cost, and failure rates when procuring multimodal coding systems. This is more informative than relying only on static image or code-generation benchmarks. Dependencies: benchmark results should be combined with privacy, cybersecurity, licensing, reproducibility, and domain-specific safety assessments.

Long-Term Applications

  • Video-to-physics digital twins — manufacturing, engineering, and industrial inspection. A mature version of the approach could reconstruct machinery, deformable components, granular materials, or fluid interactions from ordinary video and produce a parameterized simulator. Engineers could then test interventions, estimate loads, or compare alternative operating conditions. Required advances: reliable inference of hidden geometry, material parameters, contact forces, external forces, and topology changes; multi-view or depth sensing would likely be needed for high-stakes use. The current results show that dynamic reconstruction remains substantially weaker than static geometry.
  • Robotic manipulation from visual demonstrations — robotics and automation. Future agents could convert demonstrations into executable 4D world models that explain object motion, contact, deformation, and material behavior. Robots might use these models to imitate tasks, predict consequences of grasps, or plan interventions on cloth, food, fluids, or soft objects. Required advances: intervention testing, uncertainty estimation, causal identification of forces, real-time inference, and sim-to-real validation. A reconstruction that merely reproduces an observed trajectory would not be sufficient for safe manipulation under new conditions.
  • Physics-aware augmented and virtual reality — consumer software, training, and telepresence. Reconstructed dynamic scenes could be inserted into AR/VR environments as editable 3D objects with physically plausible motion. Users might pause, rotate, replay, or alter an event, such as a mechanical failure, sports action, or household demonstration. Dependencies: accurate camera calibration, low-latency reconstruction, view-consistent geometry, robust occlusion handling, and a physically meaningful model rather than frame-by-frame animation.
  • Forensic and scientific reconstruction of physical events — law enforcement, safety engineering, and science. Video-to-code reconstruction could assist with replaying collisions, material failures, laboratory phenomena, or accidents. Explicit scene programs would make assumptions inspectable and allow investigators to test alternative hypotheses. Dependencies: evidentiary use requires calibrated cameras, multiple viewpoints or independent measurements, provenance tracking, uncertainty quantification, and validation against physical constraints. The benchmark does not establish forensic reliability.
  • Predictive monitoring of deformable and flowing processes — healthcare, energy, and industrial operations. With further development, the methods could model tissue motion, fluid transport, soft materials, powders, or fracture propagation from video. Applications might include noninvasive monitoring, manufacturing quality control, pipeline inspection, battery-material analysis, or flow diagnostics. Dependencies: domain-specific sensors and constitutive models, labeled 4D data, strict validation, and compliance requirements. The current dataset excludes humans and animals and therefore does not directly support clinical deployment.
  • Generative design and inverse simulation — engineering and product development. An agent could infer a compact executable scene, modify its geometry or physical parameters, and optimize the design for desired behavior. Examples include soft grippers, packaging materials, protective structures, fluid devices, and articulated mechanisms. Required advances: differentiable or efficiently optimizable simulators, reliable material identification, counterfactual evaluation, and guarantees that the reconstructed model generalizes beyond the recorded motion.
  • Long-horizon physical prediction and intervention planning — autonomous systems. The benchmark could evolve from reconstruction toward asking whether an inferred world program predicts unseen frames, responds correctly to changed initial conditions, and handles external forces. This would provide a stronger foundation for autonomous navigation, manipulation, and planning. Dependencies: the paper explicitly identifies intervention tests and longer-horizon prediction as missing evaluations. Systems would need causal physical representations rather than prescribed trajectories that only reproduce the observed clip.
  • Large-scale synthetic training environments for embodied AI — robotics and AI research. Reconstructed real scenes could become editable simulation assets for training agents. A model might generate variations in object geometry, material properties, initial conditions, or applied forces, producing richer training data than the original video alone. Dependencies: reconstruction errors can propagate into synthetic data and create misleading training distributions. Generated environments would require uncertainty-aware filtering, simulator validation, and real-world performance checks.
  • Standardized physical-intelligence certification — academia, industry, and regulation. An expanded benchmark could support certification of multimodal agents on physically grounded capabilities, with separate thresholds for static geometry, dynamics, contact, material changes, executability, computational cost, and robustness to interventions. Dependencies: broader datasets, adversarial and out-of-distribution tests, transparent evaluation code, calibrated uncertainty, and agreed definitions of acceptable physical fidelity. The current Overall score is useful for comparison but should not alone serve as a safety or compliance standard.
  • Everyday personal scene editing and explanation — consumer applications. In the longer term, users could capture a household event with a phone and obtain an editable 4D model: replaying how an object fell, testing alternative arrangements, or generating an explanation of a physical interaction. Dependencies: privacy-preserving on-device processing, robust handling of occlusion and lighting, accurate scale and depth recovery, and clear communication that the result is an inferred hypothesis rather than a definitive record of the hidden physical event.

Glossary

  • Affine Body Dynamics (ABD): A simulation method that models bodies using affine transformations, extending rigid-body behavior to deformable or nearly rigid objects. “Affine Body Dynamics (ABD) extends related ideas to stiff and near-rigid bodies.”
  • Analysis-by-synthesis: An inference strategy that explains observations by generating and comparing candidate models. “Inverse graphics frames visual perception as inverting the forward rendering process through analysis-by-synthesis”
  • Anisotropic damage: Material damage that varies according to direction within the material. “AnisoMPM & Anisotropic damage and fracture; Drucker--Prager sand”
  • Bradley--Terry model: A statistical model for estimating rankings from pairwise comparisons. “Pairwise preferences are aggregated across scenes into a model-level Elo rating using a Bradley--Terry model”
  • Chamfer distance: A distance measure between geometric point sets based on nearest-neighbor distances. “3D Geometry measures static surface accuracy at frame~0 via Chamfer distance”
  • Co-dimensional structure: A lower-dimensional object embedded in a higher-dimensional space, such as a cloth sheet or rod in three-dimensional space. “co-dimensional structures such as cloth and rope”
  • Constitutive model: A mathematical model describing how a material responds to forces or deformation. “We extend the Genesis simulator with three new MPM constitutive models”
  • Continuum-damage fracture: A fracture model that represents distributed material damage within a continuous medium. “CD-MPM & Continuum-damage fracture”
  • Cosine similarity: A similarity measure based on the angle between two vectors. “we map the mean frame-wise cosine similarity to [0,1][0,1]”
  • Drucker--Prager model: A constitutive model commonly used to represent yielding and failure in granular materials such as soil and sand. “a Drucker--Prager model for sand”
  • Dynamic IoU: An intersection-over-union metric for measuring agreement in the spatial coverage of moving objects. “2D Dynamics measures the coverage of moving objects (Dynamic IoU).”
  • Earth Mover’s Distance (EMD): A distance between distributions measuring the minimum cost of transforming one distribution into another. “Eulerian perspective comparing frame-to-frame displacement distributions without correspondence (EMD step).”
  • Elo rating: A ranking system that estimates relative performance from pairwise outcomes. “Model pairs are sampled adaptively to shrink the widest confidence intervals, and judgments are aggregated into Elo ratings”
  • Eulerian perspective: An analysis of motion that observes changes at fixed spatial locations rather than following individual material points. “an Eulerian perspective comparing frame-to-frame displacement distributions without correspondence”
  • Executable graphics program: A program that constructs, animates, and renders a graphical scene when run. “Agents reconstruct dynamic scenes from video as executable graphics programs.”
  • Forward rendering process: The process of generating visual observations from a scene representation. “Inverse graphics frames visual perception as inverting the forward rendering process”
  • Hyperelastic material: A material whose stress–strain behavior is derived from a strain-energy function and can undergo large elastic deformations. “deformable solids, including hyperelastic, viscoelastic, and elastoplastic materials”
  • Inference-time reasoning: Computation performed during model inference to improve problem-solving or generation quality. “Effect of Inference-Time Reasoning”
  • Inverse graphics: The task of recovering a scene’s geometry, materials, lighting, and motion from visual observations. “We introduce 4D inverse graphics through code generation.”
  • Interpenetration: An invalid geometric condition in which two simulated objects occupy overlapping physical space. “unintended interpenetration, and other visible simulation artifacts”
  • Lagrangian perspective: An analysis of motion that follows persistent material points or particles through space and time. “a Lagrangian perspective tracking matter along persistent 3D paths”
  • Material Point Method (MPM): A hybrid particle-grid simulation method for modeling deformable and fluid-like materials. “The Material Point Method (MPM) combines Lagrangian material points with an Eulerian background grid”
  • Mesh rasterization: The conversion of geometric surfaces into image pixels for rendering or comparison. “Surface rasterization fails to capture background details”
  • Monocular view: A visual observation obtained from a single camera viewpoint. “real videos provide only monocular views”
  • Non-manifold edge: A mesh edge whose local neighborhood does not form a valid manifold surface structure. “open boundaries, non-manifold edges, degenerate faces”
  • Optical flow: The apparent two-dimensional motion of image brightness patterns between video frames. “On real videos, we additionally estimate optical flow”
  • Pareto frontier: The set of solutions that cannot improve one objective without worsening another. “with the Pareto frontier”
  • Perceptual metric: A measure of visual or semantic similarity intended to approximate human judgments. “The Overall score averages these five families; failed runs receive the worst score on the affected metrics”
  • Procedurally generated geometry: Geometry created algorithmically according to specified rules or parameters. “combining procedurally generated geometry with publicly available meshes”
  • Rigid body: An idealized object whose shape does not deform during motion. “rigid and articulated bodies”
  • Scene ontology: A structured classification of the entities, materials, and relationships represented in scenes. “We organize scenes hierarchically by the matter of their main dynamic objects.”
  • Smoothed Particle Hydrodynamics (SPH): A mesh-free particle-based method for simulating fluids. “Smoothed Particle Hydrodynamics (SPH) represents fluids using Lagrangian particles”
  • Spearman correlation: A rank-based statistic measuring the monotonic relationship between two variables. “human and VLM Elo correlate very strongly across models (Spearman ρ=0.980\rho=0.980”
  • Surface tension: The interfacial force that causes a liquid surface to resist deformation. “including fluid--rigid interactions and surface-tension effects”
  • Temporal correspondence: The association of the same object or material elements across different time steps. “temporal correspondence for dynamic matter”
  • Topology: The structural connectivity of a geometric object, including properties such as holes and connected components. “Flow, fracture, and deformation involve changes in shape and often topology”
  • Trajectory dynamic time warping (Trajectory DTW): A sequence-alignment measure that compares trajectories while allowing differences in temporal speed or alignment. “Trajectory DTW”
  • Vicoelastic material: A material exhibiting both viscous, rate-dependent behavior and elastic, recoverable deformation. “hyperelastic, viscoelastic, and elastoplastic materials”
  • Viscoplastic material: A material that undergoes time-dependent deformation and permanent plastic flow under stress. “a viscoplastic model to simulate materials such as shaving cream and toothpaste”
  • Visual question answering (VQA): A task in which a system answers natural-language questions about visual content. “VQA asks scene-specific questions in four categories”
  • Vision-LLM (VLM): A model trained to jointly process visual inputs and natural-language information. “We further evaluate reconstructions with a \ac{vlm} as a judge.”
  • Watertightness: The property of a three-dimensional mesh having a closed surface without gaps or open boundaries. “simple but inaccurate geometry can score well on watertightness and related measures”

Tweets

Sign up for free to view the 6 tweets with 648 likes about this paper.