4DCodeBench Benchmark: Benchmarking 4D Inverse Graphics
- 4DCodeBench is a benchmark for 4D inverse graphics through code generation, where agents must reconstruct dynamic scenes from videos by inferring scene attributes and producing executable code.
- The benchmark includes 200 scenes, combining real-world and synthetic data, and evaluates reconstructions on perceptual, dynamic, geometric, and structural integrity metrics.
- The 3D geometry of reconstructed scenes can be compared against the prototyping videos. Synthetic scenes, further, can be evaluated in terms of dynamics or motion.
4DCodeBench is a benchmark for 4D inverse graphics through code generation, in which multimodal coding agents reconstruct dynamic scenes from video as executable graphics programs. Given an RGB reference video, an agent must infer scene geometry, camera, appearance, material behavior, and temporal evolution, then produce code that regenerates the scene without accessing the reference video during reconstruction. The benchmark combines real-world footage and synthetic physics-simulated scenes, evaluates both rendered appearance and explicit 3D dynamics, and tests whether static reconstruction ability extends to deformation, fluid flow, fracture, contact, and other complex physical phenomena (Shen et al., 2 Oct 2026).
1. Task definition and research scope
4DCodeBench formulates dynamic-scene reconstruction as an inverse-graphics problem with an open representation. Let the reference video be , where is frame . An agent produces executable graphics code
The generated code constructs a geometric scene state and renders it through :
where is the reference frame rate and is the reconstructed frame. The output must match the reference video’s spatial resolution, frame count, and frame rate.
The code must specify or generate the scene’s geometry, camera, materials, lighting, and temporal evolution. No programming-language structure or simulation method is prescribed. Agents may use Blender animation and physics, analytic trajectories, custom numerical simulation, Taichi, Warp, NumPy, SciPy, PyTorch, Trimesh, or other software available in the execution container.
The benchmark is intended to evaluate more than object recognition, physics-question answering, plausible video generation, or frame-wise image imitation. Its target capability is the construction of an explicit and inspectable 4D world program. The resulting scene must expose geometric and dynamic representations that can be evaluated independently of the rendered video.
A central distinction is between static visual reconstruction and dynamic physical reconstruction. A model may produce a plausible first frame while failing to reproduce subsequent trajectories, collisions, material deformation, flow, fracture, or final configuration. The benchmark therefore treats appearance, geometry, motion, and physical-event reconstruction as related but separable capabilities.
2. Executable world representation
Each agent receives only the RGB reference video, a fixed task prompt, and output-format documentation. The prompt does not provide scene names, semantic descriptions, object lists, material labels, reference geometry, or synthetic-scene metadata.
The agent must create two directories:
6
The first contains an executable build.sh and associated source code. The second contains the reconstructed world, which may include:
7
camera.json stores camera intrinsics and extrinsics. The camera is a fixed pinhole camera with intrinsic matrix 0 and camera-to-world transform 1. Positions use world-space meters with 2 upward.
Per-frame triangle geometry is stored under meshes/. Persistent material-point positions are stored under dynamics/. For a dynamic object or particle system, the pos array has shape 3, where 4 is the number of frames, 5 is the number of tracked matter points, and the last dimension contains three-dimensional coordinates. Column 6 must represent the same physical piece of matter over time.
This correspondence requirement distinguishes a dynamic world representation from independently rebuilding a visually similar object in every frame. For liquids and other Eulerian systems without natural persistent identities, solver data may be placed under solver/, with trajectories extracted during evaluation.
The final program must regenerate the scene from scratch without reading the input video. The video is available during development, allowing the intended iterative loop of inspecting the video, writing code, rendering, comparing, and revising. The benchmark scores the final submission rather than the full development trajectory. Using the input video as a texture or copying its pixels into the scene is prohibited.
A run is executable when its rendered video decodes successfully and matches the reference resolution, frame count, and frame rate, and when its 4D world passes structural validation. A valid video with an invalid world may receive image-based scores but fails geometry-dependent metrics.
3. Dataset composition and scene ontology
The benchmark contains 200 scenes: 100 real-world video clips and 100 synthetic scenes.
Real-world scenes
The real split contains 100 clips:
- 51 from physics-video benchmarks;
- 21 from robot-manipulation datasets;
- 28 from web-sourced footage.
The sources include WISA-80K, Physics-IQ, Phys-AD, Phys101, ABC-130K, T-Rex, Robo360, AgiBot World, RoboCook, RoboCraft, ALOHA Unleashed, ManipArena, Task-Level ILC, Pexels, YouTube Creative Commons, Mixkit, and Pixabay.
Curators selected clips with diverse physical behavior, temporal continuity, minimal editing or cuts, preferably stationary cameras, limited occlusion, and visually clear physical events. Humans and animals were excluded, and visible hands were minimized because detailed human geometry and articulation were outside the benchmark’s primary focus.
Real videos provide realistic appearance and unconstrained physical phenomena but generally lack complete 3D geometry, material state, and motion ground truth. Their evaluation therefore relies more heavily on the reference video, estimated depth, optical flow, tracks, and human or VLM judgment.
Synthetic scenes
The synthetic split contains 100 scenes generated with physics simulation and rendered in Blender. Geometry includes procedural objects and public meshes such as the Stanford bunny and armadillo. Scenes vary in geometry, material type, physical parameters, initial conditions, rendering settings, number of objects, interaction structure, and dynamics regime.
The authors manually inspect synthetic simulations for numerical instability, unintended interpenetration, visible artifacts, and implausible behavior. Synthetic scenes provide ground-truth geometry and motion at every frame, enabling direct evaluation in 3D and 4D.
The simulators include Blender, libuipc, SPlisHSPlasH, MLS-MPM, CK-MPM, MPM-lite, HOT, AnisoMPM, CD-MPM, Silly Rubber, IQ-MPM, PPF, and Genesis. Their coverage includes rigid and affine-body systems, deformable solids, elastoplastic materials, cloth and rods, fluids, granular media, fracture, cutting and discontinuities, solid–fluid coupling, and robotic manipulation. Genesis is extended with constitutive models for viscoplastic materials such as shaving cream and toothpaste, snow, and Drucker–Prager sand.
Matter families and interactions
The ontology is organized around four primary matter families:
- Rigid and articulated bodies: rigid objects, articulated mechanisms, robot arms, and interacting rigid bodies.
- Deformable solids: hyperelastic, viscoelastic, elastoplastic, and damageable solids undergoing stretching, compression, bending, plastic deformation, tearing, or fracture.
- Co-dimensional structures: cloth, shells, rods, ropes, and strands embedded in three-dimensional space.
- Flowing materials: granular materials, Newtonian fluids, viscoplastic fluids, sand, snow, and particle-based liquids.
Scenes can combine multiple matter types. Eighty-three percent of scenes contain multiple dynamic objects, and 66% contain multiple material types. Interactions include rigid–deformable, rigid–fluid, fluid–deformable, deformable–co-dimensional, and rigid–co-dimensional coupling.
Motion is classified as either passive dynamics, in which motion follows from initial conditions and physical interactions, or externally driven dynamics, such as robot manipulation or prescribed rigid-object motion. Videos are capped at 60 frames per second, and real videos at 300 frames; higher-frame-rate source videos are subsampled.
4. Evaluation protocol and metrics
Each of the 18 evaluated models is run once on all 200 scenes, producing 3,600 runs. Runs execute in isolated containers with one GPU, Ubuntu 24.04, CUDA 13.0, Blender 5.2.0, FFmpeg, ImageMagick, NumPy, SciPy, PyTorch, Taichi, Warp, Trimesh, and Pillow. Offline Blender API documentation and an offline public format checker are provided.
There is no artificial step limit. Runs that stall or exceed a six-hour timeout are terminated. Proprietary systems use native agent CLIs, while open-weight models run in the Stirrup agent harness and are served using SGLang with recommended configurations.
Overall performance is the mean of five metric families:
7
The family scores are in 8, with higher values better.
Perceptual and image-space metrics
Perceptual similarity uses DINOv3 frame embeddings. Dynamic IoU compares reference and reconstructed dynamic-object masks while ignoring a thin boundary band to reduce rasterization artifacts. It measures the spatial coverage of moving objects and coarse motion trends.
For real videos, 2.5D geometry is evaluated with depth error. Reference depth is estimated using Video Depth Anything, while reconstruction depth is rasterized from the predicted 3D world. Normalized disparities are compared after removing the largest 1% of residuals per frame.
Real-video motion is evaluated with optical flow and 2D tracks. Optical flow is compared as unordered vector sets using sliced Wasserstein distance, assessing apparent-velocity distributions while reducing sensitivity to small spatial misalignments. CoTracker3 tracks reference points through the video, and corresponding reconstructed 3D points are projected through time. Dynamic-time-warping distances between paths are capped and normalized by the image diagonal.
Synthetic geometry and dynamics metrics
Synthetic reconstructions are aligned to reference scenes using trimmed ICP and similarity transforms fitted from visible frame-0 surfaces and dynamic-object trajectories. The transform that better aligns the reconstruction with the reference camera is retained.
Static 3D geometry uses a capped symmetric Chamfer distance between registered reference and reconstructed surface point clouds, normalized by the reference point cloud’s RMS radius.
Lagrangian 3D dynamics uses trajectory DTW. Persistent matter points are sampled from the reference and reconstruction, and a Hungarian assignment pairs reconstructed and reference trajectories. The metric evaluates whether the correct pieces of matter follow the correct 3D paths.
Eulerian 3D dynamics uses EMD step, which compares distributions of one-step 3D displacements through sliced Wasserstein distance. It does not require persistent correspondence and is therefore relevant to fluids and changing-topology systems.
VQA, pairwise judgment, and structural checks
The benchmark includes 1,251 hand-written binary questions:
- 371 initial-state questions;
- 354 final-state questions;
- 308 key-event questions;
- 218 contact questions.
A VLM receives 16 reconstruction frames and answers questions about object presence, location, temporal events, event order, contact behavior, and final outcomes. Gating prevents blank or objectless reconstructions from receiving favorable scores.
Pairwise VLM evaluation compares two reconstructions against the reference with respect to geometry, motion, collisions, deformation, and final state. Outcomes are aggregated using a Bradley–Terry model and reported as Elo ratings. Model-level VLM Elo agrees strongly with human Elo, with Spearman 9. On 916 comparisons judged by both humans and the VLM, agreement is 89.3%, with Cohen’s 0.
Mesh-quality diagnostics separately examine watertightness, manifoldness, degenerate faces, self-intersections, and object interpenetration. These diagnostics are not included in Overall because structural mesh integrity is not equivalent to reconstruction fidelity.
5. Results and model behavior
The benchmark evaluates 18 multimodal coding models, including GPT-6 Astra at Low, High, and Max reasoning, Claude Fable 5.1, Claude Opus 5, Claude Opus 5.5, GPT-5.6 Sol, GPT-5.6 Luna, GPT-5.6 Terra, Gemini 3.8 Flash, Qwen3.8 Flash-Next, DeepSeek v4.1 Flash, GLM 5.3 Flash, MiniMax M3, MiMo v2.5, Muse Glimmer, Gemma-4 31B, and Mistral Medium 3.5. Kimi K3 was evaluated but omitted from reported results because it used roughly 50% more cost than Fable without a commensurate quality improvement.
The leading models were GPT-6 Astra [Max], Claude Opus 5.5 [High], GPT-6 Astra [High], Claude Fable 5.1 [High], and GPT-6 Astra [Low]. Astra [Max] achieved an Overall score of 0.791 and VQA accuracy of 87.6%. Astra [High] and Astra [Low] achieved Overall scores of approximately 0.77 and 0.73. Fable [High] achieved VQA accuracy of 78.8%, whereas GLM [Max] achieved 29.5%.
The principal finding is a static–dynamic gap. Across models, perceptual and 3D Geometry scores are higher than 2D Dynamics and 3D Dynamics scores. For Astra [Max], static families averaged approximately 0.91, while dynamic families averaged approximately 0.67. Initial-state questions were generally easier than final-state and key-event questions.
Real scenes are more difficult than synthetic scenes. Relative to synthetic scenes, real scenes show standardized decreases of approximately 0.73 in VQA and 0.88 in Perceptual score. Realistic appearance, ambiguous depth, occlusion, lighting, and irregular geometry are plausible contributors, although the benchmark does not separately identify their causal effects. Astra [Max] generalized more successfully to real footage, while Opus 5.5 was nearly competitive on synthetic scenes but fell behind on real scenes.
Scene categories expose different weaknesses. Multi-material scenes have lower VQA and Perceptual scores, co-dimensional scenes have lower 2D Dynamics and 2.5D Geometry, driven scenes have lower VQA but higher 3D Geometry, and flowing scenes tend toward lower 3D Dynamics. These results implicate thin-structure geometry, material recognition, contact modeling, correspondence, and dynamic prediction as distinct sources of difficulty.
The reconstruction strategies used across solutions were:
- 67% analytic motion;
- 19% custom simulation;
- 10% Blender physics;
- 3% keyframing.
Analytic strategies included spline interpolation, prescribed vertex-displacement maps, ballistic formulas, curve-based placement, hand-written pose tables, and frame-wise pile reconstruction. Custom simulations included MPM, MLS-MPM, PBD, XPBD, DEM, PBF, mass-spring, FEM, APIC, Verlet, and custom GPU kernels.
Increasing inference-time reasoning for Astra improved Overall, 2D Dynamics, 3D Dynamics, and VQA monotonically. Astra [Low], [High], and [Max] achieved Overall scores of 0.73, 0.77, and 0.79, respectively. However, step count and token volume had negligible correlation with Overall quality, with 1. Cost had a moderate correlation of 2, but raw interaction volume did not reliably produce better reconstructions.
6. Failure modes, significance, and limitations
A recurrent failure mode is correct first-frame reconstruction followed by incorrect evolution. Agents often reproduce a plausible initial layout but fail on final configuration, event timing, collision outcomes, deformation magnitude, or object trajectories.
Many systems prioritize recognizable appearance over physical behavior. They may produce correct object categories and textures while omitting contact, using implausible motion scripts, or ignoring stiffness, elasticity, plasticity, viscosity, friction, cohesion, fracture thresholds, or fluid and granular behavior.
Analytic motion is frequently used where simulation would be necessary to reproduce interactions, contact response, deformation under changing forces, branching events, or topology changes. Nevertheless, analytic trajectories can score well because the benchmark measures reconstruction fidelity rather than causal physical correctness. Consequently, high performance does not by itself establish that an agent inferred the true physical mechanism.
Flow and fracture are particularly difficult because they involve topology changes, disappearing and appearing surfaces, ambiguous particle identity, and many interacting elements. Thin structures such as cloth, rods, ropes, and shells are also difficult because of small thickness, self-contact, bending, twisting, and long-range deformation.
Trajectory DTW reveals errors that may be hidden by rendered-frame similarity. A reconstruction can match per-frame velocity distributions while assigning motion to the wrong matter, or can produce visually plausible fluids without preserving meaningful particle paths. The trajectory metric has scene-level correlations around 0.39–0.50 with other metrics, indicating that it measures a distinct aspect of reconstruction.
The benchmark’s mesh-quality analysis shows that structural integrity is not a proxy for fidelity. Strong models can produce more defects when attempting intricate fractured, thin, or fluid geometry, while weaker models may generate clean but inaccurate primitives. Overall executability is 90.1%; proprietary models achieve at least 96.5%, while open-weight models range from 29.0% to 95.5%. Only 5.4% of runs passing the video gate fail later 4D-world validation, suggesting that formatting is not the dominant bottleneck.
The human study includes 3,587 pairwise judgments by 76 participants across 17 models and 200 scenes. Human Elo correlates with VLM Elo at 3 and with Overall score at 4. This supports the use of automated pairwise judging for benchmark expansion, although individual VLM judgments are less reliable for closely matched models.
Full evaluation requires 3,600 runs, 5,983 hours of compute, and approximately $k$52.01 per task on average, compared with $6.65 for proprietary models. Qwen uses an average of 752 steps and 1.5 million output tokens but ranks only eleventh, further indicating that raw agent activity is not a sufficient determinant of quality.
Several limitations constrain interpretation. Fidelity is evaluated more directly than physical understanding, because prescribed trajectories can perform well without encoding causal mechanisms. There is no intervention testing, no longer-horizon prediction evaluation, and no systematic measurement of how agents improve through render-and-compare iterations. Real scenes lack complete 4D ground truth and rely on estimated depth, optical flow, tracking, VLM judgment, and human preference. Synthetic scenes provide exact ground truth but inherit the assumptions of their constitutive models, discretizations, parameters, and numerical solvers. Monocular video also leaves geometry, depth, and material properties underdetermined.
4DCodeBench’s central contribution is the evaluation of executable, inspectable world representations rather than rendered videos alone. Its results indicate that current multimodal coding agents can often reconstruct static appearance and initial geometry, but remain substantially less reliable at modeling dynamic physical behavior. Deformation, fluid flow, fracture, contact, long-range trajectories, and material interactions continue to expose a pronounced gap between recognizing what a scene looks like and implementing how it evolves.