4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
Abstract: We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstractions such as physical simulations to reproduce complex behavior. To evaluate this capability, we curate a set of real-world videos and construct synthetic scenes spanning diverse physical phenomena, including deformation, fluid flow, and fracture. We perform extensive benchmarking of frontier models, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics. 4DCodeBench provides a testbed for tracking progress toward agents that can interpret the dynamics of the world through code. Our benchmark is available at https://github.com/4DCodeBench/4DCodeBench
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
The paper introduces 4DCodeBench, a test for AI systems that can look at a video and write computer code to recreate what happens in it.
The word 4D means:
- 3D space: where objects are and what they look like
- Time: how those objects move and change
For example, an AI might watch a video of a ball hitting a soft object, water flowing, cloth moving, or something breaking. It must then write a program that rebuilds the scene and shows the same events happening over time.
This is called inverse graphics. Normally, computer graphics code creates an image or video from a description. In inverse graphics, the process is reversed: the computer watches an image or video and tries to figure out the hidden 3D objects, materials, and movements that created it.
2. What questions are the researchers asking?
The researchers mainly want to know:
- Can AI coding agents understand moving scenes from video?
- Can they recreate not only what objects look like, but also how they move and change?
- Are AI systems better at copying appearance than understanding motion and physics?
- Which kinds of scenes are hardest for AI?
- Do AI systems use real physics simulations, simple mathematical movements, or other tricks?
- Do automatic computer scores agree with what human viewers think looks most accurate?
The important challenge is that the AI is not given a list of objects or a written explanation. It only receives the video and must work out what is happening by itself.
3. How did the researchers conduct the study?
Building the benchmark
The researchers created a collection of 200 scenes:
- 100 real-world videos, taken from physics videos, robot experiments, and online footage
- 100 computer-generated scenes, made using physics simulators
The synthetic scenes are useful because the researchers know the exact 3D shape and movement of every object. This is like testing a student with an answer key: they can compare the AI’s reconstruction with the true answer.
The scenes include many types of materials and actions:
- Hard objects moving or colliding
- Robots moving objects
- Cloth and ropes bending
- Soft materials stretching or squashing
- Fluids flowing
- Sand and grains moving
- Objects breaking or cracking
Asking AI models to write graphics programs
The researchers tested 18 multimodal coding models. These are AI systems that can understand images or videos and also write code.
Each model received:
- One reference video
- A task description
- A computer environment containing tools such as Blender
The model had to produce an executable program. When the program ran, it needed to:
- Build the objects in 3D
- Describe how the objects change over time
- Create a new video
- Provide the objects’ 3D positions and shapes at different moments
The program was not allowed to look at the original video while running. This prevented the model from simply copying individual frames.
Different ways of creating motion
The AI models could choose how to reproduce movement. For example, they could use:
- Analytic motion: using formulas, such as moving an object along a curved path
- Keyframing: manually specifying where an object should be at important times
- Physics simulation: using rules about forces, collisions, liquids, and materials
- Custom simulation code: writing their own program to imitate physical behavior
A useful analogy is drawing a bouncing ball. One method is to tell the ball exactly where to be in every frame. Another is to use physics rules so that gravity and collisions make the ball bounce naturally.
Measuring the results
The researchers used several kinds of measurements:
- Appearance: Does the recreated video look like the original?
- 2D motion: Do moving objects travel across the screen in the same way?
- 2.5D geometry: Does the estimated depth, or distance from the camera, look correct?
- 3D geometry: Are the reconstructed objects shaped correctly?
- 3D motion: Do the objects move through 3D space correctly?
They also used:
- A vision-LLM as an automatic judge
- Questions about the scene, such as what happened at the beginning or end
- Human comparisons in which people chose which of two reconstructions was better
In total, 76 people made 3,587 comparisons.
4. What were the main findings?
AI is better at copying appearance than motion
The clearest result is that AI models can often recreate how a scene looks, but they have much more trouble recreating how it changes.
The best-performing model, called Astra [Max] in the paper, achieved:
- About 0.91 on static aspects, such as appearance and geometry
- About 0.67 on dynamic aspects, such as motion
This means the model might correctly build something that looks like the original object, but its movement, bending, flowing, or breaking may be wrong.
For example, an AI might create a good-looking piece of cloth but fail to make it fold in the correct way. It might recreate a cup and a table but make the cup fall too slowly or bounce incorrectly.
The strongest model performed best overall
The overall ranking placed Astra [Max] first. Other strong models included Claude Opus 5.5, Astra [High], Fable, and Astra [Low].
Proprietary models generally performed better than the open-weight models tested. However, the open models showed large differences between one another.
About 90.1% of all attempts successfully produced executable results. Proprietary models had especially high success rates, while some open models often produced programs that did not run correctly or did not meet the required format.
More reasoning helped within the same model
When the researchers gave Astra more time and effort to reason, its performance improved:
| Reasoning level | Overall score | 2D motion score | 3D motion score |
|---|---|---|---|
| Low | 0.73 | 0.48 | 0.63 |
| High | 0.77 | 0.55 | 0.69 |
| Max | 0.79 | 0.60 | 0.73 |
Its ability to answer questions about the reconstructed scenes also improved, from 78.2% at the Low setting to 87.6% at the Max setting.
However, simply using more computer tokens did not always make one model better than another. This suggests that how an AI reasons may matter more than how much it writes.
AI models used different strategies
Most solutions used simple mathematical descriptions of movement:
- 67% used analytic motion
- 19% used custom physics simulations
- 10% used Blender’s physics tools
- 3% used keyframing
The models differed greatly in their choices. Some mostly guessed paths directly, while others tried to build physical simulations.
This is important because it shows that AI systems may reach similar-looking results using very different ideas about what is happening.
Real and complicated scenes were harder
Real-world videos were generally harder than synthetic scenes. Real videos contain details and physical effects that are difficult to measure exactly.
Scenes with several different kinds of materials were also harder. For example, a scene involving a hard object, cloth, and water is more difficult than a scene involving only one solid object.
Other difficult cases included:
- Cloth and rope
- Flowing materials
- Multiple objects touching each other
- Scenes with complicated interactions
- Events involving deformation, breaking, or changing shape
Automatic scores agreed well with human opinions
The automatic vision-language judge ranked the models very similarly to human viewers.
The researchers found a very strong agreement between human and automatic rankings. On shared comparisons, humans and the AI judge agreed about 89.3% of the time.
This suggests that the benchmark’s automatic tests are useful for comparing models without requiring people to watch every single reconstruction.
5. Why are these results important?
This research shows that understanding a moving world is much harder than recognizing objects or copying a single image.
To recreate a dynamic scene, an AI must work out things such as:
- Which objects exist
- What materials they are made from
- Which objects are touching
- What forces are acting
- How an object will bend, stretch, flow, or break
- What caused the movement
- How the scene will look in the future
A video does not directly reveal all of this information. For example, a ball moving across a screen could be rolling, sliding, being pushed, or following a preplanned path. The AI has to choose a reasonable explanation.
Conclusion: What could this research lead to?
4DCodeBench gives researchers a standard way to measure whether AI systems truly understand dynamic scenes rather than merely producing attractive pictures.
In the future, better systems could help with:
- Robots learning how to handle objects
- Computer animation
- Video editing and special effects
- Virtual reality and game worlds
- Scientific simulations
- Predicting how materials and objects will behave
However, the paper also shows that current AI systems still struggle with true physical understanding. They can often make something that looks right, but they may not understand why it moves that way.
The researchers suggest that future tests should ask AI systems to change the starting conditions or apply new forces. For example, instead of only recreating a ball’s original bounce, the AI could be asked what would happen if the ball were dropped from a different height. Such tests would show whether the AI learned the underlying physics or simply memorized one particular movement.
Overall, the paper presents 4DCodeBench as a useful measuring tool for progress toward AI systems that can understand, simulate, and interact with the changing physical world.
Knowledge Gaps
Knowledge Gaps, Limitations, and Open Questions
- The benchmark evaluates reconstruction fidelity but does not determine whether an agent has recovered the underlying physical mechanism; prescribed trajectories can obtain high scores without learning transferable dynamics.
- It remains unknown whether physics-based simulation produces better generalization than analytic motion, keyframing, or other kinematic approximations when scenes, initial conditions, or viewpoints change.
- The study does not test interventions such as modified initial conditions, external forces, object properties, or contact configurations, so causal physical understanding is not measured.
- Longer-horizon prediction beyond the observed video is not evaluated, leaving open whether reconstructed programs can forecast future dynamics rather than merely reproduce the input sequence.
- The real-world subset lacks complete 3D and 4D ground truth, limiting the ability to distinguish correct reconstruction from visually plausible but geometrically or physically incorrect explanations.
- The synthetic subset depends on simulator assumptions and parameterizations, which may not reflect real material behavior, contact dynamics, topology changes, sensor noise, or rendering artifacts.
- The extent to which models overfit to the particular simulators, object meshes, material categories, and procedural distributions used to create the synthetic scenes is not tested.
- The dataset contains only 200 scenes, and the paper does not establish how performance scales with substantially larger, more varied, or more systematically balanced datasets.
- Scene selection favors stationary cameras, limited occlusion, continuous footage, and the absence of humans and animals; consequently, performance under camera motion, severe occlusion, cuts, clutter, articulated humans, and animal motion remains unresolved.
- The benchmark largely excludes challenging real-world factors such as lighting changes, reflections, transparency, motion blur, sensor noise, imperfect segmentation, and dynamic backgrounds.
- The impact of monocular ambiguity is not isolated from other sources of difficulty; the study does not compare monocular input with stereo, multi-view, depth, optical-flow, or other additional observations.
- Camera and illumination parameters are left for agents to reconstruct, but the paper does not separately quantify errors in camera estimation, lighting estimation, and material appearance versus errors in object geometry and dynamics.
- The benchmark does not evaluate whether reconstructed programs remain valid under novel viewpoints, camera trajectories, lighting conditions, or rendering engines.
- The evaluation is primarily based on a single observed trajectory, so it cannot establish whether the inferred scene representation supports counterfactual rendering or manipulation.
- The overall score averages heterogeneous perceptual, geometric, and dynamic metrics, but the paper does not justify the equal weighting or analyze how alternative weightings change model rankings.
- Several evaluation signals rely on off-the-shelf vision models for depth, optical flow, point tracking, and DINO features; their biases and failure modes may systematically affect scores, especially for fluids, transparent objects, deformable materials, and large motion.
- The use of analytically exposed geometric states for some metrics may favor submissions that provide convenient internal representations, even when those representations do not correspond to visually observable or physically meaningful matter.
- The alignment procedure used before synthetic 3D evaluation may remove meaningful reconstruction errors, and its sensitivity to alignment choices is not investigated.
- The trajectory and displacement metrics may not adequately evaluate topology-changing phenomena such as fracture, splitting, merging, cutting, or fluid surface evolution; their suitability across all matter classes remains uncertain.
- Mesh-validity diagnostics identify structural defects but are excluded from the main score, leaving unresolved how geometric validity should be balanced against visual fidelity and physical accuracy.
- The benchmark does not assess conservation laws, stability, contact forces, material parameters, energy behavior, or other physical properties that could distinguish physically plausible simulations from visually matching animations.
- The analysis reports broad solution categories such as analytic motion and custom simulation, but does not measure the quality, correctness, or transferability of the underlying simulators used within those categories.
- It is unclear whether agents choose simulation or analytic motion because of genuine task-appropriate reasoning, limited tool knowledge, computational constraints, or the scoring incentives of the benchmark.
- Each model is run once per scene, so the study does not quantify stochastic variability, best-of- performance, or the reliability of repeated attempts.
- The comparison across proprietary and open-weight models is potentially confounded by differences in agent harnesses, tool interfaces, context windows, inference settings, and execution infrastructure.
- The benchmark does not provide controlled ablations of input resolution, video length, frame rate, prompt wording, available libraries, GPU resources, or token budgets, making it difficult to identify the source of performance differences.
- The reported relationship between token use and quality does not establish causal effects of computation; models may differ in tokenization, hidden reasoning, tool-call overhead, or the efficiency of their generated code.
- The reasoning-effort ablation is performed only for one model family, so the benefit of additional inference-time reasoning may not generalize to other proprietary or open-weight models.
- The evaluation focuses on final submissions and does not analyze intermediate code, renders, failed attempts, or render-and-compare iterations, leaving the effectiveness of visual feedback and iterative debugging unknown.
- The paper does not measure how often agents identify and correct specific errors, when their improvements saturate, or whether longer interaction leads to better solutions per unit of cost.
- Human evaluation covers 17 of 18 models and only a subset of comparisons is judged by multiple participants, leaving uncertainty about the robustness of human rankings, especially for closely matched systems.
- Human and VLM judgments measure overall perceived similarity but do not independently validate each metric family or determine whether judges can reliably assess 3D geometry and physical mechanisms from monocular rendered videos.
- The VQA questions are scene-specific and gated by the presence of selected objects or events; their coverage, difficulty, and sensitivity to partial or shortcut-based reconstructions are not fully established.
- The benchmark does not test whether agents can explain their inferred dynamics, identify uncertainty, distinguish observable facts from assumptions, or recognize when the video is insufficient to determine a unique reconstruction.
- Ambiguity and non-identifiability are not explicitly modeled: multiple geometries, materials, cameras, and physical programs may explain the same video, but the evaluation generally compares against one reference reconstruction.
- The paper does not investigate whether agents can represent uncertainty or produce multiple plausible hypotheses when the visual evidence underdetermines scene structure or physical parameters.
- The real and synthetic splits differ in appearance, ground-truth availability, and scene construction, so the reported performance gap cannot be cleanly attributed to realism, domain shift, or the absence of 3D supervision alone.
- The effects of individual scene properties—such as object count, material count, occlusion, deformation magnitude, topology change, and interaction type—are not fully disentangled because many are correlated within the curated dataset.
- The benchmark does not assess computational efficiency of the generated programs, including simulation time, memory usage, rendering cost, numerical stability, or scalability to longer videos and higher spatial resolution.
- The security and reliability implications of executing model-generated graphics and simulation code are not examined, including unsafe resource consumption, nondeterministic behavior, dependency failures, or malicious code generation.
- The extent to which benchmark-specific conventions, output formats, Blender APIs, and available documentation influence performance is not evaluated, limiting conclusions about general 4D inverse-graphics ability beyond this environment.
Practical Applications
Immediate Applications
- Multimodal-agent evaluation and model selection — AI/software industry, academia.
Use
4DCodeBenchas a standardized regression suite for multimodal coding agents that generate Blender, Taichi, Warp, or other simulation code from video. Organizations can compare models not only on rendered appearance, but also on executability, static geometry, 2D motion, 3D motion, and human-aligned VLM scores. Potential workflow: run candidate models on the 200 scenes, track Overall and dynamics-specific scores, inspect failure categories, and select models for downstream graphics or robotics pipelines. Dependencies: the benchmark’s 200 scenes may not represent every target domain; scores should not be treated as evidence of safe deployment in unobserved environments. - Regression testing for graphics-code agents — software engineering and digital content creation. The executable-output requirement can be incorporated into continuous integration for systems that generate Blender or simulation scripts. A submission can be automatically checked for valid video output, correct frame rate and resolution, valid per-frame geometry, non-manifold meshes, self-intersections, and object interpenetration. Potential product: a “4D code compiler” or CI service that runs generated scene programs in isolated containers and returns structural and visual diagnostics. Dependencies: secure sandboxing is essential because generated code may be unsafe, computationally expensive, or capable of accessing unintended files.
- Automated evaluation of dynamic-scene reconstruction tools — VFX, animation, and game development. Studios can use the benchmark’s perceptual, trajectory, depth, and geometry metrics to test video-to-3D and video-to-animation systems. The separate static-versus-dynamic scores are particularly useful for distinguishing systems that reproduce a convincing first frame from systems that reproduce deformation, collisions, flow, or fracture. Potential workflow: use a model for initial reconstruction, render the scene, compare it with source footage, and route low-scoring cases to an artist or simulation specialist. Dependencies: monocular real-world video is inherently ambiguous; a visually similar reconstruction may have incorrect physical parameters or hidden geometry.
- Human-aligned quality assurance for generated reconstructions — media production and research. The reported correlation between VLM rankings and human preferences supports using VLM pairwise comparison as a first-pass triage tool. For example, a studio could compare several generated reconstructions and send only close or high-value cases to human reviewers. Dependencies: VLM agreement is weaker for closely matched models, and the reported agreement was established on this task distribution. Human review remains necessary for production-critical decisions.
- Training and benchmarking curricula for physical-world reasoning — academia and education. Researchers can use the dataset’s taxonomy—rigid and articulated bodies, deformable solids, cloth and rods, fluids, grains, fracture, and multi-material interactions—to construct staged curricula for multimodal agents. Models can first learn static geometry, then externally driven motion, and finally passive multi-material dynamics. Potential research tool: a benchmark dashboard reporting separate appearance, geometry, contact, trajectory, and flow performance rather than a single aggregate score. Dependencies: the synthetic portion is generated under simulator assumptions, while the real portion lacks complete 4D ground truth; both are needed for balanced evaluation.
- Diagnosis of failure modes in physical reasoning — AI research. The consistent gap between static reconstruction and dynamic reconstruction can be used to target model improvements. Researchers can evaluate whether a new method improves material-property inference, contact handling, topology change, or long-horizon motion rather than merely improving image similarity. Dependencies: current submissions frequently use analytic or prescribed motion—67% of analyzed solutions use analytic motion—so a high visual score does not necessarily imply transferable physical understanding.
- Interactive video-to-scene prototyping for artists and engineers — design, visualization, and engineering. Today’s agents can already generate executable approximations for relatively simple rigid-body or driven-motion scenes. A user could provide a short video, obtain a Blender scene and animation script, and manually correct object geometry, camera, materials, or trajectories. Potential products: storyboard reconstruction, rapid animation blocking, educational demonstrations, and initial CAD or simulation setup. Dependencies: substantial human correction is likely for real scenes, especially those involving occlusion, transparent objects, fluids, cloth, fracture, or multiple interacting materials.
- Robotics dataset inspection and simulation bootstrapping — robotics. Robot-manipulation videos in the real subset can be used to test whether an agent recovers object geometry, contacts, and externally driven motion. Approximate reconstructions may help researchers label demonstrations, visualize failed manipulation attempts, or initialize a simulator before manual refinement. Dependencies: the paper shows that driven scenes can be easier for 3D geometry but harder for some event-level judgments; reconstructed scenes should not directly control a robot without independent state estimation and safety validation.
- Teaching and public demonstrations of physics — education and daily life. The executable representation enables learners to inspect how a scene is constructed, modify parameters, and rerun the animation. A classroom tool could compare an observed video with a student-written model of bouncing, deformation, fluid flow, or fracture. Dependencies: generated code may use kinematic shortcuts rather than correct physical laws, so educational interfaces should expose assumptions and distinguish “visual match” from “physically valid model.”
- Benchmark-informed procurement and policy evaluation of AI systems — public-sector technology governance. Government or institutional buyers can require reporting of executability, dynamic-scene performance, open-weight versus proprietary model behavior, inference cost, and failure rates when procuring multimodal coding systems. This is more informative than relying only on static image or code-generation benchmarks. Dependencies: benchmark results should be combined with privacy, cybersecurity, licensing, reproducibility, and domain-specific safety assessments.
Long-Term Applications
- Video-to-physics digital twins — manufacturing, engineering, and industrial inspection. A mature version of the approach could reconstruct machinery, deformable components, granular materials, or fluid interactions from ordinary video and produce a parameterized simulator. Engineers could then test interventions, estimate loads, or compare alternative operating conditions. Required advances: reliable inference of hidden geometry, material parameters, contact forces, external forces, and topology changes; multi-view or depth sensing would likely be needed for high-stakes use. The current results show that dynamic reconstruction remains substantially weaker than static geometry.
- Robotic manipulation from visual demonstrations — robotics and automation. Future agents could convert demonstrations into executable 4D world models that explain object motion, contact, deformation, and material behavior. Robots might use these models to imitate tasks, predict consequences of grasps, or plan interventions on cloth, food, fluids, or soft objects. Required advances: intervention testing, uncertainty estimation, causal identification of forces, real-time inference, and sim-to-real validation. A reconstruction that merely reproduces an observed trajectory would not be sufficient for safe manipulation under new conditions.
- Physics-aware augmented and virtual reality — consumer software, training, and telepresence. Reconstructed dynamic scenes could be inserted into AR/VR environments as editable 3D objects with physically plausible motion. Users might pause, rotate, replay, or alter an event, such as a mechanical failure, sports action, or household demonstration. Dependencies: accurate camera calibration, low-latency reconstruction, view-consistent geometry, robust occlusion handling, and a physically meaningful model rather than frame-by-frame animation.
- Forensic and scientific reconstruction of physical events — law enforcement, safety engineering, and science. Video-to-code reconstruction could assist with replaying collisions, material failures, laboratory phenomena, or accidents. Explicit scene programs would make assumptions inspectable and allow investigators to test alternative hypotheses. Dependencies: evidentiary use requires calibrated cameras, multiple viewpoints or independent measurements, provenance tracking, uncertainty quantification, and validation against physical constraints. The benchmark does not establish forensic reliability.
- Predictive monitoring of deformable and flowing processes — healthcare, energy, and industrial operations. With further development, the methods could model tissue motion, fluid transport, soft materials, powders, or fracture propagation from video. Applications might include noninvasive monitoring, manufacturing quality control, pipeline inspection, battery-material analysis, or flow diagnostics. Dependencies: domain-specific sensors and constitutive models, labeled 4D data, strict validation, and compliance requirements. The current dataset excludes humans and animals and therefore does not directly support clinical deployment.
- Generative design and inverse simulation — engineering and product development. An agent could infer a compact executable scene, modify its geometry or physical parameters, and optimize the design for desired behavior. Examples include soft grippers, packaging materials, protective structures, fluid devices, and articulated mechanisms. Required advances: differentiable or efficiently optimizable simulators, reliable material identification, counterfactual evaluation, and guarantees that the reconstructed model generalizes beyond the recorded motion.
- Long-horizon physical prediction and intervention planning — autonomous systems. The benchmark could evolve from reconstruction toward asking whether an inferred world program predicts unseen frames, responds correctly to changed initial conditions, and handles external forces. This would provide a stronger foundation for autonomous navigation, manipulation, and planning. Dependencies: the paper explicitly identifies intervention tests and longer-horizon prediction as missing evaluations. Systems would need causal physical representations rather than prescribed trajectories that only reproduce the observed clip.
- Large-scale synthetic training environments for embodied AI — robotics and AI research. Reconstructed real scenes could become editable simulation assets for training agents. A model might generate variations in object geometry, material properties, initial conditions, or applied forces, producing richer training data than the original video alone. Dependencies: reconstruction errors can propagate into synthetic data and create misleading training distributions. Generated environments would require uncertainty-aware filtering, simulator validation, and real-world performance checks.
- Standardized physical-intelligence certification — academia, industry, and regulation. An expanded benchmark could support certification of multimodal agents on physically grounded capabilities, with separate thresholds for static geometry, dynamics, contact, material changes, executability, computational cost, and robustness to interventions. Dependencies: broader datasets, adversarial and out-of-distribution tests, transparent evaluation code, calibrated uncertainty, and agreed definitions of acceptable physical fidelity. The current Overall score is useful for comparison but should not alone serve as a safety or compliance standard.
- Everyday personal scene editing and explanation — consumer applications. In the longer term, users could capture a household event with a phone and obtain an editable 4D model: replaying how an object fell, testing alternative arrangements, or generating an explanation of a physical interaction. Dependencies: privacy-preserving on-device processing, robust handling of occlusion and lighting, accurate scale and depth recovery, and clear communication that the result is an inferred hypothesis rather than a definitive record of the hidden physical event.
Glossary
- Affine Body Dynamics (ABD): A simulation method that models bodies using affine transformations, extending rigid-body behavior to deformable or nearly rigid objects. “Affine Body Dynamics (ABD) extends related ideas to stiff and near-rigid bodies.”
- Analysis-by-synthesis: An inference strategy that explains observations by generating and comparing candidate models. “Inverse graphics frames visual perception as inverting the forward rendering process through analysis-by-synthesis”
- Anisotropic damage: Material damage that varies according to direction within the material. “AnisoMPM & Anisotropic damage and fracture; Drucker--Prager sand”
- Bradley--Terry model: A statistical model for estimating rankings from pairwise comparisons. “Pairwise preferences are aggregated across scenes into a model-level Elo rating using a Bradley--Terry model”
- Chamfer distance: A distance measure between geometric point sets based on nearest-neighbor distances. “3D Geometry measures static surface accuracy at frame~0 via Chamfer distance”
- Co-dimensional structure: A lower-dimensional object embedded in a higher-dimensional space, such as a cloth sheet or rod in three-dimensional space. “co-dimensional structures such as cloth and rope”
- Constitutive model: A mathematical model describing how a material responds to forces or deformation. “We extend the Genesis simulator with three new MPM constitutive models”
- Continuum-damage fracture: A fracture model that represents distributed material damage within a continuous medium. “CD-MPM & Continuum-damage fracture”
- Cosine similarity: A similarity measure based on the angle between two vectors. “we map the mean frame-wise cosine similarity to ”
- Drucker--Prager model: A constitutive model commonly used to represent yielding and failure in granular materials such as soil and sand. “a Drucker--Prager model for sand”
- Dynamic IoU: An intersection-over-union metric for measuring agreement in the spatial coverage of moving objects. “2D Dynamics measures the coverage of moving objects (Dynamic IoU).”
- Earth Mover’s Distance (EMD): A distance between distributions measuring the minimum cost of transforming one distribution into another. “Eulerian perspective comparing frame-to-frame displacement distributions without correspondence (EMD step).”
- Elo rating: A ranking system that estimates relative performance from pairwise outcomes. “Model pairs are sampled adaptively to shrink the widest confidence intervals, and judgments are aggregated into Elo ratings”
- Eulerian perspective: An analysis of motion that observes changes at fixed spatial locations rather than following individual material points. “an Eulerian perspective comparing frame-to-frame displacement distributions without correspondence”
- Executable graphics program: A program that constructs, animates, and renders a graphical scene when run. “Agents reconstruct dynamic scenes from video as executable graphics programs.”
- Forward rendering process: The process of generating visual observations from a scene representation. “Inverse graphics frames visual perception as inverting the forward rendering process”
- Hyperelastic material: A material whose stress–strain behavior is derived from a strain-energy function and can undergo large elastic deformations. “deformable solids, including hyperelastic, viscoelastic, and elastoplastic materials”
- Inference-time reasoning: Computation performed during model inference to improve problem-solving or generation quality. “Effect of Inference-Time Reasoning”
- Inverse graphics: The task of recovering a scene’s geometry, materials, lighting, and motion from visual observations. “We introduce 4D inverse graphics through code generation.”
- Interpenetration: An invalid geometric condition in which two simulated objects occupy overlapping physical space. “unintended interpenetration, and other visible simulation artifacts”
- Lagrangian perspective: An analysis of motion that follows persistent material points or particles through space and time. “a Lagrangian perspective tracking matter along persistent 3D paths”
- Material Point Method (MPM): A hybrid particle-grid simulation method for modeling deformable and fluid-like materials. “The Material Point Method (MPM) combines Lagrangian material points with an Eulerian background grid”
- Mesh rasterization: The conversion of geometric surfaces into image pixels for rendering or comparison. “Surface rasterization fails to capture background details”
- Monocular view: A visual observation obtained from a single camera viewpoint. “real videos provide only monocular views”
- Non-manifold edge: A mesh edge whose local neighborhood does not form a valid manifold surface structure. “open boundaries, non-manifold edges, degenerate faces”
- Optical flow: The apparent two-dimensional motion of image brightness patterns between video frames. “On real videos, we additionally estimate optical flow”
- Pareto frontier: The set of solutions that cannot improve one objective without worsening another. “with the Pareto frontier”
- Perceptual metric: A measure of visual or semantic similarity intended to approximate human judgments. “The Overall score averages these five families; failed runs receive the worst score on the affected metrics”
- Procedurally generated geometry: Geometry created algorithmically according to specified rules or parameters. “combining procedurally generated geometry with publicly available meshes”
- Rigid body: An idealized object whose shape does not deform during motion. “rigid and articulated bodies”
- Scene ontology: A structured classification of the entities, materials, and relationships represented in scenes. “We organize scenes hierarchically by the matter of their main dynamic objects.”
- Smoothed Particle Hydrodynamics (SPH): A mesh-free particle-based method for simulating fluids. “Smoothed Particle Hydrodynamics (SPH) represents fluids using Lagrangian particles”
- Spearman correlation: A rank-based statistic measuring the monotonic relationship between two variables. “human and VLM Elo correlate very strongly across models (Spearman ”
- Surface tension: The interfacial force that causes a liquid surface to resist deformation. “including fluid--rigid interactions and surface-tension effects”
- Temporal correspondence: The association of the same object or material elements across different time steps. “temporal correspondence for dynamic matter”
- Topology: The structural connectivity of a geometric object, including properties such as holes and connected components. “Flow, fracture, and deformation involve changes in shape and often topology”
- Trajectory dynamic time warping (Trajectory DTW): A sequence-alignment measure that compares trajectories while allowing differences in temporal speed or alignment. “Trajectory DTW”
- Vicoelastic material: A material exhibiting both viscous, rate-dependent behavior and elastic, recoverable deformation. “hyperelastic, viscoelastic, and elastoplastic materials”
- Viscoplastic material: A material that undergoes time-dependent deformation and permanent plastic flow under stress. “a viscoplastic model to simulate materials such as shaving cream and toothpaste”
- Visual question answering (VQA): A task in which a system answers natural-language questions about visual content. “VQA asks scene-specific questions in four categories”
- Vision-LLM (VLM): A model trained to jointly process visual inputs and natural-language information. “We further evaluate reconstructions with a \ac{vlm} as a judge.”
- Watertightness: The property of a three-dimensional mesh having a closed surface without gaps or open boundaries. “simple but inaccurate geometry can score well on watertightness and related measures”






