LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Abstract: A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program whose execution yields a scene that can be inspected, edited, and queried. We present LEGO-Anything, an Image-to-Code framework in which a coding agent iteratively writes and executes Blender code, inspects scenes and renderings, and revises the program. To evaluate end-to-end scene recovery, we introduce LEGO-Bench, a simulator-grounded benchmark with 208 images from 104 diverse indoor and outdoor scenes. LEGO-Bench separately scores artifact validity, visible-surface geometry, and rendered appearance. Its simulator-grounded design enables extensibility and precise automatic evaluation. Among evaluated agents, GPT-6-astra achieves the strongest overall results, with 53.4% indoor and 39.6% outdoor scores, yet substantial gaps remain between delivering valid scene artifacts and faithfully recovering scene geometry and appearance. Analysis of agent construction trajectories reveals three recurring issues: weak scene initialization, regressive edits during iteration, and unreliable self-evaluation. These findings motivate LEGO-Plugin, a training-free harness plugin for more controlled iterative scene construction, which improves all six evaluated models, with relative gains of up to 62.7% in overall score. Finally, we test whether reconstructed scenes can represent natural images and support vision tasks. In LEGO-World, we derive object detections, instance masks, and relative depth as deterministic queries on scenes reconstructed by GPT-6-astra. These readouts show non-trivial performance across all three tasks but fall well short of specialized vision models, suggesting that program-constructed scenes from current coding agents are a promising but not yet sufficiently precise representation of natural images.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper introduces LEGO-Anything, a system that turns a single picture into an editable 3D scene.
Instead of creating only a picture or a finished 3D model, the system asks a coding agent to write a Blender program. Blender is software used to create and render 3D worlds. The program describes things such as:
- Which objects are in the scene
- Where the objects are located
- Their shapes and sizes
- Their colors and materials
- The camera position and lighting
Because the result is code, people can inspect it, change it, and ask questions about it later.
The paper also introduces two tools:
- LEGO-Bench, a test set for measuring how well systems rebuild 3D scenes from images
- LEGO-Plugin, a helper that makes the coding agents more reliable
The researchers also study whether these reconstructed scenes can answer computer-vision questions, such as “Where are the objects?” or “How far away are they?”
2. What questions are the researchers asking?
The paper focuses on three main questions:
- How accurately can a coding agent rebuild a 3D scene from just one image?
- Why do these agents make mistakes while building the scene?
- Can the finished 3D scene be used for other vision tasks, such as detecting objects, creating object masks, or estimating depth?
A difficult part of the problem is that one image does not show everything. For example, a picture of a chair may show its front but not its back. The agent must make educated guesses about the hidden parts and the 3D layout.
3. How did the researchers conduct the study?
Turning an image into code
The system uses a coding agent as a kind of digital builder. The process is similar to constructing a LEGO model while repeatedly checking it against a picture:
- The agent looks at the reference image.
- It writes Blender code.
- Blender runs the code and creates a 3D scene.
- The agent looks at the scene and its rendered image.
- It changes the code to fix mistakes.
- It repeats these steps until it has a final scene.
This repeated process is called an iterative code-render-inspect loop. “Iterative” simply means doing something again and again, improving it each time.
Testing the results with LEGO-Bench
The researchers created LEGO-Bench, which contains:
- 208 images
- 104 different scenes
- Indoor and outdoor environments
- 8 environment types
- 17 scene themes
- 443 different 3D assets, such as furniture and other objects
The images were made from carefully designed 3D simulator scenes. This gave the researchers the correct answers, or ground truth, for each scene. Ground truth is like an answer key that tells them the real object locations, shapes, depths, and identities.
LEGO-Bench scores the reconstructed scene in three ways:
| Measurement | Simple meaning |
|---|---|
| Validity | Does the submitted 3D file work at all? |
| Reconstruction | Are the shapes and positions of visible objects correct? |
| Appearance | Does the rendered image look like the original image? |
For example, a scene might be valid because it opens correctly, but still receive a low score if the chairs, walls, or camera are in the wrong places.
Studying mistakes and testing LEGO-Plugin
The researchers watched the agents while they worked. They looked for problems such as:
- Taking too long to create a usable first scene
- Making later edits that damage earlier improvements
- Incorrectly believing that a bad scene is good
They then created LEGO-Plugin, which works like a safety and guidance system. It has three main parts:
- Enhanced Initialization: Gives the agent a better starting arrangement based on the image
- Grounded Refinement: Uses measurable evidence about object locations and depth instead of relying only on the agent’s opinion
- Version Control: Saves good versions and allows the agent to undo harmful changes
This is similar to saving different versions of a school project so that you can return to an earlier version if a new edit makes it worse.
Testing other computer-vision tasks
Finally, the researchers asked whether the reconstructed scene could be used for:
- Object detection: Drawing boxes around objects
- Instance segmentation: Coloring the exact pixels belonging to each object
- Depth estimation: Predicting which objects are closer or farther away
The important idea is that the same 3D scene could answer all three questions without needing a separate model for each task.
4. What did the researchers find?
The agents usually produced working files, but the scenes were not always accurate
The strongest tested system, called GPT-6-astra, achieved overall scores of:
- 53.4% indoors
- 39.6% outdoors
This means it often created a usable 3D scene, but the scene was still only a partial match to the original.
A major finding was the difference between making a valid file and rebuilding the scene correctly. The agents had validity scores close to 100%, meaning their files usually opened and contained something renderable. However, their geometry and appearance scores were much lower.
In everyday terms, the agents were usually able to hand in a working project, but the project did not always look like the example they were supposed to copy.
Outdoor and complicated scenes were harder
The agents performed worse on:
- Outdoor scenes
- Scenes with many objects
- More complicated layouts
As the number of objects increased, the scenes became less accurate. Outdoor scenes were especially difficult because they often contain large spaces, more varied lighting, and objects spread across different depths.
More reasoning helped some agents
When GPT-6 agents were given more time and effort to think through the problem, their performance generally improved. For example, GPT-6-astra’s score on one test group increased from 32.3% to 61.8% when its reasoning effort was increased.
However, this improvement was not equally strong for older or weaker models. Simply giving an agent more time does not always solve its problems.
Agents sometimes made their scenes worse
The researchers found three repeated types of failure:
- Weak starting scene: The agent took too long to create a useful first version.
- Regressive edits: Later changes sometimes damaged parts that were already correct.
- Poor self-evaluation: The agent often could not reliably tell whether a new version was better or worse.
The agents were especially poor at judging the scene’s geometry. Their judgments about whether one shape arrangement was better than another were often no better than guessing.
LEGO-Plugin improved every tested model
LEGO-Plugin improved all six tested coding-agent systems.
The biggest improvements happened for weaker models, which gained about 55% to 63% relative improvement. The strongest model improved only slightly because it already avoided many of the common problems.
This suggests that better tools and safeguards can help agents work more carefully without retraining the underlying AI model.
The reconstructed scenes could support several vision tasks
The reconstructed scenes were useful for object detection, segmentation, and depth estimation. However, they were not as accurate as specialized computer-vision systems.
The results were approximately:
- Object detection: 30.1 compared with 59.9 for a specialized model
- Segmentation: 14.8 compared with 54.0
- Depth estimation: 0.155 compared with 0.078, where lower is better
So the reconstructed scenes contained useful information, but they were not yet precise enough to replace expert systems.
5. Why is this research important?
The main idea is that a 3D scene should be more than a picture. If it is represented as an executable program, people can:
- Open and inspect it
- Move or replace objects
- Change the lighting or camera
- Ask questions about object positions
- Use it in games, simulations, robotics, or virtual reality
This could eventually allow someone to photograph a room and quickly create an editable digital version of it.
However, the research also shows that this technology is still developing. Current agents are good at producing something that works, but they are not yet good enough at recovering the exact shapes, positions, and appearances shown in an image.
The paper’s tools provide useful steps forward:
- LEGO-Bench gives researchers a fair way to measure progress.
- LEGO-Plugin helps agents avoid damaging their own work.
- LEGO-World shows that one reconstructed scene can support several different vision tasks.
Overall, the paper suggests that coding agents are a promising way to build editable 3D worlds from images, but they still need better understanding, more accurate geometry, and more reliable ways to check their own work.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Real-world generalization remains untested. LEGO-Bench uses simulator-rendered scenes built from professionally authored assets, so it is unclear how well the approach handles real photographs with sensor noise, imperfect lighting, reflections, motion blur, clutter, occlusions, and unmodeled objects.
- The benchmark has limited scale and diversity. The evaluation includes 208 images from 104 scenes, eight environments, and 17 themes; broader testing is needed across more architectural styles, geographic regions, object categories, weather conditions, viewpoints, and scene layouts.
- Benchmark realism may not match natural-image distributions. Although the inputs are described as natural-image-style, the paper does not quantify the domain gap between LEGO-Bench renders and real images or evaluate whether agents exploit simulator-specific visual regularities.
- The benchmark’s ground truth is view-limited. Reconstruction is scored primarily on surfaces visible from the reference camera, leaving the fidelity of occluded, unseen, and back-facing geometry unresolved.
- Single-image 3D ambiguity is not explicitly modeled. The evaluation does not distinguish between errors that are objectively unavoidable from one view and errors caused by poor agent decisions, nor does it assess whether multiple geometrically different but image-consistent reconstructions should receive equivalent credit.
- The evaluation metrics do not fully capture scene usefulness. Visible-surface F1 and pixel-threshold appearance scores may overlook physical plausibility, object functionality, topology quality, semantic correctness, material realism, lighting consistency, and editability of the resulting programs.
- The appearance metric is potentially brittle. The pixelwise thresholded metric with a fixed may be sensitive to small camera, lighting, or rendering differences while failing to reflect perceptual similarity; comparisons with learned perceptual, structural, and human-judgment metrics are still needed.
- Metric weighting is insufficiently justified. The overall score assigns equal weight to Reconstruction and Appearance, but the paper does not establish that this weighting reflects user preferences or downstream task requirements.
- Object-level scoring may obscure instance and scene-level failures. Averaging F1 over annotated objects may treat large structural errors, small-object errors, duplicate objects, and missing room-level geometry in ways that do not correspond to practical reconstruction quality.
- Category and correspondence assumptions are underexplored. The protocol relies on private object identities, masks, and correspondences, but it remains unclear how performance changes when objects are ambiguous, semantically mislabeled, duplicated, partially visible, or absent from the asset library.
- Asset-library dependence is unresolved. The method’s ability to reconstruct scenes containing novel objects, custom furniture, deformable objects, transparent objects, vegetation, signage, or objects unavailable in Fab is not evaluated.
- The role of asset retrieval versus procedural modeling is unclear. The paper does not isolate whether performance is driven primarily by retrieving suitable assets, generating primitives, composing existing objects, or approximating appearance with camera-facing geometry.
- Camera and scale recovery are not independently evaluated. Reconstruction errors are reported jointly, so the contributions of camera pose, field of view, global scale, object placement, and object geometry remain difficult to disentangle.
- Occlusion reasoning is not systematically analyzed. The paper reports that scenes become harder with complexity, but does not quantify how performance depends on occlusion rate, object overlap, visibility fraction, or depth ordering.
- Outdoor-scene failure modes remain underspecified. Outdoor performance is consistently lower than indoor performance, but the paper does not determine whether this is caused by scale variation, vegetation, sky and lighting, sparse structure, aerial viewpoints, asset mismatch, or camera-estimation errors.
- The trajectory analysis is narrow. Detailed failure analysis focuses mainly on GPT-6-astra and GPT-5.6-sol, leaving it unclear whether weak initialization, regressive edits, and unreliable self-evaluation generalize across all models, harnesses, and scene types.
- Causal attribution of failure modes is incomplete. The paper identifies correlations between delayed initialization, score-decreasing edits, and final quality, but does not experimentally isolate how much each factor independently contributes to reconstruction failure.
- The analysis of self-evaluation uses a limited judging protocol. Pairwise judgments by vision-LLMs may themselves be poorly calibrated or insensitive to 3D errors; human judgments, specialized geometric judges, and task-based validation are needed.
- The disagreement between deterministic metrics and human quality judgments is unresolved. Although a human-validation study is mentioned, the paper does not establish whether near-chance model judgments reflect genuine inability to assess geometry or limitations of the benchmark metrics.
- The effect of agent interaction budgets is incomplete. Results vary with reasoning effort, but the paper does not provide systematic scaling curves for execution time, number of edits, tool calls, token cost, rendering cost, or total monetary cost.
- Test-time reasoning gains are not separated from computational cost. Higher reasoning effort improves GPT-6 performance, but the quality–latency–cost trade-off and the point of diminishing returns remain unspecified.
- The generality of LEGO-Plugin is uncertain. Plugin evaluation is restricted to the 42-case Office subset, so its effectiveness on outdoor scenes, complex environments, unseen themes, and real images is unknown.
- The plugin’s components lack complete ablation studies. The paper reports the combined plugin improvement but does not clearly quantify the independent and interaction effects of Enhanced Initialization, Grounded Refinement, and Version Control.
- The plugin may inherit errors from its auxiliary models. VGGT, SAM 3, and Depth Anything V2 provide the plugin’s cues, but their failure sensitivity, calibration, and contribution to downstream reconstruction quality are not evaluated.
- The plugin’s residual objectives may not reflect true 3D fidelity. Projected extents and relative depth can be satisfied by view-dependent or geometrically incorrect constructions, leaving open whether the plugin improves genuine scene structure or primarily optimizes image alignment.
- Version Control may suppress beneficial exploration. Accepting or rolling back revisions based on intermediate scores could prevent agents from making temporary degradations that enable later improvements; this exploration–preservation trade-off is not studied.
- The plugin’s compatibility with other environments is unverified. The method is implemented around Blender, Blender-MCP, and specific harnesses, so portability to other renderers, simulators, scene representations, or coding-agent interfaces remains open.
- No learned training-based alternative is compared systematically. LEGO-Plugin is training-free, but the paper does not compare it against supervised fine-tuning, reinforcement learning, preference optimization, trajectory replay, learned scene critics, or specialized image-to-code models.
- Prompt and harness sensitivity is not fully characterized. Results are reported under native harness behavior and specified configurations, but robustness to prompt wording, tool availability, context limits, memory mechanisms, and alternative execution policies is unclear.
- Reproducibility is limited by dependence on proprietary agents. The main results rely on GPT-series systems, including models and harness behavior that may be unavailable or change over time, making independent replication and longitudinal comparison difficult.
- The downstream LEGO-World evaluation has limited coverage. Detection, segmentation, and depth are tested on only 100 randomly sampled images per dataset, which may produce high uncertainty and does not establish performance across the full distributions.
- Downstream evaluation uses different datasets from LEGO-Bench. The relationship between reconstruction quality on LEGO-Bench and readout quality on COCO, LVIS, and ETH3D is not established, especially because these datasets contain real images with distributions unlike the simulator scenes used for reconstruction evaluation.
- Equal-confidence AP may underestimate or distort scene-readout performance. Because reconstructed scenes provide no calibrated detection or segmentation confidences, the evaluation imposes an equal-confidence protocol rather than testing whether confidence estimation can be derived from the scene program.
- Semantic labeling is delegated to the agent without detailed error analysis. The paper does not separate failures in object recognition and category assignment from failures in localization, geometry, visibility, or scene projection.
- The downstream tasks are restricted to 2D readouts. It remains unexplored whether the reconstructed programs support 3D detection, spatial relationships, navigation, manipulation, embodied interaction, physical simulation, question answering, or counterfactual scene editing.
- Scene editability is claimed but not evaluated. The paper motivates executable programs as editable representations, yet does not measure whether users or agents can reliably perform targeted edits while preserving unrelated scene properties.
- Program quality and maintainability are not assessed. The validity metric checks whether artifacts open and render, but not whether the generated code is modular, readable, concise, semantically structured, deterministic, or easy to debug.
- Physical plausibility is largely unmeasured. Collision and stability checks are used when constructing benchmark scenes, but reconstructed outputs are not comprehensively evaluated for collisions, support relations, gravity, articulation, watertightness, or simulation stability.
- Unseen-view consistency is unresolved. A scene can achieve good reference-view appearance using billboard-like or view-dependent shortcuts; renders from novel camera viewpoints are needed to test whether the program captures a coherent 3D world.
- Temporal and multi-view consistency are unexplored. The method operates on one image, so it is unknown whether incorporating additional views, video, depth, or user corrections would substantially improve reconstruction or reduce ambiguity.
- Failure recovery from execution errors is not systematically studied. Artifact validity is reported, but the paper does not analyze syntax errors, Blender crashes, timeout behavior, dependency failures, corrupted exports, or the agent’s ability to recover from them.
- The relationship between scene complexity and object count is confounded. Easy, medium, and hard tiers increase visible content while holding some scene factors fixed, but the effects of object count, occlusion, geometric diversity, texture complexity, spatial extent, and lighting complexity are not separately identified.
- No uncertainty estimates are provided. The system produces a single scene program despite substantial single-view ambiguity; calibrated uncertainty, multiple candidate reconstructions, or confidence-aware downstream queries remain open directions.
- User-oriented utility is not evaluated. The paper does not measure whether users prefer executable reconstructions over meshes, point maps, or image renderings, nor how much time and expertise are required to inspect, edit, or correct the generated programs.
Practical Applications
Immediate Applications
The paper’s strongest near-term value is not fully faithful 3D reconstruction, but the production of editable, executable scene artifacts, evaluation infrastructure, and more reliable agent workflows.
- Assisted 3D content creation for games, animation, and visualization
- A user could provide a single image and obtain an editable Blender scene containing approximate objects, camera placement, geometry, materials, and lighting.
- Artists could use the generated scene as a blocking or layout starting point rather than modeling an environment from scratch.
- Potential products include an
image-to-Blenderplugin, automatic scene-blocking tools, and rapid previsualization workflows for film, advertising, and game development. - Dependencies: Current reconstruction quality is substantially better for artifact validity than for geometric or appearance fidelity. Human artists would still need to correct object identity, dimensions, textures, occlusions, and outdoor layouts.
- Rapid prototyping of virtual environments
- Architecture, interior design, retail, and real-estate teams could convert reference photographs into approximate 3D mock-ups for early-stage design review.
- Generated scenes could be used to test alternative furniture arrangements, camera viewpoints, lighting conditions, or material choices.
- The executable representation is particularly useful because users can edit scene parameters rather than only view a fixed image.
- Dependencies: The method reconstructs primarily what is visible from one view; hidden structure, metric scale, and physically accurate dimensions are not guaranteed. Safety-critical architectural decisions should not rely on the output without measurement or additional views.
- Training and evaluation of coding agents for 3D tasks
LEGO-Benchcan be used by academic and industrial research teams to compare coding agents, Blender agents, scene-generation systems, and multimodal models.- Its separate measures for validity, visible-surface geometry, and rendered appearance help distinguish:
- failure to produce a usable artifact,
- incorrect spatial or geometric reconstruction,
- and visual mismatch in rendering.
- Organizations could add proprietary scenes and assets to create domain-specific evaluation suites for offices, factories, warehouses, campuses, or urban environments.
- Dependencies: Benchmark extensions require simulator-ready scenes, licensed assets, controlled cameras, and private ground-truth geometry, masks, and depth. Results may also depend on the selected simulator, rendering engine, asset library, and evaluation thresholds.
- Quality-control infrastructure for agent-generated Blender files
- The validity checks described in the paper can be integrated into asset-production pipelines to automatically verify that a generated file:
- opens successfully,
- contains renderable geometry,
- has an active camera,
- exports to a valid
.glb, - and produces a non-degenerate image.
- This is immediately useful for batch generation, reducing failures caused by malformed files or missing scene components.
- Dependencies: Validation confirms technical usability, not semantic correctness. A scene can pass validity while still having poor geometry or appearance.
- More reliable iterative scene-construction workflows
LEGO-Plugincan be deployed as a harness layer around existing coding-agent and Blender-MCP workflows.- Its three mechanisms suggest a practical production workflow:
- use image-based initialization to establish camera and scene layout,
- use explicit residuals such as projected extent and relative depth for refinement,
- maintain versioned scene states and roll back edits that reduce quality.
- This could be used in software pipelines for 3D asset generation, CAD-like prototyping, synthetic-data creation, and visual automation.
- Dependencies: The reported gains were measured mainly on a 42-case Office subset. Performance may vary across domains, image styles, asset inventories, and agent harnesses. The plugin also relies on auxiliary models such as VGGT, SAM 3, and Depth Anything.
- Human-in-the-loop 3D authoring
- The system can function as an interactive assistant rather than a fully autonomous model: a user supplies an image, reviews intermediate renders, accepts or rejects revisions, and edits the generated scene program.
- Version control is especially suitable for workflows in which users want to preserve a good intermediate scene while exploring alternatives.
- Potential applications include educational 3D modeling, concept development, museum or cultural visualization, and rapid scene reconstruction for small studios.
- Dependencies: User review remains important because the agent’s self-evaluation is unreliable, particularly for geometry. The paper shows that automatic or model-based visual judgment should not be treated as a sufficient substitute for deterministic checks or human inspection.
- Deterministic extraction of approximate visual metadata
- Once a scene is reconstructed, object projections, instance masks, and camera-space depth can be generated from the same scene artifact.
- This can support lightweight annotation assistance, rough object inventorying, approximate depth overlays, and automatic generation of scene metadata for search or asset management.
- The approach may reduce the need to run separate task-specific models for every downstream query when approximate results are acceptable.
- Dependencies: The reported downstream results are substantially below specialist models: approximately 30.1 box AP, 14.75 mask AP, and 0.155 depth AbsRel. Outputs should therefore be treated as approximate metadata, not reliable perception in high-stakes settings.
- Educational tools for executable graphics and inverse graphics
- The image-to-code formulation provides a practical teaching environment for computer graphics, Blender scripting, computer vision, and agentic programming.
- Students could inspect how natural-language or visual observations are translated into scene primitives, execute the resulting code, and study the effects of camera, geometry, lighting, and materials.
LEGO-Benchcan support assignments involving error diagnosis, agent comparison, and iterative scene improvement.- Dependencies: Educational use requires reproducible software environments and careful handling of third-party assets and model licenses.
- Policy and procurement benchmarks for generative 3D systems
- Public agencies, enterprises, and research funders could use the benchmark’s validity–geometry–appearance decomposition when evaluating vendors or models.
- This is more informative than judging only rendered screenshots, since it tests whether systems produce reusable, inspectable, and exportable artifacts.
- Procurement workflows could require reporting scene validity, reconstruction quality, failure rates, and performance across indoor and outdoor cases.
- Dependencies: Benchmark results are not automatically equivalent to real-world photographic performance. Domain-specific validation and legally licensed assets would be required.
Long-Term Applications
These applications depend on substantially improved reconstruction fidelity, broader evaluation, multi-view or sensor integration, and stronger guarantees about geometry and semantics.
- Digital twins for buildings, factories, warehouses, and infrastructure
- A mature version of the system could transform photographs, inspection images, or field-worker captures into editable digital-twin scenes.
- The resulting programs could support asset inventories, spatial queries, maintenance planning, simulation, and visualization of proposed changes.
- In industrial settings, scene programs could connect visible objects to operational metadata such as equipment IDs, maintenance histories, or sensor streams.
- Dependencies: Single-image reconstruction is currently insufficient for metric accuracy, hidden surfaces, object identity, and connectivity. Deployment would require multi-view imagery, depth sensors, calibration, semantic verification, persistent object identities, and domain-specific accuracy guarantees.
- Robotics simulation and sim-to-real training
- Reconstructed environments could provide approximate simulator scenes for robot navigation, manipulation, perception, and policy testing.
- A photograph of a room, warehouse, or outdoor area could become an initial simulation environment, with editable object placement and camera configuration.
- Executable scene programs are well suited to simulation because they can be regenerated, parameterized, and queried.
- Dependencies: Robotics requires accurate collision geometry, physical materials, object articulation, scale, friction, and dynamics. The current results focus on visible appearance and do not establish physical or interaction fidelity. Errors could produce unsafe or misleading policies.
- Augmented reality and spatial computing
- A high-fidelity system could generate editable spatial models from ordinary photographs for AR annotation, remote assistance, virtual staging, and viewpoint synthesis.
- Users might place virtual objects into reconstructed environments, inspect occlusion relationships, or navigate approximate 3D versions of locations not currently accessible.
- Dependencies: AR requires accurate camera pose, metric scale, occlusion boundaries, lighting, and often persistent spatial registration. The paper’s single-view scenes and relative-depth outputs do not yet provide these guarantees.
- Interactive visual search and 3D-aware question answering
- Executable scenes could become structured representations that answer queries such as:
- “Which objects are near the desk?”
- “What is behind the chair?”
- “How far is the door from the table?”
- “Show all objects above the floor.”
- Unlike a fixed image embedding, a scene program can expose object identities, spatial relations, camera coordinates, and editable geometry.
- Dependencies: Reliable querying requires accurate object segmentation, category labels, spatial relations, and uncertainty estimates. Current reconstruction errors would propagate directly into answers, especially for occluded or visually ambiguous objects.
- Unified multimodal perception representations
- The
LEGO-Worldresults suggest a long-term architecture in which one executable scene supports detection, segmentation, depth, pose, affordance prediction, and spatial reasoning. - Such a representation could reduce duplication among task-specific vision systems and provide consistent predictions across tasks.
- Future systems could combine learned neural features with explicit geometry and programmatic scene structure.
- Dependencies: Current readouts trail specialized models substantially. Achieving a useful universal representation will likely require joint optimization, uncertainty-aware scene programs, richer object semantics, and evaluation on diverse real-world imagery.
- The
- Automated generation of synthetic training data
- Reconstructed scenes could be edited to generate controlled variants of an observed environment: altered lighting, object placement, camera views, weather, clutter, or material properties.
- These variants could train or test perception and robotics systems while preserving a link to the original image.
LEGO-Bench’s controlled complexity tiers could support systematic studies of how scene difficulty affects model performance.- Dependencies: Synthetic data is useful only if geometry, textures, labels, and physical behavior are sufficiently realistic. Reconstruction artifacts could otherwise introduce biased or misleading training examples.
- Photogrammetry and cultural-heritage documentation
- A more accurate system could turn historical photographs, archival images, or limited field documentation into editable 3D reconstructions for preservation, education, and public access.
- Scene programs would allow curators to annotate objects, restore missing elements, and create alternative reconstructions while preserving provenance.
- Dependencies: Historical images often lack reliable scale, camera metadata, and complete visibility. The system would need explicit uncertainty representation, expert review, provenance tracking, and safeguards against presenting speculative geometry as fact.
- Urban planning and geospatial visualization
- Outdoor reconstruction could eventually support rapid modeling of streetscapes, construction sites, disaster areas, and public spaces from photographs.
- Applications include preliminary planning, visual impact analysis, infrastructure inventories, and change detection.
- Dependencies: Outdoor performance is lower than indoor performance in the paper, and urban use requires georeferencing, accurate scale, terrain modeling, weather robustness, and coverage of large scenes. A single image cannot reliably recover unseen buildings or infrastructure.
- Automated scene editing and natural-language design
- Once scenes are represented as code, users could request edits such as “move the cabinet beside the wall,” “replace the chairs,” or “create a brighter evening version.”
- Coding agents could produce reproducible, parameterized variants rather than destructive manual edits.
- This could support design iteration in games, architecture, advertising, and virtual production.
- Dependencies: Reliable editing requires stable object identities, semantic scene graphs, constraint handling, collision checking, and protection against regressive changes. The paper demonstrates the need for version control because later agent edits can damage previously correct content.
- Safety-critical inspection and emergency response
- Mature systems could reconstruct approximate scenes from inspection or disaster photographs to support remote triage, route planning, and prioritization of damaged infrastructure.
- Explicit geometry and depth could help responders reason about spatial relationships when direct access is difficult.
- Dependencies: The current accuracy levels are not sufficient for safety-critical decisions. Such applications would require calibrated sensors, multi-view confirmation, confidence intervals, human authorization, robust out-of-distribution testing, and clear separation between measured and inferred scene elements.
- Policy standards for trustworthy generative 3D
- The paper’s separation of artifact validity, geometry, and appearance could inform future standards for reporting the quality of generated 3D assets.
- A mature standard could require provenance, asset-license information, reproducibility, uncertainty, rollback history, and task-specific safety metrics in addition to visual quality.
- Dependencies: Metrics must be validated on real photographs and operational tasks, not only simulator-rendered scenes. Standards would also need to account for privacy, copyrighted assets, model-generated errors, and the risks of treating inferred 3D structure as ground truth.
Glossary
- Artifact validity: Whether a submitted digital artifact is usable, well-formed, and evaluable. “These metrics distinguish failure to deliver an evaluable scene from failures of geometric and visual fidelity.”
- Appearance metric: A measure of visual similarity between a rendered reconstruction and a reference image. “Appearance evaluates visual similarity between the reconstructed scene and the reference image.”
- Articulated asset: A 3D object composed of parts connected by joints that permit motion. “Other work targets editable indoor scenes from RGB-D scans or simulation-ready articulated assets.”
- Camera coordinates: A coordinate system whose origin and axes are defined relative to the camera viewpoint. “Both point sets are projected onto the image plane and partitioned into scored objects by the private instance masks, then compared directly in camera coordinates without alignment or rescaling.”
- Camera-space depth: The distance of a point from the camera, expressed in the camera’s coordinate system. “camera-space depth renderings provide relative depth.”
- Coding agent: An AI system that writes, executes, inspects, and revises code to accomplish a task. “Coding agents combine LLMs with tools for code editing, execution, and iterative verification.”
- Collision and stability checks: Tests that determine whether simulated objects intersect improperly or remain physically stable. “Candidate scenes undergo collision and stability checks and human review before inclusion.”
- Depth-scaled tolerance: An error threshold that changes according to an object’s distance from the camera. “precision is the fraction of submitted points within a depth-scaled tolerance of the nearest reference point ”
- Deterministic readout: A fixed procedure that derives a task-specific prediction from an existing representation without additional learned inference. “object detection, instance segmentation, and relative depth can be obtained as deterministic readouts of the same frozen scene”
- Differentiable rendering: Rendering formulated so that image outputs can be differentiated with respect to scene parameters, enabling optimization. “Rather than producing in one shot, the agent alternates between editing code, executing it, and inspecting the resulting scene and renderings”
- Executable artifact: A saved computational object that can be run to produce or inspect a result. “The agent submits the final scene, its export, and a rendered view, which together form an executable artifact that can be evaluated for fidelity and queried for downstream perception.”
- Executable scene program: A program whose execution constructs a manipulable 3D scene. “an executable scene program makes objects, geometry, layout, and camera explicit”
- F1 score: The harmonic mean of precision and recall, commonly used to evaluate retrieval or detection quality. “The Reconstruction score (F@5\%) averages object-level F1 scores”
- Field of view: The angular extent of a scene visible through a camera. “including image dimensions, horizontal field of view, a category taxonomy, and an output schema”
- Fidelity: The degree to which a reconstruction accurately matches the reference scene or image. “We study this setting as LEGO-Anything, an Image-to-Code (Image2Code) framework in which a general-purpose coding agent writes and executes Blender programs, inspects the evolving scene and its renderings, and revises the construction to match a single reference image.”
- Forward depth: The distance of a point along the camera’s viewing direction. “where is its forward depth”
- Ground truth: Authoritative reference data used to evaluate predictions. “Because every case is rendered from a fully specified simulator scene, ground truth comes for free”
- Harness: Software infrastructure that organizes an agent’s tools, execution context, and feedback. “Their harnesses organize tool access, execution context, and environmental feedback”
- Image-to-Code: A task in which an image is converted into executable code that reconstructs or represents its contents. “LEGO-Anything formulates single-image scene reconstruction as an Image-to-Code (Image2Code) problem.”
- Instance mask: A pixel-level mask identifying the region belonging to one particular object instance. “retains private geometry, object identities, camera parameters, depth, and instance masks for automatic evaluation”
- Inverse graphics: The process of inferring scene properties, such as geometry, materials, lighting, and camera parameters, from images. “SEIG reconstructs images as editable Blender programs through staged executable inverse graphics”
- MCP tool: A tool exposed through the Model Context Protocol for allowing an AI agent to interact with external software or data. “LEGO-Plugin is a training-free control layer exposed as MCP tools, workflow skills, and runtime hooks”
- Natural-image-style observation: A rendered image designed to resemble the visual appearance and diversity of a real photograph. “provides realistic, natural-image-style observations”
- Non-degenerate image: An image that contains valid, meaningful pixel data rather than being empty, constant, or otherwise unusable. “final.png must be a parseable, non-degenerate image.”
- Parametric part: A component whose geometry is controlled by numerical parameters. “P3D-Bench targets parametric parts and assemblies”
- Perceptual sufficiency: The extent to which a representation preserves enough information to support visual recognition or related tasks. “the perceptual sufficiency of frozen reconstructed scenes.”
- Point map: A representation associating image pixels with points in a 3D coordinate system. “Compared with meshes, point maps, or object sets”
- Precision and recall: Metrics measuring, respectively, the correctness of predicted elements and the completeness of retrieved reference elements. “For each object , precision is the fraction of submitted points within a depth-scaled tolerance”
- Procedural generation: The algorithmic creation of content from rules, parameters, or reusable procedures. “Procedural systems generate natural, indoor, and urban environments from structured specifications and reusable assets”
- Relative depth: Depth information that expresses ordering or proportional distance rather than absolute metric distance. “We next ask whether reconstructed scenes can serve as representations of natural images.”
- Rendered appearance: The visual characteristics of a scene as produced by a rendering process, including geometry, materials, lighting, and camera settings. “LEGO-Bench separately scores artifact validity, visible-surface geometry, and rendered appearance”
- Residual: The measured discrepancy between an observed or reconstructed value and a target value. “then compares the executed scene against these targets through explicit residuals”
- RGB-D scan: A scan containing both color data and per-pixel depth measurements. “Other work targets editable indoor scenes from RGB-D scans”
- Scene graph: A structured representation of a scene as objects and their relationships. “structured representations of natural images”
- Scene program: Executable code that specifies the objects, geometry, layout, camera, and other properties of a 3D scene. “We present LEGO-Anything, an Image-to-Code framework in which a coding agent builds such a program”
- Scene reconstruction: The recovery of a 3D scene’s objects, geometry, arrangement, and appearance from visual or other observations. “A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program”
- Segmentation: The assignment of pixels to objects, regions, or semantic categories in an image. “The resulting scene program can also serve as a representation that can be queried directly for detection, segmentation, and depth estimation.”
- Simulator-grounded benchmark: An evaluation dataset whose reference scenes and annotations originate from a simulator with known underlying state. “To evaluate how well such agents recover scenes end to end, we introduce LEGO-Bench, a simulator-grounded benchmark”
- Single-image reconstruction: The recovery of a 3D scene from one two-dimensional image. “We study this setting as LEGO-Anything, an Image-to-Code (Image2Code) framework”
- Spatial layout: The positions and arrangement of objects within a scene. “an agent must reconstruct the scene as a complete, executable 3D artifact, jointly recovering its contents, geometry, spatial layout, and appearance from one view.”
- Structured representation: An organized representation that explicitly encodes entities and their properties or relationships. “We next ask whether reconstructed scenes are precise enough to serve as structured representations of natural images”
- Test-time scaling: Improving model performance by allocating additional computation or reasoning during inference rather than training. “We run a test-time scaling experiment on a fixed 42-case Office subset”
- Trajectory analysis: Examination of the sequence of intermediate states, actions, and outcomes produced during an agent’s execution. “To diagnose why end-to-end scene reconstruction still falls short, we analyze agent construction trajectories”
- Training-free: Operating without updating model parameters using task-specific training data. “These findings motivate LEGO-Plugin, a training-free harness plugin”
- Visible-surface geometry: The shape and spatial arrangement of scene surfaces that can be seen from the reference camera. “LEGO-Bench separately scores artifact validity, visible-surface geometry, and rendered appearance”
- Volumetric or geometric fidelity: Accuracy of a reconstruction’s 3D structure relative to the reference. “yet substantial gaps remain between delivering valid scene artifacts and faithfully recovering scene geometry and appearance.”













