Papers
Topics
Authors
Recent
Search
2000 character limit reached

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Published 28 Sep 2026 in cs.CV | (2609.36380v1)

Abstract: A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program whose execution yields a scene that can be inspected, edited, and queried. We present LEGO-Anything, an Image-to-Code framework in which a coding agent iteratively writes and executes Blender code, inspects scenes and renderings, and revises the program. To evaluate end-to-end scene recovery, we introduce LEGO-Bench, a simulator-grounded benchmark with 208 images from 104 diverse indoor and outdoor scenes. LEGO-Bench separately scores artifact validity, visible-surface geometry, and rendered appearance. Its simulator-grounded design enables extensibility and precise automatic evaluation. Among evaluated agents, GPT-6-astra achieves the strongest overall results, with 53.4% indoor and 39.6% outdoor scores, yet substantial gaps remain between delivering valid scene artifacts and faithfully recovering scene geometry and appearance. Analysis of agent construction trajectories reveals three recurring issues: weak scene initialization, regressive edits during iteration, and unreliable self-evaluation. These findings motivate LEGO-Plugin, a training-free harness plugin for more controlled iterative scene construction, which improves all six evaluated models, with relative gains of up to 62.7% in overall score. Finally, we test whether reconstructed scenes can represent natural images and support vision tasks. In LEGO-World, we derive object detections, instance masks, and relative depth as deterministic queries on scenes reconstructed by GPT-6-astra. These readouts show non-trivial performance across all three tasks but fall well short of specialized vision models, suggesting that program-constructed scenes from current coding agents are a promising but not yet sufficiently precise representation of natural images.

Summary

  • The paper proposes an innovative approach to 3D scene reconstruction by encoding reconstructions as executable Blender programs, facilitating editable, inspectable, and queryable scenes .
  • LEGO-Anything demonstrates that while coding agents excel at producing valid Blender scenes, geometric and visual fidelity remains a significant challenge—the optimal executed computation had near-saturated artifact validity, but geometrics, fidelity and visual fidelity were much lower than traditional mesh/reconstruction benchmarks.
  • The LEGO-Bench evaluation framework provides a unique and important methodology for assessing both scene fidelity and representation quality in 3D scene reconstruction, and potentially facilitates exploring new strategies for self evalution and iteratively refining visual comparisons that improve understanding genericity and scene usability in reuscable scene representations.

The paper studies single-image 3D reconstruction as an executable programming problem rather than as direct prediction of a mesh, point cloud, or fixed scene representation. Its central claim is that a reconstruction is more useful when it is encoded as a Blender program whose execution produces an editable, inspectable, and queryable scene. The proposed system, LEGO-Anything, uses a general-purpose coding agent to iteratively interpret an RGB image, write and execute Blender code, inspect rendered outputs, and revise the scene program. The work introduces two evaluation resources around this formulation: LEGO-Bench, for end-to-end reconstruction, and LEGO-World, for testing whether reconstructed scenes support downstream visual tasks (2609.36380).

Executable scene reconstruction as Image-to-Code

LEGO-Anything defines reconstruction as a policy-mediated interaction between an image, a coding environment, and a scene executor. Given an image II, an agent produces a program PP through repeated code-editing and execution steps; executing PP yields a scene SS. The scene program is therefore both the output representation and the operational interface through which the agent performs reconstruction.

This formulation changes the optimization target in several ways. The agent is not required to infer a complete scene in one forward pass. It can establish a coarse scene, render it, compare the result with the reference, and apply subsequent edits. Objects, camera parameters, geometry, materials, lighting, and spatial layout remain explicit in the Blender state. The resulting artifact can also be exported and interrogated after construction.

Figure 1

Figure 1: An agent iteratively converts a single RGB image into an executable Blender scene program through planning, code editing, rendering, and inspection.

The workflow is implemented with general-purpose coding agents operating through Blender MCP in a shared execution environment. The agent receives the reference image and public task metadata but not the benchmark’s private geometry, masks, depth, or evaluator scores. Intermediate programs and scenes form a construction trajectory, allowing the paper to evaluate not only final fidelity but also initialization quality, edit behavior, and the relationship between intermediate and final states.

Figure 2

Figure 2: Agentic scene construction organized around the workspace, action space, Blender code state, and iterative validation workflow.

The choice of executable programs is consequential. A valid Blender file can be opened, rendered, edited, and exported even when its visual fidelity is poor. Consequently, the paper explicitly separates artifact deliverability from reconstruction quality rather than treating successful program execution as evidence of faithful reconstruction.

LEGO-Bench: simulator-grounded end-to-end evaluation

LEGO-Bench addresses the lack of precise ground truth for diverse single-image scene reconstruction. Its inputs are rendered from fully specified simulator scenes, combining natural-image-style visual complexity with access to private geometry, depth, camera parameters, instance masks, and object identities. The benchmark contains 208 RGB views from 104 logical scenes, spanning eight environments, 17 themes, and 443 registered assets. It includes indoor, outdoor, and Bird’s-Eye settings, with scene families whose object content increases from Easy to Medium to Hard while shared architecture, camera, lighting, materials, and object transforms remain fixed.

Figure 3

Figure 3: Representative LEGO-Bench inputs across indoor, outdoor, and Bird’s-Eye reconstruction settings.

The benchmark’s controlled construction procedure is an important methodological contribution. Difficulty is not defined solely by global object count. Instead, matched scene families add visible objects while preserving the rest of the scene, enabling comparisons that attribute performance changes more directly to scene content.

Figure 4

Figure 4: Scored-object counts increase within matched scene families, while object count alone is not used as the formal difficulty definition.

LEGO-Bench reports three principal metrics:

  • Validity checks whether the submission contains a reloadable Blender scene, a non-empty GLB export, a renderable active camera, and a parseable non-degenerate image. Invalid or unresolved submissions receive zero for the headline metrics.
  • Reconstruction measures visible-surface agreement in the reference camera frame. Reference and predicted surfaces are partitioned by private instance masks and compared without registration, rescaling, or camera fitting. The headline metric is an object-level macro-average F-score at a depth-relative tolerance of 5%.
  • Appearance compares a fresh evaluator render with the reference image under a fixed rendering protocol. It captures the joint effects of geometry, camera, materials, lighting, and color configuration.

The no-alignment reconstruction metric is deliberately stringent. It penalizes incorrect camera pose, scale, object placement, missing surfaces, and excess geometry rather than absorbing these errors through post hoc registration. Appearance is complementary rather than redundant: a scene may obtain reasonable geometric agreement but render poorly because of incorrect materials or illumination, while a visually similar render may conceal geometric errors.

Figure 5

Figure 5: LEGO-Bench separates artifact validity, visible-surface reconstruction, and evaluator-rendered appearance.

The benchmark also includes a human validation study. Across six annotators and 240 samples, compatibility agreement was 86.6% overall, with 88.9% for reconstruction judgments and 84.3% for appearance judgments. Metric–human compatibility reached 83.7% under the study’s uncertainty rule. These results support the use of the automatic metrics, although they do not establish that the metrics capture every property of scene usefulness.

Main reconstruction results

The general-purpose coding agents achieved near-saturated artifact validity but much lower geometric and visual fidelity. GPT-6-astra was the strongest evaluated configuration, reaching an overall score of 53.4% indoors and 39.6% outdoors. Its indoor Reconstruction and Appearance scores were 52.4% and 54.4%, respectively; outdoors, they fell to 34.0% and 45.5%.

Model Indoor validity Indoor score Outdoor validity Outdoor score
GPT-6-astra 100.0% 53.4% 98.0% 39.6%
GPT-6-sol 99.4% 32.3% 99.7% 24.2%
GPT-6-luna 99.7% 23.2% 99.7% 17.3%
GPT-5.6-sol 97.8% 14.8% 98.7% 15.3%
GPT-5.6-terra 99.7% 15.4% 98.0% 11.5%
GPT-5.6-luna 99.1% 14.1% 97.0% 12.1%

The principal empirical distinction is therefore between delivering a valid executable artifact and recovering the scene faithfully. The coding agents almost always produced scenes that could be opened and rendered, but the gap between validity and fidelity remained large. GPT-6-astra’s performance exceeded that of the task-specific and single-image baselines in overall score, but this result should not be interpreted as dominance on every constituent metric. Gen3DSR, for example, achieved higher indoor Reconstruction than GPT-6-astra, 65.4% versus 52.4%, but obtained zero Appearance because its output did not provide a valid appearance reconstruction under the benchmark protocol.

Scene complexity reduced Reconstruction and Appearance without materially affecting validity. Across the evaluated GPT configurations, validity remained approximately 99.5% across complexity tiers, while overall performance declined from 24.6% on Easy scenes to 21.3% on Medium and 20.5% on Hard. The decline was particularly pronounced for outdoor Reconstruction, which fell from 18.4% to 12.7%. This result indicates that additional scene content primarily increases perceptual and geometric errors rather than causing agents to fail at artifact production.

The paper also reports a test-time reasoning experiment on a 42-case Office subset. Increasing native reasoning effort improved all three GPT-6 models. GPT-6-astra increased from 32.3% at Low effort to 61.8% at XHigh effort; GPT-6-sol increased from 21.3% to 39.7%; and GPT-6-luna increased from 14.4% to 21.2%. The gains were larger for stronger models, whereas GPT-5.6 configurations showed weak or non-monotonic scaling. The result suggests that additional inference-time computation is useful in this setting, but its effectiveness depends strongly on the underlying model.

Construction trajectories and failure mechanisms

A central strength of the paper is that it does not reduce agent behavior to a final benchmark score. By rescoring intermediate renderable artifacts, it identifies three recurrent failure modes: weak initialization, regressive edits, and unreliable self-evaluation.

GPT-6-astra reached its first evaluable scene within approximately the first tenth of its available budget, whereas GPT-5.6-sol required roughly one fifth. The difference matters because delayed initialization leaves less time for render-based refinement. After initialization, GPT-6-astra improved more consistently, while GPT-5.6-sol stagnated or regressed. In the analyzed trajectories, 29.6% of GPT-5.6-sol’s edits decreased the benchmark score, and its final result was 3.2 points below its best intermediate scene.

Figure 6

Figure 6: GPT-6-astra initializes earlier relative to its runtime and improves more steadily, whereas GPT-5.6-sol exhibits more frequent regression.

The non-monotonicity of construction is a substantive finding. The final artifact is not necessarily the best artifact produced during the trajectory, even when the agent has already constructed a higher-quality scene. Late edits can damage geometry, camera settings, object placement, or appearance. The paper’s severe regression example shows that this problem is not confined to weaker systems: even the strongest model can substantially degrade a scene after reaching a better intermediate state.

Figure 7

Figure 7: A late-stage edit can substantially reduce the quality of an already stronger scene, including for the strongest evaluated model.

The self-evaluation study reinforces this diagnosis. Across 36 builder–judge pairs, self-judgments agreed with the deterministic Reconstruction direction only 45.8% of the time, below chance, and agreed with Appearance direction 62.2% of the time. Cross-model judging produced 45.4% and 63.9%, respectively. Thus, self-judging provided no consistent advantage. Geometry was especially difficult for the models to assess, plausibly because image-level similarity does not reliably expose depth, scale, occlusion, and surface-placement errors.

This result directly challenges the assumption that a coding agent can use its own visual critique as a reliable optimization signal. The paper therefore treats reference-grounded measurements, rather than free-form model judgment, as the appropriate basis for controlled refinement.

LEGO-Plugin and training-free trajectory control

LEGO-Plugin is a harness-compatible control layer designed to address the three diagnosed failures without retraining the underlying agent. It adds three modules:

  • Enhanced Initialization uses VGGT-derived camera and scene-frame estimates together with image-grounded layout cues to establish a reference-aligned starting scene.
  • Grounded Refinement uses reference object regions and relative depth estimates from SAM 3 and Depth Anything V2 to produce explicit residuals for projected extent, centroid, depth order, visibility, and related constraints.
  • Version Control wraps updates in transactions, snapshots scene state, checks protected entities and hard requirements, and commits, repairs, or rolls back candidate revisions.

Figure 8

Figure 8: LEGO-Plugin adds reference-grounded initialization, measured refinement, and transactional version control to a standard coding-agent harness.

The plugin does not access private benchmark ground truth or LEGO-Bench scores. Its evidence is derived only from the input image and current Blender state. This distinction is important because it makes the reported improvements attributable to better control and measurement rather than evaluator-specific optimization.

On the 42-case Office subset, LEGO-Plugin improved all six evaluated model configurations. The relative gains were largest for the weaker models: the GPT-5.6 variants improved by approximately 55–63%, corresponding to roughly seven to eight score points. GPT-6-luna improved by 27.5%, GPT-6-sol by 12.1%, and GPT-6-astra by only 2.1%. The inverse relationship between baseline strength and plugin benefit indicates that the plugin compensates primarily for procedural weaknesses already avoided by the strongest model.

Figure 9

Figure 9

Figure 9

Figure 9

Figure 9

Figure 9

Figure 9: LEGO-Plugin produces visibly closer reconstructions and more than doubles example scores, from 0.147 to 0.325 and from 0.109 to 0.278.

The result is strong but qualified. LEGO-Plugin improves every tested configuration under the stated protocol, yet it does not eliminate the performance gap between weaker and stronger agents. Moreover, the evaluation is concentrated on the Office subset, so the reported relative improvements do not establish equivalent gains across all environments, outdoor scenes, or Bird’s-Eye views.

Executable scenes as downstream visual representations

LEGO-World tests whether a frozen reconstructed scene can support multiple visual tasks through deterministic readouts. From the same GPT-6-astra-generated scene program, the system derives 2D object boxes, instance masks, and camera-space depth. No task-specific training is applied to the reconstructed scenes.

The results demonstrate non-trivial reuse but also substantial imprecision:

Task LEGO-Anything Specialist baseline
COCO box AP 30.14 59.88, DINO
LVIS mask AP 14.75 53.96, SAM 3
ETH3D AbsRel 0.1554 0.0783, Depth Anything 3

The scene-derived readouts achieve approximately half of the specialist detector’s box AP, but only a much smaller fraction of specialist instance-segmentation performance. Depth estimation is also substantially worse: AbsRel is 0.1554 compared with 0.0783 for Depth Anything 3. The comparison is subject to an explicit protocol limitation: scene-derived detections and masks have no calibrated confidence scores, so average precision uses equal confidence and a deterministic tie order. The reported AP therefore measures compatibility under that ranking policy rather than calibrated detection quality.

Coverage is also imperfect. Scene outputs were scored on 94 of 100 COCO and LVIS images and 99 of 100 ETH3D images, while all attempted images remained in the official evaluation denominators. These details matter because a frozen scene representation must support both semantic completeness and geometric accuracy to function as a general visual substrate.

The downstream results establish a narrower claim than universal task competence. Executable scenes already contain enough structure to support detection, segmentation, and relative depth without separate task-specific prediction heads. However, their errors remain too large for the representation to match specialized vision systems, particularly for object boundaries, instance identity, and precise geometry.

Limitations and open questions

The benchmark relies on simulator-rendered scenes rather than real photographs with independently measured 3D ground truth. This permits exact evaluation but introduces a domain assumption: performance on LEGO-Bench may not transfer directly to uncontrolled real imagery, where lighting, texture, sensor artifacts, asset identity, and physical irregularities differ from the simulator distribution.

Single-view reconstruction is intrinsically underdetermined. The benchmark evaluates visible-surface geometry and rendered appearance, not uniquely correct hidden geometry or complete physical scene structure. A scene can therefore score well while remaining incorrect behind occlusions. Conversely, appearance may be improved through view-specific construction that does not yield a generally accurate 3D model.

The appearance metric is a thresholded pixel agreement measure rather than a learned perceptual metric. It is sensitive to rendering configuration and evaluates a fresh standardized render, which is appropriate for artifact verification but does not fully characterize material realism or human perceptual similarity. The reconstruction metric is similarly view-conditioned and uses private reference masks to define object scopes; its validity for applications requiring semantic correspondence or interaction beyond the reference view remains open.

The downstream LEGO-World evaluation has additional methodological constraints. Equal-confidence AP is not directly comparable to calibrated detector rankings, and the scene-derived outputs have incomplete coverage. The experiments use GPT-6-astra without LEGO-Plugin, so they do not determine whether plugin-controlled scenes would produce better or worse downstream representations.

Finally, the trajectory analysis identifies regressions but does not fully isolate their causal sources. Model capability, harness behavior, token budget, tool latency, rendering frequency, and prompt configuration are coupled in the evaluated systems. The paper leaves open whether explicit checkpoint selection, learned value estimation, improved scene abstractions, or stronger geometric feedback would provide the most effective remedy.

Conclusion

LEGO-Anything formulates single-image 3D reconstruction as iterative construction of an executable Blender program. LEGO-Bench shows that coding agents can reliably produce valid scene artifacts, but that artifact validity substantially exceeds geometric and appearance fidelity. GPT-6-astra reaches 53.4% indoor and 39.6% outdoor overall scores, while trajectory analysis identifies delayed initialization, regressive edits, and unreliable visual self-evaluation as major sources of failure. LEGO-Plugin improves all six evaluated agents without training, with relative gains of up to 62.7%, especially for weaker configurations. LEGO-World further shows that reconstructed scenes support deterministic detection, segmentation, and depth readouts, although well below specialized models. The paper’s principal contribution is thus not only an Image-to-Code reconstruction system, but an evaluation framework that exposes the difference between executable artifact production, faithful scene recovery, and representation quality (2609.36380).

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

This paper introduces LEGO-Anything, a system that turns a single picture into an editable 3D scene.

Instead of creating only a picture or a finished 3D model, the system asks a coding agent to write a Blender program. Blender is software used to create and render 3D worlds. The program describes things such as:

  • Which objects are in the scene
  • Where the objects are located
  • Their shapes and sizes
  • Their colors and materials
  • The camera position and lighting

Because the result is code, people can inspect it, change it, and ask questions about it later.

The paper also introduces two tools:

  • LEGO-Bench, a test set for measuring how well systems rebuild 3D scenes from images
  • LEGO-Plugin, a helper that makes the coding agents more reliable

The researchers also study whether these reconstructed scenes can answer computer-vision questions, such as “Where are the objects?” or “How far away are they?”

2. What questions are the researchers asking?

The paper focuses on three main questions:

  1. How accurately can a coding agent rebuild a 3D scene from just one image?
  2. Why do these agents make mistakes while building the scene?
  3. Can the finished 3D scene be used for other vision tasks, such as detecting objects, creating object masks, or estimating depth?

A difficult part of the problem is that one image does not show everything. For example, a picture of a chair may show its front but not its back. The agent must make educated guesses about the hidden parts and the 3D layout.

3. How did the researchers conduct the study?

Turning an image into code

The system uses a coding agent as a kind of digital builder. The process is similar to constructing a LEGO model while repeatedly checking it against a picture:

  1. The agent looks at the reference image.
  2. It writes Blender code.
  3. Blender runs the code and creates a 3D scene.
  4. The agent looks at the scene and its rendered image.
  5. It changes the code to fix mistakes.
  6. It repeats these steps until it has a final scene.

This repeated process is called an iterative code-render-inspect loop. “Iterative” simply means doing something again and again, improving it each time.

Testing the results with LEGO-Bench

The researchers created LEGO-Bench, which contains:

  • 208 images
  • 104 different scenes
  • Indoor and outdoor environments
  • 8 environment types
  • 17 scene themes
  • 443 different 3D assets, such as furniture and other objects

The images were made from carefully designed 3D simulator scenes. This gave the researchers the correct answers, or ground truth, for each scene. Ground truth is like an answer key that tells them the real object locations, shapes, depths, and identities.

LEGO-Bench scores the reconstructed scene in three ways:

Measurement Simple meaning
Validity Does the submitted 3D file work at all?
Reconstruction Are the shapes and positions of visible objects correct?
Appearance Does the rendered image look like the original image?

For example, a scene might be valid because it opens correctly, but still receive a low score if the chairs, walls, or camera are in the wrong places.

Studying mistakes and testing LEGO-Plugin

The researchers watched the agents while they worked. They looked for problems such as:

  • Taking too long to create a usable first scene
  • Making later edits that damage earlier improvements
  • Incorrectly believing that a bad scene is good

They then created LEGO-Plugin, which works like a safety and guidance system. It has three main parts:

  • Enhanced Initialization: Gives the agent a better starting arrangement based on the image
  • Grounded Refinement: Uses measurable evidence about object locations and depth instead of relying only on the agent’s opinion
  • Version Control: Saves good versions and allows the agent to undo harmful changes

This is similar to saving different versions of a school project so that you can return to an earlier version if a new edit makes it worse.

Testing other computer-vision tasks

Finally, the researchers asked whether the reconstructed scene could be used for:

  • Object detection: Drawing boxes around objects
  • Instance segmentation: Coloring the exact pixels belonging to each object
  • Depth estimation: Predicting which objects are closer or farther away

The important idea is that the same 3D scene could answer all three questions without needing a separate model for each task.

4. What did the researchers find?

The agents usually produced working files, but the scenes were not always accurate

The strongest tested system, called GPT-6-astra, achieved overall scores of:

  • 53.4% indoors
  • 39.6% outdoors

This means it often created a usable 3D scene, but the scene was still only a partial match to the original.

A major finding was the difference between making a valid file and rebuilding the scene correctly. The agents had validity scores close to 100%, meaning their files usually opened and contained something renderable. However, their geometry and appearance scores were much lower.

In everyday terms, the agents were usually able to hand in a working project, but the project did not always look like the example they were supposed to copy.

Outdoor and complicated scenes were harder

The agents performed worse on:

  • Outdoor scenes
  • Scenes with many objects
  • More complicated layouts

As the number of objects increased, the scenes became less accurate. Outdoor scenes were especially difficult because they often contain large spaces, more varied lighting, and objects spread across different depths.

More reasoning helped some agents

When GPT-6 agents were given more time and effort to think through the problem, their performance generally improved. For example, GPT-6-astra’s score on one test group increased from 32.3% to 61.8% when its reasoning effort was increased.

However, this improvement was not equally strong for older or weaker models. Simply giving an agent more time does not always solve its problems.

Agents sometimes made their scenes worse

The researchers found three repeated types of failure:

  1. Weak starting scene: The agent took too long to create a useful first version.
  2. Regressive edits: Later changes sometimes damaged parts that were already correct.
  3. Poor self-evaluation: The agent often could not reliably tell whether a new version was better or worse.

The agents were especially poor at judging the scene’s geometry. Their judgments about whether one shape arrangement was better than another were often no better than guessing.

LEGO-Plugin improved every tested model

LEGO-Plugin improved all six tested coding-agent systems.

The biggest improvements happened for weaker models, which gained about 55% to 63% relative improvement. The strongest model improved only slightly because it already avoided many of the common problems.

This suggests that better tools and safeguards can help agents work more carefully without retraining the underlying AI model.

The reconstructed scenes could support several vision tasks

The reconstructed scenes were useful for object detection, segmentation, and depth estimation. However, they were not as accurate as specialized computer-vision systems.

The results were approximately:

  • Object detection: 30.1 compared with 59.9 for a specialized model
  • Segmentation: 14.8 compared with 54.0
  • Depth estimation: 0.155 compared with 0.078, where lower is better

So the reconstructed scenes contained useful information, but they were not yet precise enough to replace expert systems.

5. Why is this research important?

The main idea is that a 3D scene should be more than a picture. If it is represented as an executable program, people can:

  • Open and inspect it
  • Move or replace objects
  • Change the lighting or camera
  • Ask questions about object positions
  • Use it in games, simulations, robotics, or virtual reality

This could eventually allow someone to photograph a room and quickly create an editable digital version of it.

However, the research also shows that this technology is still developing. Current agents are good at producing something that works, but they are not yet good enough at recovering the exact shapes, positions, and appearances shown in an image.

The paper’s tools provide useful steps forward:

  • LEGO-Bench gives researchers a fair way to measure progress.
  • LEGO-Plugin helps agents avoid damaging their own work.
  • LEGO-World shows that one reconstructed scene can support several different vision tasks.

Overall, the paper suggests that coding agents are a promising way to build editable 3D worlds from images, but they still need better understanding, more accurate geometry, and more reliable ways to check their own work.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • Real-world generalization remains untested. LEGO-Bench uses simulator-rendered scenes built from professionally authored assets, so it is unclear how well the approach handles real photographs with sensor noise, imperfect lighting, reflections, motion blur, clutter, occlusions, and unmodeled objects.
  • The benchmark has limited scale and diversity. The evaluation includes 208 images from 104 scenes, eight environments, and 17 themes; broader testing is needed across more architectural styles, geographic regions, object categories, weather conditions, viewpoints, and scene layouts.
  • Benchmark realism may not match natural-image distributions. Although the inputs are described as natural-image-style, the paper does not quantify the domain gap between LEGO-Bench renders and real images or evaluate whether agents exploit simulator-specific visual regularities.
  • The benchmark’s ground truth is view-limited. Reconstruction is scored primarily on surfaces visible from the reference camera, leaving the fidelity of occluded, unseen, and back-facing geometry unresolved.
  • Single-image 3D ambiguity is not explicitly modeled. The evaluation does not distinguish between errors that are objectively unavoidable from one view and errors caused by poor agent decisions, nor does it assess whether multiple geometrically different but image-consistent reconstructions should receive equivalent credit.
  • The evaluation metrics do not fully capture scene usefulness. Visible-surface F1 and pixel-threshold appearance scores may overlook physical plausibility, object functionality, topology quality, semantic correctness, material realism, lighting consistency, and editability of the resulting programs.
  • The appearance metric is potentially brittle. The pixelwise thresholded metric with a fixed δ=30\delta=30 may be sensitive to small camera, lighting, or rendering differences while failing to reflect perceptual similarity; comparisons with learned perceptual, structural, and human-judgment metrics are still needed.
  • Metric weighting is insufficiently justified. The overall score assigns equal weight to Reconstruction and Appearance, but the paper does not establish that this weighting reflects user preferences or downstream task requirements.
  • Object-level scoring may obscure instance and scene-level failures. Averaging F1 over annotated objects may treat large structural errors, small-object errors, duplicate objects, and missing room-level geometry in ways that do not correspond to practical reconstruction quality.
  • Category and correspondence assumptions are underexplored. The protocol relies on private object identities, masks, and correspondences, but it remains unclear how performance changes when objects are ambiguous, semantically mislabeled, duplicated, partially visible, or absent from the asset library.
  • Asset-library dependence is unresolved. The method’s ability to reconstruct scenes containing novel objects, custom furniture, deformable objects, transparent objects, vegetation, signage, or objects unavailable in Fab is not evaluated.
  • The role of asset retrieval versus procedural modeling is unclear. The paper does not isolate whether performance is driven primarily by retrieving suitable assets, generating primitives, composing existing objects, or approximating appearance with camera-facing geometry.
  • Camera and scale recovery are not independently evaluated. Reconstruction errors are reported jointly, so the contributions of camera pose, field of view, global scale, object placement, and object geometry remain difficult to disentangle.
  • Occlusion reasoning is not systematically analyzed. The paper reports that scenes become harder with complexity, but does not quantify how performance depends on occlusion rate, object overlap, visibility fraction, or depth ordering.
  • Outdoor-scene failure modes remain underspecified. Outdoor performance is consistently lower than indoor performance, but the paper does not determine whether this is caused by scale variation, vegetation, sky and lighting, sparse structure, aerial viewpoints, asset mismatch, or camera-estimation errors.
  • The trajectory analysis is narrow. Detailed failure analysis focuses mainly on GPT-6-astra and GPT-5.6-sol, leaving it unclear whether weak initialization, regressive edits, and unreliable self-evaluation generalize across all models, harnesses, and scene types.
  • Causal attribution of failure modes is incomplete. The paper identifies correlations between delayed initialization, score-decreasing edits, and final quality, but does not experimentally isolate how much each factor independently contributes to reconstruction failure.
  • The analysis of self-evaluation uses a limited judging protocol. Pairwise judgments by vision-LLMs may themselves be poorly calibrated or insensitive to 3D errors; human judgments, specialized geometric judges, and task-based validation are needed.
  • The disagreement between deterministic metrics and human quality judgments is unresolved. Although a human-validation study is mentioned, the paper does not establish whether near-chance model judgments reflect genuine inability to assess geometry or limitations of the benchmark metrics.
  • The effect of agent interaction budgets is incomplete. Results vary with reasoning effort, but the paper does not provide systematic scaling curves for execution time, number of edits, tool calls, token cost, rendering cost, or total monetary cost.
  • Test-time reasoning gains are not separated from computational cost. Higher reasoning effort improves GPT-6 performance, but the quality–latency–cost trade-off and the point of diminishing returns remain unspecified.
  • The generality of LEGO-Plugin is uncertain. Plugin evaluation is restricted to the 42-case Office subset, so its effectiveness on outdoor scenes, complex environments, unseen themes, and real images is unknown.
  • The plugin’s components lack complete ablation studies. The paper reports the combined plugin improvement but does not clearly quantify the independent and interaction effects of Enhanced Initialization, Grounded Refinement, and Version Control.
  • The plugin may inherit errors from its auxiliary models. VGGT, SAM 3, and Depth Anything V2 provide the plugin’s cues, but their failure sensitivity, calibration, and contribution to downstream reconstruction quality are not evaluated.
  • The plugin’s residual objectives may not reflect true 3D fidelity. Projected extents and relative depth can be satisfied by view-dependent or geometrically incorrect constructions, leaving open whether the plugin improves genuine scene structure or primarily optimizes image alignment.
  • Version Control may suppress beneficial exploration. Accepting or rolling back revisions based on intermediate scores could prevent agents from making temporary degradations that enable later improvements; this exploration–preservation trade-off is not studied.
  • The plugin’s compatibility with other environments is unverified. The method is implemented around Blender, Blender-MCP, and specific harnesses, so portability to other renderers, simulators, scene representations, or coding-agent interfaces remains open.
  • No learned training-based alternative is compared systematically. LEGO-Plugin is training-free, but the paper does not compare it against supervised fine-tuning, reinforcement learning, preference optimization, trajectory replay, learned scene critics, or specialized image-to-code models.
  • Prompt and harness sensitivity is not fully characterized. Results are reported under native harness behavior and specified configurations, but robustness to prompt wording, tool availability, context limits, memory mechanisms, and alternative execution policies is unclear.
  • Reproducibility is limited by dependence on proprietary agents. The main results rely on GPT-series systems, including models and harness behavior that may be unavailable or change over time, making independent replication and longitudinal comparison difficult.
  • The downstream LEGO-World evaluation has limited coverage. Detection, segmentation, and depth are tested on only 100 randomly sampled images per dataset, which may produce high uncertainty and does not establish performance across the full distributions.
  • Downstream evaluation uses different datasets from LEGO-Bench. The relationship between reconstruction quality on LEGO-Bench and readout quality on COCO, LVIS, and ETH3D is not established, especially because these datasets contain real images with distributions unlike the simulator scenes used for reconstruction evaluation.
  • Equal-confidence AP may underestimate or distort scene-readout performance. Because reconstructed scenes provide no calibrated detection or segmentation confidences, the evaluation imposes an equal-confidence protocol rather than testing whether confidence estimation can be derived from the scene program.
  • Semantic labeling is delegated to the agent without detailed error analysis. The paper does not separate failures in object recognition and category assignment from failures in localization, geometry, visibility, or scene projection.
  • The downstream tasks are restricted to 2D readouts. It remains unexplored whether the reconstructed programs support 3D detection, spatial relationships, navigation, manipulation, embodied interaction, physical simulation, question answering, or counterfactual scene editing.
  • Scene editability is claimed but not evaluated. The paper motivates executable programs as editable representations, yet does not measure whether users or agents can reliably perform targeted edits while preserving unrelated scene properties.
  • Program quality and maintainability are not assessed. The validity metric checks whether artifacts open and render, but not whether the generated code is modular, readable, concise, semantically structured, deterministic, or easy to debug.
  • Physical plausibility is largely unmeasured. Collision and stability checks are used when constructing benchmark scenes, but reconstructed outputs are not comprehensively evaluated for collisions, support relations, gravity, articulation, watertightness, or simulation stability.
  • Unseen-view consistency is unresolved. A scene can achieve good reference-view appearance using billboard-like or view-dependent shortcuts; renders from novel camera viewpoints are needed to test whether the program captures a coherent 3D world.
  • Temporal and multi-view consistency are unexplored. The method operates on one image, so it is unknown whether incorporating additional views, video, depth, or user corrections would substantially improve reconstruction or reduce ambiguity.
  • Failure recovery from execution errors is not systematically studied. Artifact validity is reported, but the paper does not analyze syntax errors, Blender crashes, timeout behavior, dependency failures, corrupted exports, or the agent’s ability to recover from them.
  • The relationship between scene complexity and object count is confounded. Easy, medium, and hard tiers increase visible content while holding some scene factors fixed, but the effects of object count, occlusion, geometric diversity, texture complexity, spatial extent, and lighting complexity are not separately identified.
  • No uncertainty estimates are provided. The system produces a single scene program despite substantial single-view ambiguity; calibrated uncertainty, multiple candidate reconstructions, or confidence-aware downstream queries remain open directions.
  • User-oriented utility is not evaluated. The paper does not measure whether users prefer executable reconstructions over meshes, point maps, or image renderings, nor how much time and expertise are required to inspect, edit, or correct the generated programs.

Practical Applications

Immediate Applications

The paper’s strongest near-term value is not fully faithful 3D reconstruction, but the production of editable, executable scene artifacts, evaluation infrastructure, and more reliable agent workflows.

  • Assisted 3D content creation for games, animation, and visualization
    • A user could provide a single image and obtain an editable Blender scene containing approximate objects, camera placement, geometry, materials, and lighting.
    • Artists could use the generated scene as a blocking or layout starting point rather than modeling an environment from scratch.
    • Potential products include an image-to-Blender plugin, automatic scene-blocking tools, and rapid previsualization workflows for film, advertising, and game development.
    • Dependencies: Current reconstruction quality is substantially better for artifact validity than for geometric or appearance fidelity. Human artists would still need to correct object identity, dimensions, textures, occlusions, and outdoor layouts.
  • Rapid prototyping of virtual environments
    • Architecture, interior design, retail, and real-estate teams could convert reference photographs into approximate 3D mock-ups for early-stage design review.
    • Generated scenes could be used to test alternative furniture arrangements, camera viewpoints, lighting conditions, or material choices.
    • The executable representation is particularly useful because users can edit scene parameters rather than only view a fixed image.
    • Dependencies: The method reconstructs primarily what is visible from one view; hidden structure, metric scale, and physically accurate dimensions are not guaranteed. Safety-critical architectural decisions should not rely on the output without measurement or additional views.
  • Training and evaluation of coding agents for 3D tasks
    • LEGO-Bench can be used by academic and industrial research teams to compare coding agents, Blender agents, scene-generation systems, and multimodal models.
    • Its separate measures for validity, visible-surface geometry, and rendered appearance help distinguish:
    • failure to produce a usable artifact,
    • incorrect spatial or geometric reconstruction,
    • and visual mismatch in rendering.
    • Organizations could add proprietary scenes and assets to create domain-specific evaluation suites for offices, factories, warehouses, campuses, or urban environments.
    • Dependencies: Benchmark extensions require simulator-ready scenes, licensed assets, controlled cameras, and private ground-truth geometry, masks, and depth. Results may also depend on the selected simulator, rendering engine, asset library, and evaluation thresholds.
  • Quality-control infrastructure for agent-generated Blender files
    • The validity checks described in the paper can be integrated into asset-production pipelines to automatically verify that a generated file:
    • opens successfully,
    • contains renderable geometry,
    • has an active camera,
    • exports to a valid .glb,
    • and produces a non-degenerate image.
    • This is immediately useful for batch generation, reducing failures caused by malformed files or missing scene components.
    • Dependencies: Validation confirms technical usability, not semantic correctness. A scene can pass validity while still having poor geometry or appearance.
  • More reliable iterative scene-construction workflows
    • LEGO-Plugin can be deployed as a harness layer around existing coding-agent and Blender-MCP workflows.
    • Its three mechanisms suggest a practical production workflow:
    • use image-based initialization to establish camera and scene layout,
    • use explicit residuals such as projected extent and relative depth for refinement,
    • maintain versioned scene states and roll back edits that reduce quality.
    • This could be used in software pipelines for 3D asset generation, CAD-like prototyping, synthetic-data creation, and visual automation.
    • Dependencies: The reported gains were measured mainly on a 42-case Office subset. Performance may vary across domains, image styles, asset inventories, and agent harnesses. The plugin also relies on auxiliary models such as VGGT, SAM 3, and Depth Anything.
  • Human-in-the-loop 3D authoring
    • The system can function as an interactive assistant rather than a fully autonomous model: a user supplies an image, reviews intermediate renders, accepts or rejects revisions, and edits the generated scene program.
    • Version control is especially suitable for workflows in which users want to preserve a good intermediate scene while exploring alternatives.
    • Potential applications include educational 3D modeling, concept development, museum or cultural visualization, and rapid scene reconstruction for small studios.
    • Dependencies: User review remains important because the agent’s self-evaluation is unreliable, particularly for geometry. The paper shows that automatic or model-based visual judgment should not be treated as a sufficient substitute for deterministic checks or human inspection.
  • Deterministic extraction of approximate visual metadata
    • Once a scene is reconstructed, object projections, instance masks, and camera-space depth can be generated from the same scene artifact.
    • This can support lightweight annotation assistance, rough object inventorying, approximate depth overlays, and automatic generation of scene metadata for search or asset management.
    • The approach may reduce the need to run separate task-specific models for every downstream query when approximate results are acceptable.
    • Dependencies: The reported downstream results are substantially below specialist models: approximately 30.1 box AP, 14.75 mask AP, and 0.155 depth AbsRel. Outputs should therefore be treated as approximate metadata, not reliable perception in high-stakes settings.
  • Educational tools for executable graphics and inverse graphics
    • The image-to-code formulation provides a practical teaching environment for computer graphics, Blender scripting, computer vision, and agentic programming.
    • Students could inspect how natural-language or visual observations are translated into scene primitives, execute the resulting code, and study the effects of camera, geometry, lighting, and materials.
    • LEGO-Bench can support assignments involving error diagnosis, agent comparison, and iterative scene improvement.
    • Dependencies: Educational use requires reproducible software environments and careful handling of third-party assets and model licenses.
  • Policy and procurement benchmarks for generative 3D systems
    • Public agencies, enterprises, and research funders could use the benchmark’s validity–geometry–appearance decomposition when evaluating vendors or models.
    • This is more informative than judging only rendered screenshots, since it tests whether systems produce reusable, inspectable, and exportable artifacts.
    • Procurement workflows could require reporting scene validity, reconstruction quality, failure rates, and performance across indoor and outdoor cases.
    • Dependencies: Benchmark results are not automatically equivalent to real-world photographic performance. Domain-specific validation and legally licensed assets would be required.

Long-Term Applications

These applications depend on substantially improved reconstruction fidelity, broader evaluation, multi-view or sensor integration, and stronger guarantees about geometry and semantics.

  • Digital twins for buildings, factories, warehouses, and infrastructure
    • A mature version of the system could transform photographs, inspection images, or field-worker captures into editable digital-twin scenes.
    • The resulting programs could support asset inventories, spatial queries, maintenance planning, simulation, and visualization of proposed changes.
    • In industrial settings, scene programs could connect visible objects to operational metadata such as equipment IDs, maintenance histories, or sensor streams.
    • Dependencies: Single-image reconstruction is currently insufficient for metric accuracy, hidden surfaces, object identity, and connectivity. Deployment would require multi-view imagery, depth sensors, calibration, semantic verification, persistent object identities, and domain-specific accuracy guarantees.
  • Robotics simulation and sim-to-real training
    • Reconstructed environments could provide approximate simulator scenes for robot navigation, manipulation, perception, and policy testing.
    • A photograph of a room, warehouse, or outdoor area could become an initial simulation environment, with editable object placement and camera configuration.
    • Executable scene programs are well suited to simulation because they can be regenerated, parameterized, and queried.
    • Dependencies: Robotics requires accurate collision geometry, physical materials, object articulation, scale, friction, and dynamics. The current results focus on visible appearance and do not establish physical or interaction fidelity. Errors could produce unsafe or misleading policies.
  • Augmented reality and spatial computing
    • A high-fidelity system could generate editable spatial models from ordinary photographs for AR annotation, remote assistance, virtual staging, and viewpoint synthesis.
    • Users might place virtual objects into reconstructed environments, inspect occlusion relationships, or navigate approximate 3D versions of locations not currently accessible.
    • Dependencies: AR requires accurate camera pose, metric scale, occlusion boundaries, lighting, and often persistent spatial registration. The paper’s single-view scenes and relative-depth outputs do not yet provide these guarantees.
  • Interactive visual search and 3D-aware question answering
    • Executable scenes could become structured representations that answer queries such as:
    • “Which objects are near the desk?”
    • “What is behind the chair?”
    • “How far is the door from the table?”
    • “Show all objects above the floor.”
    • Unlike a fixed image embedding, a scene program can expose object identities, spatial relations, camera coordinates, and editable geometry.
    • Dependencies: Reliable querying requires accurate object segmentation, category labels, spatial relations, and uncertainty estimates. Current reconstruction errors would propagate directly into answers, especially for occluded or visually ambiguous objects.
  • Unified multimodal perception representations
    • The LEGO-World results suggest a long-term architecture in which one executable scene supports detection, segmentation, depth, pose, affordance prediction, and spatial reasoning.
    • Such a representation could reduce duplication among task-specific vision systems and provide consistent predictions across tasks.
    • Future systems could combine learned neural features with explicit geometry and programmatic scene structure.
    • Dependencies: Current readouts trail specialized models substantially. Achieving a useful universal representation will likely require joint optimization, uncertainty-aware scene programs, richer object semantics, and evaluation on diverse real-world imagery.
  • Automated generation of synthetic training data
    • Reconstructed scenes could be edited to generate controlled variants of an observed environment: altered lighting, object placement, camera views, weather, clutter, or material properties.
    • These variants could train or test perception and robotics systems while preserving a link to the original image.
    • LEGO-Bench’s controlled complexity tiers could support systematic studies of how scene difficulty affects model performance.
    • Dependencies: Synthetic data is useful only if geometry, textures, labels, and physical behavior are sufficiently realistic. Reconstruction artifacts could otherwise introduce biased or misleading training examples.
  • Photogrammetry and cultural-heritage documentation
    • A more accurate system could turn historical photographs, archival images, or limited field documentation into editable 3D reconstructions for preservation, education, and public access.
    • Scene programs would allow curators to annotate objects, restore missing elements, and create alternative reconstructions while preserving provenance.
    • Dependencies: Historical images often lack reliable scale, camera metadata, and complete visibility. The system would need explicit uncertainty representation, expert review, provenance tracking, and safeguards against presenting speculative geometry as fact.
  • Urban planning and geospatial visualization
    • Outdoor reconstruction could eventually support rapid modeling of streetscapes, construction sites, disaster areas, and public spaces from photographs.
    • Applications include preliminary planning, visual impact analysis, infrastructure inventories, and change detection.
    • Dependencies: Outdoor performance is lower than indoor performance in the paper, and urban use requires georeferencing, accurate scale, terrain modeling, weather robustness, and coverage of large scenes. A single image cannot reliably recover unseen buildings or infrastructure.
  • Automated scene editing and natural-language design
    • Once scenes are represented as code, users could request edits such as “move the cabinet beside the wall,” “replace the chairs,” or “create a brighter evening version.”
    • Coding agents could produce reproducible, parameterized variants rather than destructive manual edits.
    • This could support design iteration in games, architecture, advertising, and virtual production.
    • Dependencies: Reliable editing requires stable object identities, semantic scene graphs, constraint handling, collision checking, and protection against regressive changes. The paper demonstrates the need for version control because later agent edits can damage previously correct content.
  • Safety-critical inspection and emergency response
    • Mature systems could reconstruct approximate scenes from inspection or disaster photographs to support remote triage, route planning, and prioritization of damaged infrastructure.
    • Explicit geometry and depth could help responders reason about spatial relationships when direct access is difficult.
    • Dependencies: The current accuracy levels are not sufficient for safety-critical decisions. Such applications would require calibrated sensors, multi-view confirmation, confidence intervals, human authorization, robust out-of-distribution testing, and clear separation between measured and inferred scene elements.
  • Policy standards for trustworthy generative 3D
    • The paper’s separation of artifact validity, geometry, and appearance could inform future standards for reporting the quality of generated 3D assets.
    • A mature standard could require provenance, asset-license information, reproducibility, uncertainty, rollback history, and task-specific safety metrics in addition to visual quality.
    • Dependencies: Metrics must be validated on real photographs and operational tasks, not only simulator-rendered scenes. Standards would also need to account for privacy, copyrighted assets, model-generated errors, and the risks of treating inferred 3D structure as ground truth.

Glossary

  • Artifact validity: Whether a submitted digital artifact is usable, well-formed, and evaluable. “These metrics distinguish failure to deliver an evaluable scene from failures of geometric and visual fidelity.”
  • Appearance metric: A measure of visual similarity between a rendered reconstruction and a reference image. “Appearance evaluates visual similarity between the reconstructed scene and the reference image.”
  • Articulated asset: A 3D object composed of parts connected by joints that permit motion. “Other work targets editable indoor scenes from RGB-D scans or simulation-ready articulated assets.”
  • Camera coordinates: A coordinate system whose origin and axes are defined relative to the camera viewpoint. “Both point sets are projected onto the image plane and partitioned into scored objects by the private instance masks, then compared directly in camera coordinates without alignment or rescaling.”
  • Camera-space depth: The distance of a point from the camera, expressed in the camera’s coordinate system. “camera-space depth renderings provide relative depth.”
  • Coding agent: An AI system that writes, executes, inspects, and revises code to accomplish a task. “Coding agents combine LLMs with tools for code editing, execution, and iterative verification.”
  • Collision and stability checks: Tests that determine whether simulated objects intersect improperly or remain physically stable. “Candidate scenes undergo collision and stability checks and human review before inclusion.”
  • Depth-scaled tolerance: An error threshold that changes according to an object’s distance from the camera. “precision is the fraction of submitted points within a depth-scaled tolerance τ(g)=0.05 z(g)\tau(g)=0.05\,z(g) of the nearest reference point gg”
  • Deterministic readout: A fixed procedure that derives a task-specific prediction from an existing representation without additional learned inference. “object detection, instance segmentation, and relative depth can be obtained as deterministic readouts of the same frozen scene”
  • Differentiable rendering: Rendering formulated so that image outputs can be differentiated with respect to scene parameters, enabling optimization. “Rather than producing PP in one shot, the agent alternates between editing code, executing it, and inspecting the resulting scene and renderings”
  • Executable artifact: A saved computational object that can be run to produce or inspect a result. “The agent submits the final scene, its export, and a rendered view, which together form an executable artifact that can be evaluated for fidelity and queried for downstream perception.”
  • Executable scene program: A program whose execution constructs a manipulable 3D scene. “an executable scene program makes objects, geometry, layout, and camera explicit”
  • F1 score: The harmonic mean of precision and recall, commonly used to evaluate retrieval or detection quality. “The Reconstruction score (F@5\%) averages object-level F1 scores”
  • Field of view: The angular extent of a scene visible through a camera. “including image dimensions, horizontal field of view, a category taxonomy, and an output schema”
  • Fidelity: The degree to which a reconstruction accurately matches the reference scene or image. “We study this setting as LEGO-Anything, an Image-to-Code (Image2Code) framework in which a general-purpose coding agent writes and executes Blender programs, inspects the evolving scene and its renderings, and revises the construction to match a single reference image.”
  • Forward depth: The distance of a point along the camera’s viewing direction. “where z(g)z(g) is its forward depth”
  • Ground truth: Authoritative reference data used to evaluate predictions. “Because every case is rendered from a fully specified simulator scene, ground truth comes for free”
  • Harness: Software infrastructure that organizes an agent’s tools, execution context, and feedback. “Their harnesses organize tool access, execution context, and environmental feedback”
  • Image-to-Code: A task in which an image is converted into executable code that reconstructs or represents its contents. “LEGO-Anything formulates single-image scene reconstruction as an Image-to-Code (Image2Code) problem.”
  • Instance mask: A pixel-level mask identifying the region belonging to one particular object instance. “retains private geometry, object identities, camera parameters, depth, and instance masks for automatic evaluation”
  • Inverse graphics: The process of inferring scene properties, such as geometry, materials, lighting, and camera parameters, from images. “SEIG reconstructs images as editable Blender programs through staged executable inverse graphics”
  • MCP tool: A tool exposed through the Model Context Protocol for allowing an AI agent to interact with external software or data. “LEGO-Plugin is a training-free control layer exposed as MCP tools, workflow skills, and runtime hooks”
  • Natural-image-style observation: A rendered image designed to resemble the visual appearance and diversity of a real photograph. “provides realistic, natural-image-style observations”
  • Non-degenerate image: An image that contains valid, meaningful pixel data rather than being empty, constant, or otherwise unusable. “final.png must be a parseable, non-degenerate image.”
  • Parametric part: A component whose geometry is controlled by numerical parameters. “P3D-Bench targets parametric parts and assemblies”
  • Perceptual sufficiency: The extent to which a representation preserves enough information to support visual recognition or related tasks. “the perceptual sufficiency of frozen reconstructed scenes.”
  • Point map: A representation associating image pixels with points in a 3D coordinate system. “Compared with meshes, point maps, or object sets”
  • Precision and recall: Metrics measuring, respectively, the correctness of predicted elements and the completeness of retrieved reference elements. “For each object oo, precision is the fraction of submitted points within a depth-scaled tolerance”
  • Procedural generation: The algorithmic creation of content from rules, parameters, or reusable procedures. “Procedural systems generate natural, indoor, and urban environments from structured specifications and reusable assets”
  • Relative depth: Depth information that expresses ordering or proportional distance rather than absolute metric distance. “We next ask whether reconstructed scenes can serve as representations of natural images.”
  • Rendered appearance: The visual characteristics of a scene as produced by a rendering process, including geometry, materials, lighting, and camera settings. “LEGO-Bench separately scores artifact validity, visible-surface geometry, and rendered appearance”
  • Residual: The measured discrepancy between an observed or reconstructed value and a target value. “then compares the executed scene against these targets through explicit residuals”
  • RGB-D scan: A scan containing both color data and per-pixel depth measurements. “Other work targets editable indoor scenes from RGB-D scans”
  • Scene graph: A structured representation of a scene as objects and their relationships. “structured representations of natural images”
  • Scene program: Executable code that specifies the objects, geometry, layout, camera, and other properties of a 3D scene. “We present LEGO-Anything, an Image-to-Code framework in which a coding agent builds such a program”
  • Scene reconstruction: The recovery of a 3D scene’s objects, geometry, arrangement, and appearance from visual or other observations. “A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program”
  • Segmentation: The assignment of pixels to objects, regions, or semantic categories in an image. “The resulting scene program can also serve as a representation that can be queried directly for detection, segmentation, and depth estimation.”
  • Simulator-grounded benchmark: An evaluation dataset whose reference scenes and annotations originate from a simulator with known underlying state. “To evaluate how well such agents recover scenes end to end, we introduce LEGO-Bench, a simulator-grounded benchmark”
  • Single-image reconstruction: The recovery of a 3D scene from one two-dimensional image. “We study this setting as LEGO-Anything, an Image-to-Code (Image2Code) framework”
  • Spatial layout: The positions and arrangement of objects within a scene. “an agent must reconstruct the scene as a complete, executable 3D artifact, jointly recovering its contents, geometry, spatial layout, and appearance from one view.”
  • Structured representation: An organized representation that explicitly encodes entities and their properties or relationships. “We next ask whether reconstructed scenes are precise enough to serve as structured representations of natural images”
  • Test-time scaling: Improving model performance by allocating additional computation or reasoning during inference rather than training. “We run a test-time scaling experiment on a fixed 42-case Office subset”
  • Trajectory analysis: Examination of the sequence of intermediate states, actions, and outcomes produced during an agent’s execution. “To diagnose why end-to-end scene reconstruction still falls short, we analyze agent construction trajectories”
  • Training-free: Operating without updating model parameters using task-specific training data. “These findings motivate LEGO-Plugin, a training-free harness plugin”
  • Visible-surface geometry: The shape and spatial arrangement of scene surfaces that can be seen from the reference camera. “LEGO-Bench separately scores artifact validity, visible-surface geometry, and rendered appearance”
  • Volumetric or geometric fidelity: Accuracy of a reconstruction’s 3D structure relative to the reference. “yet substantial gaps remain between delivering valid scene artifacts and faithfully recovering scene geometry and appearance.”

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 2 tweets with 420 likes about this paper.