LEGO-Anything Framework
- LEGO-Anything is an Image-to-Code framework designed for reconstructing three-dimensional scenes from a single image, transforming them into executable and editable Blender code.
- The framework evaluates reconstruction accuracy using LEGO-Bench, a simulator-grounded benchmark, and introduces LEGO-Plugin for enhanced and iterative reconstruction capabilities and failure mode correction
- Executable scene programs can be regenerated, edited, validated, exported, queried for further vision tasks
- follow_up_questions_includes': [What are the practical applications of LEGO-Anything beyond image reconstruction?
- How does LEGO-Anything handle the single-image reconstruction limitations?
- Compare and contrast LEGO-Anything with other Image-to-Code and Image-to-3D techniques.
- Could you describe the main improvements brought by the LEGO-Plugin?
- Find recent papers about code-based iterative reconstruction.
LEGO-Anything is an Image-to-Code framework for reconstructing three-dimensional scenes from a single image as explicit, executable, inspectable, and editable scene programs. Rather than producing only a mesh, point map, rendering, or fixed latent representation, it uses a coding agent that iteratively writes and executes Blender code, inspects scene states and renderings, and revises the program. The resulting artifact can be regenerated, edited, validated, exported, and queried for downstream vision tasks. The framework is evaluated through LEGO-Bench, a simulator-grounded benchmark for single-image scene reconstruction, and extended with LEGO-Plugin, a training-free harness for initialization, grounded refinement, and version-controlled editing (Li et al., 28 Sep 2026).
1. Conceptual foundations and scope
The central representation is a program–scene pair:
where is an input RGB image, is an executable scene program, and is the Blender scene produced by executing that program. The program can encode object identities, geometry, transformations, camera parameters, materials, lighting, spatial layout, support relations, and potentially articulation.
The framework treats reconstruction as an iterative inverse-graphics and programming problem. A coding agent interacts with an environment , implemented using Blender:
The agent is expected to infer visible objects and their arrangement, write or modify Blender code, execute it, inspect the resulting scene and camera rendering, compare the rendering with the reference image, and revise the program. The final submission contains scene.blend, scene.glb, and final.png.
The term “LEGO-Anything” is therefore not restricted to physical LEGO bricks. In this framework, “LEGO” denotes a compositional construction process in which a complex artifact is assembled from explicit, inspectable components. This interpretation is consistent with earlier LEGO-based research that treats modular elements as programmable, physical, scientific, or generative building blocks. Bricklayer generates voxelized or “legoized” artifacts through SML programs and exports them to tools such as LEGO Digital Designer, LDraw, Minecraft, Brickr, and STL viewers (Winter et al., 2016). Image2Lego converts images into voxelized models and then into LEGO assemblies with LDraw-based instructions (Lennon et al., 2021). The present framework extends the construction metaphor to executable Blender scenes and agentic iterative reconstruction.
LEGO-Anything does not claim that executable scene programs are already geometrically faithful, physically stable, or universally applicable. Its results show a pronounced distinction between producing a valid scene artifact and recovering accurate geometry and appearance. This separation is central to both its benchmark design and its analysis of current coding agents.
2. Agent–environment interaction
A reconstruction is represented as a construction trajectory:
where is the current program, is the executed scene, and is the observation available to the agent, including scene state and renderings.
The interaction loop consists of several stages:
- Image interpretation: the agent identifies the room or environment, visible objects, approximate layout, scale relationships, camera, and depth ordering.
- Program construction: the agent writes Blender code that creates primitives or imported assets and specifies object placement, camera, materials, lighting, and rendering parameters.
- Execution: the program is run in Blender.
- Inspection: the agent examines the scene and renders the current camera view.
- Revision: the agent adds missing objects or adjusts geometry, camera, placement, materials, and lighting.
- Artifact export: the resulting Blender scene, GLB export, and PNG rendering are submitted.
The principal experimental implementation uses Blender 5.0.1, Blender-MCP, the Harbor execution framework, and the Codex harness. The agent can preserve scene state through code, making incremental modifications possible. This persistence distinguishes the framework from one-shot image-to-mesh systems.
The iterative formulation introduces failure modes that are less visible in one-shot evaluation. A scene can begin from a weak initialization, later edits can damage previously correct geometry or camera parameters, and the agent can incorrectly judge whether an edit improved the reconstruction. The paper identifies these as weak scene initialization, regressive edits, and unreliable self-evaluation.
The framework is also compatible with deterministic scene queries. Once a scene has been reconstructed, object geometry can be projected or rendered to derive detections, instance masks, and depth. The scene program thus functions both as an output representation and as an intermediate representation for downstream tasks.
3. LEGO-Bench
LEGO-Bench is a simulator-grounded benchmark for evaluating executable single-image 3D scene reconstruction. It contains 208 RGB inputs corresponding to 104 logical scenes, eight environments, 17 themes, and 443 registered assets. The benchmark includes indoor and outdoor scenes, as well as an additional NYC aerial or Bird’s-Eye stress split.
The scenes are generated in LychSim using permitted assets from Fab. They are checked for collision and stability and reviewed by humans. Within matched scene families, difficulty is nested:
0
Architecture, camera, lighting, materials, and shared object properties are held fixed while visible content is added. Difficulty therefore increases through scene-content additions rather than an arbitrary global object-count threshold.
The evaluator retains private reference geometry, depth, instance masks, object correspondences, and scene transforms. Agents receive only the RGB image and public metadata such as image dimensions, field of view, category taxonomy, and output requirements. The benchmark’s simulator-grounded construction provides realistic-looking images together with exact geometry and camera ground truth.
Coordinate conversion from LychSim’s left-handed centimeter coordinates to the evaluator’s right-handed metric coordinates is specified as:
1
Evaluation dimensions
LEGO-Bench separates three properties that are often conflated in reconstruction evaluation:
- artifact validity;
- visible-surface geometry;
- rendered appearance.
For a submission 2, validity 3 is determined by whether scene.blend opens and contains renderable geometry and an active camera, whether scene.glb is well formed and nonempty, whether final.png is decodable and nondegenerate, and whether headline evaluation can be completed. Invalid submissions receive:
4
A timed-out run can remain valid if it leaves usable artifacts, while a clean process exit does not itself guarantee validity.
Reconstruction 5 measures visible-surface agreement from the reference camera. The evaluator extracts reference-visible surfaces by back-projecting private depth, rasterizes the submitted scene under its active camera, extracts visible submitted surfaces, projects both point sets into the image plane, assigns points to scored objects using private instance masks, and compares points directly in camera coordinates. There is no registration or rescaling, so camera, translation, rotation, scale, and object-placement errors are penalized directly.
For an object 6, recall is:
7
where 8 is the reference-visible point set and 9 is the predicted visible point set. Precision is computed in the opposite direction using the nearest reference point. The principal tolerance is depth-relative:
0
where 1 is the positive camera-forward depth of the reference point. This metric is reported as F@5%. Auxiliary F@2% and F@10% variants use 2 for 3.
The object-level reconstruction score is the harmonic mean of precision and recall, and the trial-level score 4 is the macro-average over scored objects. Missing surfaces reduce recall, while extra geometry inside an instance mask reduces precision. The metric evaluates visible geometry from one camera rather than hidden geometry.
Appearance 5 is computed by re-rendering the submitted Blender scene with the evaluator’s standardized renderer and resolution. The evaluator preserves the submitted camera, geometry, materials, lights, and color settings, and compares the resulting image against the reference in 8-bit sRGB. With RGB threshold 6:
7
Appearance consequently reflects geometry, camera, materials, lighting, shadows, and color settings. It is not an independent geometric measure and is not a perceptual similarity metric.
For valid cases, the overall score is:
8
The benchmark mean is:
9
Invalid cases remain in the denominator and receive zero.
4. Results and comparative performance
The principal coding-agent results are reported as mean and standard deviation over three runs. GPT-6-astra achieves the strongest overall performance among the evaluated agents, with an indoor score of 0 and an outdoor score of 1. Its indoor validity, reconstruction, and appearance scores are respectively 2, 3, and 4; its outdoor scores are 5, 6, and 7.
| Model | Indoor validity | Indoor reconstruction | Indoor appearance | Indoor overall | Outdoor overall |
|---|---|---|---|---|---|
| GPT-6-astra | 100.0 | 52.4 | 54.4 | 53.4 | 39.6 |
| GPT-6-sol | 99.4 | 22.2 | 42.5 | 32.3 | 24.2 |
| GPT-6-luna | 99.7 | 15.8 | 30.5 | 23.2 | 17.3 |
| GPT-5.6-sol | 97.8 | 9.8 | 19.8 | 14.8 | 15.3 |
| GPT-5.6-terra | 99.7 | 9.5 | 21.3 | 15.4 | 11.5 |
| GPT-5.6-luna | 99.1 | 10.2 | 18.0 | 14.1 | 12.1 |
The main empirical pattern is that artifact validity is close to 100% for nearly all GPT configurations, whereas scene fidelity varies considerably. This indicates that executable artifact delivery is substantially easier than accurate reconstruction of camera, geometry, object placement, materials, and lighting.
Outdoor reconstruction is consistently harder than indoor reconstruction. For GPT-6-astra, reconstruction decreases from 8 indoors to 9 outdoors. The Bird’s-Eye stress split produces a validity score of 0, an F@2% reconstruction score of 1, an appearance score of 2, and an overall score of 3.
The task-specific and conventional baselines illustrate the multidimensional nature of the benchmark. Gen3DSR achieves an indoor reconstruction score of 4, higher than GPT-6-astra’s 5, but has an appearance score of zero under the evaluation setup and therefore a lower combined score. GPT-6-astra thus has the strongest overall performance, not necessarily the highest isolated score on every dimension.
Increasing scene complexity primarily reduces fidelity rather than executability. Across GPT-6 and GPT-5.6 configurations, validity remains approximately 6 on Easy, Medium, and Hard scenes, while overall score decreases from 7 on Easy to 8 on Medium and 9 on Hard.
Test-time reasoning effort improves stronger agents on the fixed 42-case Office subset. Increasing native reasoning effort from Low to XHigh changes GPT-6-astra’s score from 0 to 1, GPT-6-sol’s from 2 to 3, and GPT-6-luna’s from 4 to 5. GPT-5.6 models show weak or non-monotonic changes, indicating that additional reasoning tokens do not uniformly resolve the reconstruction problem.
5. LEGO-Plugin and controlled iterative construction
LEGO-Plugin is a training-free harness control layer designed to address the three recurrent failure modes identified through trajectory analysis. It does not replace the coding agent, introduce a new planner, or access private benchmark ground truth. Instead, it adds tools, workflow skills, and runtime hooks around the existing agent and Blender-MCP interaction.
Its modules are:
- Enhanced Initialization
- Grounded Refinement
- Version Control
Enhanced Initialization
Enhanced Initialization uses VGGT to estimate a scene frame, camera, gravity-aligned Z-up Manhattan structure, and visible-object layout cues. It constructs an initial room or landscape and camera jointly, then creates a construction plan containing, for each visible object, its normalized centroid, projected extent, relative and ordinal depth, and observability flags.
The module avoids claiming reliable absolute metric depth or absolute object size from a single image. Instead, it provides a reference-aligned starting state and image-relative layout information.
Grounded Refinement
Grounded Refinement replaces unconstrained visual self-critique with explicit image-grounded measurements. It uses SAM 3 for object regions, Depth Anything V2 for relative depth, and residuals based on projected extent, centroid, depth ordering, and rendered luminance.
Validation can check object existence, transforms and dimensions, support and containment, collision, facing direction, repeated layout, camera visibility, projected centroid and extent, relative depth, depth order, and rendered luminance. The output is a structured pass/fail/unknown report rather than a global judgment that a rendering “looks better.”
Version Control
Each bounded edit is treated as a transaction. Before an edit, the plugin snapshots transforms, meshes, materials, hierarchy, visibility, active camera, and render settings. The edit declares editable entities, protected entities, permitted change types, and acceptance requirements.
After execution, the plugin rejects undeclared or protected changes, checks hard requirements, commits accepted edits, and repairs or rolls back failed edits. The evaluation configuration permits at most eight transactions and two rollbacks per case. Reports become stale after subsequent scene changes, and runtime hooks prevent finalization when validations are unresolved, stale, or renders have not been reviewed.
On the 42-case Office subset, LEGO-Plugin improves all six evaluated models. GPT-5.6 models receive approximately 55–63% relative improvement, GPT-6-luna improves by 27.5%, GPT-6-sol by 12.1%, and GPT-6-astra by 2.1%. The largest reported relative gain is up to 62.7%. The strongest agent benefits least because it already avoids many initialization and regression failures.
Two reported GPT-5.6-sol examples improve from 6 to 7 and from 8 to 9. These results suggest that controlled initialization, explicit residuals, and transactional editing are particularly useful for agents whose unconstrained trajectories are prone to regressions.
6. LEGO-World, related systems, and open problems
LEGO-World evaluates whether an executable reconstructed scene can serve as a general-purpose visual representation. Scenes reconstructed by GPT-6-astra are queried deterministically for object detection, instance segmentation, and relative depth. The pipeline is:
0
where 1 is a deterministic scene readout for task 2.
Bounding boxes are obtained by projecting visible objects, instance masks by projecting object geometry, and depth by rendering the scene. On 100-image tracks from COCO val2017, LVIS validation, and ETH3D, LEGO-Anything obtains an AP of 30.14 for COCO bounding-box detection, an AP of 14.75 for LVIS instance segmentation, and an AbsRel of 0.1554 for ETH3D relative depth. The corresponding specialist baselines are DINO with AP 59.88, SAM 3 with AP 53.96, and Depth Anything 3 with AbsRel 0.0783.
The scene representation therefore supports all three downstream tasks without task-specific training, but remains substantially below specialized vision systems. Its predictions do not provide calibrated confidence: every predicted box or mask is assigned confidence 1.0 with deterministic tie ordering. Coverage is 94 scene outputs for COCO boxes, 94 for LVIS masks, and 99 for ETH3D depth.
Earlier LEGO-oriented systems clarify the broader design space. Image2Lego reconstructs voxelized 3D objects from single images and converts them into LEGO bricks, with multiple output resolutions and LDraw-based instructions (Lennon et al., 2021). MobileBrick uses LEGO models with known geometry as physical objects for mobile RGB-D reconstruction, providing 153 models and precise geometry annotations (Li et al., 2023). Bricklayer generates mathematical, artistic, and voxelized structures from SML programs and exports them to multiple visualization and construction back ends (Winter et al., 2016). Deep Generative Models of LEGO Graphs represent assemblies as directed labeled graphs and generate bricks and connections autoregressively (Thompson et al., 2020). Break-and-Make and InstructioNet address interactive disassembly, explicit visual instruction memory, and long-horizon reconstruction (Walsman et al., 2022, Walsman et al., 2024).
Physical construction introduces additional requirements not solved by the Blender-based framework. StableLego analyzes rigid-body equilibrium, friction capacity, weak points, and stability of LEGO layouts (Liu et al., 2024). Robotic assembly systems use human demonstrations, digital twins, custom end-effectors, force sensing, and robot-specific verification (Liu et al., 2023, Liu et al., 2023). A mathematical construction result shows that every finite face-connected union of lattice cubes can be enlarged by a factor of four and realized as a connected assembly using only upright 3 bricks, although the result does not model real studs, stability, collisions, colors, or arbitrary geometry (Gans, 5 Sep 2026).
The principal limitations of LEGO-Anything are therefore multidimensional:
- Single-view ambiguity: hidden geometry, depth, and object identity are underdetermined by one image.
- Scene fidelity: executable artifacts are frequently valid even when their geometry and appearance are inaccurate.
- Regressive editing: later changes can damage earlier improvements.
- Self-evaluation: agent judgments align poorly with deterministic geometric evaluation, particularly for reconstruction.
- Viewpoint dependence: image-grounded comparisons are sensitive to camera changes.
- Simulator and asset bias: LEGO-Bench is simulator-grounded and depends on its asset library, rendering process, and scene priors.
- Physical feasibility: Blender validity does not guarantee structural stability, legal LEGO connectivity, constructability, or robot reachability.
- Articulation: the pilot articulation score is zero for the evaluated frozen Office submission, indicating that articulated scene understanding remains unresolved.
- Downstream precision: deterministic scene queries are useful but substantially weaker than specialist detection, segmentation, and depth systems.
- Training and execution cost: large-scale iterative agents require substantial reasoning, environment interaction, and computational resources.
LEGO-Anything’s main methodological contribution is the treatment of scene reconstruction as executable program synthesis under visual feedback. Its benchmark distinguishes validity from geometry and appearance; its trajectory analysis exposes initialization, regression, and self-evaluation failures; and LEGO-Plugin introduces grounded, transactional iteration without retraining the underlying agents. The resulting representation is promising because it is editable, inspectable, regenerable, and queryable. Its current limitation is equally clear: producing an executable scene is substantially easier than producing one that faithfully recovers the geometry, appearance, hidden structure, articulation, and physical feasibility of the scene depicted in the input image.