LEGO-Anything: Coding Agents for 3D Scene Reconstruction
This presentation examines a novel approach to single-image 3D reconstruction that treats the problem as executable programming rather than direct geometric prediction. LEGO-Anything uses general-purpose coding agents to iteratively write and refine Blender scene programs from a single RGB image. The work introduces LEGO-Bench, a simulator-grounded benchmark with precise ground truth, and demonstrates that while coding agents reliably produce valid executable artifacts, geometric and appearance fidelity remain substantially lower. The presentation explores construction trajectories, identifies key failure modes including regressive edits and unreliable self-evaluation, and shows how training-free control modules can improve weaker models by over 60 percent.Script
What if reconstructing a 3D scene from a photograph meant writing a program, not predicting a mesh? LEGO-Anything treats single-image reconstruction as an executable coding problem, where agents iteratively author Blender programs that can be opened, edited, and queried long after construction.
The agent doesn't generate geometry in one pass. Instead, it establishes a coarse scene, renders it, inspects the result against the reference, and applies targeted edits. Objects, cameras, materials, and lighting remain explicit in the Blender state, making the scene both the reconstruction output and the operational interface for refinement.
Construction trajectories reveal three failure modes. Weak initialization delays the first evaluable scene. Regressive edits damage geometry or camera settings even after reaching higher quality. And unreliable self-evaluation means agents cannot consistently judge whether an edit improved the reconstruction, especially for depth and occlusion.
LEGO-Plugin addresses these failures without retraining the agent. It adds reference-grounded initialization using camera estimates and layout cues, grounded refinement with measured depth and extent constraints, and version control that rolls back destructive edits. Weaker models improved by over 60 percent, while the strongest model gained only 2 percent, meaning the plugin compensates for procedural weaknesses already avoided by capable agents.
The same reconstructed scene can support detection, segmentation, and depth through deterministic readouts. LEGO-World demonstrates that executable programs already contain enough structure for downstream tasks, reaching 30 box average precision on COCO and an absolute relative error of 0.155 on depth, though well below specialized vision models.
Coding agents reliably deliver valid executable artifacts, but validity substantially exceeds reconstruction fidelity. The central finding is the separation between producing a scene that can be opened and rendered versus recovering geometry and appearance faithfully. Visit EmergentMind.com to explore this work further and create your own research videos.