- The paper introduces SEIG, a staged generator–verifier framework that decomposes single-image reconstruction into geometry, materials, composition, and lighting, enabling a vision-language model to produce editable Blender code.
- SEIG outperforms monolithic VIGA baselines on five of six metrics across NeRF Synthetic and Edit3D benchmarks, including DINO similarity scores of 0.7188 and 0.6293.
- The resulting structured scenes support relighting, object-level editing, and physics simulation, although greedy stage ordering, initialization errors, inference cost, and single-view ambiguity remain important limitations.
Overview
"Thinking in Blender: Staged Executable Inverse Graphics with Vision-LLMs" (2606.02580) addresses a long-standing inverse graphics problem: reconstructing an image as an editable, physically grounded 3D scene. The authors introduce SEIG (Staged Executable Inverse Graphics), an agentic framework that takes a single reference image and produces executable Blender code whose render matches the input. The central claim is that pretrained vision-LLMs (VLMs) already encode sufficient latent priors about geometry, materials, composition, and lighting to perform executable inverse graphics—provided the reconstruction is decomposed into sequential, individually verifiable stages—and that this decomposition matters more than access to specialized 2D or 3D foundation models.
The work positions itself against two lines of prior art. Neural scene representations such as NeRF and 3D Gaussian Splatting reconstruct scenes effectively but entangle geometry, reflectance, and illumination in latent representations that are not directly editable as structured programs (2606.02580). Agentic approaches such as VIGA formulate reconstruction as a monolithic write–render–compare–revise loop, which the authors argue leaves tightly coupled scene factors to be solved simultaneously and creates a search space in which errors in one factor obscure corrections to another.
Method
SEIG uses Blender as its reconstruction engine because of its unified Python API for scene editing and rendering. The pipeline mirrors the workflow of professional 3D artists through four sequential stages—geometry, materials, composition, and lighting—each formulated as an agentic function that depends only on earlier stage outputs and maintains stage-specific context.
Initialization. The VLM first decomposes the scene into a hierarchical scene graph with a root node for the environment and object/part nodes storing visual descriptions, approximate geometry, material appearance, spatial relations, and a per-node Blender reconstruction strategy. A coarse primitive-based scaffold is then instantiated with stable object names shared across all later stages. Because initialization errors propagate, the authors sample multiple independent scene graphs and scaffolds and apply a rollout selector favoring complete object coverage and plausible structure.
Staged refinement. The geometry stage performs local shape edits, geometric transforms, and structural edits per object, supported by tools for alternative-viewpoint rendering, object isolation, and edit reversion. The material stage replaces placeholder textures with PBR materials via shader nodes, restricted by a material-only tool so it cannot alter earlier results. The composition stage adjusts object transforms and optionally the target-view camera, but not geometry or materials. The lighting stage infers light direction, shadow softness, color temperature, and exposure cues while keeping all other factors fixed, with instructions toward conservative edits and reversion of over- or under-exposed changes.
Generator–verifier loops. Each stage runs a multi-round loop in which the generator writes code via tool calls and renders the result; a verifier scoped to the active factor compares render against reference and returns an explicit approval checklist—a concrete todo list injected into the generator's context. Once the checklist is satisfied, the verifier must approve; if a stage-specific round budget is exhausted (five rounds for geometry, three each for materials and composition, two for lighting), the verifier selects the best attempt to advance. This checklist mechanism addresses noisy free-form critiques that would otherwise give the generator inconsistent targets.
Throughout, the same pretrained VLM (Claude Opus 4.7 via API) serves as both generator and verifier at every stage, with no fine-tuning, prompt-tuning, or task-specific supervision—so observed quality differences are attributable to harness design rather than model capability.
Quantitative evaluation
Evaluation covers 35 synthetic scenes: five rendered views from seven NeRF synthetic scenes (the metallic-spheres "materials" scene is excluded due to specular reflections) and 13 object-centric scenes from VoxHammer ("Edit3D"). Six metrics span pixel-level (PSNR, SSIM), learned perceptual (LPIPS, DreamSim), and semantic similarity (DINOv2 ViT-L/14 CLS-token cosine, CLIP ViT-B/32 embedding cosine). When reference meshes exist, reconstructions are registered via NDP and ICP (keeping the smaller Chamfer distance) before rendering from the reference camera, avoiding conflation of reconstruction quality with camera estimation error.
SEIG achieves the best score on five of six metrics on both benchmarks:
| Method |
PSNR ↑ |
SSIM ↑ |
LPIPS ↓ |
DreamSim ↓ |
DINO ↑ |
CLIP ↑ |
| VIGA VLM-only (NeRF syn.) |
12.33 |
0.7122 |
0.3506 |
0.3693 |
0.6221 |
0.8451 |
| VIGA full (NeRF syn.) |
11.18 |
0.6647 |
0.3944 |
0.3624 |
0.5545 |
0.7986 |
| SEIG (NeRF syn.) |
13.58 |
0.6881 |
0.3493 |
0.3021 |
0.7188 |
0.8830 |
| VIGA VLM-only (Edit3D) |
11.52 |
0.6776 |
0.3931 |
0.3847 |
0.5606 |
0.8366 |
| VIGA full (Edit3D) |
12.48 |
0.6743 |
0.4466 |
0.4441 |
0.4832 |
0.7883 |
| SEIG (Edit3D) |
12.65 |
0.6737 |
0.3823 |
0.3433 |
0.6293 |
0.8446 |
Two comparisons carry distinct implications. First, SEIG outperforms full VIGA despite using no specialist foundation models (SAM, SAM-3D), indicating the gains stem from harness design rather than tool access. Second, SEIG outperforms the VLM-only ablation, isolating the contribution of staged decomposition itself. Notably, full VIGA underperforms its own VLM-only ablation on most metrics—the authors attribute this to the VLM agent frequently overwriting textures of SAM-3D-generated objects, producing fragmented, mis-colored meshes. This is consistent with findings from BlenderGym and IR3D-Bench that visual precision, not tool orchestration, is the dominant bottleneck in current agentic 3D pipelines.
Qualitative analysis
Qualitatively, SEIG recovers global geometry, materials, and composition more reliably than either VIGA configuration. Two failure-mode analyses are instructive. On a humanoid character, full VIGA exhibits the Janus artifact—frontal facial features duplicated onto the back of the head—a characteristic single-view lifting failure of mesh-prior generators like SAM-3D; SEIG avoids this by composing the figure from primitives rather than lifting a single-view mesh. Conversely, on a heavily occluded bread-basket scene, SEIG produces rounded loaves rather than bread sticks—an interpretation equally consistent with the visible silhouette. The authors present this candidly as evidence of the inherent underdetermination of single-view reconstruction, while noting that both VIGA variants fail even to recover a coherent basket structure, indicating a compositional-discipline failure rather than view ambiguity.
A property distinguishing the staged pipeline from monolithic alternatives is that every intermediate scene is itself a coherent, editable Blender program, since each stage commits its output before the next begins. This makes intermediate decisions inspectable and selectively reusable at any stage boundary.
Downstream applications
Because geometry, materials, lights, and objects are committed as separate named stage outputs, reconstructed scenes support standard graphics operations without retraining or post-processing. Relighting reduces to adding or reconfiguring light sources in Blender; per-object editing (part duplication, texture editing, shape manipulation, rearrangement) operates directly on scene-graph nodes; and physics simulation runs in Blender's built-in engine without remeshing or watertighting, since the pipeline yields object-decomposed meshes rather than a fused implicit representation requiring conversion into discrete physical entities. Rigid-body tabletop dynamics and soft-body cushion deformation are demonstrated directly on recovered scenes.
Limitations
The authors identify two principal limitations. First, the greedy staged formulation means early-stage errors propagate: inaccurate geometry can constrain subsequent material, lighting, and compositional reasoning, creating local minima from which later stages cannot easily recover. Global refinement passes revisiting earlier factors could mitigate this but at substantially increased computational cost. Second, repeated inference calls across multiple generator–verifier stages incur substantially higher runtime and API cost than single-pass generation pipelines. An additional implicit limitation is the reliance on a rollout selector at initialization; since initialization determines the scaffold used by all later stages, poor coverage or bad object association remains difficult to recover from local refinement alone.
Conclusion
SEIG demonstrates that a single off-the-shelf VLM, given a carefully staged generator–verifier harness, can reconstruct editable Blender scenes from one image without task-specific training, specialist foundation models, or differentiable rendering—and that staged decomposition outperforms monolithic agentic pipelines even when those pipelines are augmented with strong 2D/3D specialists. The result supports the claim that task decomposition, rather than toolkit richness, is the critical factor for executable inverse graphics with current VLMs. Open questions left by the paper include whether global multi-factor refinement passes can be made computationally tractable, how the framework extends to multi-image and dynamic-scene reconstruction, and how to handle the inherent ambiguity of occluded content in single-view settings.