LEGO-Combined Benchmarks: Evaluations for 3D Environment and Reconstruction
- LEGO-Bench evaluations encompass tools for assessing text-conditioned 3D scene synthesis (
- Holistic Success Rate
- Partial Success Rate
- And judgment strategies like LEGO-Eval and Image to Code frameworks for accurate 3D
LEGO-Bench is the name of two distinct benchmark resources in the supplied research record. In “LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environments with Tool Augmentation,” it evaluates whether generated 3D indoor environments satisfy long, fine-grained natural-language specifications (Hwangbo et al., 4 Nov 2025). In “LEGO-Anything: Coding Agents for 3D Scene Reconstruction,” it evaluates executable Blender programs reconstructed from single RGB images, separately measuring artifact validity, visible-surface geometry, and rendered appearance (Li et al., 28 Sep 2026). The two benchmarks are unrelated in task definition, data, evaluation protocol, and output representation; the shared name is therefore ambiguous without citation or additional context.
1. Name disambiguation and conceptual scope
The LEGO-Bench introduced with LEGO-Eval targets text-conditioned 3D scene synthesis. Its input is a detailed natural-language instruction describing rooms, architectural components, objects, attributes, materials, spatial relations, and object placements. A generation system must produce a scene satisfying all specified constraints. The benchmark evaluates both scene-generation systems and alignment evaluators that judge whether a generated scene conforms to its instruction (Hwangbo et al., 4 Nov 2025).
The LEGO-Bench introduced with LEGO-Anything targets single-image, scene-level, executable reconstruction. A coding agent receives one RGB image and iteratively writes and executes Blender code, inspects intermediate scenes and renderings, and revises the program. The output is an editable scene artifact rather than merely an image, mesh, point cloud, or fixed 3D prediction. The benchmark evaluates whether the artifact is valid and whether it recovers visible geometry and rendered appearance (Li et al., 28 Sep 2026).
| Dimension | LEGO-Bench for scene synthesis | LEGO-Bench for scene reconstruction |
|---|---|---|
| Primary task | Generate a 3D indoor environment from a fine-grained instruction | Reconstruct a 3D scene from one RGB image |
| Input | Natural-language scene specification | Single RGB image |
| Output | Simulator-compatible 3D scene | Executable Blender artifact |
| Main evaluation | Constraint satisfaction and instruction–scene alignment | Validity, visible geometry, and appearance |
| Principal paper | (Hwangbo et al., 4 Nov 2025) | (Li et al., 28 Sep 2026) |
The name should not be conflated with other LEGO-related resources. “LEGO-Puzzles” evaluates multimodal models on LEGO-based spatial and sequential reasoning (Tang et al., 25 Mar 2025). “StableLego” evaluates static structural stability of block-stacking assemblies (Liu et al., 2024). “LEGOBench: Scientific Leaderboard Generation Benchmark” is unrelated to physical or virtual LEGO and evaluates scientific leaderboard generation from arXiv and PapersWithCode data (Singh et al., 2024).
2. LEGO-Bench for fine-grained 3D scene synthesis
The first LEGO-Bench is designed to expose failures that coarse scene-generation evaluations can conceal. A generated environment may look plausible while omitting required objects, assigning incorrect attributes, placing objects in the wrong rooms, blocking architectural elements, or violating specified spatial relationships. The benchmark therefore evaluates whether all constraints in a detailed instruction are jointly satisfied.
Each instruction is decomposed into constraints
A scene is valid only if every constraint is satisfied:
The constraints are divided into four categories:
- Floor Layout: spatial layout of rooms, walls, doors, and windows.
- Material Selection: visual appearance of floors and walls.
- Object Selection: object identity and appearance, including doors and windows.
- Object Placement: object locations and rotations.
This formulation emphasizes multi-hop grounding. For example, an instruction such as “the cup is on the red table” requires identifying the table, verifying its color, identifying the cup, and validating their relation. Similarly, “the bed in the room with the red floor is white” requires resolving the room context before evaluating the bed attribute.
The benchmark contains 130 natural-language instructions, 130 manually annotated scenes, and 1,250 total constraints, with an average of 9.6 constraints per instruction. Instructions contain an average of 2.27 rooms and 98.54 words. Approximately 55% of constraints concern objects and 39% concern architectural components; approximately 40% concern material and object selection, while approximately 60% concern floor layout and object placement. Most instructions contain between 9 and 11 constraints.
The reference scenes were constructed to fully satisfy their corresponding instructions. Annotators received a two-hour training session, wrote approximately 30–50 instructions each, identified and classified constraints, linked constraints to text spans, and associated prerequisite constraints where required. Holodeck initially generated textual scene representations, including asset IDs, object identities, positions, rotations, and other scene metadata. Annotators then manually edited these representations until the rendered scenes aligned with the instructions. The scenes and annotations underwent mutual review and two verification iterations.
No conventional train, validation, and test split is reported. The resource is used primarily as an evaluation benchmark rather than as a supervised-learning dataset with official partitions.
3. Scene-generation evaluation and results
The scene-synthesis benchmark reports two principal metrics. Holistic Success Rate is the proportion of instructions for which every constraint is satisfied. Partial Success Rate is the proportion of individual constraints that are satisfied, reported overall and by constraint category.
Four LLM-based scene-synthesis systems are evaluated: I-Design, LayoutGPT, Holodeck, and LayoutVLM. Holodeck generates complete scenes, including object selection, attributes, and placement. I-Design selects and places objects. LayoutGPT and LayoutVLM primarily position a given set of objects and are augmented with Holodeck to produce complete scenes suitable for evaluation.
| Method | Holistic SR | Floor Layout | Material Selection | Object Selection | Object Placement | Average Partial SR |
|---|---|---|---|---|---|---|
| I-Design | 3.8 | 92.7 | 63.7 | 11.0 | 4.1 | 34.2 |
| LayoutGPT | 6.9 | 96.0 | 65.3 | 40.9 | 37.3 | 55.2 |
| Holodeck | 8.4 | 96.3 | 61.6 | 46.3 | 43.7 | 58.5 |
| LayoutVLM | 10.0 | 95.6 | 65.3 | 49.8 | 46.0 | 60.6 |
LayoutVLM obtains the highest reported holistic success rate, but only 10.0% of instructions are fully satisfied. Floor-layout performance is substantially higher, ranging from 92.7% to 96.3%, whereas object selection and placement are considerably weaker. The gap between partial and holistic performance indicates that systems can satisfy many individual requirements while still violating at least one constraint in most complete scenes.
Instructions are grouped by complexity: simple instructions contain 2–7 constraints, moderate instructions contain 8–12, and complex instructions contain more than 12. Holistic success decreases as constraint density increases. The paper reports that existing methods consistently fail on complex instructions. A human survey found an average of 18.2 constraints per room, indicating that the benchmark’s difficult regime corresponds to the level of detail found in natural descriptions of indoor environments.
4. LEGO-Eval as a tool-augmented evaluator
LEGO-Eval is the evaluator associated with the scene-synthesis benchmark. It decomposes the instruction into constraints, plans tool calls, selects tool arguments, executes the tools, and validates each constraint using retrieved evidence. Its formal evaluation output is
where is the instruction, is the generated scene, is a binary alignment judgment, and is an interpretable explanation.
The evaluator uses 21 tools in three groups.
Environment-interaction tools retrieve visual evidence, including top-down scene views, room views, object views, wall views, material images, multiview object renderings, and spatial-relation visualizations.
Textual-reasoning tools retrieve structured scene metadata, including room, wall, door, window, and object lists and information. These tools expose room polygons, floor materials, wall geometry and orientation, object IDs, asset IDs, room assignments, exact 3D positions, rotations, and geometric representations.
Multimodal-reasoning tools identify objects and verify attributes such as color, shape, material, texture, and surface pattern.
The evaluator supports multi-hop grounding by reusing previously identified entities. Once a red table has been identified as table-2, later constraints can refer to that object without repeating the search. Tool plans may include parallel calls, and previous constraint results are reused to reduce redundant retrieval.
For comparison with other alignment metrics, the authors use the 130 compliant instruction–scene pairs and 130 additional manually curated scenes that intentionally violate the instructions. The resulting 260 pairs are evaluated with SceneEval, CLIPScore, VLM-as-a-judge, and LEGO-Eval. The reported metrics are F1, precision, recall, and Cohen’s kappa.
| Evaluator | Holistic F1 | Holistic Recall | Holistic Precision | Holistic | Partial F1 | Partial |
|---|---|---|---|---|---|---|
| SceneEval, full dataset | 0.33 | 0.50 | 0.25 | 0.00 | 0.28 | 0.00 |
| CLIPScore, threshold 20 | 0.49 | 0.51 | 0.51 | 0.02 | 0.46 | 0.00 |
| GPT-4.1 judge | 0.40 | 0.53 | 0.67 | 0.05 | 0.68 | 0.35 |
| LEGO-Eval, GPT-4.1 | 0.81 | 0.82 | 0.84 | 0.63 | 0.83 | 0.66 |
| LEGO-Eval, GPT-4.1-mini | 0.70 | 0.72 | 0.78 | 0.43 | 0.78 | 0.56 |
| LEGO-Eval, Qwen2.5VL-32B | 0.64 | 0.66 | 0.70 | 0.32 | 0.72 | 0.44 |
LEGO-Eval with GPT-4.1 achieves a holistic F1 of 0.81 and Cohen’s kappa of 0.63, compared with 0.40 and 0.05 for GPT-4.1 used as a conventional VLM judge. The paper reports an improvement of 0.41 F1 over VLM-as-a-judge at the holistic level. SceneEval cannot evaluate 41% of the constraints in the dataset; its full-dataset setting treats unevaluable constraints as incorrect.
Tool ablations show that textual reasoning and environment interaction are important. Removing textual reasoning reduces holistic F1 by 5.05%, while removing environment interaction and multimodal reasoning reduces it by 24.90%. The evaluator requires explicit grounding because conventional global image–text similarity or unstructured VLM judgment can reason about nonexistent objects, confuse visually similar entities, or validate relations before establishing object identity.
5. LEGO-Bench for executable single-image reconstruction
The second LEGO-Bench evaluates LEGO-Anything, an Image-to-Code framework in which a coding agent constructs an executable Blender scene from one RGB image (Li et al., 28 Sep 2026). The agent interacts with Blender through an editing environment and produces a program 0:
1
where 2 is the input image, 3 is the coding agent, and 4 is the resulting 3D scene.
The agent may iterate through intermediate programs, scenes, and observations:
5
The benchmark is simulator-grounded. Its images are rendered in LychSim using scene and asset packs from Fab, while private evaluator-side information includes complete scene geometry, object identities, camera parameters, depth, instance masks, object correspondences, and scene transforms. Candidate scenes undergo collision checks, stability checks, and human review. The benchmark contains 208 RGB inputs rendered from 104 logical scenes across 8 environments and 17 themes, using 443 registered assets.
Indoor and outdoor scenes are organized into matched Easy, Medium, and Hard tiers. Within a scene family, architecture, materials, lighting, camera, shared asset identities, transforms, scales, and bounding boxes remain fixed, while additional visible objects are added at higher difficulty levels:
6
An additional unpaired NYC Bird’s-Eye split contains 10 overhead-view cases per complete run. The benchmark does not report a conventional public train, validation, and test split. It instead uses indoor and outdoor evaluation groups, paired complexity tiers, a 42-case Office subset, a 10-case Bird’s-Eye stress split, and repeated runs.
Submissions provide:
scene.blend,scene.glb, andfinal.png.
Artifact validity requires a reloadable, non-empty Blender scene, at least one mesh, an active camera, a non-empty well-formed GLB, and a decodable non-degenerate render. Validity does not imply reconstruction fidelity. If an artifact is invalid or headline evaluation cannot be completed, validity, reconstruction, and appearance are all set to zero.
6. Reconstruction metrics, results, and limitations
The reconstruction metric evaluates visible geometry from the reference camera. It does not evaluate hidden geometry and does not perform post-hoc alignment, global rescaling, camera fitting, or per-object transformation fitting.
For each scored object, reference-visible points 7 are compared with submitted visible points 8. The depth-relative tolerance is
9
where 0 is the positive forward depth of reference point 1. Recall measures the fraction of reference points matched by submitted points, while precision measures the fraction of submitted points close to reference geometry. Per-object F1 is
2
The trial-level reconstruction score is the equal-weight macro-average over eligible objects:
3
The headline metric is F@5%; auxiliary variants include F@2% and F@10%.
Appearance is evaluated by independently rerendering the submitted Blender scene. If 4 is the reference image and 5 the evaluator rerender, a pixel is counted as correct when the largest absolute RGB-channel difference is at most 30 on the 0–255 scale:
6
The validity-gated overall score is
7
The strongest evaluated coding agent is GPT-6-astra with Codex, which obtains an Indoor overall score of 53.4% and an Outdoor overall score of 39.6%. Its artifact validity is 100.0% indoors and 98.0% outdoors, while reconstruction and appearance remain substantially lower. The results distinguish producing a technically usable Blender artifact from faithfully recovering the input scene.
| Model and harness | Indoor validity | Indoor reconstruction | Indoor appearance | Indoor overall | Outdoor validity | Outdoor reconstruction | Outdoor appearance | Outdoor overall |
|---|---|---|---|---|---|---|---|---|
| GPT-6-astra + Codex | 100.0 | 52.4 | 54.4 | 53.4 | 98.0 | 34.0 | 45.5 | 39.6 |
| GPT-6-sol + Codex | 99.4 | 22.2 | 42.5 | 32.3 | 99.7 | 19.8 | 28.7 | 24.2 |
| GPT-6-luna + Codex | 99.7 | 15.8 | 30.5 | 23.2 | 99.7 | 13.2 | 21.4 | 17.3 |
| GPT-5.6-sol + Codex | 97.8 | 9.8 | 19.8 | 14.8 | 98.7 | 12.8 | 17.7 | 15.3 |
Increasing scene complexity harms fidelity but not artifact deliverability. Across GPT-6 and GPT-5.6 configurations, overall score decreases from 24.6% on Easy scenes to 21.3% on Medium scenes and 20.5% on Hard scenes. On the Bird’s-Eye split, GPT-6-astra obtains 35.3% overall with 24.0% reconstruction under F@2%, demonstrating the difficulty of large-scale overhead reconstruction.
Trajectory analysis identifies weak scene initialization, regressive edits, and unreliable self-evaluation. For GPT-5.6-sol, 29.6% of edits decrease the score, and the final submission is 3.2 points below its best intermediate scene. Across builder–judge pairs, self-judgment agrees with reconstruction direction only 45.8% of the time and with appearance direction 62.2% of the time.
LEGO-Plugin addresses these issues without training. Its modules provide enhanced initialization with VGGT, grounded refinement using SAM 3 and Depth Anything V2, and transactional version control with snapshots, validation, commits, and rollbacks. It limits evaluation to at most eight transactions and two rollbacks per case, with at most two correction rounds for a candidate. On the 42-case Office subset, it improves all six model settings, with relative gains of up to 62.7%; the gain for GPT-6-astra is 2.1%.
The reconstruction benchmark remains limited by single-view ambiguity. A single RGB image cannot uniquely determine hidden geometry, metric depth, absolute scale, occluded objects, back-side appearance, or physically correct structure. The headline metric therefore evaluates visible surfaces from the reference view rather than complete unseen-world correctness. Appearance is entangled with geometry, camera, materials, lighting, shadows, and color configuration. The benchmark is primarily an evaluation resource rather than a supervised dataset with a conventional split.
The two LEGO-Bench resources consequently represent different methodological priorities. The scene-synthesis benchmark measures whether detailed language specifications are grounded into complete, relationally correct environments and whether evaluators can verify that grounding. The executable-reconstruction benchmark measures whether coding agents can convert a single image into an inspectable, editable, simulator-grounded scene program whose visible geometry and appearance are faithful to the reference. Neither benchmark is a physical LEGO-construction benchmark, and neither should be identified with LEGO-Puzzles, StableLego, or the scientific leaderboard benchmark despite their overlapping names.