DistScene: Building 3D Worlds with Objects and Environments
DistScene tackles single-image 3D scene generation by jointly reconstructing objects and their surrounding environments in a shared coordinate system. Unlike cascaded pipelines that suffer from error propagation or object-only models that ignore environmental context, DistScene generates the environment as an explicit geometric anchor, enabling better object placement, scale, and spatial relationships. The method uses a novel Object-to-Scene Distillation process to create 125,000 synthetic training scenes from a pretrained object generator, then employs a two-stage generation pipeline: scene-frame synthesis for global layout followed by object-centric refinement for high-frequency detail.Script
Reconstructing a complete 3D scene from a single photograph means recovering not just individual objects, but the room or environment that holds them together. DistScene solves this by generating objects and their surrounding space jointly, so placement, scale, and support relationships emerge naturally rather than being patched together afterward.
The authors needed large-scale training data with every object and environment cleanly separated. They built it by distilling from a pretrained object generator, using an Large Language Model to describe assets, generating 70,000 objects and 46,000 empty environments, then composing 125,000 synthetic scenes with collision checks and physical placement constraints.
Generation happens in a shared scene frame where one latent slot represents the environment and separate slots represent each object. A sparse transformer lets all components attend to each other during denoising, so the model learns object placement by conditioning on the environmental geometry that anchors the entire scene.
After scene-frame generation, each object is extracted, normalized, and refined at high resolution in its own local coordinate system while staying conditioned on the full scene context. This two-stage design solves the resolution problem: global layout uses coarse voxels spread across the entire environment, while refinement recovers fine geometry without losing the object's position or scale in the scene.
On MIDI test, the method reduces scene Chamfer Distance by 32 percent and raises F-score by 6.5 percentage points compared to the strongest baseline. More importantly, it improves scene geometry, object geometry, and bounding box alignment simultaneously, confirming that environmental context improves both local detail and global spatial coherence.
The method inherits biases from its training corpus. Indoor scenes may omit ceilings because the environment generator was trained on images of open-top rooms, and outdoor reconstructions can show holes when the fixed scene-frame resolution spreads too thin over large areas. These limitations point to representation capacity and data diversity as the next frontiers. To explore this work further and generate your own explanatory videos, visit EmergentMind.com.