DistScene: Object-to-Scene Distillation for 3D Scene Generation

Kunming LuoHongyu YanKen DengChengcheng ZhouTianyu LiuHaipeng LiHaibin HuangXuelong LiPing Tan

Method

Object-to-Scene Distillation: Automatic and Scalable Training Data Generation

DistScene self-distilled data curation: Trellis.2 generates objects and environments, physics checks guide scene assembly, and rendering provides conditioning images.
Left: an overview of the complete self-distillation process. Right: the data curation stage. We adopt Trellis.2 as the object generator from which training data are distilled. A VLM first generates descriptions of 3D assets, including objects and environments; these descriptions are then used to generate the corresponding input images for Trellis.2. After obtaining the 3D assets, we place them into the generated environments with physics checks and render conditioning images for each scene.

Unified Scene Representation: Scene Frame Generation with Object-Centric Refinement

DistScene inference pipeline: scene frame generation jointly synthesizes the environment and objects, followed by object-centric refinement and placement back into the scene.
Left: the environment and objects are jointly generated in a shared scene frame through two stages, sparse structure generation and geometry-latent generation. Right: each generated object is transformed to a local voxel support for scene-conditioned refinement and then placed back into the scene using the inverse transformation.

Interactive Demo

Select an input image, then open the preview when needed.

Note that this 3D scene is a compressed preview version, approximately 10% of the original generated result. Please refer to results for the detailed result.

Rendered Demo

Select an input image, then show the video when needed.