DistScene: Object-to-Scene Distillation for 3D Scene Generation
DistScene generates compositional 3D scenes from single images via object-to-scene distillation, improving spatial coherence by modeling environments as explicit components with shared coordinate frames.
Published Oct 3, 2026▲ 4 on Hugging FacearXiv ↗

Only vote on papers you've read. Sign in with GitHub to vote.
DistScene advances single-image 3D scene coherence by modeling environments as explicit geometric frames rather than object collections, though its synthetic-scene distillation still awaits real-world deployment validation.
Abstract
We present DistScene, a framework for single-image compositional 3D scene generation by jointly modeling the environment and individual objects. Unlike existing methods that represent scenes primarily as collections of objects, we model the environment as an explicit scene component to provide geometric context for object placement. Specifically, we introduce Scene-Frame Generation, which jointly generates separate environment and object components in a shared coordinate frame, allowing their geometry and relative placement to be learned together. Then we introduce Object-Centric Refinement to refine each object in a local frame with scene context. Finally, we develop Object-to-Scene Distillation to transfer pretrained object-generation priors to scene generation through automatically composed and rendered synthetic scenes. Evaluations on indoor and outdoor benchmarks demonstrate improved scene-level spatial coherence over the evaluated baselines. Project page: https://coolbeam.github.io/DistScene/