Good Papers

Dream-Cubed: Controllable Generative Modeling in Minecraft by Training on Billions of Cubes

Dream-Cubed introduces a billion-scale Minecraft voxel dataset and diffusion models that generate interactive 3D worlds directly from block tokens with inpainting and outpainting support.

Tim Merino, Sam Earle, Ryunosuke Iwai, Julian Togelius, Edoardo Cetin

Published 2026Sydney Poster Session 4 · Wed, Dec 9, 5:00 PM–8:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

76%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel10/20reviewers recommend it
lenient 5/5
medium 3/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
Dream-Cubed's compositional cube diffusion and fully open dataset set a new bar for interactive voxel generation, though its block-level FID and human preference studies leave buildability and occlusion as unresolved blind spots.

Abstract

We introduce Dream-Cubed, a large-scale dataset of Minecraft worlds at voxel resolution, and a family of models using cubes as powerful compositional units for efficient generation of interactive 3D environments. Dream-Cubed comprises tens of billions of tokens from a carefully curated mixture of procedural biome terrain and high-quality human-authored maps. We use this dataset to conduct the first large-scale study of 3D diffusion models for voxel generation, analyzing discrete and continuous diffusion formulations, data compositions, and architectural design choices. Our models operate directly in the space of blocks, enabling efficient and semantically grounded generation while supporting interactive user workflows such as inpainting and outpainting from user-authored blocks. To quantitatively evaluate our models, we adapt the FID metric to assess semantic differences between real and generated world renderings, and validate generation quality through a human preference study. We release the full dataset, code, and all our pretrained models, which we hope will provide a foundation for future research in efficient generative modeling for structured, interactive 3D environments.