Good Papers

Octrees as an Explicit 3D Language

OctLLM represents 3D geometry via sparse octree tokens and trains separate 3D branches to achieve state-of-the-art multimodal 3D generation and understanding without degrading language ability.

Ran Dan, Si‐Tong Wei, Pengfei Xiong, Wei Zhang, Yadong Mu, Peng-Shuai Wang

Published Oct 1, 2026▲ 11 on Hugging FaceCode ★ 12arXiv ↗

78%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel11/20reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
OctLLM delivers striking image-to-3D gains and frozen-language preservation via sparse octree tokens and branched parameters, but its 3D-language premise, thin baselines, and missing ablations leave the reasoning gains unproven.

Abstract

Existing 3D large language models (LLMs) compromise on two fronts: they compress shapes into latent codebook indices or coordinate text, which removes spatial structure from what the model observes, and they acquire the 3D modality by fine-tuning the backbone, which overwrites its general language ability. We present OctLLM, which addresses both limitations. Geometry enters as an explicit 3D sequence of octree occupancy tokens. However, full octree sequences grow rapidly with depth; OctLLM therefore randomly empties penultimate-level nodes and omits descendants while preserving shape, yielding a shorter coordinate- and depth-anchored Sparse Octree (S-Octree) for position-aware mask-modeling generation and 3D understanding. On the other front, existing methods introduce a new modality with full fine-tuning or LoRA, but full fine-tuning is costly, LoRA limits 3D capacity, and both modify the language pathway. OctLLM instead adds 3D capacity in parameters separate from the pretrained ones: mesh tokens are routed through independent trainable branches in a subset of blocks while text and image tokens retain the frozen vision-language pathway, and the two streams interact through shared self-attention. It trains far fewer parameters than full fine-tuning, yet sets a new state of the art among unified multimodal LLMs, lowering image-to-3D FID by $17.4\%$ and raising render-grounded captioning by $28.7$ points over ShapeLLM-Omni, while matching the backbone on general language benchmarks.