Octrees as an Explicit 3D Language
OctLLM represents 3D geometry via sparse octree tokens and trains separate 3D branches to achieve state-of-the-art multimodal 3D generation and understanding without degrading language ability.
Published Oct 1, 2026▲ 11 on Hugging FaceCode ★ 12arXiv ↗

Only vote on papers you've read. Sign in with GitHub to vote.
OctLLM delivers striking image-to-3D gains and frozen-language preservation via sparse octree tokens and branched parameters, but its 3D-language premise, thin baselines, and missing ablations leave the reasoning gains unproven.
Abstract
Existing 3D large language models (LLMs) compromise on two fronts: they compress shapes into latent codebook indices or coordinate text, which removes spatial structure from what the model observes, and they acquire the 3D modality by fine-tuning the backbone, which overwrites its general language ability. We present OctLLM, which addresses both limitations. Geometry enters as an explicit 3D sequence of octree occupancy tokens. However, full octree sequences grow rapidly with depth; OctLLM therefore randomly empties penultimate-level nodes and omits descendants while preserving shape, yielding a shorter coordinate- and depth-anchored Sparse Octree (S-Octree) for position-aware mask-modeling generation and 3D understanding. On the other front, existing methods introduce a new modality with full fine-tuning or LoRA, but full fine-tuning is costly, LoRA limits 3D capacity, and both modify the language pathway. OctLLM instead adds 3D capacity in parameters separate from the pretrained ones: mesh tokens are routed through independent trainable branches in a subset of blocks while text and image tokens retain the frozen vision-language pathway, and the two streams interact through shared self-attention. It trains far fewer parameters than full fine-tuning, yet sets a new state of the art among unified multimodal LLMs, lowering image-to-3D FID by $17.4\%$ and raising render-grounded captioning by $28.7$ points over ShapeLLM-Omni, while matching the backbone on general language benchmarks.