Thinking with Visual Primitives
Thinking with Visual Primitives interleaves spatial markers into chain-of-thought reasoning to close the reference gap, achieving frontier-level visual QA with extreme token efficiency.
Published 2026Paris Poster Session 1 · Wed, Dec 9, 12:30 PM–2:30 PM local time · Paris Poster HallOpenReview ↗
Only vote on papers you've read. Sign in with GitHub to vote.
Abstract
Despite the remarkable progress in Multimodal Large Language Models (MLLMs), the pre- vailing Chain-of-Thought (CoT) paradigms remain predominantly confined to the linguistic space. While recent advancements have focused on bridging the Perception Gap through high- resolution cropping (e.g., Thinking with Images), they overlook a more fundamental bottleneck: the Reference Gap. The inherent ambiguity of natural language often fails to provide precise, unambiguous pointers to complex spatial layouts, leading to logical collapse in tasks requiring rigorous grounding. In this work, we introduce Thinking with Visual Primitives, a novel reasoning framework that elevates spatial markers—such as points and bounding boxes—to “minimal units of thought”. By interleaving these visual primitives directly into the thinking process, our model can “point” while it “reasons”, effectively grounding its cognitive trajectory in the physical coordinates of the image. Notably, our framework is built on a highly optimized architecture with extreme visual token efficiency. Despite its compact model scale and signifi- cantly lower image-token budget, our model achieves frontier-competitive performance on a focused suite of challenging visual QA tasks, matching or exceeding models such as GPT-5.4, Claude-Sonnet-4.6, and Gemini-3-Flash. This demonstrates a path toward more efficient and scalable System-2-like multimodal intelligence.