RATS! Patches Talk Through Registers: Emergent Parts in Register Attention Transformers
RATS decomposes vision transformers' classification token into learnable register tokens that spontaneously specialize into object parts, improving segmentation by up to 12 mIoU.
Published 2026Atlanta Poster Session 2 · Wed, Dec 9, 4:30 PM–7:30 PM local time · Hall C1arXiv ↗OpenReview ↗

Only vote on papers you've read. Sign in with GitHub to vote.
Abstract
When humans see a bird, they recognize far more than just "bird" -- they see a head, wings, and talons, a structured assembly of reusable parts that can be identified across every bird they have ever seen. We ask whether a self-supervised visual model can discover the same compositional structure on its own. To this end, we propose RATS (Register Attention Transformers), which decomposes the classification token into N learnable register tokens that route patch information through an L->N->N->L bottleneck via a three-step compress-communicate-broadcast attention. The N registers are partitioned across the H attention heads, so that registers assigned to different heads do not interact with each other. Without auxiliary losses or part annotations, each register spontaneously specializes into a proto-semantic region whose emerging structure resembles object parts. RATS surpasses all baselines by +12 mIoU on average across five segmentation benchmarks, with consistent gains on ADE20K (+1.11 mIoU) and COCO (+0.2 AP^m). Its register dictionary further exhibits part-level consistency and semantic proximity across related categories. Our results suggest that RATS may provide a useful architectural prior for structured and interpretable visual representation learning.