LoopVL: Recurrent Visual Intelligence
LoopVL applies recurrent loop transformers to vision-language models via iterative shared-module updates, outperforming larger non-recurrent models and exhibiting visual aha moments.
Published Sep 29, 2026▲ 469 on Hugging FaceCode ★ 128arXiv ↗
Only vote on papers you've read. Sign in with GitHub to vote.
LoopVL's recurrent vision-language loops yield genuine visual attention shifts and rigorous scratch training, though its "aha moments" need deeper non-recurrent ablation to prove they are not mere attention oscillation.
Abstract
We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through language pre-training, multimodal training, and post-training. LoopVL outperforms a range of similarly sized and larger non-recurrent models on multimodal understanding and visual reasoning benchmarks. We also observe Visual Aha Moments in LoopVL, characterized by pronounced shifts in visual attention across loops. LoopVL provides practical evidence for recurrent vision-language modeling and offers an intuitive perspective on how shared parameters can support deeper multimodal computation over continuously evolving visual-language states.