Good Papers

Showing Vision backbones Show all papers

74%Highly rated
?Highly ratedVote to see the score

VisionHOPE: Visual Backbones as Self-Modifying Learning Systems

VisionHOPE formulates visual backbones as self-modifying learning systems with coupled co-evolving memories and proves stable non-expansive dynamics, achieving competitive results on ImageNet-1K, COCO, and ADE20K.

Siran Peng, Tianshuo Zhang, Tianyu Fu, Weisong Zhao and 7 more

Published Sep 27, 2026 · 0 citations · ▲ 323 on Hugging Face · Code ★ 880

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 2/5
medium 7/10
strict 0/5
86%Must read
?Must readVote to see the score

GTR: Gated Token Recurrence for Efficient Dense Prediction

GTR replaces softmax attention with gated token recurrence for efficient high-resolution dense prediction, achieving 58.9 COCO box AP with low latency.

Zhe Feng, Longfei Liu, Wei Liu, Kai Chen and 6 more

Published Sep 22, 2026 · 0 citations · ▲ 12 on Hugging Face · Code ★ 50

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
93%Must read
?Must readVote to see the score

Swin Transformer: Hierarchical Vision Transformer using Shifted Windows

Swin Transformer is a hierarchical vision transformer using shifted windows for efficient local self-attention and linear image complexity, achieving state-of-the-art results on ImageNet, COCO, and ADE20K.

Ze Liu, Yutong Lin, Yue Cao, Han Hu and 4 more

Published Oct 1, 2021 · 33,075 citations

100% Readers1 of 1 upvoted
18/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 21 reviewers recommend it
lenient 5/5
medium 10/11
strict 3/5
57%Worth a look
?Worth a lookVote to see the score

Training Vision Transformers to Focus: Emergence of Attention Head Specialization

Yuji Nozawa, Yu-Chieh Lin, Youyang Ng

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Lost in the Slots: Revisiting Object-Centric Representations in the era of Foundation Models

Priyam Dey, Aditya Sahdev, Omkar M Kashyap, Venkatesh Babu Radhakrishnan

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

SOLA: A Structured Operator Library for Attention in Pretrained Vision Transformers

Enzo Tartaglione, Vito Paolo Pastore

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

OrbitLoRA: Learning Rotation-Aware Low-Data Adaptation of Vision Foundation Models

Yunlu Chen, Dominik Engel, Peter Wonka, Ivan Viola

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Patch4Patch: Restoring Structural Connectivity in Patch-based Vision Encoders

Yaqi Zhang, Shuntian Yao, Niantai Qu, Runguo Chen and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

SpRePE: A Spherical Geometry-Aware Position Embedding scheme for Vision Transformers

Qilong Jia, Yi Xiao, Wei Xue

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

COHE: Auditing Non-Transitivity in Sample Difficulty Proxies for Vision Models

Dayena Jeong, Sunglok Choi

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

Channel Mixer: A Pretrainable Tokenizer for Scalable Multi-Channel Vision Transformers

Lukas Miklautz, Lucas Miranda, Nikola Bulat, Armin Lambacher and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Prototypes of the Mind: A Unified Framework for Probing the Visual Brain

Shi Chen, Seoyoung Ahn, Doris Tsao

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Virtual Task Prompting for Multi-Task Scene Understanding

Zedong WANG, Yucheng Wang, Dan Xu

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Disentangling Channel Semantics in Vision Transformers via Token Decorrelation and Composition-Aware Modulation

Daeun Kim, Hyejin Park, Hyesong Choi, Dongbo Min

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Evaluating Deployable Inference-Time Error Prediction in Vision MoEs

Aristeidis Tsaris, JangHyeon Lee, Abhishek Potnis, Philipe Dias and 4 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

C-LoRA: Continual Low-Rank Adaptation for Pre-trained Visual Models

Xin Zhang, Liang Bai, Xian Yang

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Vision Correlators: Correlation-Driven Visual Understanding with Hypergraphs

Mengqi Lei, Guohuan Xie, Siqi Li, Shihui Ying and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows

Albert Gao, Bing Xue, Andrea Zanette

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

On Making $SE(2)$-Invariant Networks Optimal

Tomas Karella, Emily Shinkle, Alice Allen, Pieter J Swart and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

RoSeViT: Role-Separated Vision Transformers for ARC Visual Reasoning

Zhiguang Liu, Yi Shang

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
89%Must read
?Must readVote to see the score

Not Too Generative, Not Too Discriminative: The Human Alignment Sweet Spot

Joint Energy-Based Models show human visual alignment peaks at intermediate generative-discriminative training, not either objective alone.

Jorge Chang Ortega, Bastien Le Lan, Thomas Serre, Victor Boutin

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 3/5
78%Highly rated
?Highly ratedVote to see the score

Multidimensional Observer Model and Perceptual Dimensions of Human Image Quality Assessment

A multidimensional observer model representing images as latent perceptual distributions reveals that human image quality judgments rely on extremely low-dimensional, task-dependent perceptual spaces.

Sheng Zhao, Weikai Lin, Yuhao Zhu

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
80%Must read
?Must readVote to see the score

The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation

Cross-domain event-to-RGB distillation induces color invariant, shape-biased networks by suppressing texture dependence for edge-based robustness.

Shunsuke Yasuki, Masato Taki, Soshun Kihara

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

DC-ViT: Modulating Spatial and Channel Interactions for Multi-Channel Images

DC-ViT decouples spatial and cross-channel attention pathways via DSA and learns task-specific channel weights via DAG, outperforming existing multi-channel vision transformers.

Umar Marikkar, Syed S Husain, Muhammad Awais, Sara Atito

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Vision Transformers Learn Gestalt-Like Figure-Ground Cues from Natural Images

Vision transformers learn Gestalt-like figure-ground cues, surroundedness, convexity, and symmetry for uniform regions, from natural images, with linear probes generalizing zero-shot to artificial stimuli.

Matthias Tangemann, Benjamin Lo, Zygmunt Pizlo, Kaleem Siddiqi and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 2/5
86%Must read
?Must readVote to see the score

Task-Induced Riemannian Metrics for Vision Transformer Feature Spaces

Task-induced Riemannian metrics define ViT feature geometry via decoder Jacobians, and a learnable low-rank approximation enables accurate geometric token pruning without fine-tuning.

Andrew Bond, Ege E Özlü, Tuna Çimen, Ilkin U Melanlioglu and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 3/5
medium 9/10
strict 2/5
91%Must read
?Must readVote to see the score

Attention Transfer Is Not Universally Effective for Vision Transformers

Attention transfer fails for four ViT families due to architectural mismatch, and adding the teacher's native components to students fully restores its effectiveness.

Huaiyuan Qin, Muli Yang, Gabriel James Goenawan, Peng Hu and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 4/5