Good Papers

Showing Vision-language models Show all papers

91%Must read

DEPICT: Scoring Text-to-Image Alignment by Answer Agreement

DEPICT scores text-to-image alignment via expected agreement between image and caption answers, boosting negation accuracy to 88% and exceeding fine-tuned evaluators on human correlation benchmarks.

Vasco Ramos, Sandra Godinho Silva, Joao Magalhaes, Ricardo Rei and 1 more

Published Oct 2, 2026 · 0 citations · ▲ 3 on Hugging Face

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Foresight: planning future perception in streaming VLMs without retraining

FORESIGHT uses dual-stream anticipatory planning in frozen streaming VLMs to dynamically configure future perception, improving online benchmarks by up to 18.7 points without retraining.

Ashok Prasad Neupane, Dipan Bartaula, Ankit Belbase, Saugat Adhikari and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 3 on Hugging Face · Code

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

LoopVL: Recurrent Visual Intelligence

LoopVL applies recurrent loop transformers to vision-language models via iterative shared-module updates, outperforming larger non-recurrent models and exhibiting visual aha moments.

Zhe Qian, Ziyang Gong, Zhongxing Xu, Hehan Li and 8 more

Published Sep 29, 2026 · 0 citations · ▲ 470 on Hugging Face · Code ★ 172

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training

PaddleOCR-VL-1.6 applies region-aware data optimization and progressive reinforcement learning post-training to achieve 96.33% on OmniDocBench v1.6.

Zelun Zhang, Hongen Liu, Suyin Liang, Yubo Zhang and 11 more

Published Jun 2, 2026 · 0 citations · ▲ 26 on Hugging Face · Code ★ 90,719

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

VisPlay: Self-Evolving Vision-Language Models from Images

VisPlay enables vision-language models to self-evolve via reinforcement learning on unlabeled images by having a questioner and reasoner generate and train on diverse visual reasoning tasks, improving reasoning, generalization, and reducing hallucinations across benchmarks.

Yicheng He, Chengsong Huang, Li, Zongxia, Jiaxin Huang and 1 more

Published Nov 19, 2025 · 0 citations · ▲ 45 on Hugging Face · Code ★ 79

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

MinerU2.5 decouples global layout analysis from local content recognition via coarse-to-fine parsing, achieving state-of-the-art document parsing accuracy with low computational overhead.

Junbo Niu, Zheng Liu, Zhuangcheng Gu, Bin Wang and 36 more

Published Sep 26, 2025 · ▲ 180 on Hugging Face · Code ★ 81,202

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Otter: A Multi-Modal Model With In-Context Instruction Tuning

Otter is a multi-modal model instruction-tuned with visual and textual in-context examples via the MIMIC-IT dataset, improving convergence and generalization on complex video and multi-image tasks.

Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang and 5 more

Published May 20, 2025 · 75 citations · Code ★ 3,443

– ReadersNo votes yet
10/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

SmolDocling is a 256M-parameter vision-language model for end-to-end document conversion using DocTags to capture content, structure, and spatial layout, matching models up to 27 times larger.

Ahmed Nassar, Andres Marafioti, Matteo Omenetti, Maksym Lysak and 9 more

Published Mar 14, 2025 · ▲ 177 on Hugging Face · Code ★ 68,484

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

Diffusion Instruction Tuning

Lavender aligns vision-language model attention with Stable Diffusion during supervised fine-tuning, boosting accuracy up to 30% with minimal training data.

Chen Jin, Ryutaro Tanno, Amrutha Saseendran, Tom Diethe and 1 more

Published Feb 4, 2025 · 0 citations · ▲ 2 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Personalized Visual Instruction Tuning

PVIT introduces a framework that curates personalized visual instruction data to cure multimodal models' face blindness, significantly boosting personalized dialogue performance.

Renjie Pi, Jianshu Zhang, Tianyang Han, Jipeng Zhang and 2 more

Published Oct 9, 2024 · 0 citations · ▲ 70 on Hugging Face · Code ★ 34

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding

ERNIE-Layout enhances document pre-training with layout knowledge for improved visually-rich document understanding.

Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo and 11 more

Published 2022 · 65 citations · Code ★ 106

– ReadersNo votes yet
8/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Hierarchical Semantic Tree Anchoring for CLIP-Based Class-Incremental Learning

Tao Hu, Lan Li, Zhenhao Wen, Da-Wei Zhou

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Revitalizing Medical Time Series with Vision-Informed Retrieval: A Vision-Language Perspective

Guoqi Yu, Juncheng Wang, Shujun Wang

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

C2FT: Enhancing Fine-Grained Perception in MLLMs via Confuse-then-Contrast Fine-Tuning

Shaoxuan He, Benlei Cui, Shikai Qiu, Yuwen Zhai and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Not Every Image Teaches Vision: Visual-Necessity-Gated Continual Learning for Multimodal Large Language Models

Jiamu Xue, Haidong Kang, Wen Luo, Jubo Chen and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Despa: Resolving Spatial Collapse in VLMs via Depth-Grounded Geometry

Yujing Lou, Pingyi Chen, Shen Cao, Lubin Fan and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Pixels over Symbols: Sensory Realism Improves Behavioral Alignment in Models of Cognition

Mark Bai, Xiaoxuan Lei, Zihan Weng, MOTAHAREH POURRAHIMI and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

GaitLingo: Self-Supervised Gait Representation Learning with Language Priors

Chenye Wang, Zhengxiang Lan, Saihui Hou, Zhikang Liu and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Diffusion Thinking for Fast Long-Form Spatial Reasoning in Vision--Language Models

Zhitao Zeng, weitao Du, Yueming Jin

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context

Zhaowei Wang, Lishu Luo, Haodong Duan, WeiWeiLiu and 8 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features

Dae Ung Jo, Jongin Lim, YoungJoon Yoo, Daeho Um

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Render Structure Uncertainty for HTML Repair in MLLM-based UI-to-Code Generation

Haoran Ma, Jiechao Gao, Shisong Tang, bing han and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Toward Multimodal Sheet Music Recognition and Understanding

Guang Yang, Brian Zheng, Victoria Ebert, Noah Smith

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

LAMP: Language-Modulated Geometric Preservation for Multi-Modal Object Re-Identification

Shuying Li, Chao Su, Yongxiang Li, Peng Hu and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Concentrated Gradients Amplify Forgetting: Dominant-direction Projection for Continual Multimodal Learning

Chengxiang Huang, Haopeng Zhang, Yuzhe Han, Rui Dai and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Omni-SpikeDet: A Spiking Open-World Detector with Dynamic Text–Image Alignment

Ziqi Li, Tao Gao, Ting Chen, Xin Zhang and 3 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

VisEditBench: A Benchmark for Vector-Format Diagram Editing with Visual Instructions

Akito Taneguchi, Itsumi Saito, Haruto Yoshida, Jun Suzuki

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

What Drives Test-Time Adaptation for CLIP? A Controlled Empirical Study from an Update Perspective

Jiazhen Huang, Xiao Chen, Zhiming Liu, Yaru Sun and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

VLS: A Vision-Language-Shape Model for Open-Vocabulary Partonomic 3D Reconstruction

Xiaoqian Ruan, Pei Yu, Dian Jia, Hyeonjeong Park and 2 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

UGGRH: Unsupervised Generative Completion and Graph-attention Refinement for Incomplete Cross-modal Hashing

Yunfei Chen, Yuchen Zhang, Hongyu Lin, Peng Liu and 1 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Decoupled Prototype Contrastive Alignment Hashing for Cross-Modal Retrieval

Yunfei Chen, Renwei Xia, Shangchong Gao, Zhan Yang

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

ALAM: Algebraically Consistent Latent Transitions for Vision-Language-Action Models

Zuojin Tang, Haoyun Liu, Xinyuan Chang, Changjie Wu and 10 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

MedVTok: A General-Purpose Medical Visual Tokenizer

Chenglong Ma, Yuanfeng Ji, Junzhi Ning, Jiyao Liu and 17 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Sophon: A Procedural Diagnostic for Spatial Reasoning in Vision-Language Models

Zach Gazak, Ryan Swindle, Justin Fletcher

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Spectral Rank Calibration for Continual LoRA Merging in Multimodal Large Language Models

Chenrui Wu, Haishuai Wang, Jiajun Bu, Jiangchuan Liu

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Cross-Modal Prior-Guided Training with Visual Foundation Models for Unsupervised LiDAR Point Cloud Registration

KeZheng Xiong, Shiyun Xu, Sheng Ao, Mingming Wang and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Beyond Global Alignment: Structured Compositional Reasoning for Vision-Language Models

Zhoujun Ye, Yiwei Fu, Qiyun Huang, Jie Yang and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Reading, Not Thinking: Bridging the Modality Gap When Text Becomes Pixels

Kaiser Sun, Xiaochuang Yuan, Hongjun Liu, Chen Zhao and 3 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models

Xurui Song, Shuo Huai, Jingjing Jiang, Jiayi Kong and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

VLSplat: Vision-Language Guided Object-Centric 3D Gaussian Splatting via Scene Graph

Yuntae Jeon, Younho Jeon, Sujin Jin, Sungho Jo and 1 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

BlenderFORGE: Framework for Optimizing Reactive 3D-Graphics Editing Ability of MLLMs

Zilin Guo, Tianrui ZHANG, Yi RONG, Yichen Liu and 5 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

MindVLM: Neural-Grounded Visual Captioning via Subject-Aware Semantic Evidence Selection

Zixiang Yin, Yu-Ping Wang, Zhengming Ding

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Where to Connect? Boosting MLLMs via Dynamic Gated Pathways across ALL ViT and LLM Layers

Yingying Yan, Jiaqi Tang, Wei Wei, Qianzhou Wang and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Mitigating Asymmetric Boundary Encroachment in Continual Learning of Vision-Language Models

Mingfeng Li, Xinyang Chen, Xiucheng Li, Weili Guan and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

ViLo: LiDAR Localization with Vision-Language Priors

Minghang Zhu, Jianshi Wu, Yuxin Guo, Wen Li and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Geometry is an Operator: Lie-Algebraic Space Routing for View-Robust 3D MLLMs

Jingjun Yi, Chen Hu, Qi Bi, Hao Zheng and 6 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

TPRL: Adaptive Visual Token Pruning in LVLMs via Language-Guided Reinforcement Learning

Sihan Cao, Jianwei Zhang, Pengcheng Zheng, Jiaxin Yan and 6 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

MedVIGOR: Visual Evidence Internalization for Observation-Driven Reasoning in Medical VLMs

Yuan Wu, Jiayu Qian, Sipeng Wu, Songpan Gao and 5 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Unlocking Fine-Grained Perception in CLIP via Structurally-Aware Latent Masked Modeling

Juntong Li, Lingwei Dang, Haomin Wu, Ziyan Qiu and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Contextualized Evaluation of Vision Language Models through Dynamic Interviews

Yijiang Li, Huiqi Zou, Bingyang Wang, Ziang Xiao

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Lost on Campus: Evaluating Embodied Spatial Reasoning of Vision-Language Models in the Wild

Zehan Zheng, Yanyuan Chen, Deming Li, Yutao Tang and 4 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

RobustPruner: Decoupled Relevance and Uncertainty for Efficient Visual Token Pruning in MLLMs

Hanwei Zhu, Junhan Fu, Xi Zhang, Jiamang Wang and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

MMLongCite: A Benchmark for Evaluating Faithfulness of Long-Context Vision-Language Models

Keyan Zhou, Zecheng Tang, Lingfeng Ming, Qiguang Chen and 8 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

In-Context Learning Can Help Vision Language Models Overcome Training Prior

Kun Wang, Xindi Wu, Sanghyuk Chun, Olga Russakovsky and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Heads That Write, Not Just Point: Image Retrieval Heads in Vision-Language Models

Junsung Park, Uiwon Hwang, Donghun Kang, Yeongtak Oh and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

From Perception to Punchline: Empowering VLM with the Art of In-the-wild Memes

Xueyan Li, Yingyi Xue, Mengjie Jiang, Qingzi Zhu and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Counterfactual Instruction Grounding for Vision-Language-Action Models

Shiyu Liu, Yu Zhou, Meng Liu, Xuanming Guo and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

MM-SCALE: Evaluating Evidence-Grounded Moral Judgment in Vision-Language Models

Eunkyu Park, Wesley Deng, Cheyon Jin, Matheus Kunzler Maldaner and 7 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Rethinking Contrastive Targets in Radiology Language-Image Pretraining

Yunhe Gao, Ashwin Kumar, Jiaming Liu, Chong Wang and 7 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

BoxTuning: Object-Aware Visual Prompting for Multimodal Model Fine-Tuning

Zekun Qian, Ruize Han, Wei Feng

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Anchoring LLM-based Chest X-ray Report Generation via Diffusion Language Planning

Jiechao Gao, Chang Liu, Yuandong Pan, Ying Liu and 1 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

CompJudge: Fine-Grained Comparative Evaluation using Multimodal LLM for Subject-Driven Generation

Nam Hyeon-Woo, Wenbin Ouyang, Ciprian A Corneanu

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

FedIBS: Federated Vision-Language Adaptation via Intrinsic Bias Selection

Xiaoming Wu, Wei Wenyu, Xin Wang, Ming Yang

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

TwinPrune: Density-Aware Two-Phase Token Pruning for Vision-Language Models

Kai Liu, Anqi Li, Junxian Li, Zhixin Wang and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Aligning MLLMs with the Latent Structure of Human Cognition via Behavior-Derived Semantic Dimensions

Ning E, Changde Du, Yizhuo Lu, Huiguang He

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Strong Helps Weak: Directional Cross-Modal Alignment Transfer in Multi-modal LLMs

Hoigi Seo, Byung Hyun Lee, Minjun Kim, Dohyun Mah and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

AIR: Rethinking Image-Text Offset Alignment in Multimodal Contrastive Representation Space

Guimeng Liu, Milad Abdollahzadeh, Ngai-Man (Man) Cheung

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Are Multimodal Benchmarks Really Useful? Item-Level Multimodal Benchmark Diagnosis via Structure-Response Co-Calibration

Shiqi Zhang, Weixin Zeng, Ziheng Zhang, Jiuyang Tang and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

MMCompass: Diagnosing Position Bias in Generative Multimodal Reward Models

Hongbo Zhao, Mingkun Yang, Yantao Liu, Yichang Zhang and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

LEGO: Sizing Rules for Budget-Aware Dense-to-MoE Conversion of Vision-Language Models

Qishen Yin, ZiangWu, Juntong Wu, Peng Jin and 5 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Learning to See Through Language: An Exploration of Language Modeling's Effect on Visual Representations

Mark Endo, Shiye Su, Yuhui Zhang, Serena Yeung-Levy

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

From Views to Worlds: Active Exploration over 3D Worlds for Vision-Language Models

Qijian Tian, Jiayu Ying, Ke Fan, Lizhuang Ma and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Logit-Conditioned Diffusion Decoding for Frozen Discrete-Token VLMs

Ji Woo Hong, Hee Suk Yoon, Gwanhyeong Koo, Eunseop Yoon and 4 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

When Medical VLMs Stop Understanding: MedTEC-Bench for Probing Semantic Specificity

Ahsan H Akash, Alina Devkota, Donald A Adjeroh, Binod Bhattarai and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Hear, Localize, and Reason: Spatially Aware Scene Understanding for Audio-visual LLMs

Sung-Bin Kim, Lee Jung-Mok, Jinwoo Jung, Oh Hyun-Bin and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

RUBRIC-MME: Real-User Behavior-grounded Rubric for Multimodal Interaction Capability Evaluation

Jiajie Teng, Jianping Jiang, Huiyu Duan, Jingdong Chen and 7 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

GEAR-Align: Grounding-Evidence-Aware Gradient Routing for Multimodal Alignment

Yu Yongkang, Haobo Wang, Meng Chen, Han Fang and 6 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

E0: Expressive Fine-Grained Discrete Action Prediction for Vision-Language-Action Models via Tweedie Discrete Diffusion

Zhihao Zhan, Jiaying Zhou, Likui Zhang, Qinhan Lyu and 9 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

V-LUMEN: Visual Lookup Memory for Embedding Scaling in Vision-Language Models

Miso Choi, Daekeun Kim, Eunji Kim, Jungbeom Lee

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Mission Impossible: Diagnosing and Fixing Non-Operative Instruction Following in Image Editing

Guoyizhe Wei, Feng Wang, Alan Yuille, Rama Chellappa

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

GIVLA: Deep Geometry Internalization for A Lightweight VLA via Geometry Instruction and Gradient-Informed Training

YUNHE LI, Qiming Liu, Haoyuan Wang, Hesheng Wang

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

A General Concept-based Decomposition for Vision–Language Embeddings

Simone Alberto Peirone, Ortal Senouf, Francesca Pistilli, Giuseppe Averta and 1 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Evaluating Spatiotemporal Reasoning of Vision-Language Models in Atari Gameplay

Mingjia Huo, Yao Fu, Bo Chang, Yaqing Wang and 7 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Too Aligned to be Real: Detecting AI-Generated Images via Cross-modal Alignment Shift

Haifeng Zhang, Qinghui He, Xiuli Bi, Bo Liu and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Spatial Representation Distillation and Knowledge Routing for Vision-Language-Action Models

Jungin Park, Chaoran Zhu, Changjae Oh, Kwanghoon Sohn

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

MLLM Makes Strong Backbone for Multi-Modal Object Detection

Yifeng Yang, Yutong Li, Jubo Feng, Qinying Gu and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Does the Question Really Matter? Training-Free Data Selection for Vision-Language SFT

Peng Sun, Yi Yang, Huawen Shen, Yi Ban and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

What Sketches Tell Us about LVLMs: Conventions, Grounding, and Localisation

Chaitat Utintu, Ahmed Bourouis, ARKAPRABHA BASU, Yi-Zhe Song

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

DRAMA: Dissecting Attention Redundancy for Accelerating Multimodal Diffusion Large Language Models

Yi-Hsin Hung, Fangfu Liu, Shao-Yuan Lo, Chu-Song Chen

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

RECAP: Looking Once Is Not Enough for Vision-Language Reasoning

Zhaolu Kang, Tailong Luo, Chenxin Li, Zhenyu Yu and 12 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

RankAlign: Unsupervised Vision-Language Representation Alignment via Rank Transformation

Enzhe Zhao, Marco Fiorucci, Lamberto Ballan

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Do Composed Image Retrieval Benchmarks Require Multimodal Composition?

Matteo Attimonelli, Alessandro De Bellis, Aryo Gema, Rohit Saxena and 7 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Exploiting Textual Semantics for Robust Cross-View Object Correspondence

Bing Fan, Yunhe Feng, Yan Huang, Heng Fan

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Position: Telic Errors Make VLMs Unreliable Annotators in Sensitive Contexts

Kokil Jaidka, Insyirah Mujtahid, Chebrolu Niranjan, Sahajpreet Singh

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Training-Free Active Test-Time Adaptation for Vision-Language Models

Jihwan Bang, Sumyeong Ahn, Hwanjun Song, Jae-Gil Lee

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

RADAR: Text-Guided Medical Image Segmentation via Residual Aggregation and Dense Alignment Representations

Jie Gui, Hang Tu, Yunfeng Dou, Wen Sha and 2 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

CATS: Acceptance-Oriented Critical Token Adaptive Selection for Multimodal Speculative Decoding

Kaiwen Liu, Yangkai Xie, Shuxia Lin, Liu Chonghan and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Reranking with Intra-modal Visual Association for Text-to-Image Person Re-Identification

Junhong Wang, Changxing Ding, Yusha Peng, Wentao Tan and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

SearchV: Evolutionary Fine-Grained Visual-Token Skipping for Efficient Vision-Language Models

Xinrui Chen, Zhen Huang, Shuwei Li, Fanyi Zeng and 7 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

ConfDet: Learning Reliable Confidence for MLLM-based Detection

Xuanjie Mao, Peng Ye, Ziteng Ma, Lin Zhang and 5 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Test-Time Adaptation via Self-Reinforced Optimal Transport for Zero-Shot OOD Detection with Vision–Language Models

MingCai Chen, Heng-yang Lu, Yuntao Du, Baoming Zhang and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Dual-Pronged LoRA: Achieving Near-Zero Forgetting and High-Performance Adaptation for MLLMs

Jingqi Ye, haonan he, Minglei Li, Fujun Han and 2 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Listening to the Retriever: Perturbation-Sensitive Question Selection for Interactive Person Retrieval

Yunhui Shao, Yang Bai, Shuai You, Bin Yang and 3 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Mitigating Saliency Collapse: Robust Saliency-Aware Long-Text Image-Text Alignment

Qiuyu Kong, Zanxi Ruan, Marco Cristani, Yiming Wang

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Modality-Depth Routing for Visual Reasoning in VLM Post-Training

Yiming Ren, Yiran Xu, Chufan Shi, Yu Qiao and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

MMDiff: Multimodal Model Diffing for Feature Discovery and Control

Lachin Naghashyar, Hunar Batra, Ashkan Khakzar, Philip Torr and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study

Hao Dong, Hongzhao Li, Shupan Li, Muhammad Haris Khan and 2 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Tune-Up Open-Weight CLIP: Optimization Framework for Self-Supervised Fine-tuning of CLIP

Anant Mehta, Xiyuan Wei, Xingyu Chen, Tianbao Yang

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

SP$^2$ec: Adaptive Self-Speculative Decoding for Vision-Language Models

Yuqi Huang, Xingyao Li, Yunlong Hou, Fengzhuo Zhang and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder

Zidan Wang, Yaqian Li, xiaokai zhang, Kun He and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Saliency-Aware Multi-Route Thinking: Grounding and Reasoning on Vision-Language Agents

Mingjia Shi, Yinhan He, Yaochen Zhu, Cassie Dong and 1 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

When Depth Lies: Benchmarking Vision-Language Models on Mirror-Induced RGB-D Ambiguity

HAO YIN, Tianchen Guo, Heming Du, Yan Ke and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Anatomy-Activated Mixture-of-Experts for 3D Medical Vision-Language Pre-training

Szymon Płotka, Gizem Mert, Pedro R. A. S. Bassi, Wenxuan Li and 8 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Phase-Adaptive Fusion: Spatio-Temporal Modulation for VLA Models

Lueke Ni, Junhao Zhu, Xinyu Zhang

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Diagnosing and Repairing Visual Collapse in Compact Medical Multimodal LLMs

Jinjie Xie, Baihua Li, Qinggang Meng

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

PRISM: Spectral Pruning and Reconstruction for Parameter-Efficient Model Merging of MLLMs

Wuxuan Shi, Haotian Chen, He Li, Qiang Yang and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Permute-then-Adapt: Weak-to-Strong Contrastive Image--Text Adaptation

Jinhao Li, Sarah Erfani, Lei Feng, Guangrui Li and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

CCDiff: Inverse Canonical Correlation Analysis for Discovering Visual Differences in Natural Language

Neelesh Bisht, Xingjian Li, Zihan Li, Bo Jiang and 4 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

CrossWeave: Emergent Cross-Modal Scene and Instance Retrieval from Sparse 2D-3D Alignment

Aadith Warrier, Gnana Prakash Punnavajhala, Siddharth Tourani, Muhammad Haris Khan and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

PolyVision: Conditional Visual Scaling Via Dynamic Expert Routing For Vision-Centric MLLMs

Tiehan Fan, Chen Zhao, Nikai Du, Zili Yi and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AesGI-Bench: Benchmarking and Evaluating the Aesthetic Quality of AI-Generated Images via Large Multimodal Models

Wang Jiarui, Xubo Su, Huiyu Duan, WeizeSun and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

CURE: Visual Reprogramming of Vision-Language Models under Limited Supervision

Lingzhi Wang, Wei Wang, Xinyang Chen

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

V-GIFT: Boosting Visual Instruction Tuning with Self-Supervised Guidance

Sophia Sirko-Galouchenko, Monika Wysoczańska, Andrei Bursuc, Nicolas THOME and 1 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Perturb, Repair, Verify: Self-Play Vision-Language Verifiers for Compositional Understanding

Zijun Huang, Yingjun Du, Cees Snoek

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Lang-SVG: Hierarchical Image Vectorization with Language Priors

Xi Liu, Chaoyi Zhou, Run Wang, Jiaang Li and 3 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Learning When to Think: Dual-Reference Offline Optimization for Adaptive VLM Reasoning

Hongbin Lin, Sizhe Zou, Juangui Xu, Xinyue Xu and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Concept-Aware Wasserstein Routing with Vision-Language Guidance for Few-Shot WSI Classification

Ankit Kumar, Shounak Das, Sandeep Kumar, Kaustubh Atey and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Batch-Conditioned Semantic Anchors for Robust Transductive Adaptation of Vision--Language Models

Mohammed Rahman Sherif Khan Mohammad, Ardhendu Behera, Sandip Pradhan, Swagat Kumar and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

SD-LoRA: Training-Time Structural Distillation into LoRA for Few-Shot Vision-Language Adaptation

Mohammed Rahman Sherif Khan Mohammad, Ardhendu Behera, Sandip Pradhan, Swagat Kumar and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Identifiable alignment of unpaired representations

Stas Syrota, Albert Lopez i Serrano, Johanne Franck, Søren Hauberg

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

SkillCIR: Intent-Guided Skill Composition for Training-Free Composed Image Retrieval

Yuanmin Tang, Lin Li, Yang Du, Yuan Gao and 7 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Mixture of Attribute-Aware Attention Experts for Fine-grained E-Commerce Composed Image Retrieval

Yufei Ma, Zihan Liang, Zhipeng Qian, Huangyu Dai and 4 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

From Structural Feedback to Prompt Policies: Learning Faithful Text-to-Image Prompt Editors

Wentao Ye, Yali Ye, Zhiqing Xiao, Ru Peng and 6 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

TAMEing the Open-World Personalization: Towards Open-Set Personalized MLLM Assistant

Rongpei Hong, Jian Lang, Ting Zhong, Fan Zhou

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

BRIDGE: Brain-Vision Representation Integration through Depth and Granularity Encoding

Chenyuan Hong, Binghao Ye, Yufei Guo, Guoqi Li

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Scaling Laws for Multimodal Data Mixtures

Aditi Khandelwal, Ayush Kumar Tarun, Yixuan Xu, Imanol Schlag and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding

CRS unites geometric road graphs with open vocabulary semantics to generate structured reasoning data, showing small models trained on few scenes surpass large vision-language models at structured road reasoning.

Lena Wild, Katie Luo, Marco Pavone

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

TokenSwap benchmarks and reduces MLLMs' modality gap by interleaving visual tokens with text, finding reasoning models have smaller gaps and training with TokenSwap mitigates it.

Andong Hua, Colton Bishop, Igor Mordatch, Arian Hosseini and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

MulCLIP: A Multi-level Alignment Framework for Enhancing Fine-grained Long-context CLIP

MulCLIP aligns images with long captions via multi-level token and patch strategies, improving fine-grained vision-language understanding without region proposals.

Chau Truong, Hieu Ta, Zhenzhen Liu, Dung Le

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Visual Instruction Tuning Aligns Modalities through Abstraction

Visual instruction tuning embeds image features into LLM intermediate semantic layers, aligning them with text abstractions to drive multimodal processing.

Luis Palacios, Lorenzo Basile, Diego Doimo, Alberto Cazzaniga

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity

Multimodal LLMs outperform pathology foundation models in cross-institution histological similarity by avoiding shortcut acquisition features tied to learning objectives rather than scale.

Yishu Zhang, Yun Li, David Zhang

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

3D-PLOT-LLM: Part-Level Object Tokens for 3D Large Language Models

3D-PLOT-LLM inserts learnable part tokens into frozen point features to enable part-level reasoning in 3D LLMs with under 1M parameters, outperforming prior part-aware models on part-QA and grounded description benchmarks.

Jintang Xue, Xinyu Wang, Yixing Wu, Jingwen Chen and 1 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

AMPS: Adaptive Modality Preference Steering via Functional Entropy

AMPS uses instance-aware functional entropy to adaptively steer multimodal model modality preferences, improving control while minimizing inference errors.

Zihan Huang, Xintong Li, Rohan Surana, Tong Yu and 4 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization

MM-IssueLoc benchmarks multimodal repository-level issue localization using visual evidence across 652 instances, showing current systems achieve under 39% file accuracy and text-only scores do not transfer.

Shaoxiong Zhan, Shi Hu, Hai Lin, BoyuFeng and 6 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration

Instruction tokens structurally anchor modality arbitration via shallow-layer multimodal buffering and deep-layer selective subspace resolution by sparse attention heads, with targeted intervention validating functional specificity.

Yu Zhang, Mufan Xu, Xuefeng Bai, Kehai Chen and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

A systematic evaluation of vision-language models for observational astronomical reasoning tasks

AstroVLBench evaluates VLMs across five astronomical modalities, finding accuracy depends on physical grounding and raw numerical data improves results, yet all models lag behind domain-specialized methods.

Wenke Ren, Hengxiao Guo, Wenwen Zuo, Xiaoman Zhang

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

CoMet: Context and Multiplicity Decomposition for Multimodal Uncertainty Estimation

CoMet decomposes multimodal LLM uncertainty into context and multiplicity terms via a lightweight module, improving calibration without generation or sampling.

Sanghyuk Chun, William Yang, Amaya Dharmasiri, Olga Russakovsky

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

AVIS: Adaptive Test-Time Scaling for Vision–Language Models

AVIS introduces a per-query adaptive policy that jointly scales visual token pruning and reasoning rollouts to improve vision-language model accuracy-compute trade-offs.

Ahmadreza Jeddi, Minh Le, Amirhossein Kazerouni, Hakki Karaimer and 7 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder

ProCLIP progressively aligns CLIP image encoders with LLM-based embedders via curriculum distillation and contrastive tuning to support long multilingual texts without disrupting pretrained vision-language alignment.

Xiaoxing Hu, Kaicheng Yang, Ziqi Ye, Ziyang Gong and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 11 on Hugging Face · Code ★ 27

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop

PhysVista benchmarks physical intelligence in vision-language models via a perception-reasoning-assessment loop, exposing major gaps in physical reasoning and plausibility assessment.

Xinge Peng, Yiting Lu, tianwu zhi, Wen Wen and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 30 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Direct Product Flow Matching: Decoupling Radial and Angular Dynamics for Few-Shot Adaptation

Direct product flow matching decouples radial and angular dynamics via constant-speed geodesic transport and hidden-state conditioning for state-of-the-art few-shot vision-language adaptation.

Hongxu Chen, Yanghao Wang, Bowei Zhu, Hongxiang Li and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation

DocPTBench introduces 1,300 photographed documents for parsing and translation, showing MLLMs drop 18% parsing and 12% translation accuracy versus digital-born documents.

Yongkun Du, Pinxuan Chen, Xuye Ying, Zhineng Chen

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · Code ★ 18

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

VVTRec: Radio Interferometric Reconstruction through Visual and Textual Modality Enrichment

VVTRec uses visual and textual visibility enrichment with vision-language models to reduce artifacts and improve radio interferometric image reconstruction.

Kai Cheng, Ruoqi Wang, Qiong Luo

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

AdaM-Rec: Adaptive Modality Routing for Multimodal Recommendation

AdaM-Rec adaptively routes between textual and visual modalities via LLM-based proxy recall tasks to improve multimodal recommendation accuracy.

Honghao Fu, Jiacheng Chen, Manxi Lin, Junjun Zheng and 6 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models

Multimodal LLMs suffer spurious cross-modality interference that distorts decisions, and a unified finetuning framework with perturbation augmentation and consistency regularization improves robustness and generalization.

Rui Cai, Bangzheng Li, Xiaofei Wen, Muhao Chen and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face · Code ★ 8

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Multi-view Relational Distillation for Spatial Reasoning with Vision-Language Models

Multi-view relational distillation improves vision-language model spatial reasoning by distilling cross-view patch similarities rather than features, preserving language alignment with minimal overhead.

Kiet Nguyen, Hanbo Shim, Jinwoo Kim, Seunghoon Hong

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

M3-AD: Reflection-aware Multi-modal, Multi-category, and Multi-dimensional Benchmark and Framework for Industrial Anomaly Detection

M3-AD proposes a reflection-aware multimodal benchmark and RA-Monitor framework that improves industrial anomaly detection via learnable self-correction, outperforming several MLLMs.

Chao Huang, Yanhui Li, Hongxi Huang, Yunkang Cao and 5 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation

AgenticOCR transforms OCR into query-driven, on-demand extraction to selectively parse document regions, improving visual RAG efficiency and accuracy over page-level chunking.

Zhengren Wang, Dongsheng Ma, Huaping Zhong, Jiayu Li and 5 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Allocentric Perceiver: Disentangling Allocentric Reasoning from Egocentric Visual Priors via Frame Instantiation

Allocentric Perceiver recovers 3D geometry and transforms it into query-aligned target frames to offload mental rotation from VLMs, boosting allocentric spatial reasoning by ~10% without training.

Hengyi Wang, RuiQiang Zhang, Chang Liu, Guanjie Wang and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

Text as Partial Constraint: Core–Residual Alignment for Robust Vision–Language Learning

TPC treats captions as partial constraints, aligning vision-language representations to a consensus semantic core while penalizing dependence on unsaid residuals, yielding robust zero-shot recognition and improved LVLM grounding.

Chengzhen Yu, Canran Xiao, SiYuan Ma, Yang Liu

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation

Animation2Code benchmarks video-to-code generation for web animations, showing state-of-the-art vision-language models struggle with temporal consistency despite high appearance fidelity.

Anya Ji, Abhijith Varma Mudunuri, David Chan, Alane Suhr

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs

OmniSpace improves autonomous vehicle MLLM spatial reasoning via camera pose injection, multi-view epipolar attention, and 3D geometric distillation without auxiliary 3D models, surpassing existing methods across planning, risk detection, and language benchmarks.

Anh Hao Vo, Phu Loc Nguyen, Khoa Vo, Sieu Tran and 6 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

LensVLM: Selective Context Expansion for Compressed Visual Representation of Text

LensVLM lets VLMs scan compressed rendered text and selectively expand only relevant regions via learned tools, maintaining near-full accuracy at 4.3x compression and outperforming baselines up to 10.1x across text QA benchmarks.

Roy Xie, Dan Friedman, Donghan Yu, Bowen Pan and 6 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Same Concept, Different Directions: Cross-Modal Feature Heterogeneity in Sparse Autoencoders

Cross-modal sparse autoencoders exhibit feature heterogeneity where shared concepts activate different latents across image and text modalities, and training modality-specific autoencoders with post-hoc alignment improves reconstruction, retrieval, and steering.

Chungpa Lee, Jihoon Kwon, Kyle Min, Jy-yong Sohn

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

dRAE: Representation Autoencoder with Hyper-Spherical Codes

Hyper-spherical quantization decouples semantics from magnitude via angular routing to prevent codebook collapse, enabling scalable discrete representation autoencoders with full codebook usage and high-fidelity reconstruction.

Tianren Ma, Lin Long, Chuyan Chen, Mu Zhang and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 10 on Hugging Face · Code ★ 10

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

ITO: Multi-View Alignment and Training-Time Fusion for Image-Text Pretraining

ITO improves image-text pretraining via multi-view cross-modal alignment and discarded training-time fusion, beating CLIP at 100M-1B scale on classification and retrieval.

Hanpeng Liu, Yaqian Li, Zidan Wang, Shuoxi Zhang and 5 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

THEIA: A Multimodal Dataset and Benchmark for Vision-Language Analysis of Layout

THEIA introduces a vision-language dataset and benchmark for analog circuit layout analysis via GDSII images, showing fine-tuned models outperform general VLMs by up to 73%.

Giuseppe Chiari, Michele Piccoli, Federico Viola, Davide Zoni

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

MARS: Harmonizing Multimodal Convergence via Adaptive Rank Search

MARS introduces adaptive rank search balancing multimodal convergence dynamics via dual scaling laws to optimize low-rank fine-tuning of multimodal large language models.

Minkyoung Cho, Insu Jang, Shuowei Jin, Zesen Zhao and 4 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Closing the Indexing-Decoding Gap in Multimodal Generative Retrieval via Prefix Retention Optimization

PRO closes the indexing-decoding gap in multimodal generative retrieval via prefix ranking distillation, vocabulary scheduling, and geometric score fusion to improve beam search retention and retrieval accuracy.

Yufei Chen, Zihan Wang, Yubao Tang, Yukun Zhao and 2 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Binding Visual Features Point by Point

Pointing via text induces internal visual search routines that eliminate binding errors, enabling compositional generalization and solving vision-language binding via serial processing.

Udith Haputhanthri, Declan Campbell, Rim Assouel, Jonathan D Cohen and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Topology-Aware Representation Alignment for Semi-Supervised Vision-Language Learning

ToMA uses persistent homology to align cross-modal manifold edges via image-text pairs, improving semi-supervised vision-language learning in specialized domains.

Junwon You, Mihyun Jang, Sangwoo Mo, Jae-Hun Jung

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

SALART-VQA: Diagnosing Whether VLMs Understand Salient Artifacts in Generated Images

SalArt-VQA evaluates VLM artifact understanding via fine-grained questions, revealing high detection recall but only 53% fully correct reasoning and a sensitivity-calibration tradeoff.

Xiaoxiao Sun, Ruotian Zhang, Junzhe Huang, James Burgess and 1 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

OneVision-Encoder applies codec-aligned sparsity to video, processing only high-entropy regions to outperform dense backbones with fewer tokens. It achieves 4.1% higher video accuracy than Qwen3-ViT across 16 benchmarks.

feilong tang, Xiang An, Yunyao Yan, Yin Xie and 14 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026 · ▲ 52 on Hugging Face · Code ★ 403

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Tango3D: Towards Alignment for Global and Local 2D-3D Correspondence

Tango3D unifies global retrieval and dense pixel-to-point correspondence via shared 2D-3D alignment with progressive training.

Zebin He, Mingxin Yang, Shuhui Yang, Hanxiao Sun and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

A dataset of multi-aspect human visual similarity judgments benchmarks vision-language models and yields the TPIPS metric, which aligns with human perception and enables text-guided image retrieval and generative evaluation.

Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

JMed48k introduces a Japanese medical licensing benchmark with 48,862 questions showing proprietary vision-language models gain substantially from images while medical-specific systems ignore visual evidence.

Yue Xun, Junyu Liu, Qian Niu, Xinyi Wang and 9 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

FiLoRA: Focus-and-Ignore LoRA for Controllable Feature Reliance

FiLoRA is an instruction-conditioned LoRA framework that modulates multimodal model reliance on internal feature pathways via gated low-rank modules, enabling controllable amplification or suppression of feature groups without changing task semantics.

Hyunsuk Chung, Caren Han, Seungyeon Ji, Jinwoo Kim and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

FinMTM: A Multi-Turn Multimodal Benchmark for Financial Reasoning and Agent Evaluation

FinMTM introduces a bilingual multi-turn multimodal benchmark with 11,133 financial visual QA pairs to evaluate vision-language models on reasoning and agent tasks, revealing significant limitations in fine-grained perception and complex workflows.

Chenxi Zhang, Ziliang Gan, Liyun Zhu, Youwei Pang and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Neural‑Visual Decoding via Cognitive‑guided Adaptive Blurring and Information‑Constrained Alignment

CAIA improves EEG visual decoding by cognitively guided adaptive blurring and information-constrained alignment, significantly boosting zero-shot brain-to-image retrieval accuracy.

Fan Yin, Chuhang Zheng, Peiliang Gong, Donghai Guan and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

OmniGF: A Dual-Branch Vision-Language Framework for Unified Gaze Following

OmniGF unifies multi-person gaze following via dual-branch vision-language decoding with head embeddings, achieving state-of-the-art spatial, semantic, and social gaze reasoning.

Qiaomu Miao, Haoyu Wu, Jingyi Xu, Minh Hoai and 1 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
91%Must read
?Must readVote to see the score

CardioLens: Revealing the Clinical Reality Gap of MLLMs via Multi-Sequence Cardiac MRI Evaluations

CardioLens evaluates MLLMs on multi-sequence cardiac MRI, revealing poor clinical workflow performance and category-collapse failures despite reasoning prompts and slice selection.

Zixian Su, Hongkai Zhang, Fan Gao, Encheng Su and 11 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 5/5
89%Must read
?Must readVote to see the score

The Alignment Illusion in Multimodal Large Language Models

Standard similarity metrics show an alignment illusion in MLLMs because shared language-model pathways create weight-induced visual-text similarity; the proposed PA gap better tracks actual visual content integration via multi-directional structure.

Hong-Han Wang, Yuntao Wang, Hu Ding

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 3/5
medium 9/10
strict 4/5
76%Highly rated
?Highly ratedVote to see the score

Language-Assisted Image Clustering Guided by Discriminative Relational Signals and Adaptive Semantic Centers

A language-assisted clustering framework uses cross-modal relational signals and adaptive semantic centers to improve clustering accuracy by 2.6% over state-of-the-art methods.

Jun Ma, Xu Zhang, Zhengxing Jiao, Yaxin Hou and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios

MDPBench introduces a 3,400-image multilingual document parsing benchmark across 17 languages revealing open-source models suffer severe performance drops on photographed and non-Latin script documents.

Zhang Li, Lin Zhibo, Qiang Liu, Ziyang Zhang and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 9 on Hugging Face · Code ★ 895

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
88%Must read
?Must readVote to see the score

Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation

PU-DPO treats unmentioned radiology findings as unlabeled rather than negative, using edited contrastive pairs to prevent omission noise from corrupting preference optimization and improving hidden finding recovery.

Yuta Kobayashi, Pradyun Ramesh, Muhammad Ahmed Chaudhry, Vincent Jeanselme and 4 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
86%Must read
?Must readVote to see the score

GLINT: Sparsely Gated Vision-Language Alignment for Fine-Grained Radiology Representations

GLINT introduces sparse gating and dense feature regularization to learn fine-grained radiology vision-language representations that outperform baselines on classification, grounding, and zero-shot 3D CT segmentation.

Jonggwon Park, Seongeun Lee, Junhyun Park, Hannah Yun and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

SwiftVLM: Efficient Vision-Language Model Inference via Cross-Layer Token Bypass

SwiftVLM introduces cross-layer token bypass to preserve visual tokens across pruning stages, enabling training-free vision-language model acceleration with superior accuracy-efficiency trade-offs.

Qian, Xinran Yu, Xinan Wang, Danyang Li and 4 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Boosting Brain-to-Image Decoding with TRIBE v2 Data Augmentation

TRIBE v2 synthetic fMRI augmentation improves brain-to-image decoding by up to 68%, though optimal synthetic-to-real ratios vary by dataset, and synthetic-only training achieves above-chance zero-shot decoding.

Yohann Benchetrit, Marlene Careil, Simon Dahan, Hubert Banville and 2 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
80%Must read
?Must readVote to see the score

NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning

NeuronEye improves vision-language reasoning by selectively activating query-relevant sparse visual concept clusters and suppressing dominant cues during inference in frozen VLMs. It raises CV-Bench accuracy by +3.1 and BLINK Multi-view by +8.3 on Qwen2.5-VL-7B without retraining.

Ruiyu Yan, Bowen Chen, Shaowen Wan, Lin Zhao

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution

ViCO minimizes vision tokens via consistency training across compression ratios, cutting tokens up to 50% while preserving capabilities through semantic-based routing.

Long Cui, Weiyun Wang, Jie Shao, Zichen Wen and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

BayesRAG: Probabilistic Mutual Evidence Corroboration for Multimodal Retrieval-Augmented Generation

BayesRAG applies Bayesian inference and Dempster-Shafer theory to fuse multimodal retrieval evidence, improving retrieval confidence via cross-modal mutual corroboration. It significantly outperforms state-of-the-art methods on challenging multimodal benchmarks.

Xuan Li, Yining Wang, Haocai Luo, Shengping Liu and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · Code ★ 6

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5
88%Must read
?Must readVote to see the score

Half-Truths Break Similarity-Based Retrieval

CLIP-style dual encoders often prefer half-true image descriptions with incorrect added details over correct shorter ones due to weak part-level supervision; CS-CLIP improves half-truth accuracy to 69.3% through component-level contrastive fine-tuning.

Bora Kargi, Arnas Uselis, Seong Joon Oh

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026 · ▲ 6 on Hugging Face · Code ★ 15

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
88%Must read
?Must readVote to see the score

When a Zero-Shooter Cheats: Improving Age Estimation via Activation Steering

Vision-language models use celebrity identity shortcuts rather than visual age cues, and activation steering suppresses this to cut mean absolute error by up to 25%.

Erik Imgrund, Pia Hanfeld, Klim Kireev, Konrad Rieck

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

ReToken: One Token to Improve Vision–Language Models for Visual Retrieval

ReToken introduces one learnable retrieval token that selects sparse visual tokens from long contexts, improving vision-language models by up to 13.4 points on visual retrieval while fitting on a single GPU.

Yao Xiao, Reuben Tan, Zhen Zhu, Yuqun Wu and 2 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 8 on Hugging Face · Code ★ 18

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
80%Must read
?Must readVote to see the score

EasyLens: A Training-Free Plug-and-Play Subtle-Lesion Representation Amplifier for Medical Vision-Language Models

EasyLens is a training-free plug-and-play module that amplifies subtle lesion representations in frozen medical vision-language models via prototype-based patch selection and morphology-guided residual enhancement, improving detection across datasets.

QIWEI ZENG, Hao Wang, Jinghao Lin, Shuchang Ye and 5 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
88%Must read
?Must readVote to see the score

PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition

PoseBridge extracts pose-anchored semantic cues from human pose estimation to recover upstream visual context lost in skeletons, improving zero-shot skeleton action recognition by up to 17.4 points.

Sanghyeon Lee, Jinwoo Kim, Jong Taek Lee

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

DriveSpatial: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving

DriveSpatial benchmarks vision-language models' spatiotemporal autonomous driving intelligence, finding a 28.4-point human gap with cognitive scene construction as the key bottleneck.

Anh Hao Vo, Khoa Vo, Phu Loc Nguyen, Sieu Tran and 9 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
74%Highly rated
?Highly ratedVote to see the score

From Denoising to Refining: A Corrective Framework for Vision-Language Diffusion Model

ReDiff reframes vision-language diffusion as active refining via error-correction training and online self-correction loops, breaking error cascades to improve coherence, factual accuracy, and parallel generation.

Yatai Ji, Teng Wang, Yuying Ge, Zhiheng Liu and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 30 on Hugging Face · Code ★ 45

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

LithoBench: Benchmarking Large Multimodal Models for Remote-Sensing Lithology Interpretation

LithoBench evaluates large multimodal models on remote-sensing lithology interpretation via expert-annotated multi-level tasks, revealing substantial limitations in geological reasoning.

jun wang, Fengpeng Li, Tianjin Huang, Hang Dong and 1 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 2/5
89%Must read
?Must readVote to see the score

GT-Free OCR Metrics: A Reference-Free Evaluation Framework for Document OCR Systems

A render-and-compare framework benchmarks 147 reference-free visual metrics for document OCR, with the best composite achieving ρ = 0.494 correlation to ground-truth quality. Silent region misclassification artificially preserves visual similarity, suppressing correlation when masking is omitted.

Kshitij Singh

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Attention Alignment Between Humans and Vision-Language Models

Vision-language model decoder architecture dominates human attention alignment, with LSTM decoders reaching 85, 87% of the human noise ceiling but remaining diffuse, while transformer decoders show sharper task differentiation despite lower alignment; encoder effects are secondary and neural predict

Isaac Christian, Udith Haputhanthri, Declan Campbell, Samuel Nastase and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 3/5
72%Highly rated
?Highly ratedVote to see the score

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models

EyeVLM benchmarks vision-language models on gaze following and social gaze prediction, finding they lack precise gaze understanding despite training improvements.

Hengfei Wang, Anshul Gupta, Pierre Vuillecard, Jean-marc Odobez

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5
80%Must read
?Must readVote to see the score

COVD: Continual Open-Vocabulary Object Detection with Novel Concept Injection

COVD enables continual open-vocabulary detection via novel concept injection, and NoIn-Det freezes visual encoders to align text representations with new concepts without extra parameters, outperforming existing continual methods on Novel-114.

Yupeng Zhang, Ruize Han, Yuzhong Feng, Zixin Ren and 2 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
86%Must read
?Must readVote to see the score

Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding

CAFT learns local text-region alignments before global image-text matching via hierarchical encoders, achieving state-of-the-art long-caption retrieval without region supervision.

Byeongju Woo, Zilin Wang, Byeonghyun Pak, Sangwoo Mo and 1 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 0/5
83%Must read
?Must readVote to see the score

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation

Vision-OPD distills a crop-conditioned teacher into a full-image student via on-policy self-distillation to improve fine-grained visual understanding without external teachers or tools. It achieves competitive or superior performance on fine-grained benchmarks against larger open-source, closed-sour

Qianhao Yuan, Jie Lou, XingYu Li, Hongyu Lin and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 0/5