Good Papers

Showing Vision-language models Show all papers

91%Must read

DEPICT: Scoring Text-to-Image Alignment by Answer Agreement

DEPICT scores text-to-image alignment via expected agreement between image and caption answers, boosting negation accuracy to 88% and exceeding fine-tuned evaluators on human correlation benchmarks.

Vasco Ramos, Sandra Godinho Silva, Joao Magalhaes, Ricardo Rei and 1 more

Published Oct 2, 2026 · 0 citations · ▲ 3 on Hugging Face

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
86%Must read
?Must readVote to see the score

Foresight: planning future perception in streaming VLMs without retraining

FORESIGHT uses dual-stream anticipatory planning in frozen streaming VLMs to dynamically configure future perception, improving online benchmarks by up to 18.7 points without retraining.

Ashok Prasad Neupane, Dipan Bartaula, Ankit Belbase, Saugat Adhikari and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 3 on Hugging Face · Code

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

LoopVL: Recurrent Visual Intelligence

LoopVL applies recurrent loop transformers to vision-language models via iterative shared-module updates, outperforming larger non-recurrent models and exhibiting visual aha moments.

Zhe Qian, Ziyang Gong, Zhongxing Xu, Hehan Li and 8 more

Published Sep 29, 2026 · 0 citations · ▲ 469 on Hugging Face · Code ★ 128

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training

PaddleOCR-VL-1.6 applies region-aware data optimization and progressive reinforcement learning post-training to achieve 96.33% on OmniDocBench v1.6.

Zelun Zhang, Hongen Liu, Suyin Liang, Yubo Zhang and 11 more

Published Jun 2, 2026 · 0 citations · ▲ 26 on Hugging Face · Code ★ 90,719

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

VisPlay: Self-Evolving Vision-Language Models from Images

VisPlay enables vision-language models to self-evolve via reinforcement learning on unlabeled images by having a questioner and reasoner generate and train on diverse visual reasoning tasks, improving reasoning, generalization, and reducing hallucinations across benchmarks.

Yicheng He, Chengsong Huang, Li, Zongxia, Jiaxin Huang and 1 more

Published Nov 19, 2025 · 0 citations · ▲ 45 on Hugging Face · Code ★ 79

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

MinerU2.5 decouples global layout analysis from local content recognition via coarse-to-fine parsing, achieving state-of-the-art document parsing accuracy with low computational overhead.

Junbo Niu, Zheng Liu, Zhuangcheng Gu, Bin Wang and 36 more

Published Sep 26, 2025 · ▲ 180 on Hugging Face · Code ★ 81,154

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Otter: A Multi-Modal Model With In-Context Instruction Tuning

Otter is a multi-modal model instruction-tuned with visual and textual in-context examples via the MIMIC-IT dataset, improving convergence and generalization on complex video and multi-image tasks.

Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang and 5 more

Published May 20, 2025 · 75 citations · Code ★ 3,443

– ReadersNo votes yet
10/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 21 reviewers recommend it
lenient 5/5
medium 4/11
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

SmolDocling is a 256M-parameter vision-language model for end-to-end document conversion using DocTags to capture content, structure, and spatial layout, matching models up to 27 times larger.

Ahmed Nassar, Andres Marafioti, Matteo Omenetti, Maksym Lysak and 9 more

Published Mar 14, 2025 · ▲ 177 on Hugging Face · Code ★ 68,465

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

Diffusion Instruction Tuning

Lavender aligns vision-language model attention with Stable Diffusion during supervised fine-tuning, boosting accuracy up to 30% with minimal training data.

Chen Jin, Ryutaro Tanno, Amrutha Saseendran, Tom Diethe and 1 more

Published Feb 4, 2025 · 0 citations · ▲ 2 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

Personalized Visual Instruction Tuning

PVIT introduces a framework that curates personalized visual instruction data to cure multimodal models' face blindness, significantly boosting personalized dialogue performance.

Renjie Pi, Jianshu Zhang, Tianyang Han, Jipeng Zhang and 2 more

Published Oct 9, 2024 · 0 citations · ▲ 70 on Hugging Face · Code ★ 34

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 4/5
medium 2/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding

ERNIE-Layout enhances document pre-training with layout knowledge for improved visually-rich document understanding.

Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo and 11 more

Published 2022 · 65 citations · Code ★ 106

– ReadersNo votes yet
8/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 21 reviewers recommend it
lenient 4/5
medium 4/11
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Hierarchical Semantic Tree Anchoring for CLIP-Based Class-Incremental Learning

Tao Hu, Lan Li, Zhenhao Wen, Da-Wei Zhou

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Revitalizing Medical Time Series with Vision-Informed Retrieval: A Vision-Language Perspective

Guoqi Yu, Juncheng Wang, Shujun Wang

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

C2FT: Enhancing Fine-Grained Perception in MLLMs via Confuse-then-Contrast Fine-Tuning

Shaoxuan He, Benlei Cui, Shikai Qiu, Yuwen Zhai and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Not Every Image Teaches Vision: Visual-Necessity-Gated Continual Learning for Multimodal Large Language Models

Jiamu Xue, Haidong Kang, Wen Luo, Jubo Chen and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Despa: Resolving Spatial Collapse in VLMs via Depth-Grounded Geometry

Yujing Lou, Pingyi Chen, Shen Cao, Lubin Fan and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Pixels over Symbols: Sensory Realism Improves Behavioral Alignment in Models of Cognition

Mark Bai, Xiaoxuan Lei, Zihan Weng, MOTAHAREH POURRAHIMI and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

GaitLingo: Self-Supervised Gait Representation Learning with Language Priors

Chenye Wang, Zhengxiang Lan, Saihui Hou, Zhikang Liu and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer?

Xueqi Cheng, Yushun Dong

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

HEED: Density-Weighted Residual Alignment for Hybrid Vision-Language Model Distillation

Yihao Liang, Niraj Jha

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
Show 20 more papers