Good Papers

Showing Multimodal reasoning Show all papers

80%Must read
?Must readVote to see the score

CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG

CANOPY adaptively compresses multimodal evidence via hierarchical region scoring and targeted retrieval, improving QA accuracy while reducing input tokens by up to 27.7%.

Hyojeong Yun, Jueun Kim, Wook-Shin Han

Published Oct 1, 2026 · 0 citations · ▲ 16 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
86%Must read
?Must readVote to see the score

AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models

AURAL uses adaptive latent reasoning with joint chunk prediction to match chain-of-thought performance while cutting first-token latency 11.8x versus explicit reasoning.

Yuxiang Wang, Kunyu Feng, Yuancheng Wang, Zihang Liu and 11 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
83%Must read
?Must readVote to see the score

EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos

EgoTools introduces a dataset and benchmark for egocentric tool-use reasoning, showing current models struggle with visual grounding while training improves performance.

Shulin Tian, Junsu Kim, Shuai Liu, Hao Li and 16 more

Published Sep 30, 2026 · 0 citations · ▲ 67 on Hugging Face · Code ★ 9

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
80%Must read
?Must readVote to see the score

OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning

OmniReasoning introduces a benchmark, data engine, and self-distillation method to improve audio-visual joint reasoning, boosting Qwen3-Omni-30B-A3B-Thinking by up to 12.8 points.

Junming Lin, Yuxuan Wang, Zhenxin Lei, Yuxin Liu and 10 more

Published Sep 30, 2026 · 0 citations · ▲ 17 on Hugging Face · Code ★ 1

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
78%Highly rated

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Imagine3D-LLM learns compact scene representations via learnable summary tokens decoded into 3D Gaussian splatting with photometric loss, improving multi-view spatial reasoning benchmarks.

Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han and 9 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published Sep 29, 2026 · ▲ 72 on Hugging Face · Code ★ 37

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning

WM-VLM adds a world-model branch to vision-language models to generate intermediate visual states for spatial reasoning, outperforming baselines by up to 39.25 points on mental rotation tasks.

Yuheng Zha, Yilei Wang, Qiyue Gao, Junrong Chen and 4 more

Published Sep 28, 2026 · 0 citations · ▲ 11 on Hugging Face · Code ★ 1

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
86%Must read
?Must readVote to see the score

Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

ReaLVR fixes latent reasoning's weak visual grounding by supervising latent tokens with visual evidence, improving reasoning across scales up to 235B.

Xi Xiao, Tianchen Zhao, Youngeun Kim, Zhuowei Li and 9 more

Published Sep 28, 2026 · 0 citations · ▲ 217 on Hugging Face · Code ★ 40

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
80%Must read
?Must readVote to see the score

BabyVision: Visual Reasoning Beyond Language

BabyVision benchmarks core visual reasoning without language and finds top MLLMs score far below human children.

Liang Chen, Weichu Xie, Yiyan Liang, Hongfeng He and 26 more

Published Jan 10, 2026 · 1 citation · ▲ 201 on Hugging Face · Code ★ 257

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
67%Highly rated
?Highly ratedVote to see the score

ACM SIGMM Multimodal Reasoning Workshop

The ACM SIGMM Multimodal Reasoning Workshop at IIT Patna gathered 108 participants to discuss multimodal reasoning, generative intelligence, and trustworthy AI through talks, tutorials, and hands-on sessions.

Sriparna Saha

Published 2026 · 0 citations

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 1/5
86%Must read

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Video generation enables unified multimodal reasoning via Sora-2, which matches vision-language models and exceeds GPT-5 on spatial tasks while scoring 92% on MATH.

Jingqi Tong, Yurong Mou, Hangcheng Li, Mingzhe Li and 10 more

Published Nov 6, 2025 · 0 citations · ▲ 242 on Hugging Face · Code ★ 319

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
91%Must read

CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models

CollabVR pairs vision-language models with video generation models in closed-loop step-level planning and verification, reducing drift and simulation errors for major video reasoning gains.

Joowon Kim, Seungho Shin, Joonhyung Park, Eunho Yang

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 71 on Hugging Face · Code ★ 10

100% Readers1 of 1 upvoted
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 2/5
67%Highly rated
?Highly ratedVote to see the score

Single-Pass Evidence Measurement for Interpretable and Uncertainty-Aware Multimodal Face Anti-Spoofing

Yingjie Ma, Haonan Wang, Xun Lin, Ruixin Zhang and 7 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Learning to Learn from Multimodal Experience

Xingyu Sui, Weixiang Zhao, Yongxin Tang, Yanyan Zhao and 3 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions

Junho Kim, Xu Cao, Houze Yang, Bikram Boote and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

CMI-Trans: Cross Modal Inconsistency-aware Transport for HSI-LiDAR Classification

Yanli Li, Xuan Tan, Ding Qi, XINYANG JIANG

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

EDMA: Entropy-Driven Multimodal Answering

Emanuele Mezzi, Gertjan Burghouts, Fabio Massacci, Mengyuan Zhang

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AeroMosaic:Transport-Aware Multimodal Evidence Fusion for Atmospheric Pollution Risk Inference

Muyang Zheng, Jiaming Ma, Zongyu Zhang, Qingsong Wen

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

Do More Modalities Always Help? A Geometric Perspective on Missing-Modality Robustness

Songyuan Sui, Zhen Tan, Mohan Zhang, Rana M Khan and 2 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

Intervene3D: Intervention-Based Controlled Inference for Multimodal Perception under Partial Observability

Constantino Msigwa, Denis Bernard, Jaeseok Yun

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?

Ziyi Wang, Li Li, Aolin Zhou, Yankun Shen and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
Show 20 more papers