Good Papers

Showing Multimodal reasoning Show all papers

80%Must read
?Must readVote to see the score

CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG

CANOPY adaptively compresses multimodal evidence via hierarchical region scoring and targeted retrieval, improving QA accuracy while reducing input tokens by up to 27.7%.

Hyojeong Yun, Jueun Kim, Wook-Shin Han

Published Oct 1, 2026 · 0 citations · ▲ 16 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models

AURAL uses adaptive latent reasoning with joint chunk prediction to match chain-of-thought performance while cutting first-token latency 11.8x versus explicit reasoning.

Yuxiang Wang, Kunyu Feng, Yuancheng Wang, Zihang Liu and 11 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos

EgoTools introduces a dataset and benchmark for egocentric tool-use reasoning, showing current models struggle with visual grounding while training improves performance.

Shulin Tian, Junsu Kim, Shuai Liu, Hao Li and 16 more

Published Sep 30, 2026 · 0 citations · ▲ 67 on Hugging Face · Code ★ 9

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning

OmniReasoning introduces a benchmark, data engine, and self-distillation method to improve audio-visual joint reasoning, boosting Qwen3-Omni-30B-A3B-Thinking by up to 12.8 points.

Junming Lin, Yuxuan Wang, Zhenxin Lei, Yuxin Liu and 10 more

Published Sep 30, 2026 · 0 citations · ▲ 17 on Hugging Face · Code ★ 1

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Imagine3D-LLM learns compact scene representations via learnable summary tokens decoded into 3D Gaussian splatting with photometric loss, improving multi-view spatial reasoning benchmarks.

Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han and 9 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published Sep 29, 2026 · ▲ 72 on Hugging Face · Code ★ 37

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning

WM-VLM adds a world-model branch to vision-language models to generate intermediate visual states for spatial reasoning, outperforming baselines by up to 39.25 points on mental rotation tasks.

Yuheng Zha, Yilei Wang, Qiyue Gao, Junrong Chen and 4 more

Published Sep 28, 2026 · 0 citations · ▲ 11 on Hugging Face · Code ★ 1

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

ReaLVR fixes latent reasoning's weak visual grounding by supervising latent tokens with visual evidence, improving reasoning across scales up to 235B.

Xi Xiao, Tianchen Zhao, Youngeun Kim, Zhuowei Li and 9 more

Published Sep 28, 2026 · 0 citations · ▲ 217 on Hugging Face · Code ★ 40

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

BabyVision: Visual Reasoning Beyond Language

BabyVision benchmarks core visual reasoning without language and finds top MLLMs score far below human children.

Liang Chen, Weichu Xie, Yiyan Liang, Hongfeng He and 26 more

Published Jan 10, 2026 · 1 citation · ▲ 201 on Hugging Face · Code ★ 257

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

ACM SIGMM Multimodal Reasoning Workshop

The ACM SIGMM Multimodal Reasoning Workshop at IIT Patna gathered 108 participants to discuss multimodal reasoning, generative intelligence, and trustworthy AI through talks, tutorials, and hands-on sessions.

Sriparna Saha

Published 2026 · 0 citations

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

Video generation enables unified multimodal reasoning via Sora-2, which matches vision-language models and exceeds GPT-5 on spatial tasks while scoring 92% on MATH.

Jingqi Tong, Yurong Mou, Hangcheng Li, Mingzhe Li and 10 more

Published Nov 6, 2025 · 0 citations · ▲ 242 on Hugging Face · Code ★ 319

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models

CollabVR pairs vision-language models with video generation models in closed-loop step-level planning and verification, reducing drift and simulation errors for major video reasoning gains.

Joowon Kim, Seungho Shin, Joonhyung Park, Eunho Yang

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 71 on Hugging Face · Code ★ 10

100% Readers1 of 1 upvoted
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Single-Pass Evidence Measurement for Interpretable and Uncertainty-Aware Multimodal Face Anti-Spoofing

Yingjie Ma, Haonan Wang, Xun Lin, Ruixin Zhang and 7 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Learning to Learn from Multimodal Experience

Xingyu Sui, Weixiang Zhao, Yongxin Tang, Yanyan Zhao and 3 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions

Junho Kim, Xu Cao, Houze Yang, Bikram Boote and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

CMI-Trans: Cross Modal Inconsistency-aware Transport for HSI-LiDAR Classification

Yanli Li, Xuan Tan, Ding Qi, XINYANG JIANG

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

EDMA: Entropy-Driven Multimodal Answering

Emanuele Mezzi, Gertjan Burghouts, Fabio Massacci, Mengyuan Zhang

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AeroMosaic:Transport-Aware Multimodal Evidence Fusion for Atmospheric Pollution Risk Inference

Muyang Zheng, Jiaming Ma, Zongyu Zhang, Qingsong Wen

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

Do More Modalities Always Help? A Geometric Perspective on Missing-Modality Robustness

Songyuan Sui, Zhen Tan, Mohan Zhang, Rana M Khan and 2 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

Intervene3D: Intervention-Based Controlled Inference for Multimodal Perception under Partial Observability

Constantino Msigwa, Denis Bernard, Jaeseok Yun

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?

Ziyi Wang, Li Li, Aolin Zhou, Yankun Shen and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

GeLVR: Geometry-Consistent Latent Visual Reasoning in Multimodal LLMs

Jiaqi Wang, Yi Feng, Xian Wu

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

IntentLens: Grounding Underspecified Multimodal Queries for Recommendation via Tool-Augmented Reasoning

xiao chen, Dong Fang, Haitao Li, Qing Li

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

AVID: A 5T fMRI Dataset for Benchmarking Auditory-induced Visual Mental Imagery Decoding

Shiqi Shen, Shurui Li, Yuanning Li, Xilin Zhang

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

TIGER: Bridging the Multimodal Reasoning-Access Gap via Modality Counterfactuals

Gregory Kang Ruey Lau, Huynh Minh Nguyen, Bryan Kian Hsiang Low

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

MIRA: Reinforcing Multimodal Reasoning via Deceptive Contextual Augmentation

Zhihan Yin, Jianxin Liang, Yifeng Yao, Nonghai Zhang and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Thinking with Images as Continuous Policy: Numerical Visual Chain-of-Thought

Kesen Zhao, Beier Zhu, Junbao Zhou, Xingyu Zhu and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation

Zhenlong Yuan, Yue Wang, Jing Tang, Rui Chen and 8 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score
67%Highly rated
?Highly ratedVote to see the score

Forgery Evidence Peaks Mid-Stack: Forensic Evidence Relay for Multimodal Forgery Detection

Yingxin Lai, xinyuan Wang, Yufei Liu, Jialin Guo and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

ANCHOR: Audio-Visually Grounded Chain-of-Thought Reasoning Benchmark

Joel Julin, Souraja Kundu, Liza Dahiya, George Z Wei and 5 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Verifier Choice is a Benchmark Design Variable: Auditing Structural Counting Evaluation in Text-to-Image Models

Shurun Li, Michael H Wang, Haibo Zeng, Long Wang

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

TopoGraphRAG-Bench: Evaluating Multimodal GraphRAG on Layout-Grounded Evidence Reasoning

Ruochi Li, peter lin, Haoxuan Zhang, Haihua Chen and 3 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

SeoulMMOD: A Large-Scale Multimodal Origin-Destination Flow Benchmark

Taeyoung Yu, Seonbin Jo, Jiwon Kim, Junyoung Byun

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Who Called? V33DA: A Physically Verified Multimodal Benchmark for Vocal Attribution in Zebra Finch Groups

Maris Basha, Yuhang Wang, Xiaoran Chen, Longbiao Cheng and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Look Before You Reason: Implicit Visual Thinking for Efficient Multimodal Reasoning

Zerui Chen, Changrui Chen, Fei Ni, Jiankang Deng

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

MUST: Stage-Adaptive Stability Control for Test-Time Scaling in Multimodal Reasoning

YunFeng Deng

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Visual Reinforcement Fine-Tuning via Bootstrapped Medical Reasoning

Yequan Bie, Peng Xie, Zhixuan CHEN, Yihui Wang and 7 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Prior Text-Informed Gate Attention Framework for Multimodal Psychiatric Disorder Diagnosis

Wensheng Zhai, Kaizhong Zheng, Difei Mei, Badong Chen and 1 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning

jiayi lei, Yuandong Pu, Xingyu Han, Rongpeng Zhu and 7 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Semantic Freedom Bottleneck for Domain-Generalized Multimodal Face Anti-Spoofing

Yingjie Ma, Haonan Wang, Xun Lin, Hui Ma and 8 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Learning Actionable Information Landscapes for Multimodal Active Sensing in Hawkmoths

Abdelrahman Sharafeldin, Yaqing Wang, Simon Sponberg, Hannah Choi

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Fusion or Confusion? Multimodal Complexity Is Not All You Need

Tillmann Rheude, Roland Eils, Benjamin Wild

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Reasoning Gestures: LLM-Inferred Communicative Functions for Co-Speech Gesture Generation

Pinxin Liu, Haiyang Liu, Yunlong Tang, Luchuan Song

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

When Are Multimodal Predictions Biologically Supported? A Diagnostic Evaluation Framework

Dylan Steiner, Gustavo Arango, Gerald J Sun, Etai Jacob

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

CRISP: Compositional Reasoning over Images via Stackable Programs for VLMs

Arnas Uselis, Yujin Jeong, Yanpeng Zhao, Alexander Rubinstein and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

AIM: Adaptive Interaction in Multi-Agent Debate for Multimodal LLM Inference

Wei Fan, JinYi Yoon, Bo Ji

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

MotorSense: A Video-EMG Dataset of Motoric Representations for Action Understanding

Eadom T Dessalene, Michael Maynord, Amir Hossein Shahidzadeh, Botao He and 5 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

Jinghui Lu, Jiayi Guan, Zhijian Huang, Jinlong Li and 36 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

SynGeo: Synergizing Seeing and Proving through Revisable Geometric States

Tianyi Xu, Liu Yang, Wenjun GAO, Junyu Ou and 4 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

CausalConflictBench: Can Multimodal Models Follow Local Mechanisms That Conflict with Commonsense?

Bo Tian, Jianfeng Qu, Peng-Fei Zhang, Siyu Li and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Triggering Generalist Reasoning via Predictive Uncertainty for Dual-System VLA

Hyemin Yang, Wooseong Jeong, Giwon Lee, Kuk-Jin Yoon

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

GeoDial: A Multimodal Dialog Tutoring Dataset for Geometry Problem-Solving with Visual Tutor Turns

Sankalan Pal Chowdhury, Junling Wang, Donya Rooein, April Wang and 1 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Learning Where to Look: Observation Policy Optimization for Thinking with Images

Junfeng Wang, Jiawei Liu, Yongchao Xu, Tao Jiang and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
83%Must read
?Must readVote to see the score

Unveiling Fine-Grained Visual Traces: Evaluating MultiModal Interleaved Reasoning Chains in Multimodal STEM Tasks

StepSTEM introduces 283 graduate-level STEM problems with strictly complementary visual-textual inputs to evaluate cross-modal reasoning via step-level alignment, revealing current MLLMs achieve only 38.29% accuracy due to heavy reliance on text.

Jing Jin, Hao Liu, Yan Bai, Yihang Lou and 8 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

MAEB: Massive Audio Embedding Benchmark

MAEB benchmarks 30 audio tasks across 100+ languages, finding no single model dominates and acoustic and linguistic skills trade off.

Adnan E Assadi, Isaac Chung, Chenghao Xiao, Roman Solomatin and 14 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 25 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Reinforcing Multimodal Reasoning Against Visual Degradation

ROMA improves multimodal reasoning robustness to visual corruption via dual-pass RL optimization that avoids reward poisoning while preserving clean accuracy.

Rui Liu, Dian Yu, Haolin Liu, Yucheng Shi and 4 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

CAVE: A Structured Credit Assignment Approach for Fragmented Visual Evidence Reasoning

CAVE uses structured process rewards to improve vision-language models' integration of nonlocal visual evidence, boosting fragmented reasoning benchmarks.

Tengda Guo, Jie Leng, Hanlei Li, Yaoyuan Liang and 5 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Touch-R1: Reinforcing Touch Reasoning in MLLMs

Touch-R1 trains a tactile reasoning MLLM via GRPO with tactile-grounded rewards, outperforming Octopi-13B and GPT-4o by 18.4% and 24.7% on TouchReason-Bench.

Yingxin Lai, Yafei Zhou, Fucai Zhu, Siyu Zhu and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning

CausalSpatial benchmarks object-centric causal spatial reasoning, revealing MLLMs score 54% versus human 84% due to ungrounded textual reasoning, fixed by video-simulation framework COW.

Wenxin (Wendy) Ma, Chenlong Wang, Ruisheng Yuan, Hao Chen and 9 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains

SR-MCR aligns multimodal reasoning via intrinsic process rewards and a critic-free GRPO objective, achieving 81.4% average accuracy across visual benchmarks.

jusheng zhang, Ningyuan Liu, Kaitong Cai, Sidi Liu and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

AgroOmni: A Large-Scale Multi-view Agricultural Dataset for Cross-Scale Multimodal Reasoning

AgroOmni introduces 288K multi-view agricultural VQA pairs across scales, and AgroNVILA achieves state-of-the-art 62.32% on AgroMind while demonstrating strong cross-scale generalization.

Jiarui Zhang, Junqi Hu, Zurong Mai, Yang Liu and 9 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

OpenView: Empowering MLLMs with Out-of-view VQA

OpenView introduces out-of-view VQA via panoramic synthesis, boosting MLLM accuracy from 48.6% to 64.1%.

Qixiang Chen, Cheng Zhang, Chi-Wing Fu, Jingwen Ye and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
80%Must read
?Must readVote to see the score

RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark

RAIL introduces a CHC-based benchmark evaluating LALMs across five auditory cognitive abilities, revealing highly uneven performance among 26 state-of-the-art models.

Hongyu Jin, Siyi Wang, Yang Xiao, Jiaheng Dong and 9 more

Published 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Leveraging Latent Visual Reasoning in Silence

Latent visual reasoning enhances multimodal training despite being largely unused at inference; attention-based reinforcement learning preserves its benefits by promoting latent-text interaction during training.

Dongyao Zhu, Zhen Wang, Xi Xiao, Han Jiang and 6 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

92%Must read
?Must readVote to see the score

DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding

DocScope benchmarks verifiable long-document reasoning via structured trajectory evaluation, finding correct answers rarely include complete evidence chains and region grounding remains weakest.

Xiang Feng, Jiawei Zhou, Zhangfeng Huang, Kewei Wang and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning

Visual replay and routing depth scaling fix visual under-optimization and gradient instability in latent reasoning, achieving state-of-the-art results with faster inference.

Yudong Han, Yong Wang, Zaiquan Yang, Zhen Qu and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

A Benchmark for Omni-Modal Reasoning in Long Videos

LongShOTBench evaluates long-form omni-modal video reasoning via rubric-scored open-ended questions, and LongShOTAgent achieves 66.64% as the top training-free system.

Mohammed Irfan Kurpath, Jaseel M Kaithakkodan, Jinxing Zhou, Sahal Shaji Mullappilly and 11 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 7 on Hugging Face · Code ★ 26

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning

GeoSym Engine automates symbolically-verifiable geometric reasoning data synthesis, and models trained on GeoSym127K achieve large gains on diagram-dependent geometry benchmarks.

Jinhao Jing, Zheng Ma, Jinwei Liang, Qiannian Zhao and 8 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

On the Nature of Attention Sink that Shapes Decoding Strategy in Omni-LLMs

Attention sinks in Omni-LLMs serve as global representation biases rather than redundant heads, and the proposed OutRo decoding method improves video reasoning with minimal overhead.

Suho Yoo, Youngjoon Jang, Joon Son Chung

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency

VCBENCH benchmarks multimodal elementary math reasoning requiring multi-image visual dependencies, finding top models score under 50%.

Zhikai Wang, Jiashuo Sun, Wenqi Zhang, Zhiqiang Hu and 2 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 13 on Hugging Face · Code ★ 11

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Cephalonauts One: A deep fMRI dataset for decoding naturalistic speech in the human brain

Cephalonauts One provides 30 hours per subject of whole-brain fMRI during naturalistic speech, paired with audio, transcripts, and embeddings, plus a brain decoding benchmark showing continuous performance gains with more training data.

Antoine Collas, Louis Jalouzot, Géraud Ilinca, Corentin Caris and 10 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 2/10
strict 1/5
89%Must read
?Must readVote to see the score

Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests

Spatial-IQ hierarchically decomposes spatial reasoning into perceptual and cognitive sub-tasks, showing models use shortcuts and that hierarchical chain-of-thought training improves consistency and accuracy.

Patrick Rim, Tom Long, Ekta Prashnani, Ruth Rosenholtz and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

TAC: Timestamped Audio Captioning

TAC generates temporally grounded audio captions via synthetic training, reducing hallucinations and outperforming competitors in detection and dense captioning; cascading it with LLMs achieves state-of-the-art audio and audio-visual reasoning.

Sonal Kumar, Prem Seetharaman, Ke Chen, Oriol Nieto and 7 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

92%Must read
?Must readVote to see the score

CausalDriveBench: Evaluating Causal Reasoning in Vision-Language-Action Models for Autonomous Driving

CausalDriveBench evaluates causal reasoning in autonomous driving vision-language-action models via structured QA and counterfactual trajectories, finding weak causal understanding despite fluent reasoning and accurate baseline predictions.

Narendiran Chembu, Navvrat Rao, Shreedhar Kodate, Gayatri S Banda and 9 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 5/5
86%Must read
?Must readVote to see the score

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning

RAPO uses information-theoretic lower bounds to select high-entropy reflection anchors and optimize visual dependence via GRPO, substantially improving long-chain multimodal reasoning with reduced visual fading.

Xuan Gong, HanBo Huang, Hao Zheng, Yiran Zhang and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models

GAP fixes visual-latent reasoning instability via feature, context, and capacity alignment, improving Qwen2.5-VL 7B perception and reasoning performance.

Yanting Miao, Yutao Sun, Dexin Wang, Pascal Poupart and 7 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation

A reasoning-prefix masking framework distills think-answer visual reasoning into compact VLMs by masking salient reasoning cues to force visual anchoring, improving multimodal benchmarks over prior distillation methods.

Seonghoon Yu, Dongjun Nam, Byung-Kwan Lee, Jeany Son

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

LogSTOP: Temporal Scores over Prediction Sequences for Matching and Retrieval

LogSTOP computes Linear Temporal Logic temporal property scores over noisy local prediction sequences, outperforming language models and retrieval baselines by at least 16% on matching and retrieval.

Avishree Khare, Hideki Okamoto, Bardh Hoxha, Georgios Fainekos and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 2/5
83%Must read
?Must readVote to see the score

VocalCoachBench: Benchmarking Audio-Language Models on Expert Feedback for Singing

VocalCoachBench benchmarks audio-language models on expert singing feedback, revealing they identify broad vocal issues but fall below baselines on fine-grained diagnosis and strict alignment.

Hayeon Bang, Hounsu Kim, Wonil Kim, Juhan Nam

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 3/5
91%Must read
?Must readVote to see the score

DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models

DRScaffold improves lightweight vision-language model reasoning via structured four-stage supervision, surpassing a frozen 32B model on DRBench.

Xinrui Shi, Kai Liu, Ziqing Zhang, Jianze Li and 2 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
76%Highly rated
?Highly ratedVote to see the score

Measure Less, Know More: Self-Supervised Test-Time Feature Acquisition

ECHO-k uses pretrained representations as proxy targets for self-supervised reinforcement learning to sequentially acquire informative modalities at test time, improving budgeted downstream performance across diverse backends.

Eeshaan Jain, Linus Bleistein, Bart Deplancke, Charlotte Bunne

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing

VLRS-Bench introduces a remote sensing vision-language reasoning benchmark spanning cognition, decision, and prediction tasks that exposes major bottlenecks in current multimodal models.

Zhiming Luo, Hebaixu Wang, Haonan Guo, Jing Zhang and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

Neural-Behavioral Representation of Natural Whole-body Movement in Monkeys

A neural-behavioral framework decodes natural whole-body monkey movement from large-scale epidural cortical signals via an autoregressive model without physical constraints.

Jieshi He, Puzhe Li, Yanan Sui, Mu-ming Poo

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
89%Must read
?Must readVote to see the score

PitchBench: Measuring Pitch Hearing in Audio-Language Models

PitchBench evaluates pitch hearing in audio-language models via 28 experiments, finding their pitch perception remains highly unreliable across instruments and acoustic conditions.

Milan Liessens Dujardin, Song-Ze Yu, Craver C Thomas-Smith, David Chan and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 4/5
74%Highly rated
?Highly ratedVote to see the score

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs

UniVLR unifies text and visual reasoning into a shared visual workspace, using compressed visual latent tokens to outperform prior methods with fewer reasoning tokens.

Houcheng Jiang, Jiajun Fu, Junfeng Fang, Chen Gao and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
83%Must read
?Must readVote to see the score

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning

LatentOmni replaces text chain-of-thought with interleaved audio-visual latent reasoning states to preserve dense sensory signals, improving joint reasoning over explicit text baselines.

Yifan Dai, zhenhua wu, Bohan Zeng, Daili Hua and 17 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 44 on Hugging Face · Code ★ 24

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 0/5
86%Must read
?Must readVote to see the score

Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs

ProCauEval reveals LMMs perceive video but ignore it for causal reasoning, and ADPO reduces textual shortcuts via negative teacher alignment.

Jiafeng Liang, Zhihao Zhu, Zihan Zhang, Baoqi Ren and 6 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
91%Must read
?Must readVote to see the score

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning

StemBind introduces a shared-stem benchmark diagnosing MLLM abstract visual reasoning, finding a persistent rule-to-instance binding gap where models identify patterns but fail to apply them correctly.

Xixiang He, Baiqi Wu, Xingming Li, Ao Cheng and 3 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 5/5
83%Must read
?Must readVote to see the score

OpenMedReason: Scientific Reasoning Supervision for Medical Vision–Language Models

OpenMedReason is a 450K-instance open medical reasoning dataset derived from scientific articles that improves LVLM diagnostic accuracy by 20% and enhances perception, knowledge, and reasoning.

Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi and 5 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 3/5
89%Must read
?Must readVote to see the score

When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning

AVIC adaptively scales test-time visual imagination via world models for spatial reasoning, matching fixed strategies with fewer calls while exceeding GPT-4o.

Shoubin Yu, Yue Zhang, Zun Wang, Jaehong Yoon and 3 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 9 on Hugging Face · Code ★ 20

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Thinking with Visual Primitives

Thinking with Visual Primitives interleaves spatial markers into chain-of-thought reasoning to close the reference gap, achieving frontier-level visual QA with extreme token efficiency.

Ruijie Lu, Yiyang Ma, Xiaokang Chen, Lingxiao Luo and 4 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
89%Must read
?Must readVote to see the score

OmniTraffic: A Controllable Generation Pipeline and Benchmark for Spatio-Temporal Traffic Reasoning

OmniTraffic introduces a controllable 3D traffic generation pipeline and benchmark with 8M VQA samples for spatio-temporal reasoning, revealing large model gaps and improved real-world performance via simulated fine-tuning.

Maonan Wang, Zhengyan Huang, Kemou Jiang, Yuhang Fu and 12 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
86%Must read
?Must readVote to see the score

ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation

A new benchmark reveals current multimodal AI fails at multi-step ECG reasoning, achieving near-zero completion in linking clinical criteria to visual signal evidence.

Jungwoo Oh, Hyunseung Chung, Junhee Lee, Min-Gyu Kim and 5 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 6 on Hugging Face · Code ★ 18

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

Towards Explainable Industrial Anomaly Detection via Knowledge-Guided Latent Reasoning

Reason-IAD improves explainable industrial anomaly detection via knowledge-guided latent reasoning and dynamic visual injection, outperforming state-of-the-art methods across tasks.

Peng Chen, Chao Huang, Yunkang Cao, Chengliang Liu and 6 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models

TWNM formalizes audio scene analysis as a three-level hierarchy and equips audio-language models with physically grounded spatial representations via ambisonic simulation and progressive training, achieving 70.8% overall accuracy on controlled spatial benchmarks.

Yuhuan You, Lai Wei, Tianshu Qu

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
89%Must read
?Must readVote to see the score

ARK: A Dual-Axis Multimodal Retrieval Benchmark along Reasoning and Knowledge

ARK introduces a dual-axis multimodal retrieval benchmark spanning knowledge domains and reasoning skills, revealing persistent bottlenecks in fine-grained visual and spatial reasoning.

Yijie Lin, Guofeng Ding, Haochen Zhou, Haobin Li and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 3/5