Good Papers

Showing Visual grounding Show all papers

45%Niche pick
?Niche pickVote to see the score

Diffusion Fine-Tuning: Iterative Refinement for Advanced Grounding with Diffusion Large Language Models

Zhangyang Qi, Jinsong Li, Jiaqi Wang, Hengshuang Zhao

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Learning to Synergize Textual and Visual Prompts for Fine-Grained Traffic Element Detection in HD Maps

Xiaoyang Bi, Haowen Guo, Caoshengzhe Xue, Siyuan Li and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model

Hanqing Wang, Mingyu Liu, Xiaoyu Chen, Chengwei MA and 9 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

HiLoc: A Hierarchical Representation Method for Spatial Localization in Multimodal Large Language Models

Evelyn Zhang, Fufu Yu, Hanjun Li, Aoqi Wu and 7 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Visual Grounding First, Multimodal In-context Learning Follows

Minhyuk Seo, Minjae Lee, Chaeeun Lee, Wei Lin and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

LaSA-Net: A Language-Guided Network for Outdoor Generalized 3D Referring Expression Segmentation

Lingfei Ma, Bin Liu, Wen Li, Wentao Sun and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

CSI-TextBench: A Dataset and Benchmark for Language-Grounded Ambient Sensing Perception

Guozhen Zhu, Yuqian Hu, Sakila S Jayaweera, Wei-Hsiang Wang and 2 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

P$^3$-VLM: A Point-based Alternative for Grounded 3D Vision-Language Models

Anna-Maria Halacheva, Jan-Nico Zaech, Sombit Dey, Luc V Gool and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Shift-Aware Identity-Guided Latent Refinement for Referring Audio–Visual Segmentation

Kun Li, Sami S Brandt, Michael Yang

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Hierarchical Graph Alignment for Cross-Modal 3D Scene Grounding

Tianyi Shang, Zhenyu Li

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

QGround: Condition-Wise Evidence Aggregation for 3D Grounding with 2D VLMs

Fengyun Wang, Jian Wang, Dingwei Zhang, Jinhui Tang and 1 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Referring and Reasoning Camouflaged Object Segmentation in Audio-Visual Scenes

Tianxin Han, Qing Dong, Xingwei Wang, Jie Jia

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
88%Must read
?Must readVote to see the score

Evidence-RL: Towards Evidence-intensive Visual Reasoning

Counterfactual Evidence Disentanglement (CED) audits vision-language model grounding by comparing evidence-region and non-evidence-region support drops inside GRPO, improving visual reasoning across benchmarks without inference overhead or evidence annotations.

Haojie Huang, Xinlei Yu, Chengming Xu, Zhangquan Chen and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 16 on Hugging Face · Code ★ 4

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 3/5
89%Must read
?Must readVote to see the score

SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring

SIEVES improves selective prediction for visual question answering by scoring visual evidence quality, boosting out-of-distribution coverage up to three times across open and closed models without requiring internal weights.

Hector G. Rodriguez, Marcus Rohrbach

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
91%Must read
?Must readVote to see the score

Grounding Driving VLA via Inverse Kinematics

Reformulating driving VLA as inverse kinematics with future visual prediction and diffusion-based decoding recovers visual grounding, letting a 0.5B model match 7B-8B planning performance.

Junsung Park, Hyunjung Shim

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs

SSR3D-LLM introduces latent spatial reasoning steps to refine 3D object rankings step-by-step, improving fine-grained grounding across benchmarks while preserving unified language tasks.

Jiawei LI, Ziyi Liu, Weijie Shi, Long Chen and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Segment and Select: Vision-Language Segmentation in 3D Scenarios

SEGA3D performs 3D vision-language segmentation via fine-grained mask candidates and LLM-guided selection, surpassing prior methods by up to 8.3 mIoU.

Yulin Chen, Zhihang Zhong, Yuenan Hou

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
80%Must read
?Must readVote to see the score

Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding

GUI-SD uses on-policy self-distillation with privileged visual contexts and entropy-guided distillation for GUI grounding, outperforming GRPO methods in accuracy and efficiency.

Yan Zhang, Daiqing Wu, Huawen Shen, Can Ma and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
80%Must read
?Must readVote to see the score

Vision-Language Grounding as Bidirectional Concept Correspondence

Grounding is formulated as bidirectional concept correspondence to recover all image-text span correspondences without prespecified phrases via ConCor-1, improving F1 by 48% and 29% over baselines.

Jieyu Zhang, Ziqi Gao, Luke Zettlemoyer, Ranjay Krishna

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 6 on Hugging Face · Code ★ 8

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
88%Must read
?Must readVote to see the score

TIGER-FG: Text-Guided Implicit Fine-Grained Grounding for E-commerce Retrieval

TIGER-FG uses text-guided implicit fine-grained grounding and dual distillation to improve cropped-query e-commerce retrieval, boosting Recall@1 by up to 34.4 points without object detection.

Xinyu Sun, Huangyu Dai, Lingtao Mao, Zexin Zheng and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
Show 20 more papers