Good Papers

Showing Embodied agents Show all papers

78%Highly rated
?Highly ratedVote to see the score

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning

VeriFine co-evolves policies, curricula, and judges via adaptive diagnosis and coactive calibration to sustain embodied reasoning self-improvement. Experiments on driving and navigation show continuous gains in both policy and judge capability.

Zewei Zhou, Rachel Luo, Yulong Cao, Chaowei Xiao and 9 more

Published Oct 6, 2026 · ▲ 4 on Hugging Face

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents

Attacca trains visual goal-conditioned policies on complete search-to-interact trajectories with decoupled goal images and behavioral-phase conditioning to improve long-horizon embodied task success by up to 7x.

Gyusik Seo, Jaehong Yoon

Published Oct 6, 2026 · ▲ 3 on Hugging Face · Code ★ 3

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
89%Must read
?Must readVote to see the score

ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context

Embodied progress reward models fail at long tasks due to missing context, but ProgressCompass supplies needed context to cut progress estimation errors by up to 82%.

Jianshu Zhang, Keyi Wu, Chengxuan Qian, Xiyuan Yang and 5 more

Published Sep 29, 2026 · 0 citations · ▲ 11 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Video2Skill: From Streaming Experience to Reusable Embodied Skills

Video2Skill benchmarks streaming embodied skill discovery, showing VLMs group manipulation events poorly and rarely expand skill libraries despite supervised fine-tuning.

Jianshu Zhang, Ce Zhang, Xiyuan Yang, Chenwei Xu and 5 more

Published Sep 29, 2026 · 0 citations · ▲ 11 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 3/5
76%Highly rated
?Highly ratedVote to see the score

Adaptive Latent Capacity for World Models

ALeWM learns adaptive-width JEPA world models via MixSIGReg and prefix sampling, concentrating predictive information in early latent coordinates to improve planning with lower capacity.

Idan Achituve, Lior Dikstein, Idit Diamant, Arnon Netzer and 1 more

Published Sep 26, 2026 · 0 citations · ▲ 6 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
86%Must read
?Must readVote to see the score

Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding

DMM refines multi-agent action intents iteratively via communication to fix incompatible joint actions, achieving near-perfect success on 1,600 MovingAI tasks and scaling to over one million agents.

Valeriy Vyaltsev, Anton Andreychuk, Taisia Zlotnikova, Konstantin Yakovlev and 2 more

Published Sep 25, 2026 · 0 citations · ▲ 62 on Hugging Face · Code ★ 4

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

Lumine: An Open Recipe for Building Generalist Agents in 3D Open Worlds

Lumine is an open recipe for generalist agents that completes hours-long 3D open-world missions via end-to-end vision-language modeling with adaptive reasoning and strong cross-game zero-shot generalization.

Tan, Weihao, Li, Xiangyang, Fang, Yunhao, Yao, Heyuan and 10 more

Published Nov 12, 2025 · 0 citations · ▲ 218 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

RECIPE: Procedural Planning via Grounding in Instructional Video

Luigi Seminara, Antonino Furnari, Lorenzo Torresani

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

AndroidReality: How Far Are Mobile Agents from the Real World?

Xiaoou Liu, Longchao Da, Hanyang Chen, Yuan Ling and 1 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving

Xingtai Gui, Yucheng Zhou, Dongqian Guo, jiahao gong and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

A Latent World-Action Model with Jointly Aligned Reasoning

Hao Luo, Wanpeng Zhang, Yicheng Feng, Sipeng Zheng and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Decoupling Action from Egocentric Observation for World Simulation

Yue Ma, Pengjie Song, Xinyu Wang, Yi He and 9 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Embedded-Arena: Building Hardware-in-the-Loop Coding Agents to Run AI on Microcontrollers

Zhihan Zhang, Alexander Le Metzger, Jiuyang Lyu, Chun-Cheng Chang and 9 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score
NeurIPS 2026Embodied agents

Causal-Geo: Neuro-Symbolic Spatial Grounding for Situated Agent Planning

Sanbi Luo

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Learning Agentic World Vision-Language-Action Models for Autonomous Driving

Guoqing Wang, Pin Tang, Xiangxuan Ren, Chao Ma

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

REINS: A Self-Evolving Agent Harness for Real-Time Trajectory Planning

Zhihong Cui, Hengyu Liu, Haoran Tang, shijun liu and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Say the Same, Act Differently: Text-Orthogonal Action Subspaces in Reasoning Vision-Language-Action Models

Zihao Feng, Qingzhao Zhang, Chunyu Xia, Bo Yu and 2 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

GNES: Neural-Guided Evolutionary Program Search for Interpretable Multi-Agent Control

Chen Wang, Minfang Lu, Xiangke Wang, Cheng Zhu and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

ASH: Agents that Self-Hone via Embodied Learning

Benjamin Schneider, Xavier Schneider, Victor Zhong, Sun Sun

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

LongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoE

Lei Jin, Yu Shang, Yiding Ma, Xinhao Jin and 4 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

MHWA: Multi-timescale Hierarchical World-Action Model

Pengcheng Pan, Guoqing Ma, Yuhan Zhang, Yang Chen and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Toward Embodied World Agents via Embodied-Planning Dataset and Interactive World Models

Xiaokun Feng, Junshu Tang, zeyi lin, Ling-Hao Chen and 3 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Endowing Your Vision-Language-Action Model with a Predictive Mind

Pengxiang Ding, Haoying Wang, Minghui Lin, Qishen Wang and 8 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

The limbic navigation system as a hierarchical RNN

Zilong Ji, Huiwen Zhang, Krishna Gorantla, Neil Burgess

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

SCOUT: Planning under Occlusion via Object-Centric World Model Rollouts

Amir Hossain Raj, Xuesu Xiao

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Benchmarking Fine-Grained Spatio-Temporal Awareness in Embodied Brain Models

Minghao Zhu, Zhikai Wang, Ronghao Dang, Bohan Hou and 9 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Think Densely, Act Sparsely: Latent Expert Cognitive Chains for Vision-Language-Action Autonomous Driving

Jie Wang, Guang Li, Zhijian Huang, Jinlong Li and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

3A-VLA: Abstraction-Aligned Action Learning for Vision-Language Agents in 3D Game Worlds

Zheyuan Zhou, Liang Du, Zixun Sun, Xiaoyu Zhou and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

WORD: Diffusion-Based Posterior Inference for Online Goal Recognition

Almog Anschel, Sarah Keren

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

MAPS: Margin-Aware Priors and Verifier-Guided Search for Embodied Planning

XIN Li, Junquan Huang, Xujia Li, Lei Chen

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
80%Must read
?Must readVote to see the score

ChainFlow-VLA: Causal Flow Planning with Vision-Language Models

ChainFlow-VLA unifies causal trajectory generation and global diffusion refinement via vision-language-conditioned residual distributions, scoring 94.85 on NAVSIM v1.

Xiyang Wang, Xinlin Wang, Tingguang Zhou, Gong Chen and 6 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
80%Must read
?Must readVote to see the score

DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning

DriveDreamer-Policy unifies depth generation, video prediction, and motion planning via geometry-aware world representations, achieving 89.2 PDMS on Navsim v1 and 88.7 EPDMS on v2.

Yang Zhou, Xiaofeng Wang, Hao Shao, Letian Wang and 7 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 2 on Hugging Face · Code ★ 60

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search

Imagining in 360° decouples humanoid visual search into an Imaginator predicting spatial priors and an Actor using them to improve search efficiency without costly trajectory annotations.

Jingdong Zhang, Yizhou Wang, Zhengzhong Tu, Xin Li and 2 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation

WorldDrive unifies vision and motion representations to couple scene generation with planning, achieving leading vision-only autonomous driving performance with high-fidelity future video generation.

Xingtai Gui, Meijie Zhang, Tianyi Yan, Wencheng Han and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
83%Must read
?Must readVote to see the score

ReflectDrive-2: Reinforcement-Learning-Aligned Self-Editing for Discrete Diffusion Driving

ReflectDrive-2 is a discrete diffusion planner that uses reinforcement learning to train self-editing trajectory tokens, boosting NAVSIM PDMS to 91.0 with 31.8 ms latency.

Huimin Wang, Yue Wang, Bihao Cui, Pengxiang Li and 6 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 8 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 2/5
86%Must read
?Must readVote to see the score

Resilience Matters for Embodied Agents System: New Metrics, Systematic Evaluation, and Optimization

Resilience metrics for embodied agents expose hidden recovery and stability differences beyond success rates, guiding deployment-specific optimization.

Yapeng Liu, Yuanzhao Zhai, Xudong Gong, Feng Dawei and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Vision to Geometry: 3D Spatial Memory for Sequential Embodied MLLM Reasoning and Exploration

3DSPMR leverages 3D spatial memory with field-of-view geometric priors to reuse exploration knowledge across sequential embodied tasks, significantly improving reasoning and navigation performance on the SEER-Bench benchmark.

Zhongyi Cai, Yi Du, Chen Wang, Yu Kong

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

CrafterDojo: A Suite of Foundation Models for Building Open-Ended Embodied Agents in Crafter

CrafterDojo introduces CrafterVPT, CrafterCLIP, and CrafterSteve-1 alongside toolkits to unlock Crafter as a lightweight testbed for open-ended embodied agents.

Junyeong Park, Hyeonseo Cho, Sungjin Ahn

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 1/5
83%Must read
?Must readVote to see the score

AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving

AlignDrive conditions longitudinal planning on the lateral path via anchor-based 1D displacement prediction and safety-critical augmentation, achieving state-of-the-art Bench2Drive results.

Yanhao Wu, Haoyang Zhang, Fei He, Rui Wu and 5 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
86%Must read
?Must readVote to see the score

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators

Astra enhances vision-language model spatial reasoning by letting agents generate imagined simulator views via RL, improving MMSI-Bench scores over direct answering.

Chenming Zhu, Jingli Lin, Yilin Long, Peizhou Cao and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 16 on Hugging Face · Code ★ 31

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving

LaST-VLA replaces explicit chain-of-thought reasoning with a physically grounded latent spatio-temporal framework, achieving record NAVSIM scores and improved reasoning.

Yuechen Luo, Fang Li, Shaoqing Xu, Yang Ji and 9 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 5/10
strict 1/5
86%Must read
?Must readVote to see the score

World–Value–Action Model: Implicit Planning for Vision–Language–Action Systems

WAV introduces a latent-space planning framework for vision-language-action models that predicts future states and evaluates trajectory values to enable efficient long-horizon decision-making.

Runze Li, Hongyin Zhang, Junxi Jin, Qixin Zeng and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Intend, Reflect, Refine: An Adaptive Multimodal Reflection Framework for Autonomous Driving

IRR-Drive uses adaptive multimodal text and BEV reflection to self-correct driving intentions before trajectory generation, achieving state-of-the-art NAVSIM results.

Zisheng Chen, Yuping Qiu, Jianhua Han, Tao Tang and 5 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
83%Must read
?Must readVote to see the score

Generative Scenario Rollouts for End-to-End Autonomous Driving

GeRo enables vision-language-action models to generate language-grounded future traffic scenes via autoregressive rollouts, improving Bench2Drive driving scores by 15.7 and success rates by 26.2.

Rajeev Yasarla, Deepti Hegde, Shizhong Han, Hsin-Pai Cheng and 10 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
80%Must read
?Must readVote to see the score

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving

CoWorld-VLA embeds multi-expert world tokens into vision-language-action models and couples diffusion planning with scene context to generate continuous ego trajectories, improving autonomous driving performance.

Jingqi Wang, minqing huang, Zihan Liang, Yujiao Xiang and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
86%Must read
?Must readVote to see the score

AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models

AlloSpatial is an agentic framework that converts egocentric observations into allocentric spatial priors via cognitive mapping and reasoning harnesses, improving spatial reasoning by 5%-18% and outperforming larger general-purpose models.

Shouwei Ruan, Bin Wang, Zhenyu Wu, Qihui Zhu and 4 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 4 on Hugging Face · Code ★ 21

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 0/5
91%Must read
?Must readVote to see the score

TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning

TaskGround grounds full household scenes into task-relevant slices to infer executable task structures, improving compact open-weight models' success rates by large margins over direct prompting while cutting token costs up to 18x.

ZhiYuan Feng, Yu Deng, Ruichuan An, Zhenhua Liu and 10 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
76%Highly rated
?Highly ratedVote to see the score

DiffeoMorph: Learning to Morph 3D Shapes Using Differentiable Agent-Based Simulations

DiffeoMorph learns agent-based 3D shape morphogenesis via differentiable attention-based graph networks and a rotation-aligned 3D Zernike shape-matching loss.

Seong Ho Pahng, Guoye Guan, Benjamin Fefferman, Sahand Hormoz

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

MineEvolve: Self-Evolution with Accumulated Knowledge for Long-Horizon Embodied Minecraft Agents

MineEvolve converts Minecraft execution feedback into structured skills and remedies via Monitor, Inducer, Curator, and Adaptor, improving long-horizon agent performance across planners.

Zhengwei Xie, Zhisheng Chen, Ziyan Weng, Jinhan Li and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
86%Must read
?Must readVote to see the score

WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation

WorldVLN autoregressively predicts short-horizon world-state transitions to generate waypoint actions for aerial vision-language navigation, achieving over 12% success-rate gains and real-world drone transfer.

Baining Zhao, jiacheng xu, Weicheng Feng, Xin Zhang and 12 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

RAO-Nav: Probing Omni-Language Models for Zero-shot Semantic Audio-Visual Navigation

RAO-Nav applies omni-language models to zero-shot audio-visual navigation via a reasoning pipeline with latent navigation reasoning, surpassing trained state-of-the-art baselines without training data.

Qilang Ye, Meng Liu, Yu Zhou

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 1/5