Good Papers

RAO-Nav: Probing Omni-Language Models for Zero-shot Semantic Audio-Visual Navigation

RAO-Nav applies omni-language models to zero-shot audio-visual navigation via a reasoning pipeline with latent navigation reasoning, surpassing trained state-of-the-art baselines without training data.

Qilang Ye, Meng Liu, Yu Zhou

Published 2026Sydney Poster Session 1 · Tue, Dec 8, 10:00 AM–1:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

76%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel10/20reviewers recommend it
lenient 3/5
medium 6/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

We explore whether Omni-Language Models (OLMs) can be directly applied to zero-shot Semantic Audio-Visual Navigation (SAVN). Recent work demonstrates that even state-of-the-art specialized models still struggle to achieve generalist multimodal navigation, despite extensive task-specific training. In this paper, we introduce RAO-Nav, short for Reasoning All-in-One OLM, a deployment pipeline for zero-shot SAVN. By leveraging the rich implicit audio-visual knowledge encoded in OLMs, the embodied agent is enabled to ``hear'', ``see'', ``reason'', and ``act'' in the environment. To further elicit the built-in thinking ability of OLMs, we propose a test-time Latent Navigation Reasoning (LNR) module that can be seamlessly integrated into the decoding space. LNR encourages the model to retrieve more target-relevant observations and make effective navigation decisions. Through comprehensive experiments, we show that our framework surpasses existing state-of-the-art baselines on public SAVN benchmarks without using any training data. Moreover, we introduce a new \emph{Global Navigation Instruction} setting to further evaluate the ability of OLMs to serve as embodied navigation agents. Code: https://github.com/rikeilong/OmniAV\_Nav.