Good Papers

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving

CoWorld-VLA embeds multi-expert world tokens into vision-language-action models and couples diffusion planning with scene context to generate continuous ego trajectories, improving autonomous driving performance.

Jingqi Wang, minqing huang, Zihan Liang, Yujiao Xiang, JiaJie Huang, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang, Gong Chen

Published 2026Sydney Poster Session 2 · Tue, Dec 8, 5:00 PM–8:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

80%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel12/20reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
CoWorld-VLA earns its "must read" status by forcing diffusion planners to treat explicit multi-expert world tokens as conditioning rather than hidden backbone reasoning, though NAVSIM-only validation and unclear out-of-domain durability leave structural claims partly unproven.

Abstract

Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing reasoning mechanisms still struggle to provide planning-oriented intermediate representations: textual Chain-of-Thought (CoT) fails to preserve continuous spatiotemporal structure, while latent world reasoning remains difficult to use as a direct condition for action generation. In this paper, we propose CoWorld-VLA, a multi-expert world reasoning framework for autonomous driving, where world representations serve as explicit conditions to guide action planning. CoWorld-VLA extracts complementary world information through multi-source supervision and encodes it into expert tokens within the VLA, thereby providing planner-accessible conditioning signals. Specifically, we construct four types of tokens: semantic interaction, geometric structure, dynamic evolution, and ego trajectory tokens, which respectively model interaction intent, spatial structure, future temporal dynamics, and behavioral goals. During action generation, CoWorld-VLA employs a diffusion-based hierarchical multi-expert fusion planner, which is coupled with scene context throughout the joint denoising process to generate continuous ego trajectories. Experiments on NAVSIM v1 & v2 demonstrate future-scene modeling capability and strong planning performance, including collision avoidance and trajectory accuracy. Ablation studies further validate the complementarity of expert tokens and their effectiveness as planning conditions for action generation. Code will be available at https://github.com/AFARI-Research/CoWorld-VLA.