Good Papers

Showing papers from Megvii Technology Inc. Show all papers

67%Highly rated
?Highly ratedVote to see the score

Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving

Xingtai Gui, Yucheng Zhou, Dongqian Guo, jiahao gong and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

A 3D Scene is Worth 1K Tokens: 3D-Grounded Representation for Scene Generation at Scale

Dongxu Wei, QiXu, Zhiqi Li, Hangning Zhou and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
89%Must read
?Must readVote to see the score

PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models

PriorVLA freezes a prior expert and trains an adaptation expert via expert queries to preserve pretrained vision-language-action priors, updating only 25% of full fine-tuning parameters while outperforming baselines on OOD and few-shot robot manipulation.

Xinyu Guo, Bin Xie, Wei Chai, Xianchi Deng and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
80%Must read
?Must readVote to see the score

ChainFlow-VLA: Causal Flow Planning with Vision-Language Models

ChainFlow-VLA unifies causal trajectory generation and global diffusion refinement via vision-language-conditioned residual distributions, scoring 94.85 on NAVSIM v1.

Xiyang Wang, Xinlin Wang, Tingguang Zhou, Gong Chen and 6 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation

WorldDrive unifies vision and motion representations to couple scene generation with planning, achieving leading vision-only autonomous driving performance with high-fidelity future video generation.

Xingtai Gui, Meijie Zhang, Tianyi Yan, Wencheng Han and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Realtime-VLA FLASH: Speculative Inference Framework for Diffusion-based VLAs

Realtime-VLA FLASH uses a draft model and parallel verification to replace most full diffusion-based VLA inference rounds with faster speculative ones, cutting average latency 3.04x to 19.1 ms.

Jiahui Niu, Kefan Gu, Yucheng Zhao, shengwen Liang and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
80%Must read
?Must readVote to see the score

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving

CoWorld-VLA embeds multi-expert world tokens into vision-language-action models and couples diffusion planning with scene context to generate continuous ego trajectories, improving autonomous driving performance.

Jingqi Wang, minqing huang, Zihan Liang, Yujiao Xiang and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5