
CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models
CollabVR pairs vision-language models with video generation models in closed-loop step-level planning and verification, reducing drift and simulation errors for major video reasoning gains.
Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 71 on Hugging Face · Code ★ 10
Only vote on papers you've read. Sign in with GitHub to vote.