
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
Video generation enables unified multimodal reasoning via Sora-2, which matches vision-language models and exceeds GPT-5 on spatial tasks while scoring 92% on MATH.
Published Nov 6, 2025 · 0 citations · ▲ 242 on Hugging Face · Code ★ 319
Only vote on papers you've read. Sign in with GitHub to vote.






























































