FAST-Brain unifies flow-aligned generation, graph convolutions, and transformers to model rs-fMRI directly, with approximation error scaling by intrinsic dimension rather than ambient dimension, achieving state-of-the-art functional and effective connectivity recovery.
VTS frames grounded long-video QA as self-correcting search over an adaptive temporal tree with explicit backtracking, improving grounding and answer accuracy across benchmarks.
AVSD separates cross-view consensus from privileged residuals in multi-view self-distillation to adaptively supervise reasoning models, improving math and code benchmarks over single-view methods and GRPO.
PhyMotion evaluates human video motion via physics-simulated 3D trajectory rewards across kinematics, contact, and dynamics, improving RL post-training realism by +68 Elo.
StreamGaze introduces a benchmark for evaluating gaze-guided temporal and proactive reasoning in streaming videos, revealing large performance gaps between state-of-the-art MLLMs and humans.
AVIC adaptively scales test-time visual imagination via world models for spatial reasoning, matching fixed strategies with fewer calls while exceeding GPT-4o.