VTS frames grounded long-video QA as self-correcting search over an adaptive temporal tree with explicit backtracking, improving grounding and answer accuracy across benchmarks.
PhyMotion evaluates human video motion via physics-simulated 3D trajectory rewards across kinematics, contact, and dynamics, improving RL post-training realism by +68 Elo.
AVIC adaptively scales test-time visual imagination via world models for spatial reasoning, matching fixed strategies with fewer calls while exceeding GPT-4o.
MetaCanvas enables multimodal LLMs to plan directly in diffusion latent spaces, outperforming global-conditioning baselines across six precise visual generation tasks.
SkillRL evolves agents via recursive skill-augmented reinforcement learning with automatic skill discovery and hierarchical library co-evolution, cutting token use while achieving state-of-the-art results across complex tasks.