DriveHierarchy hierarchically benchmarks VLM driving across four ranks from perception to closed-loop execution, linking open-loop understanding to embodied performance for diagnosing 15 models.
AVIC adaptively scales test-time visual imagination via world models for spatial reasoning, matching fixed strategies with fewer calls while exceeding GPT-4o.
SkillRL evolves agents via recursive skill-augmented reinforcement learning with automatic skill discovery and hierarchical library co-evolution, cutting token use while achieving state-of-the-art results across complex tasks.