PAGER closes the semantic-execution gap for point-precise geometric GUI control via dependency-structured planning and pixel-level execution, achieving 4.1x higher task success than general baselines.
VisHarness trains a visual agent to orchestrate heterogeneous experts for multi-turn reasoning, achieving strong results on segmentation, detection, and counting tasks.
StraTA introduces trajectory-level strategies into agentic reinforcement learning via hierarchical rollout training, improving long-horizon decision-making and reaching 93.1% on ALFWorld.