RobotUse: Allocating Computation, Context, and Decisions
RobotUse organizes robot computation, context, and decisions around revisable physical actions via visual target selection and persistent playbooks, achieving 45% RoboLab success and real-world learning.
Published Oct 4, 2026▲ 14 on Hugging FaceCode ★ 4arXiv ↗
Only vote on papers you've read. Sign in with GitHub to vote.
RobotUse delivers a practical harness that replaces raw execution history with persistent subgoal playbooks, yielding a real 6.7-point RoboLab lift and real-world transfer, though critics remain skeptical that its 45% success and unquantified feedback justify the…
Abstract
Robot agents must connect their intended actions to observed outcomes while retaining the context needed to revise their choices over repeated attempts. Existing interfaces often leave these choices inside predefined tools or require agents to manage detailed execution code and its growing history. We introduce RobotUse, a robot agent harness that organizes computation, context, and decisions around specifying and revising physical actions. Agents visually select targets and poses, while the backend handles geometry, motion planning, and control. Subagents retain detailed interactions within each subgoal and return the information needed for subsequent decisions. Continual harnessing lets agents learn from execution by updating a persistent playbook. On RoboLab, RobotUse achieves 45% task success, outperforming CaP-X by 6.7 percentage points while maintaining compact decision contexts and reducing reliance on predefined action abstractions. Furthermore, we show that RobotUse learns from real-world execution despite imperfect feedback and transfers what it learns to subsequent tasks. Project page is available at https://robotuse-team.github.io/.