Good Papers

Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents

Attacca trains visual goal-conditioned policies on complete search-to-interact trajectories with decoupled goal images and behavioral-phase conditioning to improve long-horizon embodied task success by up to 7x.

Gyusik Seo, Jaehong Yoon

Published Oct 6, 2026▲ 3 on Hugging FaceCode ★ 3arXiv ↗

78%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel11/20reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
Attacca delivers striking long-horizon Minecraft gains through decoupled goal masking and phase conditioning, but without isolated ablations, visible-target baselines, or released code, its 7x improvement remains unverified outside this toy world.

Abstract

A central capability of embodied agents is to accomplish complex objectives through sequences of interdependent tasks. Yet existing visual goal-conditioned policies underlying these agents are typically evaluated on isolated interactions where the target is already visible, and thus do not capture the conditions that arise during continuous long-horizon task execution. In such settings, each task begins from the state left by the previous one: the agent may end at a different position and orientation, the world may have been modified, and the next interaction target may lie outside the current field of view. As a result, agents relying on such policies may struggle to proceed to the next task when they cannot ground their target in the current observation. To address this challenge, we propose Attacca, a new approach that trains visual goal-conditioned policies on complete search-to-interact trajectories using goal images decoupled from the execution environment. Attacca uses context-decoupled goal sampling to pair each demonstration with a class-compatible masked goal image from another world, removing direct scene and pose correspondence. It learns dense current-view grounding through a target-mask prediction head, providing auxiliary supervision beyond action imitation. We further introduce behavioral-phase conditioning that teaches the policy to distinguish Search, Approach, and Interact stages and adapt its control as execution progresses. We evaluate Attacca on multiple short- and long-horizon embodied tasks in Minecraft. Our method achieves 39.0-47.5% clean success, improving over the strongest baseline by 1.7-2.4x. On long-horizon tasks, it attains 54%, 30%, and 28% completion, yielding up to a 7x improvement.