Good Papers

Beyond the Current Scene: Event-Referential Grasping with Active View Selection

BeyondSCe enables zero-shot event-referential grasping via active view selection, achieving 76% and 77% success on visible and occluded targets versus 40% and 55% baselines.

Hyunjoon Lee, Haebeom Jung, Eunsung Cha, Daeun Lee, Yu-Chiang Frank Wang, Jaesung Choe, Jaesik Park

Published Sep 30, 2026▲ 51 on Hugging FaceCode ★ 8arXiv ↗

74%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel9/20reviewers recommend it
lenient 5/5
medium 3/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
BeyondSCe delivers impressive real-robot gains for event-referential grasping with zero-shot models and smart view selection, but its reliance on frozen VLMs, untested event-prior robustness to moved objects, and unclear scene selection leave its broader transferability unresolved.

Abstract

A robot that observes people interacting with objects should be able to carry out later requests that refer back to those interactions. Such requests may specify a grasp target by the role it played in a past event rather than by its name or appearance. Moreover, the target may no longer be visible when the robot is asked to act. We present BeyondSCe, a zero-shot robotic grasping system for this event-referential setting. Given the event history and the current scene, the system identifies the requested object or part and localizes it for grasping. If the target is occluded, it combines an event prior recovered from the history with current scene geometry to select camera viewpoints likely to reveal the target. The system uses pretrained models without additional task-specific training. In real-robot experiments with a single wrist-mounted RGB-D camera, it achieves grasp success rates of 76% and 77% for initially visible and occluded targets, respectively, compared with 40% and 55% for the strongest baseline in each condition. On four additional scenes with heavy occlusion, it increases grasp success rates from 75% to 95% while reducing the mean number of views from 3.35 to 2.20, compared with an active-perception baseline given the target's ground-truth 3D bounding box.