Good Papers

RelationVGGT : Visual Geometry Transformers for 3D Spatial Relation Segmentation

RelationVGGT enables feed-forward 3D spatial relation segmentation across multi-view images without camera poses or category names by combining visual semantics with geometry-aware representations via a relation transformer.

Minsu Kim, Jaesung Choe, Jiwoo Lee, Frank Wang, Seon Joo Kim

Published 2026Sydney Poster Session 5 · Thu, Dec 10, 10:00 AM–1:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

71%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel7/20reviewers recommend it
lenient 4/5
medium 3/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit -- yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We formulate 3D spatial relation segmentation in a feed-forward, pose-free multi-view setting: given a visually specified subject and a relational text query, the model segments the target across views without receiving its category name. To this end, we propose RelationVGGT, a novel feed-forward framework that integrates semantic features from a visual foundation model with geometry-aware representations from a 3D geometry foundation model and leverages a relation transformer for subject-conditioned, cross-view relation prediction -- requiring neither per-scene optimization nor known camera poses. We additionally provide a fully automated annotation pipeline built on ScanNet++ with VLMs and LLMs, enabling scalable training data generation for this new task.