Feed-forward framework decomposes scenes into instance-structured 3D token groups from unposed multi-view images to enable reconstruction, segmentation, and direct object editing.
RelationVGGT enables feed-forward 3D spatial relation segmentation across multi-view images without camera poses or category names by combining visual semantics with geometry-aware representations via a relation transformer.