ComPose tracks 6DoF object pose in RGB video by using hand motions as complementary cues, achieving robust accuracy under severe occlusion without external priors.
MolmoMotion predicts goal-conditioned 3D point trajectories from visual history and language, outperforming baselines on PointMotionBench and improving robot manipulation and video synthesis.
GenCOPE achieves synthetic-to-real generalized category-level object pose estimation via 2D/3D semantic consistency and cross-modality fusion using only global features, outperforming prior methods on REAL275 and Wild6D.
Self-supervised training on 160,000 in-the-wild videos via a shared coarse mesh yields emergent canonical object frames without pose labels, matching supervised category-level pose estimation accuracy.