ComPose tracks 6DoF object pose in RGB video by using hand motions as complementary cues, achieving robust accuracy under severe occlusion without external priors.
MolmoMotion predicts goal-conditioned 3D point trajectories from visual history and language, outperforming baselines on PointMotionBench and improving robot manipulation and video synthesis.
GenCOPE achieves synthetic-to-real generalized category-level object pose estimation via 2D/3D semantic consistency and cross-modality fusion using only global features, outperforming prior methods on REAL275 and Wild6D.
Self-supervised training on 160,000 in-the-wild videos via a shared coarse mesh yields emergent canonical object frames without pose labels, matching supervised category-level pose estimation accuracy.
SimpliHuMoN is a simple transformer that predicts human pose and trajectory together, achieving state-of-the-art results across standard benchmarks without task-specific modifications.
ViDiHand leverages pretrained video diffusion models to reconstruct 4D hand poses directly from full egocentric video without detectors, substantially outperforming prior methods on ARCTIC, HOT3D, and HOI4D.
ProxyPose recasts 6-DoF pose tracking as video-to-video translation using a diffusion model to generate proxy videos for classical pose estimation, achieving state-of-the-art accuracy without 3D models or masks.
A primitive-fitting framework recovers articulated object kinematics from single casual videos via joint optimization of part segmentation and joint parameters under occlusions.
AnyHand provides 6.6M synthetic RGB-D hand images with occlusions and aligned depth, significantly improving 3D hand pose estimation benchmarks and showing data diversity rivals scale.
StableHand estimates world-space dual-hand motion from egocentric video via quality-aware flow matching, cutting W-MPJPE by 20-25% over baselines on occluded benchmarks.
COMPOSE recasts multi-view 3D pose estimation as hypergraph exact-cover optimization, improving training-free average precision by up to 31 points over prior optimization methods.
MoSE3 predicts dense per-pixel world-space SE(3) motion from monocular video via point tracks and rigidity embeddings, achieving state-of-the-art 6-DoF estimation and 3D tracking.