Self-supervised training on 160,000 in-the-wild videos via a shared coarse mesh yields emergent canonical object frames without pose labels, matching supervised category-level pose estimation accuracy.
Semantic motion anchors discretize gesture motion into verbalized primitives to align text and gestures, improving retrieval and generation by capturing communicative intent over low-level kinematics.