Good Papers

UniMate: One Unified Model to Animate Diverse Skeletons

UniMate is a unified diffusion transformer that synthesizes motion for arbitrary skeletons from text and rigged assets without test-time optimization, using topology-aware attention and a new dataset to outperform specialized animators.

Linzhan Mou, Lei, Jiahui, Zhiyang Dou, Chenyue Cai, Chaoyue Song, Adam Finkelstein, Szymon Rusinkiewicz

Published Sep 4, 2026▲ 23 on Hugging FaceCode ★ 1,553arXiv ↗

80%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel12/20reviewers recommend it
lenient 5/5
medium 5/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
UniMate delivers a genuinely slick topology-aware diffusion transformer and impressive zero-shot cross-topology transfer, but its synthetic-rig pipeline, canonicalization losses, and sharp failure modes on exotic skeletons leave real-world animator and biomechanist adoption still unproven.

Abstract

Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion to drive them remains a bottleneck. Existing learned animators are topology-constrained: they rely on category-specific templates or require per-skeleton fine-tuning and reference motions at inference. We present UniMate, a unified foundation model that synthesizes articulated motion for arbitrary skeletons from a rigged 3D asset and a text prompt, with no test-time optimization or per-skeleton retraining. UniMate introduces a topology-aware diffusion transformer, which integrates skeletal topology into attention via three mechanisms: (1) a graph-aware attention bias from pairwise joint relations and geodesic distances; (2) a spectral rotary position embedding generalizing RoPE to arbitrary kinematic trees via the graph Laplacian; and (3) a global topological conditioner attention-pooled from the rest-pose skeleton. We also curate UniML3D, 13,006 motion sequences spanning bipedal, quadrupedal, avian, marine, insectoid, serpentine, and articulated rigid objects with unified canonicalization and text pairing. Trained on this dataset, UniMate outperforms state-of-the-art baselines in quality, generalization, and efficiency, and supports zero-shot cross-topology transfer, in-betweening, expansion, and text-guided editing. Our project page is available at https://linzhanmou.com/unimate/.