Good Papers

MoSE3: Learning World-Space SE(3) at Every Pixel

MoSE3 predicts dense per-pixel world-space SE(3) motion from monocular video via point tracks and rigidity embeddings, achieving state-of-the-art 6-DoF estimation and 3D tracking.

Jiahuan Cheng, Zhiyi Li, Tian Xia, Ruojin Cai, Yilun Du, Qianqian Wang

Published 2026Atlanta Poster Session 5 · Fri, Dec 11, 10:00 AM–1:00 PM local time · Hall C1▲ 1 on Hugging FacearXiv ↗OpenReview ↗

86%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel14/20reviewers recommend it
lenient 4/5
medium 7/10
strict 3/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
MoSE3 delivers a pioneering dense world-space SE(3) predictor with elegant rigidity embeddings and state-of-the-art tracking, though synthetic-only training and unproven real-world rigidity clustering leave clinical validation and true generalization unresolved.

Abstract

Dense 3D point tracking has been a prominent paradigm for modeling motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel: it captures where pixels go, not the rotation of the underlying part, nor which pixels move together as one body. We propose MoSE3, the first feed-forward model that predicts dense SE(3) motion from monocular RGB video, producing full 6-DoF rigid transforms at every pixel in world space. Per-pixel SE(3) motion offers a richer view of how a scene moves: rotation, translation, and grouping all at once. Directly predicting SE(3) is challenging: rotations lie on a curved manifold that is ill-suited to Euclidean regression, and annotations for SE(3) are particularly difficult to acquire. To address these challenges, MoSE3 predicts per-pixel SE(3) through two jointly learned intermediates, 3D point tracks and rigidity embeddings, and recovers SE(3) by differentiably fitting transforms within each soft rigid cluster, enabling end-to-end prediction and supervision. To close the data gap, we introduce Art-Kubric, a large-scale synthetic dataset with dense SE(3) and rigidity labels for articulated objects with rich physical interactions. MoSE3 achieves state-of-the-art SE(3) estimation at pixel, part, and object levels on both rigid and articulated benchmarks, and state-of-the-art average 3D point tracking accuracy across three datasets, while showing strong generalization to real-world videos despite being trained solely on synthetic motion data.