Good Papers

Masked Visual Actions for Unified World Modeling

Masked Visual Actions expresses robot and object motion as revealed pixel trajectories to unify forward dynamics, planning, and inverse modeling in video world models with minimal finetuning.

Hadi Alzayer, Wenlong Huang, Haonan Chen, Christopher Luey, Lvmin Zhang, Maneesh Agrawala, Gordon Wetzstein, Fei-Fei Li, Yilun Du, Jiajun Wu, Jia-Bin Huang

Published 2026Sydney Poster Session 1 · Tue, Dec 8, 10:00 AM–1:00 PM local time · Hall 1-4▲ 9 on Hugging FaceCode ★ 110arXiv ↗OpenReview ↗

76%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel10/20reviewers recommend it
lenient 5/5
medium 4/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
Masked visual actions elegantly unify forward dynamics and inverse modeling through pixel-space trajectory masks with minimal data, though "unified" remains unproven without embodiment ablations, inverse dynamics error bars, and stress-tested contact grounding.

Abstract

Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video. Revealing robot motion makes the model act as a forward dynamics model that predicts the scene's response to low-level robot actions, while revealing desired object motion makes the same model recover robot behavior consistent with that outcome. Finetuned with only 15 hours of masked examples from real videos and simulation, a single checkpoint achieves strong visual fidelity and controllability across diverse scenes and multiple embodiments. In downstream manipulation settings, the model produces imagined rollouts whose outcomes correlate with real-world execution for policy evaluation, improves decision making by ranking candidate futures in model-based planning, and supports inverse modeling by synthesizing robot motion from desired object motion.