RefDecoder conditions video VAE decoders on reference images via attention to recover lost detail, boosting reconstruction PSNR by up to 2.1 dB and improving consistency across video generation tasks without retraining.
MolmoMotion predicts goal-conditioned 3D point trajectories from visual history and language, outperforming baselines on PointMotionBench and improving robot manipulation and video synthesis.
ADUCA is a parameter-free cyclic algorithm for Minty variational inequalities that uses delayed operator updates to avoid line searches and achieves near-optimal global oracle complexity.
Grounding is formulated as bidirectional concept correspondence to recover all image-text span correspondences without prespecified phrases via ConCor-1, improving F1 by 48% and 29% over baselines.