MolmoMotion predicts goal-conditioned 3D point trajectories from visual history and language, outperforming baselines on PointMotionBench and improving robot manipulation and video synthesis.
Continuous flow language models outperform discrete diffusion in quality and speed, and distilling their unique flow map enables one-step generation surpassing eight-step discrete diffusion.
PACE predicts agentic benchmark scores from small, selected non-agentic test subsets via regression, achieving under 4% error and over 0.80 correlation at under 1% evaluation cost.
Monarch-RT factorizes video diffusion attention via Monarch matrices to reach 95% sparsity without quality loss, enabling 16 FPS real-time generation on one GPU with 1.4-11.8x kernel speedups.
AstraFlow is a dataflow-oriented RL system for agentic LLMs that decouples rollout, dataflow, and training to enable multi-policy collaborative training with 2.7x faster training.
PROSPER defines the Maximum Entropy Blackwell Winner and applies it to multi-objective preference fine-tuning, outperforming baselines on instruction following and chat benchmarks.
Concept modulation models unify conditional latent variable model identifiability and extrapolation via attribute potentials and algebraic criteria for unseen attributes.
Spatiotemporal Noise-Contrastive Estimation learns energy-based models via joint spatiotemporal differences to avoid failure modes of spatial or temporal methods alone, matching state-of-the-art density estimation.
FineVLA introduces fine-grained action-aligned supervision for steerable vision-language-action policies, yielding up to 86.8% simulation and 62.7 real-world success and boosting steerable control over coarse instructions.
RL4F introduces an offline RL benchmark for tokamak plasma control using DIII-D dynamics, finding model-based methods perform best but no method dominates all tasks.
OdysSim trains 8B behavioral foundation models via SOUL taxonomy and multi-stage recipes, ranking first on eight human simulation benchmarks while nearly matching real-user reaction alignment.
SCALLOP introduces a Hutchinson-free likelihood distillation objective for few-step Boltzmann generators, reducing training variance and time while achieving up to 10x inference speedup.
Point4D infers dense 3D point trajectories across multi-hundred-frame videos via a decoupled query-based motion decoder, outperforming prior short-window feed-forward 4D methods.
MoLF dynamically routes optimizer updates between full fine-tuning and LoRA to match or beat the stronger static method across tasks, and its efficient variant surpasses AdaLoRA and AdaMix by up to 11.70 points.