ROMA improves multimodal reasoning robustness to visual corruption via dual-pass RL optimization that avoids reward poisoning while preserving clean accuracy.
UniT unifies human and humanoid actions via visual anchoring into shared latent tokens for scalable policy learning and world modeling. It achieves state-of-the-art data efficiency, zero-shot transfer, and cross-embodiment dynamics alignment.
AnyMo introduces OmniHuMo dataset with 5,000 hours of multimodal motion data and proposes a masked modeling framework for scalable any-modality conditional motion synthesis with flexible spatial and stylistic control.
RepoMirage evaluates code agents via repository perturbations, revealing severe repository context reasoning gaps and exploration drift, while RepoAnchor improves performance through structure-first scaffolding.
FaithfulFaces improves identity-preserving video generation via pose-shared identity alignment and achieves state-of-the-art consistency across pose changes and occlusions.
OpenSearch-VL introduces an open-source recipe training multimodal deep search agents via curated data, diverse tools, and multi-turn fatal-aware GRPO, achieving over 10-point benchmark gains comparable to proprietary models.
CM2 replaces verifiable outcome rewards with checklist rewards for multi-turn tool-use RL, improving 8B models by 8, 12 points on agent benchmarks using simulated environments.
PAMod models cyclical non-stationary shifts via phase-amplitude modulation in normalized space to achieve state-of-the-art forecasting with lower cost and broad plug-and-play gains.
ReDiff reframes vision-language diffusion as active refining via error-correction training and online self-correction loops, breaking error cascades to improve coherence, factual accuracy, and parallel generation.
Unify-Agent reframes image synthesis as an agent pipeline with search and recaptioning, improving generation of long-tail factual concepts via 143K curated trajectories.
RLVR training causes reasoning outputs to structurally converge on seen prompts, and Min-kNN Distance detects this collapse via black-box sampling to identify contamination.