LLM RL suffers from training-inference policy mismatch; MIPU optimizes monotonic inference policy improvement to stabilize training and boost reasoning performance.
ForceFlow uses force-aware flow matching with asymmetric multimodal fusion and vision-to-force handover to achieve robust contact-rich manipulation with 37% higher success and stronger zero-shot generalization.
Looped-MoE models scale better than standard transformers via routing divergence that recovers expressivity, and loop boundaries enable efficient early exits with minimal quality loss.
Embodied-R1.5 is an 8B-parameter embodied foundation model achieving state-of-the-art results on 16 of 24 embodied VLM benchmarks via multi-task balanced RL and a closed-loop planner-grounder-corrector framework.