LLM RL suffers from training-inference policy mismatch; MIPU optimizes monotonic inference policy improvement to stabilize training and boost reasoning performance.
Embodied-R1.5 is an 8B-parameter embodied foundation model achieving state-of-the-art results on 16 of 24 embodied VLM benchmarks via multi-task balanced RL and a closed-loop planner-grounder-corrector framework.