RACES recursively composes verifiable environments as LEGO bricks to scale RL reasoning training, boosting model performance on unseen benchmarks with far fewer base environments.
PAPO-VLA improves vision-language-action reliability by identifying planning actions via action variation and trajectory outcomes, weighting them by causal importance in policy optimization, and boosting benchmark performance.
Dirichlet-Guided Group Forecasting reduces time-series over-smoothing by modeling multi-modal predictive distributions with Dirichlet-guided sampling, improving accuracy, diversity, and dynamical consistency.
Vision-OPD distills a crop-conditioned teacher into a full-image student via on-policy self-distillation to improve fine-grained visual understanding without external teachers or tools. It achieves competitive or superior performance on fine-grained benchmarks against larger open-source, closed-sour