RACES recursively composes verifiable environments as LEGO bricks to scale RL reasoning training, boosting model performance on unseen benchmarks with far fewer base environments.
Vision-OPD distills a crop-conditioned teacher into a full-image student via on-policy self-distillation to improve fine-grained visual understanding without external teachers or tools. It achieves competitive or superior performance on fine-grained benchmarks against larger open-source, closed-sour