Increasing rollouts fails in group-based RL because normalized gradients cancel; SALT adaptively reweights updates via subspace decomposition to recover effective learning.
OASES co-trains a search policy and adaptive evaluator to provide outcome-aligned process rewards, outperforming RL baselines on multi-hop QA benchmarks.
xHC expands Transformer hyper-connections beyond four streams via sparse updates and temporal augmentation, improving scaling efficiency. It boosts 18B MoE downstream scores by 4.0 points over mHC with lower compute and reduced memory traffic via xHC-Flash.
Knowledge-graph paths provide intermediate supervision for self-evolving search agents, improving question validity via relational context and solver rewards via waypoint coverage, boosting multi-hop QA across benchmarks.
Vision-OPD distills a crop-conditioned teacher into a full-image student via on-policy self-distillation to improve fine-grained visual understanding without external teachers or tools. It achieves competitive or superior performance on fine-grained benchmarks against larger open-source, closed-sour