Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution
On-policy power distillation trains models to generate sharpened answers directly, improving single-sample math reasoning by up to 27.3 points and outperforming multi-candidate sampling and reward-based methods.
Published Oct 5, 2026 · ▲ 4 on Hugging Face · Code ★ 1
Only vote on papers you've read. Sign in with GitHub to vote.






