On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
In controlled strong-to-weak distillation, rollout policy is less central than token-level KL direction and learning rate, though on-policy data can improve generalization on harder reasoning tasks.
Published Sep 28, 2026 · 0 citations · ▲ 194 on Hugging Face
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.




























