Good Papers

On-Policy Distillation Teaches New Skills but Not New Knowledge

On-policy distillation teaches compositional reasoning across unseen structures but transfers minimal new factual knowledge, instead organizing existing knowledge.

Yixuan Tang, Yi Yang

Published Oct 7, 2026arXiv ↗

88%
OverallMust read
?
OverallMust readVote to see the score
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel15/20reviewers recommend it
lenient 4/5
medium 8/10
strict 3/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown. We separate these capabilities using a controlled synthetic framework that measures the student's initial capabilities and independently controls the teacher's additional facts, compositional skill, or both. Across four models from three families, reverse-KL OPD reliably transfers compositional skill across unseen reasoning structures, but transfers minimal factual knowledge. Decoupling the distillation recipe reveals the source of this asymmetry: replacing reverse KL with forward KL restores factual transfer, whereas student rollouts specifically improve the execution of multi-step reasoning. Experiments on recent factual QA and competition mathematics show a similar asymmetry under reverse-KL OPD, yielding notable reasoning gains without factual memory expansion. Together, these results demonstrate that on-policy distillation does not expand a model's parametric knowledge, but instead teaches it to organize and compose the knowledge it already possesses.