Interpolated Policy Distillation: A Controllable Continuum Between Off-Policy and On-Policy Distillation
Interpolated Policy Distillation mixes student and teacher token distributions to balance trajectory quality and learnability, outperforming off-policy and on-policy distillation across reasoning benchmarks.
Published Sep 29, 2026arXiv ↗
Only vote on papers you've read. Sign in with GitHub to vote.
IPD delivers a rigorous token-level continuum between off- and on-policy distillation that produces genuinely interleaved reasoning trajectories and strong benchmark gains, yet its practical value hinges entirely on an unverified speculative-decoding rule, missing trajectory-length ablations, and…
Abstract
Off-policy and on-policy distillation have traditionally been formulated as separate paradigms, each favoring a different property of distillation trajectories. Teacher-generated (off-policy) traces are typically high-quality but lie far from the student's distribution, whereas student-generated (on-policy) rollouts are more learnable but often contain erroneous reasoning. We view these paradigms as the endpoints of a policy continuum and posit that a more effective rollout policy may lie in between. We introduce \textbf{Interpolated Policy Distillation (IPD)}, which defines the next-token distribution at every decoding step as an explicit linear interpolation between the student and teacher distributions. The interpolation operates at the distribution level, token by token, and its coefficient provides direct control over the balance between trajectory quality and student learnability. Naively sampling from this policy would require sequentially querying the teacher at every token and is thus expensive. To make IPD practical, we accelerate it with a new speculative-decoding rule while exactly preserving the interpolated next-token distribution.At the trajectory level, the resulting rollouts naturally interleave student- and teacher-generated segments. Unlike recent heuristic segment-interleaving methods, however, this interleaving is induced by an exactly realized token-level interpolated policy rather than by hand-designed switching rules. Across text-only and multimodal reasoning benchmarks, IPD consistently outperforms both endpoint policies (SFT and OPD), their conventional two-stage combination (SFT-then-OPD), and recent heuristic segment-interleaving methods, demonstrating that token-level policy interpolation better balances trajectory quality and student learnability.