
The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation
RIDE extrapolates RL-induced representation residuals for stable on-policy distillation, surpassing output-space methods and matching or exceeding RL teachers.
Published Sep 29, 2026 · 0 citations · ▲ 579 on Hugging Face · Code ★ 6
Only vote on papers you've read. Sign in with GitHub to vote.












































