
On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training
On-policy methods continuously adjust parameter update directions, unlike consistent SFT updates; constraining SFT to these directions via OPSFT transfers on-policy generalization advantages to supervised fine-tuning.
Published Sep 29, 2026 · 0 citations · ▲ 81 on Hugging Face
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.




