Latent-MOPD: Latent Multi-Teacher On-Policy Distillation
Latent-MOPD distills multi-teacher LLM specialists via hidden-state and prediction-level on-policy supervision, outperforming token-only and representation-only baselines across math, code, and logic benchmarks.
Published Oct 1, 2026 · 0 citations · ▲ 63 on Hugging Face
Only vote on papers you've read. Sign in with GitHub to vote.







