Good Papers

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

Latent-MOPD distills multi-teacher LLM specialists via hidden-state and prediction-level on-policy supervision, outperforming token-only and representation-only baselines across math, code, and logic benchmarks.

Zhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li, Koyoshi Shindo, Brandon Joseph Lwowski, Jing Li

Published Oct 1, 2026▲ 63 on Hugging FacearXiv ↗

78%
OverallHighly rated
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel11/20reviewers recommend it
lenient 3/5
medium 7/10
strict 1/5
Panel consensus
Latent-MOPD achieves sweeping same-family and cross-family gains by fusing hidden-state and token-level multi-teacher supervision, though its precise channel contributions and cross-family margins remain insufficiently ablated.

Abstract

On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method for LLMs. It integrates existing specialists through both their predictions and the hidden states used to compute them, without additional teacher training. To coordinate representation supervision from multiple specialists, we select late-layer targets according to the teacher-student relationship, bridge unequal hidden widths with a shared projection, and group updates by domain. Each teacher's supervision gradually shifts from hidden states to token predictions, with both channels using the same routed specialist. In our main same-family setting, Latent-MOPD outperforms the token-only, representation-only and uniform-averaging baselines on all nine benchmarks across math, code and logic. With the same parameter count as each teacher, the student also surpasses the per-benchmark best teacher on a majority of these benchmarks. With larger, separately developed cross-family teachers, Latent-MOPD outperforms both single-channel baselines on all benchmarks. A same-family all-layer representation-only control remains stable with domain-pure updates but collapses when teacher domains are interleaved within an update. Our results show that a single student can integrate capabilities from several specialists through both their output distributions and internal representations.