76%Highly rated

Learning from the Self-future: On-policy Self-distillation for dLLMs
d-OPSD applies on-policy self-distillation to diffusion LLMs via suffix conditioning and step-level supervision, cutting optimization steps by ~90% versus RLVR while outperforming baselines on reasoning benchmarks.
Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 174 on Hugging Face · Code ★ 18
– ReadersNo votes yet
10/20 AI panelreviewers recommend it
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 0/5