Good Papers

Showing papers from Tsinghua Univ, Tsinghua University Show all papers

76%Highly rated

Learning from the Self-future: On-policy Self-distillation for dLLMs

d-OPSD applies on-policy self-distillation to diffusion LLMs via suffix conditioning and step-level supervision, cutting optimization steps by ~90% versus RLVR while outperforming baselines on reasoning benchmarks.

Yifu Luo, Zeyu Chen, Haoyu Wang, Xinhao Hu and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 174 on Hugging Face · Code ★ 18

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 0/5