Good Papers

DiLaDiff: Distilled Latent-augmented Diffusion for Language Modeling

DiLaDiff proposes a latent-augmented masked diffusion language model with consistency distillation that improves quality and accelerates inference by generating continuous latents in negligible time.

Jean-Marie Lemercier, Tomas Geffner, Morteza Mardani, Karsten Kreis, Arash Vahdat, Ante Jukić

Published 2026Paris Poster Session 5 · Fri, Dec 11, 11:30 AM–1:30 PM local time · Paris Poster HallarXiv ↗OpenReview ↗

69%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel3/20reviewers recommend it
lenient 2/5
medium 1/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Diffusion language models intrinsically fail to capture correlations between decoded tokens, which leads to a harsh trade-off between sampling quality and throughput. To solve this issue, we propose DiLaDiff, a variant of masked diffusion language models with three components: (1) a continuous latent space with semantic capabilities, learned by an auto-encoder fine-tuned from an existing masked diffusion language model; (2) a latent diffusion model learning the prior over the encoder distribution; (3) a consistency model distilling the learned prior into a few-step latent generative model. We show that, even without distillation, our latent-guided diffusion model outperforms the masked diffusion baseline while significantly accelerating inference. Consistency distillation further lowers the computational overhead of continuous diffusion, such that the latent is generated in negligible time compared to discrete decoding.