GDSD improves diffusion language models via advantage-guided denoiser self-distillation that bypasses ELBO likelihood surrogates, boosting test accuracy up to 19.6%.
Post-hoc confidence remasking in masked diffusion language models offers little benefit under standard decoding and worsens diversity collapse under stochastic sampling, showing setting-dependent gains.