TUBE: Tangent Upper Bound on Evidence for Discrete Diffusion Language Models
TUBE introduces a variational upper bound with unbiased Monte Carlo estimation to evaluate discrete diffusion model log-likelihoods, revealing that block diffusion and any-order autoregressive models remain below exact autoregressive baselines.
Published 2026Paris Poster Session 2 · Wed, Dec 9, 5:00 PM–7:00 PM local time · Paris Poster HallarXiv ↗OpenReview ↗
Only vote on papers you've read. Sign in with GitHub to vote.
TUBE delivers the first unbiased upper bound for diffusion LLMs, rigorously exposing ELBO gaming and confirming autoregressive dominance, though its unreported variance, missing large-scale benchmarks, and thin baseline coverage leave critical empirical gaps.
Abstract
Log-likelihood is a standard metric for evaluating generative models. Unfortunately, in contrast to autoregressive models (ARMs), discrete diffusion models generally do not admit exact computation of this quantity. Existing evaluations, therefore, rely on the evidence lower bound (ELBO), leaving unclear how much higher the true value may be. We address this by introducing the Tangent Upper Bound on Evidence (TUBE), a variational upper bound on log-likelihood that admits an unbiased Monte Carlo estimator. Our TUBE extends across latent-variable models, including masked diffusion models (MDMs), any-order ARMs (AO-ARMs), and block variants of both. Applied to block MDMs and block AO-ARMs, TUBE reveals our key empirical finding that these models lie strictly below the exact ARM baseline, showing that ARMs still dominate in likelihood.