Good Papers

Discrete Flow Matching: Convergence Guarantees Under Minimal Assumptions

Discrete Flow Matching achieves non-asymptotic KL and total variation convergence bounds under minimal approximation error assumptions with improved scaling in vocabulary size and dimension.

Le-Tuyet-Nhi PHAM, Giovanni Conforti, Zhenjie Ren, Alain Durmus

Published 2026Paris Poster Session 6 · Fri, Dec 11, 2:30 PM–4:30 PM local time · Paris Poster HallarXiv ↗OpenReview ↗

70%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel4/20reviewers recommend it
lenient 2/5
medium 1/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Flow Matching has recently emerged as a popular class of generative models for simulating a target distribution $μ_1$ from samples drawn from a source distribution $μ_0$. This framework relies on a fixed coupling between $μ_0$ and $μ_1$, and on a deterministic or stochastic bridge to define an interpolating process between the two distributions. The time marginals of this process can then be approximately sampled by estimating the transition rates, or more generally the generator, of its Markovian projection. This framework has recently been extended to the case of discrete source and target distributions, under the name Discrete Flow Matching (DFM). However, theoretical guarantees for such models remain scarce. In this paper, we study two DFM models on $\mathbb{Z}_m^d = \{0,\ldots,m-1\}^d$, sampled through time discretization, and derive non-asymptotic associated bounds for both of them. In contrast to previous work, we establish non-asymptotic bounds in Kullback--Leibler divergence for the early-stopped version of the target distribution. We also derive explicit convergence guarantees in total variation distance with respect to the true target distribution. Importantly, these bounds rely only on an approximation error assumption, relaxing standard score assumptions used in earlier works, while also yielding improved dependence on the vocabulary size $m$ and the dimension $d$.