
MEND: RL For Flow Models via Proximal Velocity Matching
MEND uses proximal velocity matching to cap rewards and accept only cost-effective sample moves, outperforming prior flow-model RL methods in far fewer updates without KL penalties or reference models.
Published Oct 5, 2026 · ▲ 2 on Hugging Face · Code ★ 1
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
























