Good Papers

iADD: Improving Alignment and Diversity in Diffusion Policy Optimization

iADD analyzes diffusion policy optimization to show only-latter-timestep updates harm diversity, then proposes incremental Feynman-Kac training that improves alignment-diversity tradeoffs across tasks.

Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel

Published Oct 1, 2026▲ 4 on Hugging FaceCodearXiv ↗

76%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel10/20reviewers recommend it
lenient 2/5
medium 7/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
iADD delivers rigorous theory and strong ablations that correct DDPO's diversity collapse through incremental Feynman-Kac training, though its proof remains confined to toy tasks without ImageNet-scale or real-world validation.

Abstract

Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that \emph{only-latter timestep} updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.