RepFusion: Leveraging Multimodal Priors for Denoising in Representation Space
RepFusion conditions a diffusion transformer on multimodal LLM outputs to denoise visual representations, outperforming comparable newly initialized denoisers.
Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 17 on Hugging Face
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.