Good Papers

RepFusion: Leveraging Multimodal Priors for Denoising in Representation Space

RepFusion conditions a diffusion transformer on multimodal LLM outputs to denoise visual representations, outperforming comparable newly initialized denoisers.

Xichen Pan, Satya Narayan Shukla, Aashu Singh, Shlok K Mishra, Saining Xie

Published 2026Atlanta Poster Session 4 · Thu, Dec 10, 4:30 PM–7:30 PM local time · Hall C1▲ 17 on Hugging FacearXiv ↗OpenReview ↗

71%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel7/20reviewers recommend it
lenient 3/5
medium 4/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Large language models (LLMs) are widely used in text-to-image (T2I) systems, but they are typically limited to text encoding, while denoising is handled by newly trained generative backbones. The emergence of representation autoencoders (RAEs) shifts the generation target toward semantically structured visual representations, creating a latent space that is more compatible with pretrained LLM priors. Inspired by multimodal LLMs (MLLMs), where an MLP projector is sufficient to align clean visual representations with a pretrained LLM, we repurpose the MLLM itself as a noisy representation encoder, extending this mechanism from clean to noisy inputs. We present RepFusion, which uses the resulting MLLM outputs as the conditioning signal for a diffusion transformer. In controlled comparisons at similar inference budgets, RepFusion outperforms baselines that devote comparable capacity to newly initialized denoisers. These results demonstrate that MLLMs provide strong priors for denoising visual representations and that, by conditioning on evolving noisy representations, test-time compute can be productively spent on repeated MLLM conditioning in modern T2I systems.