Using summed last-k encoder layers and combining RAE with REPA, RAEv2 achieves state-of-the-art gFID of 1.06 in 80 epochs with 10x faster convergence and free guidance.
A dataset of multi-aspect human visual similarity judgments benchmarks vision-language models and yields the TPIPS metric, which aligns with human perception and enables text-guided image retrieval and generative evaluation.
UNITE unifies tokenization and latent diffusion via a shared generative encoder, enabling single-stage joint training from scratch without adversarial losses or pretrained encoders to reach near state-of-the-art FID scores.