Good Papers

LiFT: Loop Flow Transformers

Loop Flow Transformers loop a shared diffusion transformer with depth-indexed regression targets, improving generation with more inference compute and fewer parameters than dense models.

Mohammad Mahdi Derakhshani, Pedro M. P. Curvo, Gertjan J. Burghouts, Jan-Willem van de Meent, Cees G. M. Snoek

Published Oct 4, 2026▲ 9 on Hugging FacearXiv ↗

74%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel9/20reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
LiFT delivers a striking ImageNet FID gain with far fewer parameters via continuous-depth looped DiT cores, but its evaluation lacks error bars, COCO or text-to-image tests, and wall-clock proof that longer rollouts justify the inference cost.

Abstract

We introduce Loop Flow Transformers (LiFT), a family of looped generative models that scales computation by repeatedly applying a shared Diffusion Transformer (DiT) core, with only light changes to the standard architecture. Rather than asking every recurrent step for the final prediction, LiFT trains each step with a single regression target: a point on a straight path from the model's initial estimate to the flow-matching target. Because we index these targets by a continuous depth coordinate, a trained model can loop far beyond its training depth with no retraining, early exits, or other modifications. In our experiments, these longer rollouts improve generation, so inference computation can grow without adding parameters. On ImageNet at 256x256, LiFT-L/2 achieves an FID 3.34 points lower than our dense DiT-XL/2 baseline while using approximately 60% fewer parameters, 32% fewer training FLOPs, and 52% fewer inference FLOPs.