Good Papers

Learning Discriminative Geometry for Drifting Models

Drifting models suffer from poor pixel-space performance because representation geometry controls KDE sample weighting; persistent representation learning learns discriminative geometry from raw pixels, cutting FID by 82, 95% without pretrained encoders.

Doudou Zhang, Wenwen Hou, Yilin Chen, Qi Chen

Published Oct 3, 2026▲ 4 on Hugging FaceCode ★ 3arXiv ↗

86%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel14/20reviewers recommend it
lenient 3/5
medium 9/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
A striking 82, 95% FID improvement and rigorous geometry analysis make this a major advance, though its gradient equivalence framing and persistent feature-learning dynamics invite deeper scrutiny.

Abstract

Recently proposed Drifting Models shift iterative distribution refinement from inference to training, enabling effective one-step generation. However, their performance on complex image datasets depends strongly on the representation used to construct the drifting field: pixel-space drifting performs poorly, whereas pretrained feature spaces substantially improve sample quality for reasons that remain unclear. We trace this gap to the discriminative geometry of the representation, which determines sample weighting in kernel density estimation (KDE) and, consequently drift. We introduce persistent representation learning, which continuously learns a more discriminative representation geometry as the generator evolves across batches. We further establish a current-step gradient equivalence between the KDE ratio loss and drift regression loss under matched conditions, connecting density-ratio-based generator optimization to empirical drifting and motivating direct control of the drifting velocity. Across multiple datasets, our method learns effective discriminative representations directly from pixels and reduces FID by approximately $82-95\%$ over the original pixel-space Drifting Models, without pretrained encoders. Adapting pretrained representations and applying velocity clipping provide further gains.