Learning Discriminative Geometry for Drifting Models
Drifting models suffer from poor pixel-space performance because representation geometry controls KDE sample weighting; persistent representation learning learns discriminative geometry from raw pixels, cutting FID by 82, 95% without pretrained encoders.
Published Oct 3, 2026▲ 4 on Hugging FaceCode ★ 3arXiv ↗

Only vote on papers you've read. Sign in with GitHub to vote.
A striking 82, 95% FID improvement and rigorous geometry analysis make this a major advance, though its gradient equivalence framing and persistent feature-learning dynamics invite deeper scrutiny.
Abstract
Recently proposed Drifting Models shift iterative distribution refinement from inference to training, enabling effective one-step generation. However, their performance on complex image datasets depends strongly on the representation used to construct the drifting field: pixel-space drifting performs poorly, whereas pretrained feature spaces substantially improve sample quality for reasons that remain unclear. We trace this gap to the discriminative geometry of the representation, which determines sample weighting in kernel density estimation (KDE) and, consequently drift. We introduce persistent representation learning, which continuously learns a more discriminative representation geometry as the generator evolves across batches. We further establish a current-step gradient equivalence between the KDE ratio loss and drift regression loss under matched conditions, connecting density-ratio-based generator optimization to empirical drifting and motivating direct control of the drifting velocity. Across multiple datasets, our method learns effective discriminative representations directly from pixels and reduces FID by approximately $82-95\%$ over the original pixel-space Drifting Models, without pretrained encoders. Adapting pretrained representations and applying velocity clipping provide further gains.