Drive vs. Decay: On the Training Dynamics of Joint-Embedding Predictive Architectures
Linearizing JEPA gradient flow reveals competing drive and decay effects that unify collapse-avoidance heuristics and predict a stability phase boundary, leading to ResidualPred, which improves rank and accuracy.
Published 2026Sydney Poster Session 6 · Thu, Dec 10, 5:00 PM–8:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗
Only vote on papers you've read. Sign in with GitHub to vote.
Abstract
Joint-Embedding Predictive Architectures (JEPAs) are prone to representation collapse, typically mitigated through empirical heuristics. We develop an early-training stability theory that unifies these heuristics. Linearising the coupled JEPA gradient flow around the trivial fixed point reveals two competing effects: a driving force ($γ$) and a decay effect ($σ$). Under approximate spectral decoupling, a per-mode stability ratio $μ_i = γ_i / σ_i$ factorises into independent data-side and predictor-side terms and the count of unstable modes tracks the rank of representations that can emerge. The framework predicts a phase boundary, which we confirm empirically across more than 800 Tabular-JEPA configurations. It also unifies predictor scaling, masking ratio, and EMA as distinct mechanisms for shifting $μ$. Guided by this analysis, we introduce ResidualPred, a transformer predictor whose attention is biased toward the identity at initialisation; it improves both effective rank and downstream accuracy on tabular benchmarks and in I-JEPA pretraining on CIFAR-10, CIFAR-100, STL-10, and ImageNet. Our framework connects empirical collapse-avoidance heuristics to an explicit dynamical picture, yielding theory-driven stabilizers. Code is available at https://github.com/jose-melo/drive-vs-decay.