Good Papers

Adaptive Latent Capacity for World Models

ALeWM learns adaptive-width JEPA world models via MixSIGReg and prefix sampling, concentrating predictive information in early latent coordinates to improve planning with lower capacity.

Idan Achituve, Lior Dikstein, Idit Diamant, Arnon Netzer, Hai Victor Habi

Published Sep 26, 2026▲ 6 on Hugging FacearXiv ↗

76%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel10/20reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
ALeWM's learned prefix ordering and MixSIGReg's cheap structural prior earn praise as a clever fix for tighter latent planning, though reservations over missing open-source code, untested scaling, and whether lower capacity hides overfitting linger.

Abstract

We introduce Adaptive LeWorldModel (ALeWM), a world model based on a joint-embedding predictive architecture (JEPA) that learns to concentrate predictive information in compact prefixes of a wide latent representation. To encourage this ordering, ALeWM learns a sequence-conditioned distribution over prefix lengths and trains the predictor to estimate the full next embedding from a sampled input prefix. As standard anti-collapse objectives encourage variation across latent coordinates and do not organize them by predictive importance, we also introduce MixSIGReg. MixSIGReg regularizes the masked embeddings against a prior-weighted mixture with Gaussian active prefixes and zeros in the remaining coordinates. As a result, the ALeWM objective encourages early coordinates to retain information useful for prediction and recursive planning. Our analysis shows that the mixture distribution used by MixSIGReg assigns higher variance to earlier coordinate blocks and lower variance to later ones. In addition, we show that, under specified assumptions, prediction error is minimized by placing the information most useful for prediction in earlier blocks. Empirically, we study the behavior of ALeWM in a controlled dynamical system with known state variables and in goal-conditioned visual control. We show that ALeWM consistently achieves higher mean success rates than tuned fixed-width LeWM, with lower planning capacity on average.