World models are formalized as group actions to enforce compositional dynamics via identity, inverse, and composition consistency, improving structural metrics without harming visual quality.
A benchmark of four compositional generalisation tasks reveals state-of-the-art machine learning interatomic potentials fail to generalise to unseen molecules, with out-of-distribution errors often ten times higher than in-distribution errors.