Value-based agents trained on diverse reward functions implicitly encode world models, extractable via P-learning, with sufficient conditions for exact dynamics recovery and cross-goal generalization.
Generative models learn rules at τ_rule and memorize at τ_mem, defining an innovation window that widens with dataset size but narrows with rule complexity across diffusion and autoregressive architectures.