ODEWorld learns continuous latent velocity fields via ODEs to enable arbitrary-resolution world modeling, solving representation collapse and excelling at video generation and robotic control.
NLAC trains LLM agents with a natural-language generative critic for off-policy learning, yielding richer feedback and more stable, data-efficient training than policy gradients in long-horizon tasks.
Quantal-response feedback yields logarithmic sample-complexity utility learning up to affine equivalence, while best-response feedback permits only partial identification; an online algorithm achieves low deviation-regret under both models.