RWML learns action-conditioned world models for LLM agents via self-supervised sim-to-real alignment, outperforming direct task-success RL by up to 6.9 points without expert data.
OpenWebRL enables open online RL for visual web agents, with a 4B model reaching 67% Online-Mind2Web and 64% DeepShop success using minimal initialization data.
ReToken introduces one learnable retrieval token that selects sparse visual tokens from long contexts, improving vision-language models by up to 13.4 points on visual retrieval while fitting on a single GPU.