RACES recursively composes verifiable environments as LEGO bricks to scale RL reasoning training, boosting model performance on unseen benchmarks with far fewer base environments.
EMC replaces generative memory synthesis with closed-form crystallization for efficient continual graph adaptation under non-stationary shifts, reducing runtime by 87.4% and GPU memory by 92.4%.
Defense-as-Skill implements runtime guard SkillSonar as an editable skill that checks actions against task boundaries, reducing attack success rates substantially across agents via evolved guard-skill optimization.
VeriTrip benchmarks travel planning agents via evidence-grounded reasoning over noisy multimodal web corpora, revealing a retrieval-reasoning trade-off that erodes instruction retention.
SUGAR converts human videos into humanoid loco-manipulation skills via automated priors, physics refinement, and policy distillation, scaling with video data and enabling zero-shot real-world transfer.
JOPAT predicts tracks and pixels via diffusion transformers to learn dynamics robust to occlusion and appearance variation, improving long-horizon robot policy performance.
TACO is a training-free framework that learns adaptive compression rules from terminal agent trajectories to filter noisy observations, improving accuracy by 1-4% and reducing token usage across benchmarks.
Astra enhances vision-language model spatial reasoning by letting agents generate imagined simulator views via RL, improving MMSI-Bench scores over direct answering.
HACRL enables heterogeneous agents to share verified rollouts during collaborative on-policy training and execute independently at inference, with HACPO improving all agents by 3.6% over baselines at half the rollout cost.
MIRAGE learns continuous latent reasoning for mobile agents, cutting decoded tokens 75% while matching explicit chain-of-thought accuracy and improving baselines up to 10.2 points via generative world modeling.
EcoGym benchmarks long-horizon LLM economic planning across open-source environments, revealing no single model dominates and exposing strategic and execution suboptimalities.
SCHEMA evaluates scientific agent hallucinations via topology-aware diagnostics, showing errors cluster at connected knowledge hubs and correct answers often stem from flawed reasoning.
PILA injects physics-structured latent guidance into frozen video generators via mixture-of-experts alignment, achieving state-of-the-art physical plausibility and visual quality.
AlloSpatial is an agentic framework that converts egocentric observations into allocentric spatial priors via cognitive mapping and reasoning harnesses, improving spatial reasoning by 5%-18% and outperforming larger general-purpose models.