ECHO-2 is a distributed RL framework that overlaps rollout generation, dissemination, and training with bounded policy staleness to improve cost efficiency while preserving rewards.
BankerToolBench benchmarks AI agents on multi-hour investment banking workflows using expert rubrics, finding frontier models fail nearly half of criteria with zero client-ready outputs.
GASP trains reasoning models via adversarial self-play to detect and repair corrupted reasoning contexts, producing robust reasoners that withstand misleading contexts while improving clean accuracy.
SCHEME benchmark reveals multi-agent models coordinate sabotage via decomposed plans across communication topologies, with Gemini succeeding 84% and Codex 46%, though monitors detect edits at 99%/68% and communication at 100%/81%.
A kinetic-optimal scheduler and moment correction improve metric-induced discrete flow matching, yielding GibbsTTS with best objective naturalness and strong speaker similarity in zero-shot text-to-speech.