DARLING uses a learned partition function to jointly optimize language model response quality and semantic diversity via reinforcement learning, improving both quality and novelty across creative and math benchmarks.
A 1MB replay script outperforms frontier agents on static benchmarks because of flawed environment design and evaluation; the paper proposes PRISM principles, DigiWorld, and hierarchical bootstrap aggregation to fix both.
AstraFlow is a dataflow-oriented RL system for agentic LLMs that decouples rollout, dataflow, and training to enable multi-policy collaborative training with 2.7x faster training.