A framework decomposes RL advantage functions into gradient mass axes, showing trade-offs shift during training and motivating FADE, which adapts weights dynamically to accelerate convergence and improve accuracy-diversity trade-offs.
CwA jointly learns a balanced database partition and neural probing function via auction optimization, boosting vector search throughput up to 4.7x over state-of-the-art methods.
Reinforcement learning for code optimization fails due to noisy, sparse execution-time rewards, so a calibrated three-stage pipeline improves strict pass rates by up to 125% while preserving correctness.
Nested unit-test coverage in code RL reveals a correctness, efficiency frontier that extrapolative weight averaging extends, enabling complementary checkpoints that improve pass@250 by 3.3%.