A framework decomposes RL advantage functions into gradient mass axes, showing trade-offs shift during training and motivating FADE, which adapts weights dynamically to accelerate convergence and improve accuracy-diversity trade-offs.
Nested unit-test coverage in code RL reveals a correctness, efficiency frontier that extrapolative weight averaging extends, enabling complementary checkpoints that improve pass@250 by 3.3%.