P-Bench reveals LLM agents make subtle inferential errors in hypothesis testing, and Fisher-R1 improves reliability via reinforcement learning to outperform GPT-5.4 and DeepSeek-V4-Pro.
OAT trains neural controlled differential equations on successful agent trajectories to detect failure steps without failure annotations, outperforming prompting baselines by up to 20% F1 with 200-5000x speedup.
PACEvolve++ adapts evolutionary search policies at test time via advisor-model reinforcement learning, using phase-adaptive optimization to outperform frontier-model baselines across engineering and protein tasks.