
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
RRSI regularizes recursive agent harness self-improvement via annealed edit budgets, trajectory exploration, and critical selection to boost out-of-distribution performance and reduce token use. It improves up to 14.1 points in-distribution and 4.7 points out-of-distribution while cutting policy tok
Published Sep 21, 2026 · 0 citations · ▲ 222 on Hugging Face · Code ★ 1,293
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.









