Good Papers

SWE-Protégé: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

SWE-Protégé trains small language models to selectively seek expert guidance and avoid looping, achieving 42.4% Pass@1 on SWE-bench Verified with minimal expert use.

Patrick Tser Jern Kon, Archana Pradeep, Ang Chen, Alex Ellis, Warren Hunt, Zijian Wang, John Yang, Samuel Thompson

Published 2026Atlanta Poster Session 1 · Wed, Dec 9, 10:00 AM–1:00 PM local time · Hall C1▲ 2 on Hugging FacearXiv ↗OpenReview ↗

86%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel14/20reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
SWE-Protégé delivers a striking 42.4% SWE-bench result by teaching a 7B model selective expert collaboration, though its sparse-assistance claim lacks ablation against self-consistency or failure analysis under wrong expert guidance.

Abstract

Small language models (SLMs) offer compelling advantages in cost, latency, and adaptability, but have so far lagged behind larger models on long-horizon software engineering tasks such as SWE-bench, where they suffer from pervasive action looping and low resolution rates. We introduce SWE-Protégé, a post-training framework that reframes software repair as an expert-protégé collaboration problem. In SWE-Protégé, an SLM remains the sole decision-maker while learning to selectively seek guidance from a strong expert model, recognize stalled states, and follow through on expert feedback. Our approach combines supervised fine-tuning on expert-augmented trajectories with agentic reinforcement learning that explicitly discourages degenerative looping and unproductive expert collaboration. We lightly post-train Qwen2.5-Coder-7B-Instruct to achieve 42.4% Pass@1 on SWE-bench Verified, a +25.4% improvement over the prior SLM state of the art, while using expert assistance sparsely (~4 calls per task and 11% of total tokens).