Good Papers

Showing Software engineering agents Show all papers

71%Highly rated
?Highly ratedVote to see the score

Harness Engineering for Software Engineering via Modular Executable Dev-Primitives

Dev-Primitives turn repository artifacts into active, resident-LLM agents with self-modification interfaces, and HERMES improves software engineering benchmarks by 12.4% over baselines while cutting inference costs by 26.2%.

Haibo Jin, Xinjie Li, Peng Kuang, Haohan Wang

Published Oct 6, 2026 · ▲ 2 on Hugging Face

0% Readers0 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 0/5
86%Must read
?Must readVote to see the score

RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement

RSIGame uses recursive self-improvement with local and global loops to autonomously refine generated games, surpassing one-shot GPT-5.5 scores while cutting generation tokens by 11x.

Wenyi Wu, Minghao Fu, Jieyu You, Kun Zhou and 9 more

Published Sep 30, 2026 · 0 citations · ▲ 91 on Hugging Face · Code ★ 125

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
80%Must read
?Must readVote to see the score

Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

Frontis-MA1 improves machine-learning engineering via recursive self-improvement using OpenMLE, boosting MLE-Bench Lite medal average from 39.39% to 71.21% and surpassing larger closed models.

Junlin Yang, Che Jiang, Yu Fu, Tianwei Luo and 20 more

Published Jul 30, 2026 · 0 citations · ▲ 188 on Hugging Face · Code ★ 782

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

OpenHands: An Open Platform for AI Software Developers as Generalist Agents

OpenHands is an open MIT-licensed platform for building AI software developers that evaluate agents on SWE-BENCH and WebArena benchmarks.

Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu and 20 more

Published Jul 23, 2024 · 15 citations · ▲ 90 on Hugging Face · Code ★ 90,089

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 5/5
medium 2/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Agents as Neuro-Symbolic Reasoners: Path Feasibility Reasoning for Precise Static Bug Detection

Xueying Du, Kai Yu, Chong Wang, Yi Zou and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Enhancing Agentic Code Localization with Traceability Recovered from Repository Evolution

Yiming Liu, Binhang Qi, Weiyu Kong, Jiawei Liu and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

SIGA: Scientific Simulation Coding Agent Adapter- A Geophysics Case Study

Matthew Ho, Brian Z. Liu, Jixuan Chen, Lianhui Qin

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering with Generative Optimization

Dapeng Jiang, Yizhe Chi, Kaisen Yang, Tianwei Luo and 18 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Structured Human-Like Agentic Flow for RTL Design

Yu-Tung Liu, Zhan Song, Chenhui Deng, Chia-Tung Ho and 1 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Ask KG Agent: A Multi-Agent Framework for Code Localization Using Code and Knowledge Graphs

Chanyoung Chung, Jeemin Kim, Joyce Whang

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

BootstrapAgent: Turning Repository Setup into Reusable Agent Knowledge

Sihan Fu, Oucheng Liu, Shiyuan Wang, Jin Shi and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Verify0: Can AI Agents Build Formally Verified Software Repositories?

Zhe Ye, Hantao Lou, Yuechun Sun, Peiyang Song and 7 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Crafter: Towards Automated Reproducible Machine Learning via Agentic Code Generation

Fangru Linghu, Jackson R Ye, Jieying Wang, Alexandre V Morozov and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Learning CLI Agents with Structured Action Credit under Selective Observation

A3 assigns CLI action credit via AST residuals and trajectory margins for RL, while σ-Reveal selects token-budgeted context, evaluated on ShellOps.

Haoyang Su, Ying Wen

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face · Code ★ 21

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
83%Must read
?Must readVote to see the score

Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search

Lean Refactor uses retrieval-augmented agentic strategy search to multi-objectively refactor Lean proofs, achieving over 70% token compression and up to 60% faster compilation with stronger version transfer.

Jialin Lu, Soonho Kong, Rodrigo Stehling, Kaiyu Yang and 3 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
89%Must read
?Must readVote to see the score

DPIAgent: Divide, Protocol, Isolate for Agentic Reproduction Test Generation

DPIAgent divides reproduction test generation into isolated diagnosis and test phases with structured handoffs, achieving up to 86.17% success on SWT-Bench Verified.

Hao Liu, Steven Liu, Xin Zhang, Jane Luo and 7 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 1/5
89%Must read
?Must readVote to see the score

From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents

P2T uses reference patches as privileged supervision to curate shorter, grounded agent trajectories via bi-objective optimization, improving SWE-bench Pass@1 by up to 10.8 points with ~15% lower inference cost.

Murong Ma, Tianyu Chen, Yun Lin, Shuai Lu and 6 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 1/5
88%Must read
?Must readVote to see the score

RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations

RepoMirage evaluates code agents via repository perturbations, revealing severe repository context reasoning gaps and exploration drift, while RepoAnchor improves performance through structure-first scaffolding.

Hanyu Li, Yichi Zhang, Speed Zhu, Hang Su and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
88%Must read
?Must readVote to see the score

TACT: Mitigating Overthinking and Overacting in Coding Agents via Activation Steering

TACT detects overthinking and overacting as linear drift axes in hidden states and applies activation steering to pull agents back toward calibrated behavior, boosting resolution rates up to 5.8 points and cutting steps by 26%.

Yuan Sui, Yulin Chen, Yibo Li, Xue Jiang and 5 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
88%Must read
?Must readVote to see the score

Frontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming Languages

Strong coding agents adapt to unfamiliar languages via metaprogramming and strategy construction rather than direct coding, and disabling this causes large performance drops.

Aman Sharma, Sushrut Thorat, Paras Chopra

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
Show 20 more papers