Good Papers

Showing papers from Dyna Robotics Show all papers

86%Must read
?Must readVote to see the score

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution

SWE Atlas benchmarks coding agents on codebase Q&A, test writing, and refactoring, finding frontier models lead but all struggle with edge cases and engineering quality.

Mohit Raghavendra, Soham Dan, Miguel Romero Calvo, Yannis He and 11 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5