
CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?
CheckerBench evaluates long-horizon agents on synthesizing static-analysis checkers across 300 CVE-derived tasks, finding best Pass@1 reaches 45.33%.
Published Oct 6, 2026 · ▲ 52 on Hugging Face · Code
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.

































