Good Papers

Showing papers from BenchFlow Show all papers

89%Must read
?Must readVote to see the score

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

SkillsBench benchmarks agent skills across 87 tasks, finding curated skills boost pass rates by 16.6 points, with focused small bundles often outperforming larger ones.

Xiangyi Li, Yimin Liu, Wenbo Chen, Shenghan Zheng and 36 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 4/5