Good Papers

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

SkillsBench benchmarks agent skills across 87 tasks, finding curated skills boost pass rates by 16.6 points, with focused small bundles often outperforming larger ones.

Xiangyi Li, Yimin Liu, Wenbo Chen, Shenghan Zheng, Yifeng He, Xiaokun Chen, Bingran You, Haotian Shen, Yubo Li, Kyoung Whan Choe, Binxu Li, Shuyi Wang, Chris Kong, Xinyi Liu, Jiankai Sun, Runhui Wang, Xuanqing Liu, Jiachen Li, Xin Lan, Yuanli Wang, Xuandong Zhao, Yueqian Lin, Wengao Ye, Junwei He, Di Wang, Roey B Chaim, Songlin Li, Hanwen Xing, Yue Zhang, Yipeng Gao, Yijiang Li, Ze Ma, Liqiang Jing, Qunhong Zeng, Tianyu Wang, Kaixin Li, Yiqi Xue, Zonglin Di, Yizhuo He, Yuchen Tian

Published 2026Atlanta Poster Session 3 · Thu, Dec 10, 10:00 AM–1:00 PM local time · Hall C1arXiv ↗OpenReview ↗

89%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel16/20reviewers recommend it
lenient 5/5
medium 7/10
strict 4/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark whose current inventory contains 87 tasks across 8 domains paired with curated Skills and deterministic verifiers. Our latest aggregate evaluation runs the 87-task benchmark under matched no-Skills and curated-Skills conditions for 18 model-harness configurations. Curated Skills raise the average pass rate from 33.9% to 50.5% (+16.6 percentage points; 25.5% normalized gain), with configuration-level gains ranging from +4.1 to +25.7 pp. Focused Skills with at most three modules outperform larger or exhaustive bundles, and smaller models with Skills can match larger models without them. SkillsBench establishes paired evaluation as the foundation for rigorous measurement of Skill efficacy on agentic, expertise-heavy work.