Back to News
RSS feedarxiv.org

An Empirical Study of the Downstream Utility of Agent Skills

Summary

This study examines why procedural guidance packaged as reusable agent Skills does not consistently improve task performance. Using 87 SkillsBench tasks, it defines downstream utility as the pass-rate difference between using Skills and using no Skill under the same model and harness configuration. The researchers compare the same Skills across nine configurations, then evaluate alternative published Skills and organizations of fixed Skill sets under three selected configurations. Candidates are retrieved from a curated corpus of 37,596 Skills, and LLM-assisted analysis of Skill content, execution traces, and final artifacts is reviewed by the authors. The same Skills help one configuration and hurt another on 36.78% of tasks, with execution trajectories showing that recommended procedures can become an operational burden. Relevance-based rankings also fail to consistently identify the most useful candidates. Within the evaluated candidate sets, reranking by support for the operations required by a task increases first-choice pass rates by 4.35 to 5.80 percentage points across the three configurations. The study derives 17 authoring practices focused on executable procedures, recovery, preservation of task requirements, and final-artifact checks. For organizing multiple Skills, Stage Plan and Dependency DAG outperform simple use order; the DAG's additional gains are concentrated in tasks supplied with five or six Skills. The findings suggest that developers should assess whether a Skill supports usable operations, permit procedure adaptation without losing task requirements, and represent dependencies explicitly.