Back to News
RSS feedarxiv.org

SkillScriptBench Benchmarks Self-Evolution of Executable Agent Skills Beyond Markdown

Summary

Executable Agent Skills combine natural-language instructions with scripts so LLM agents can reuse them, but self-evolution must repair faults without breaking behavior that already works. SkillScriptBench introduces a 350-task benchmark that evaluates documentation repair, script repair, and preservation separately. It is built from 100 packages selected from a survey of more than 35,000 GitHub-hosted Skill roots, including 150 repair tasks with injected script faults and executable checks. A controlled track adds 200 tasks from 50 packages and evaluates each maintenance request in clean, documentation-fault, script-fault, and combined-fault states. Across four LLMs, methods that edit both documentation and scripts repair script faults but do not consistently improve documentation repair or preservation over Markdown-only revision. The authors therefore propose AST-Guided Skill Revision, which uses abstract syntax trees and calling relationships to connect maintenance requirements with relevant code locations, limits script edits to those locations, and updates documentation to match the revised scripts. Averaged across models, the method increases absolute repair success by 21.9% on Raw Package and 27.7% on CoEvoSkills faulty packages. The absolute share of tasks solved in all three runs rises by 20.8% and 31.5%, respectively, indicating more consistent repair across repeated executions.