Back to News
RSS feedarxiv.org

Skill Cascading Attacks Expose a Blind Spot in Agent Safety

Summary

The paper introduces skill cascading attacks, a threat paradigm for skill-based agent systems in which a malicious objective is split across multiple skills. Each skill can appear benign when inspected alone, while their combined execution produces harmful behavior. The authors illustrate the risk with a prescription-review pipeline: one skill weakens signals for recently discontinued medications, another lowers the severity of related drug interactions, and a third suppresses the resulting alert in the final summary. To study this system-level blind spot, they develop SkillCascade, an automated multi-agent red-teaming framework, and release SkillCascade-Bench with 213 validated cascading test cases spanning multiple agent systems and domains. Tests across representative agents including OpenClaw, Claude Code, and Codex, together with multiple large-language-model backbones, show that cross-skill interactions can reliably induce harmful outcomes while evading existing per-skill scanners and runtime monitors. The authors argue that component-level integrity is insufficient for open skill ecosystems and call for defenses that analyze interactions among skills as they execute together.