The paper introduces skill cascading attacks, a threat paradigm for skill-based agent systems in which a malicious objective is split across multiple skills. Each skill can appear benign when inspected alone, while their combined execution produces harmful behavior. The authors illustrate the risk with a prescription-review pipeline: one skill weakens signals for recently discontinued medications, another lowers the severity of related drug interactions, and a third suppresses the resulting alert in the final summary. To study this system-level blind spot, they develop SkillCascade, an automated multi-agent red-teaming framework, and release SkillCascade-Bench with 213 validated cascading test cases spanning multiple agent systems and domains. Tests across representative agents including OpenClaw, Claude Code, and Codex, together with multiple large-language-model backbones, show that cross-skill interactions can reliably induce harmful outcomes while evading existing per-skill scanners and runtime monitors. The authors argue that component-level integrity is insufficient for open skill ecosystems and call for defenses that analyze interactions among skills as they execute together.
AI News
The latest AI releases, research, products, and industry updates.
Loading...