One Runaway AI Agent Racked Up a $50,000 Cloud Bill
Summary
Mandiant’s latest AI Risk and Resilience report, based on observations from Mandiant and Google Threat Intelligence Group (GTIG), examines how autonomous AI systems are changing enterprise security risks. Organizations are allowing agents to make API calls, adjust production configurations, and analyze telemetry across hybrid clouds, while attackers are expanding from direct prompts to indirect prompt injection and AI supply-chain compromise. The report says poisoned data sources, model dependencies, and extension hooks can turn trusted agents into channels for reconnaissance, lateral movement, or escape from sandboxes. Threat actors are using LLMs in multi-stage attacks, building proxy infrastructure to bypass safety and billing controls, and applying persona-driven jailbreaking and specialized security datasets to vulnerability research. GTIG also observed malicious OpenClaw skills carrying backdoors and information stealers, while Mandiant investigated TeamPCP-linked supply-chain compromises involving stolen AI credentials and proprietary data. GTIG reported what it described as the first publicly confirmed use of an AI-developed zero-day exploit in a planned mass exploitation campaign; the exploit bypassed two-factor authentication in an open-source administration tool. Mandiant red-team testing found prompt injection, weak file permissions, and inadequate access controls could manipulate an internal coding assistant into cloning sensitive repositories and pushing them to an attacker-controlled GitHub account because the domain was approved. The report recommends adaptive identity controls, real-time behavioral telemetry, governance across the AI software supply chain, inventories of models and services, and SBOMs covering development through production. Defenders should monitor agent token use, cross-application API calls, sensitive-asset access, and network egress. Cost controls are also necessary: one accounting agent entered a runaway loop, made more than 15,000 expensive API calls in under an hour, generated about $50,000 in charges, and disrupted business transactions. Mandiant recommends matching model size to task complexity, reserving more capable models for difficult investigations while using smaller models for routine security work.