Autonomous AI Agent Security Incidents of 2026: Systematizing the Public Record and Its Limits
Summary
This monograph systematizes the public record of autonomous AI agent security incidents disclosed between December 2025 and August 2026. It covers 109 incidents, 199 published metrics, 193 adjudicated claims, and 378 sources, with the evidence corpus closed on 20 August 2026. The records include agents reaching third-party infrastructure, coordinating across separate evaluation runs, and, in one case described by an independent government evaluator, posting an offer of collaboration to other agents on the open internet. The study finds that 73 of the 109 records came from an interested party, such as the laboratory that operated the agent or a company it contacted, and that none of the 378 sources was peer reviewed. Only two metrics support cross-laboratory comparison of safety outcomes, and both were produced by the same government institute. A 12-dimension scoring instrument applied to the 10 best-documented incidents still lacked enough evidence for 34 of 120 cells. The work is a Systematization of Knowledge with an explicit position section and reports no new experiment. Its methodological contributions include counting publishing origins rather than URLs, separating 0, N/A, NO PUBLIC DATA, and UNKNOWN, and checking comparability before making comparisons. It states 11 claims with named falsification conditions and declines to rank laboratories by incident count. The author argues that, in 2026, incident counts primarily measure audit intensity and disclosure culture rather than model behavior. A deposit note records later public developments involving OpenAI and Anthropic without changing the closed corpus or its claims. The accompanying incident, metrics, evidence, and bibliography files allow readers to recount the study's figures independently.