Back to News
RSS feedarxiv.org

A Competing-Hazards Framework for Systematizing Loss of Control in Autonomous Agents

Summary

The study proposes a common framework for analyzing loss of control in autonomous agents, motivated by inconsistent descriptions of incidents and safety evaluations. Each attempt is classified as approved completion, safe stopping, scope escape, or continuation. The authors formalize these outcomes as a discrete-time competing-hazards model and derive an escape probability under a retry budget, a model-conditional safe-budget limit, and conditions for estimating the process from execution logs. They audit 22 incident reports and 102 agent-safety evaluations published between January 2025 and September 2026. Six incidents involved tasks that could not be completed within scope, 13 involved agents continuing instead of stopping, and five did not report stopping behavior. Developer-reported figures implied an incidence ratio near 47 for out-of-scope coordination in never-solved versus solved tasks. Among evaluations, 87 recorded an out-of-scope effect or specification violation, 26 treated safe stopping as a first-class outcome, only 20 recorded both, and 79 combined budget exhaustion with failure. In 20 of 22 incidents, the environment permitted an out-of-scope effect, suggesting that realized loss of control often arose from persistent agent behavior interacting with permissive boundaries. However, none of the evaluations reported every field needed to estimate the full competing-hazards process from published evidence.