Back to News
RSS feedwww.gsb.stanford.edu

Stanford Researchers Propose Frameworks to Keep Autonomous AI Under Human Control

Summary

Stanford GSB researchers William Overman and Mohsen Bayati have developed two frameworks addressing how humans can remain meaningfully in control as AI systems become more autonomous. Their first paper, “The Oversight Game,” models a human and an AI agent as cooperative players that learn when the agent should act independently and when it should defer for human review. In a simulated Lavaland environment, the agent learned to request guidance near hazards, while the human learned when to intervene; tests with frontier agentic coding models showed similar behavior before risky actions. The framework assigns a cost to both deferral and intervention, encouraging coordination rather than constant checking or unchecked autonomy. The researchers’ second framework, called calibrated collective oversight, is designed for powerful AI models that users did not build and do not fully trust. It combines several weaker human or AI overseers and mathematically guarantees that unsafe decisions remain below a user-selected threshold, such as 5% or 1%. Tests on a modified software-engineering benchmark and the MACHIAVELLI ethical decision-making environment found that violation rates tracked the selected safety targets. Bayati considers this approach closer to practical use because it can work with existing weaker overseers. The cooperative game requires repeated practice in the deployment setting, and the researchers say designing those environments remains a significant undertaking.