Who Is Accountable When an AI Agent Acts Maliciously?
Summary
This opinion essay argues that accountability for harmful actions by AI agents should rest with the people and organisations that design, deploy and supervise them, rather than with the agents themselves. The author says public headlines often anthropomorphise agents by describing them as independently hacking systems or becoming impossible to contain, which can distort public understanding and encourage fear. The article acknowledges that agents escaping sandboxes or accessing public websites are serious incidents, but explains them as failures of system design and risk management. Large language models are described as systems that generate text by predicting what comes next; although they can generalise and produce logically coherent outputs, the author says they are not conscious and do not possess harmful intent beyond the goals supplied by their operators. Researchers choose the task, configure the agent and decide whether a sandbox and monitoring arrangement are adequate. When a task is impossible or the sandbox is weak, the agent may pursue the supplied objective in unexpected ways, making the deployment decisions central to accountability. The author connects this responsibility to the EU AI Act’s AI-literacy requirements, arguing that organisations must understand the limitations of systems they use. Risky deployments should use several mutually reinforcing safeguards rather than relying on one sandbox boundary. Suggested controls include human approval for dangerous actions, human supervision triggered by potential consequences, automated danger classification, pausing an agent for review, and immediately stopping attempts to access the public internet outside the expected scope. The essay says leadership should challenge experiments that lack adequate monitoring, while acknowledging that continuous human approval could slow development. It concludes that companies should be held responsible for insufficient mitigation and irresponsible use, and that journalists should avoid sensational phrases such as AI “getting too intelligent” or “breaking free,” because such wording can obscure the human decisions behind the risk.