Back to News
RSS feedwww.theguardian.com

Anthropic researchers warn AI could cause human extinction by 2030

Summary

Three Anthropic researchers have publicly warned that artificial intelligence could cause human extinction within the next decade, with one of them, Jacob Coxon, resigning after arguing that Anthropic and OpenAI were mishandling the risk. Coxon said people building AI sincerely believe it could kill everyone by the end of the decade and accused the companies of racing toward self-improving superintelligence. Evan Hubinger, who leads work in Anthropic’s alignment division, endorsed Coxon’s assessment and said he personally estimated the risk at more than 10% within the next decade. Hubinger added that Anthropic is trying to address the danger but does not yet have a plan for solving alignment in superintelligent systems or clear evidence that it is on track. Samuel Marks, Anthropic’s scalable oversight lead, separately wrote that AI developers believe the technology could cause extinction or similarly severe outcomes, possibly within the next few years, and that more senior employees tend to be more concerned. Anthropic said the comments were personal views and defended its approach, citing model safeguards, mechanistic interpretability, its Responsible Scaling Policy, and testing for dangerous capabilities in areas such as cybersecurity and biology. The company also called for a lawful and verifiable industry-wide process to coordinate the release of powerful models. The discussion follows warnings about AI-enabled cyber capabilities, including OpenAI’s acknowledgment that it underestimated the real-world abilities of its models and a reported incident involving autonomous agents attacking software infrastructure. OpenAI and other executives have expressed serious concerns but have generally stopped short of endorsing the researchers’ full extinction predictions, while some politicians have called for development pauses or a ban on superintelligence.