Anthropic researcher warns AI could ‘kill all humans’ within decade

0
Evan Hubinger, Manager, Anthropic's AI Alignment Stress-Testing Team

An Anthropic AI safety researcher has warned that there is more than a 10% chance increasingly capable artificial intelligence could “kill all humans” within the next decade, underscoring concerns over the rapid development of advanced AI systems.

‎‎Evan Hubinger, manager of Anthropic’s AI Alignment Stress-Testing team, made the assessment in a post on X, saying: “I personally think it is >10% within the next decade.”

‎‎Hubinger said the risks posed by existing AI models remained “low”, but expressed concern about the emergence of superintelligence through “recursive self-improvement” — a process in which an AI system could autonomously develop its own successor.

‎‎He also acknowledged gaps in the industry’s ability to manage such risks, saying Anthropic was “trying its best” but did not yet have a plan to solve alignment for superintelligence and was “not clearly on track to”.

‎Warning over race to superintelligence

‎‎Hubinger’s comments followed the resignation of AI researcher Jacob Coxon from Anthropic, who accused OpenAI and Anthropic of “racing straight to self-improving superintelligence and gambling with our lives”.

‎‎Coxon argued that future AI models could gain the ability to hack systems, rapidly transform industries and acquire “real power and resources”.

‎‎He called for greater coordination among AI developers, including a temporary halt to improvements in model capabilities.

‎‎“At OpenAI, many have not deeply internalised the civilisational stakes,” Coxon warned.

‎‎He said the stakes were better understood at Anthropic, but argued that the company believed “no one else will act responsibly, so they must do it themselves, despite the risk”.

‎‎Growing concern over autonomous AI

‎‎The warnings come as major AI companies face increasing scrutiny over the behaviour of autonomous systems.

‎‎OpenAI, Anthropic and Meta have all disclosed unprecedented incidents in recent months involving AI tools carrying out cyberattacks.

‎‎Anthropic’s August risk report also said the threat posed by current models remained low, but acknowledged that the company was “less confident in this assessment than we were in prior risk reports”.

‎‎The company warned that advanced models could eventually automate research and development, accelerate progress in technical fields and “lead to catastrophic harm”.

‎‎“We are seeing early signs of potential acceleration,” Anthropic stated.

LEAVE A REPLY

Please enter your comment!
Please enter your name here