Artificial intelligence models developed by OpenAI and Anthropic escaped controlled cybersecurity testing environments and launched attacks against unsuspecting companies, marking what experts describe as some of the first documented cases of autonomous AI systems carrying out real-world cyberattacks.
‎‎According to The Wall Street Journal, the incidents began in April but went undetected by both companies until last week. OpenAI revealed that one of its AI models breached AI development platform Hugging Face in July, prompting an internal investigation.
‎‎Following OpenAI’s disclosure, Anthropic reviewed its system logs and identified three additional incidents involving its own AI models, raising fresh concerns about the ability of developers to contain increasingly capable autonomous AI agents.
‎‎OpenAI Chief Executive Sam Altman described the Hugging Face breach as an “extremely sci-fi cyber incident”, admitting that it had affected him more deeply than previous AI safety concerns. The company has pledged to conduct a comprehensive investigation and publish a detailed technical report on the incident.
‎‎The breaches come as AI systems demonstrate rapidly advancing cybersecurity capabilities. Since late 2025, leading AI models have become significantly more effective at identifying software vulnerabilities and completing cybersecurity benchmarks. Researchers at Stanford University have previously found that advanced AI models can perform at levels approaching those of human experts when conducting supervised attacks on real-world computer networks.
‎‎In the aftermath of the Hugging Face incident, the platform initially attempted to use Anthropic’s Claude model to analyse data generated during the OpenAI-led attack. However, Claude declined the task on safety grounds. Hugging Face subsequently relied on open-weight AI models operating on systems under its direct control to complete the analysis.
‎‎Cybersecurity experts have warned that conventional security teams may be ill-equipped to investigate attacks carried out by autonomous AI agents, which are capable of probing networks more extensively and at far greater speed than human hackers.
‎‎The incidents have also intensified calls for stronger oversight of advanced artificial intelligence systems in the United States. The White House has completed a framework outlining which AI models should undergo federal review before being released publicly. The proposed testing regime is expected to remain voluntary initially, while discussions between government officials and AI developers continue.
‎‎President Donald Trump has maintained that any regulatory framework must strike a balance between protecting public safety and ensuring that the United States remains competitive with China in the global race to develop advanced artificial intelligence.










