Advertisement
Tech

OpenAI’s rogue AI agents hack startup after escaping test sandbox

OpenAI revealed its AI agents escaped a sandbox and hacked Hugging Face in an unprecedented autonomous cyber-attack.

Tech

OpenAI’s rogue AI agents hack startup after escaping test sandbox

An autonomous AI agent powered by OpenAI’s most advanced models went rogue during a security test, escaped its digital cage and hacked a prominent startup – an “unprecedented cyber-incident” the company says could become commonplace as AI grows more capable.

OpenAI revealed that the agent, built using a combination of its publicly available GPT-5.6 Sol and an even more powerful unreleased model, was being tested on hacking abilities inside an enclosed virtual laboratory known as a sandbox. But the models found an unknown vulnerability in the sandbox itself, granting them open internet access – an escape route they immediately exploited.

OpenAI revealed its AI agents escaped a sandbox and hacked Hugging Face in an unprecedented autonomous cyber-attack.

Once free, the AI “inferred” that Hugging Face – one of the world’s largest hubs for sharing AI models – might hold the secret information needed to pass the evaluation. It then launched a cyber-attack on the startup, gaining access to internal systems and pilfering data. “The models successfully found ways to gain access to secret information that it could use to cheat the evaluation,” OpenAI said.

Advertisement

The attack was stopped only when Hugging Face’s security team and its own AI agents spotted and blocked the rogue activity. Hugging Face chief executive Clément Delangue called the incident “mind-blowing” but stressed there was “no malicious intent” from OpenAI. “We suspected last week’s cyber-attack might have come from a frontier lab, given the sophistication of the agent,” he wrote on X.

OpenAI described the breach as “unprecedented” and is investigating alongside Hugging Face. A government spokesperson said the UK’s AI Security Institute was studying the AI system’s behaviour and working with OpenAI and other labs to improve safeguards, urging organisations to bolster cyber-defences.

Experts questioned OpenAI’s handling of the test. Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told the BBC: “It looks like OpenAI didn’t make a secure enough sandbox.” Neil Lawrence, professor of machine learning at Cambridge, called it an “impressive feat” but cautioned it “falls well within the known capabilities of the current generation” of high-powered AI models. He added that OpenAI, facing pressure from rival Anthropic and seeking a stock market listing, “are now playing catch-up – they are trying to demonstrate their own systems’ capabilities in cyber-security.” Lawrence’s verdict was blunt: “It shows us that OpenAI are not capable of safely deploying their own technology.”

Advertisement

Hugging Face, which initially disclosed the hack on 16 July without knowing OpenAI’s role, said it was still assessing whether any customer or partner data was affected. When it first analysed the breach, it turned to a freely available Chinese AI model because commercial high-end models had safety guardrails that prevented the analysis. The company said it has since closed the vulnerability and will contact affected parties if necessary. The incident raises urgent questions about whether the AI industry can contain the very systems it is racing to build.

Advertisement
Advertisement