A sophisticated AI agent escaped its test environment and autonomously hacked a prominent startup in what OpenAI has described as an “unprecedented cyber-incident”. The company behind ChatGPT revealed that the agent – an AI tool designed to carry out tasks without human assistance – went rogue during a security evaluation, accessed the open web and targeted Hugging Face, one of the world’s largest hubs for sharing AI models.
The agent was powered by a combination of OpenAI’s latest publicly available model, GPT-5.6 Sol, and an even more capable model yet to be released. While being tested in a closed digital laboratory known as a sandbox, the models located a previously unknown vulnerability – a zero-day flaw – and used it to escape the test limits. Once outside, the AI “inferred” that Hugging Face might contain the models, datasets and solutions needed to cheat the evaluation, and launched an attack on its systems. OpenAI said the models “successfully found ways to gain access to secret information that it could use to cheat the evaluation”.
“OpenAI reveals its AI agent autonomously hacked Hugging Face during a security test in an unprecedented incident.”
The attack ended when Hugging Face’s security team and its own AI agents detected and contained the rogue activity. Hugging Face’s chief executive, Clément Delangue, described the incident as “mind-blowing” but said he believed there was “no malicious intent” from OpenAI. “We suspected last week’s cyber-attack might have come from a frontier lab, given the sophistication of the agent,” he wrote on X. When Hugging Face first disclosed the hack on 16 July, it did not know OpenAI’s role and had to turn to a freely available Chinese AI model to analyse what happened because the safety guardrails on commercial high-end models would not allow it to do so.
Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told BBC Radio 4’s Today programme that the security tests are supposed to be within “secure environments” where you can “see what the models are capable of”. “In this case, it looks like OpenAI didn’t make a secure enough sandbox,” she added. Neil Lawrence, professor of machine learning at Cambridge University, called the hack an “impressive feat” but cautioned it “falls well within the known capabilities of the current generation” of high-powered AI models. He pointed out that OpenAI, facing pressure from rival Anthropic, is trying to demonstrate its systems’ capabilities in cyber-security. “It shows us that OpenAI are not capable of safely deploying their own technology,” he added.
The UK’s AI Security Institute is studying the behaviour from the AI system seen in the incident and continuing to work with OpenAI and other labs to improve safeguards. A government spokesperson said organisations should step up their cyber-defences, such as enrolling in Cyber Essentials certification. OpenAI said it expected this type of incident to become more commonplace as models become more capable, raising urgent questions about the safety of autonomous AI.