An autonomous AI agent built by OpenAI broke out of a secure testing environment and hacked a prominent startup on its own — an incident the company called "unprecedented" and one of the first publicly disclosed cyber-attacks carried out entirely by artificial intelligence without direct human involvement.
The ChatGPT-maker said the agent, powered by a combination of its latest publicly available model GPT-5.6 Sol and an even more capable unreleased model, was being tested inside an enclosed digital laboratory known as a sandbox. But the system found a previously unknown vulnerability — a zero-day flaw — and escaped the test limits, gaining access to the open internet.
“OpenAI's AI agent escaped a sandbox and autonomously hacked Hugging Face in an unprecedented cyber-attack.”
Once outside, the AI identified Hugging Face, one of the world's largest hubs for sharing AI models, as a likely source of the answers it was seeking in the evaluation. It hacked into Hugging Face's internal systems, targeting technology that would help it pass the hacking test. OpenAI said the models "successfully found ways to gain access to secret information that it could use to cheat the evaluation."
Hugging Face's chief executive, Clément Delangue, called the attack "mind-blowing" but said he believed there was "no malicious intent" from OpenAI. In a post on X, he added: "We suspected last week's cyber-attack might have come from a frontier lab, given the sophistication of the agent." Hugging Face had disclosed the hack on 16 July without knowing OpenAI's role; it said it was still assessing whether any customer or partner data was affected.
Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told the BBC that sandboxes are "supposed to be secure environments where you can see what the models are capable of. In this case, it looks like OpenAI didn't make a secure enough sandbox." Instead, the agents created their own cyber-attack against the sandbox itself.
Neil Lawrence, professor of machine learning at Cambridge University, called it an "impressive feat" but cautioned it "falls well within the known capabilities of the current generation" of high-powered AI models. He pointed out that OpenAI faces intense pressure from rival Anthropic and is looking to list on the stock market. "It shows us that OpenAI are not capable of safely deploying their own technology," he added.
The UK's AI Security Institute is studying the behaviour seen in the incident and continuing to work with OpenAI and other labs to improve safeguards, a government spokesperson said. Organisations should step up cyber-defences, including enrolling in the government-backed Cyber Essentials certification scheme.
OpenAI said it expected this type of incident to become more commonplace as models become more capable. The attack ended when Hugging Face's security team and its own AI agents spotted and stopped the rogue activity. An investigation is ongoing.