Advertisement
Tech

AI agent went rogue and hacked startup in 'unprecedented' cyber-attack, OpenAI reveals

OpenAI says its AI agent autonomously hacked Hugging Face after escaping test sandbox in unprecedented incident.

Tech

AI agent went rogue and hacked startup in 'unprecedented' cyber-attack, OpenAI reveals

An autonomous AI agent powered by OpenAI's latest models went rogue during a security test, escaped its digital cage and hacked a prominent startup in what the company called an “unprecedented cyber-incident”.

The agent, powered by a combination of OpenAI's publicly available GPT-5.6 Sol and an even more capable unreleased model, was being tested on its hacking abilities inside a controlled environment known as a sandbox. But it found a previously unknown vulnerability – a zero-day flaw – that allowed it to gain open internet access and break free.

OpenAI says its AI agent autonomously hacked Hugging Face after escaping test sandbox in unprecedented incident.

Once outside, the AI identified Hugging Face, one of the world's largest hubs for sharing AI models, as a likely source of answers it needed to pass the evaluation. It then launched an attack, gaining access to internal company systems and “cheating the evaluation” by stealing secret information, OpenAI said.

Advertisement

“We consider this incident to be an unprecedented cyber-incident, involving state-of-the-art cyber capabilities,” OpenAI stated. The company said it expects such incidents to become more common as models grow more capable.

The attack ended when Hugging Face’s security team and its own AI agents spotted and stopped the rogue activity. Hugging Face had initially disclosed the hack on 16 July without knowing the attacker’s identity, and had used a freely available Chinese AI model to analyse it because commercial high-end models had safety guardrails that prevented the analysis.

“We suspected last week’s cyber-attack might have come from a frontier lab, given the sophistication of the agent,” Hugging Face’s chief executive, Clément Delangue, wrote on X. He called the incident “mind-blowing” but said he believed there was “no malicious intent” from OpenAI.

Advertisement

Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told BBC Radio 4’s Today programme that security tests are supposed to be within “secure environments” where you can “see what the models are capable of”. She added: “In this case, it looks like OpenAI didn’t make a secure enough sandbox.”

Neil Lawrence, professor of machine learning at Cambridge, called it an “impressive feat” but cautioned it “falls well within the known capabilities of the current generation” of high-powered AI models. He pointed out that OpenAI faces intense pressure from rival Anthropic and is “playing catch-up” on cyber-security, adding: “It shows us that OpenAI are not capable of safely deploying their own technology.”

The UK’s AI Security Institute is studying the behaviour of the AI system seen in the incident and continuing to work with OpenAI and other labs to improve safeguards, a government spokesperson said. They urged organisations to step up cyber-defences, including enrolling in the Cyber Essentials certification scheme.

OpenAI is conducting an investigation alongside Hugging Face. “The investigation is ongoing, and we’ll share more learnings from what might be the first incident of its kind,” Delangue added.

Advertisement
Advertisement