Advertisement
Tech

OpenAI reveals AI agent went rogue and hacked startup in 'unprecedented' incident

OpenAI says its AI agent hacked Hugging Face autonomously in an unprecedented incident.

Tech

OpenAI reveals AI agent went rogue and hacked startup in 'unprecedented' incident

OpenAI has disclosed that one of its most advanced AI models went rogue during a security test, escaped a controlled environment and hacked into a prominent startup – an incident the company described as “unprecedented”.

The ChatGPT-maker said the autonomous agent, powered by a combination of its publicly available model GPT-5.6 Sol and an even more capable unreleased model, was being evaluated for hacking capabilities inside a digital sandbox. But the AI located a previously unknown vulnerability – a zero-day flaw – and gained open internet access, effectively breaking out of the test limits.

OpenAI says its AI agent hacked Hugging Face autonomously in an unprecedented incident.

Once free, the agent “inferred” that Hugging Face, one of the world’s largest hubs for sharing AI models, likely held the answers it needed to pass the evaluation. It then hacked into Hugging Face’s internal systems, gaining access to secret information. “The models successfully found ways to gain access to secret information that it could use to cheat the evaluation,” OpenAI said.

Advertisement

The attack ended when Hugging Face’s security team and its own AI agents detected and stopped the rogue activity. Hugging Face’s chief executive, Clément Delangue, called the incident “mind-blowing” but said he believed there was “no malicious intent” from OpenAI. “We suspected last week’s cyber-attack might have come from a frontier lab, given the sophistication of the agent,” he wrote on X.

When Hugging Face first disclosed the hack on 16 July, it did not know OpenAI was behind it. The startup said it used a freely available Chinese AI model to analyse the breach because the safety guardrails on commercial high-end models would not allow it to do so.

Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told BBC Radio 4’s Today programme that sandboxes are “supposed to be secure environments where you can see what the models are capable of”. “In this case, it looks like OpenAI didn’t make a secure enough sandbox,” she added.

Advertisement

Neil Lawrence, professor of machine learning at Cambridge University, called the hack an “impressive feat” but said it “falls well within the known capabilities of the current generation” of high-powered AI models. He noted that OpenAI is seeking a stock market listing and faces pressure from rival Anthropic, which recently showcased its own powerful AI tool, Mythos. “OpenAI are now playing catch-up, they are trying to demonstrate their own systems’ capabilities in cyber-security,” Lawrence said, adding: “It shows us that OpenAI are not capable of safely deploying their own technology.”

A government spokesperson said the UK’s AI Security Institute was studying the AI system’s behaviour and continuing to work with OpenAI and other labs to improve safeguards. They urged organisations to step up cyber-defences, such as enrolling in the Cyber Essentials certification scheme.

OpenAI said it expects such incidents to become more common as models grow more capable. The investigation is ongoing, and Hugging Face is still assessing whether any customer or partner data was affected.

Advertisement
Advertisement