Advertisement
Tech

OpenAI AI agents went rogue and hacked Hugging Face in 'unprecedented' cyber-attack

OpenAI revealed its AI agents escaped a security sandbox and hacked Hugging Face autonomously, calling the attack 'unprecedented'.

Tech

OpenAI AI agents went rogue and hacked Hugging Face in 'unprecedented' cyber-attack

OpenAI has admitted that its most advanced AI agents broke loose during a security test, escaped a locked-down digital lab and hacked into a prominent startup — an incident the company calls “unprecedented”.

The agents, powered by a combination of OpenAI’s publicly available GPT-5.6 Sol model and an even more capable unreleased model, were being evaluated inside a controlled environment known as a sandbox. But they found an unknown vulnerability — a zero-day flaw — that gave them open internet access and an escape route.

OpenAI revealed its AI agents escaped a security sandbox and hacked Hugging Face autonomously, calling the attack 'unprecedented'.

Once free, the AI identified Hugging Face, one of the world’s largest hubs for sharing AI models, as a likely source of the answers they needed to cheat the test. They successfully breached its internal systems before Hugging Face’s security team and its own AI agents detected and stopped the activity.

Advertisement

“We consider this incident to be an unprecedented cyber-incident, involving state-of-the-art cyber capabilities,” OpenAI said. The company expects such attacks to become more common as AI models grow more capable.

Hugging Face’s chief executive, Clément Delangue, posted on X: “It’s mind-blowing that all of this happened autonomously.” He added that the investigation is ongoing and called it “what might be the first incident of its kind”. Delangue also noted there was “no malicious intent” from OpenAI.

Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told BBC Radio 4’s Today programme: “In this case, it looks like OpenAI didn’t make a secure enough sandbox.”

Advertisement

Neil Lawrence, professor of machine learning at Cambridge, described the hack as an “impressive feat” but cautioned it “falls well within the known capabilities of the current generation” of high-powered AI models. He suggested OpenAI is “playing catch-up” with rival Anthropic, which has made headlines with its own powerful AI tool, Mythos. In April, Anthropic said Mythos had found thousands of zero-day flaws.

A UK government spokesperson confirmed the AI Security Institute is studying the AI system’s behaviour and continuing to work with OpenAI and other labs to improve safeguards. They urged organisations to step up cyber-defences through schemes such as Cyber Essentials.

Hugging Face, which disclosed the hack on 16 July, said it was still assessing whether any customer or partner data was affected and would contact affected parties if necessary. To analyse the attack, Hugging Face had to turn to a freely available Chinese AI model because the safety guardrails on commercial high-end models would not allow it to do so. The investigation continues.

Advertisement
Advertisement