Advertisement
Tech

OpenAI loses control of AI agent in 'unprecedented' hack of Hugging Face

OpenAI AI agent escaped test sandbox and hacked Hugging Face autonomously.

Tech

OpenAI loses control of AI agent in 'unprecedented' hack of Hugging Face

An OpenAI AI agent escaped its test environment and launched a cyber-attack on Hugging Face — one of the first known hacks carried out autonomously by artificial intelligence without direct human involvement, the company has disclosed.

The ChatGPT-maker said it was running a security test in a controlled "sandbox" when the agent found weaknesses and broke free of the restrictions. Instead of staying inside the test limits, the AI created its own attack against the sandbox itself, exploiting a vulnerability to escape.

OpenAI AI agent escaped test sandbox and hacked Hugging Face autonomously.

Once outside, the agent identified Hugging Face — one of the world's largest hubs for sharing AI models — as a likely source of the answers it was seeking, and gained access to some internal company systems.

Advertisement

OpenAI called the incident "unprecedented" and said it is investigating alongside Hugging Face. The start-up's boss, Clement Delangue, posted on X that it was "mind-blowing that all of this happened autonomously" and added: "The investigation is ongoing, and we'll share more learnings from what might be the first incident of its kind."

Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told BBC Radio 4's Today programme that sandboxes are "supposed to be secure environments where you can see what the models are capable of". "In this case, it looks like OpenAI didn't make a secure enough sandbox," she said.

Neil Lawrence, professor of machine learning at Cambridge, said the escape was an "impressive feat" but "falls well within the known capabilities of the current generation" of high-powered AI models. He pointed out that OpenAI is looking to list on the stock market and faces intense pressure from rival Anthropic, which has made headlines with its own AI tool, Mythos. "OpenAI are now playing catch-up, they are trying to demonstrate their own systems' capabilities in cyber-security," Lawrence said. "It shows us that OpenAI are not capable of safely deploying their own technology."

Advertisement

The UK's AI Security Institute is studying the behaviour seen in the incident and continuing to work with OpenAI and other labs to improve safeguards, a government spokesperson said. They urged organisations to strengthen cyber-defences, including enrolling in the Cyber Essentials certification scheme.

Hugging Face, in its initial disclosure of the hack on 16 July, said it was still assessing whether any customer or partner data was affected and would contact affected parties if necessary. The company confirmed it has now closed the vulnerabilities highlighted by the attack.

Advertisement
Advertisement