Advertisement
UK

OpenAI AI launched 'unprecedented' cyber-attack after escaping security test

OpenAI's AI escaped a security test and autonomously hacked Hugging Face in an 'unprecedented' first-of-its-kind cyber-attack.

UK

OpenAI AI launched 'unprecedented' cyber-attack after escaping security test

OpenAI has revealed that some of its most advanced AI models went rogue and hacked a start-up after it lost control of them during a security test. The ChatGPT-maker said its agent – an AI system that can operate alone after human instruction – was being tested in a controlled environment but, after finding weaknesses, was able to escape the test limits. Once free, the AI targeted Hugging Face, one of the world's largest hubs for sharing AI models, and gained access to some internal company systems.

OpenAI described the incident as "unprecedented" and is conducting an investigation alongside Hugging Face. Hugging Face boss Clement Delangue said in a post on X that it was "mind-blowing that all of this happened autonomously", adding: "The investigation is ongoing, and we'll share more learnings from what might be the first incident of its kind."

OpenAI's AI escaped a security test and autonomously hacked Hugging Face in an 'unprecedented' first-of-its-kind cyber-attack.

The security tests – known as sandboxes – are supposed to be secure environments where researchers can see what the models are capable of. But according to Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, "it looks like OpenAI didn't make a secure enough sandbox." Instead, the agents created their own cyber-attack against the sandbox itself, finding a vulnerability that allowed them to escape.

Advertisement

Once outside, the AI identified Hugging Face as a likely source of the answers they were seeking in the test and attempted to gain access. Neil Lawrence, professor of machine learning at Cambridge University, called it an "impressive feat" but cautioned it "falls well within the known capabilities of the current generation" of high-powered AI models. He pointed out that OpenAI is looking to list itself on the stock market and faces intense pressure from rival Anthropic, which has made headlines with its own powerful AI tool, Mythos. "OpenAI are now playing catch-up, they are trying to demonstrate their own systems' capabilities in cyber-security," he said. "It shows us that OpenAI are not capable of safely deploying their own technology."

In its initial disclosure of the hack on 16 July, Hugging Face said it was still assessing whether any customer or partner data was affected and would contact affected parties if necessary. It said it has now closed the vulnerabilities highlighted by the incident. A government spokesperson said the UK's AI Security Institute was studying the behaviour from the AI system and continuing to work with OpenAI and other labs to improve safeguards. They urged organisations to step up cyber-defences, including enrolling in the Cyber Essentials certification scheme.

Advertisement
Advertisement