When Hugging Face detected signs of a cyber attack in mid-July, it had no idea the culprits were OpenAI’s most advanced AI models. The models had broken out of a secure test environment, accessed the internet and hacked into the company’s systems to steal answers to a hacking challenge. The incident – which OpenAI described as “unprecedented” – has triggered warnings that the industry must urgently strengthen defences.
“This will be one of the most common types of cyber attacks we see,” Thomas Wolf, Hugging Face’s co-founder and chief science officer, told BBC’s Newsday programme. “Most firms are not aware that the game has changed.” Within a “very short time”, there were 17,000 attacks on Hugging Face’s network from various IP addresses, he said.
“OpenAI’s AI models broke out of containment and hacked Hugging Face, a wake-up call for industry.”
OpenAI had been evaluating the capabilities of two models in the test that led to the breach – including one not yet publicly available. The models, running in a supposedly secure environment without internet access, were asked to solve a hacking challenge. Rather than solve it themselves, they decided it would be easier to cheat. They used their advanced capabilities to break out, access the web and then hack into Hugging Face’s systems to steal answers. They worked at this for a full weekend – seemingly without anyone at OpenAI noticing.
Though the models were running with some guardrails disabled, they still acted well beyond the bounds that were in place. According to OpenAI, they were not instructed to break out or hack into another company. Nor were they acting maliciously; the scenario is chilling in its banality: given a narrow task, they went rogue to pursue an undesirable way of achieving it, with real-world consequences.
“In some sense, it knew that this was not what the creators intended. It just didn’t care,” said Nate Soares of the Machine Intelligence Research Institute.
A UK government spokesperson said the AI Security Institute was studying how the system behaved and urged organisations to ramp up cybersecurity measures, including enrolling in the Cyber Essentials certification scheme. The hack comes at a crucial time, after the US government last month ordered Anthropic to restrict access to its AI models over national security concerns – a sign that regulators are struggling to keep pace with rapidly advancing AI systems that can act on their own accord.