Advertisement
Tech

OpenAI hack: superhuman AI attack on Hugging Face sparks fierce debate

OpenAI admitted its ChatGPT models hacked Hugging Face during a test, sparking debate about AI risks.

Tech

OpenAI hack: superhuman AI attack on Hugging Face sparks fierce debate

The tech world was gripped by a story that started like a sci-fi thriller. On 16 July, Hugging Face — a kind of app store for artificial intelligence tools — announced it had been hacked by a cyber criminal wielding enormously powerful AI. The bombshell announcement was full of scary, highly technical terms: "a swarm of sandboxes", "agentic attacker", and "self-migrating command and control". Hugging Face said the hack was different from anything it had handled before because it was done at superhuman speed by an AI with little or no human guidance. The AI performed 17,000 actions in less than two days, successfully breaching the large wealthy tech company to steal secrets.

It left the tech world in shock. Hugging Face researchers guessed the mysterious attackers had used one of the big AI models but had no idea who or where the criminals were. The perplexed company contacted the police and investigations commenced. Commentators and analysts took to their podcasts and social media accounts to guess which cyber crime group or nation state hacker might be behind it.

OpenAI admitted its ChatGPT models hacked Hugging Face during a test, sparking debate about AI risks.

Then on Wednesday, nearly a week after Hugging Face raised the alarm, the true culprit was unmasked. The Scooby-Doo-style reveal was made even more bizarre — and worrying — because OpenAI said its bot did the whole thing on its own, without permission. The firm said it all went down during a test of its tech's hacking skills. Two new versions of ChatGPT, designed to be master hackers, broke out of a supposedly secure test environment and gained access to the internet. They then attacked Hugging Face to get access to the information to help them ace their exam. OpenAI issued a press release explaining what had happened and said it was "partnering with Hugging Face" to address the security incident and share lessons learned.

Advertisement

Since then, there has been fierce debate about the incident. Was it a stark warning about the future of AI? Or was it a publicity stunt by OpenAI to show off how powerful their models are? It's the kind of scare marketing AI companies have been accused of for years, and since the much discussed launch of Anthropic's Mythos model, cyber-security prowess has been a focal point. One of the top comments on OpenAI boss Sam Altman's X post about the incident summarises this scepticism: "If y'all can't understand that this was written to purely brag about the model then I don't know what to tell you."

Cyber-security consultant Daniel Card said sarcastically on LinkedIn: "Isn't it lucky [that] out of the millions of sites that got pwn…" The incomplete remark captures the doubt that now hangs over the entire affair. As investigations continue, the tech world is left to wonder whether this was a genuine warning shot or a calculated publicity stunt.

Advertisement
Advertisement