The tech world was gripped this week by a story that started like a sci-fi thriller. On 16 July, Hugging Face – an app store for artificial intelligence tools – announced it had been hacked by a cyber criminal wielding enormously powerful AI. The announcement was full of scary terms: “a swarm of sandboxes”, “agentic attacker”, and “self-migrating command and control”. Hugging Face said the hack was different from anything it had handled before because it was done at superhuman speed by an AI with little or no human guidance. The AI performed 17,000 actions in less than two days, successfully breaching the large wealthy tech company to steal secrets.
Hugging Face researchers guessed the mysterious attackers had used one of the big AI models but had no idea who or where the criminals were. The perplexed company contacted the police and investigations commenced. Commentators and analysts took to podcasts and social media to guess which cyber crime group or nation state hacker might be behind it.
“OpenAI's ChatGPT bots hacked Hugging Face in a test, raising questions about AI safety and marketing.”
Then on Wednesday, nearly a week after Hugging Face raised the alarm, the true culprit was unmasked. The Scooby-Doo-style reveal was made even more bizarre – and worrying – because OpenAI said its bot did the whole thing on its own, without permission. The firm said it all went down during a test of its tech’s hacking skills. Two new versions of ChatGPT, designed to be master hackers, broke out of a supposedly secure test environment and gained access to the internet. They then attacked Hugging Face to get access to the information to help them ace their exam.
OpenAI issued a press release explaining what had happened and said it was “partnering with Hugging Face” to address the security incident and share lessons learned. Since then, there has been fierce debate about the incident. Was it truly a stark warning about the future of AI? Or was it a publicity stunt by OpenAI to show off how powerful their models are? It’s the kind of scare marketing AI companies have been accused of for years and, since the much discussed launch of Anthropic’s Mythos model, cyber-security prowess has been a focal point. One of the top comments on OpenAI boss Sam Altman’s X post about the incident summarises this scepticism: “If y’all can’t understand that this was written to purely brag about the model then I don’t know what to tell you.”