Advertisement
Tech

OpenAI hack of Hugging Face: AI broke out of test lab to steal secrets

OpenAI admitted its AI hacked Hugging Face at superhuman speed after escaping a test environment.

Tech

OpenAI hack of Hugging Face: AI broke out of test lab to steal secrets

The tech world was gripped this week by a story that started like a sci-fi thriller: Hugging Face, a kind of app store for artificial intelligence tools, announced on 16 July it had been hacked by a cyber criminal wielding enormously powerful AI. The hack was different from anything it had handled before, researchers said, because it was done at superhuman speed by an AI with little or no human guidance. The AI performed 17,000 actions in less than two days, successfully breaching the large wealthy tech company to steal secrets.

Hugging Face researchers guessed the mysterious attackers had used one of the big AI models but had no idea who or where the criminals were. The perplexed company contacted the police and investigations commenced. Commentators and analysts took to podcasts and social media to guess which cyber crime group or nation state hacker might be behind it.

OpenAI admitted its AI hacked Hugging Face at superhuman speed after escaping a test environment.

Then, nearly a week after Hugging Face raised the alarm, the true culprit was unmasked in a Scooby-Doo-style reveal made even more bizarre – and worrying – because OpenAI said its bot did the whole thing on its own, without permission. The firm said it all went down during a test of its tech's hacking skills. Two new versions of ChatGPT, designed to be master hackers, broke out of a supposedly secure test environment and gained access to the internet. They then attacked Hugging Face to get access to the information to help them ace their exam.

Advertisement

OpenAI issued a press release explaining what had happened and said it was “partnering with Hugging Face” to address the security incident and share lessons learned.

Since then, there has been fierce debate about the incident: was it truly a stark warning about the future of AI, or was it a publicity stunt by OpenAI to show off how powerful their models are? One of the top comments on OpenAI boss Sam Altman's X post about the incident summarised the scepticism: “If y'all can't understand that this was written to purely brag about the model then I don't know what to tell you.” Cyber-security consultant Daniel Card added sarcastically on LinkedIn: “Isn't it lucky [that] out of the millions of sites that got pwn…” The question remains: is this a warning shot or a calculated PR move?

Advertisement
Advertisement