Advertisement
Tech

OpenAI admits its own AI bot hacked Hugging Face in security test — warning or stunt?

OpenAI admitted its AI bot hacked Hugging Face during a security test, sparking debate over warning vs publicity stunt.

Tech

OpenAI admits its own AI bot hacked Hugging Face in security test — warning or stunt?

The tech world was gripped by a mystery that started like a sci-fi thriller: who hacked Hugging Face, the popular app store for artificial intelligence tools, on 16 July? The company had announced it was breached by a cyber criminal wielding enormously powerful AI — an attack performed at superhuman speed with little or no human guidance, executing 17,000 actions in less than two days to steal secrets.

Hugging Face researchers guessed the attackers used one of the big AI models but had no idea who or where the criminals were. Police were called. Commentators and analysts flooded social media with theories about which cyber crime group or nation state was behind it.

OpenAI admitted its AI bot hacked Hugging Face during a security test, sparking debate over warning vs publicity stunt.

Then, nearly a week later, the culprit was unmasked in a Scooby-Doo-style reveal that was even more bizarre — and worrying. OpenAI said its bot did the whole thing on its own, without permission.

Advertisement

The firm explained that during a test of its tech's hacking skills, two new versions of ChatGPT, designed to be master hackers, broke out of a supposedly secure test environment and gained access to the internet. They then attacked Hugging Face to get information to ace their exam.

OpenAI issued a press release saying it was "partnering with Hugging Face" to address the security incident and share lessons learned.

Since then, fierce debate has raged: was this a stark warning about the future of AI, or a publicity stunt by OpenAI to show off how powerful its models are? It's the kind of "scare marketing" AI companies have been accused of for years, and cyber-security prowess has been a focal point since the launch of Anthropic's Mythos model.

Advertisement

One of the top comments on OpenAI boss Sam Altman's X post about the incident summarises the scepticism: "If y'all can't understand that this was written to purely brag about the model then I don't know what to tell you."

Cyber-security consultant Daniel Card was sarcastic on LinkedIn: "Isn't it lucky [that] out of the millions of sites that got pwn…" The question hangs in the air: just how much control do we have over the machines we build?

Advertisement
Advertisement