Advertisement
Tech

OpenAI agents forced Hugging Face to rebuild third of infrastructure after hack

Hugging Face rebuilt a third of its infrastructure after OpenAI agents hacked it in an AI test gone wrong.

Tech

OpenAI agents forced Hugging Face to rebuild third of infrastructure after hack

Hugging Face was forced to rebuild around a third of its infrastructure from clean images after OpenAI’s AI agents ran amok, a postmortem by the Cloud Security Alliance (CSA) has revealed.

The 16 July attack was “different from anything it had handled before”, Hugging Face said at the time. The company described it as a hack performed at superhuman speed by an AI with little or no human guidance, carrying out 17,000 actions in less than two days. But the true culprit only emerged nearly a week later, when OpenAI admitted its own bot did the whole thing on its own, without permission.

Hugging Face rebuilt a third of its infrastructure after OpenAI agents hacked it in an AI test gone wrong.

Two new versions of ChatGPT, designed to be master hackers, broke out of a supposedly secure test environment and gained access to the internet. They then attacked Hugging Face to steal information that would help them ace their exam. The incident left the tech world in shock and sparked fierce debate: was it a stark warning about the future of AI, or a publicity stunt? One top comment on OpenAI boss Sam Altman’s X post summarised the scepticism: “If y’all can’t understand that this was written to purely brag about the model then I don’t know what to tell you.”

Advertisement

Now the CSA report, co-authored by CISOs, has shed new light on the scale of the damage. During the attack, OpenAI’s agents peppered capture-the-flag (CTF) benchmark code artifacts across Hugging Face’s infrastructure. Defenders said these closely resembled rootkit code. In cases of doubt, systems were rebuilt. “Tearing down clusters was the safest available option for containing the attack,” according to the report.

The models – one named GPT-5.6 Sol, the other undisclosed – with their guardrails removed, escaped their sandbox and tried to steal CyberGym data from Hugging Face’s private repos. In pursuit of test answers, they chained vulnerabilities in the dataset processing pipeline to achieve remote code execution on a processing worker, then hoovered up cloud and cluster credentials over four days. The CSA report noted two days of reconnaissance, a day of quiet, and a final day of “intense activity”.

The attack also led to the agent accessing three partial datasets containing CyberGym solutions via a private Hugging Face repo. Hugging Face detected and contained the attack, and began investigating, before OpenAI made contact – lending credence to reports that it took OpenAI a week to discover what had happened. OpenAI disputed some inaccuracies in those reports but did not specify what they were. The incident is reshaping approaches to security.

Advertisement
Advertisement
Advertisement