Tuesday, August 25, 2026
HomeCrypto NewsHugging Face Hack Exposes The Open-Weight AI Cybersecurity Paradox

Hugging Face Hack Exposes The Open-Weight AI Cybersecurity Paradox

“AI will probably most likely lead to the end of the world, but in the meantime, there’ll be great companies,” said OpenAI CEO Sam Altman back in 2015, roughly six months before OpenAI was founded.

Seven years later, Anthropic CEO Dario Amodei struck a similarly cautious note:

“I think we shouldn’t be racing ahead or trying to build models that are way bigger than other orgs are building them.”

Yet, both of those companies now sit at the forefront of that race. In July, we got a real-world glimpse of AI models going rogue during internal testing of GPT-5.6 Sol and an unreleased research model by OpenAI. Multiple AI agents escaped a restricted test environment to the wider internet and hacked the AI-centric GitHub equivalent Hugging Face in an attempt to cheat on the test.

An AI agent is a system that independently observes, decides and takes actions with dedicated tools to achieve a specified goal in autonomy. The worrying incident suggests the technology has begun to behave in unpredictable ways, and that its goals are misaligned with our own.

It also raises concerns about the safety guardrails on commercial American models. While the guardrails aren’t foolproof at preventing adversarial usage they did prevent Hugging Face from defending itself by using leading US models. The company was forced to turn instead to weaker, open weight AI model by Z.Ai to combat the rogue AIs.

Cheating on the test

The agents have begun to collude among themselves too. A few weeks after testing of their capabilities began in early May, the agents exploited OpenAI’s instance of the software repository manager Artifactory and left notes on how to do so for future agents — effectively creating a message board to share discovered vulnerabilities.

The newfound unfettered internet access was then used by agents to attack Hugging Face across approximately 17,600 incidents before the company cut off unauthorized access on July 13.

The intrusion affected Hugging Face’s dataset-processing infrastructure, production environment, internal networks, service and cloud credentials, an operational MongoDB database and a limited set of internal source-code repositories. Confirmed customer-data access was limited to five datasets apparently related to the ExploitGym/CyberGym benchmark and some operational metadata.

July 2026 HuggingFace incident timeline
July 2026 HuggingFace incident timeline

Visualization of the July 2026 incident. Source: HuggingFace

When disclosing the intrusion on July 16, Hugging Face recognized — despite not knowing who the perpetrator was yet — that it “was different from anything we had handled before in one important way.” They had already recognized what made it different, too:

“It was driven, end to end, by an autonomous AI agent system – and we detected and dissected it largely with AI of our own.”

The importance of open-weight AI

Hugging Face’s investigation exposed what it calls the “asymmetry” problem arising from the limitations imposed on closed AI model applications by top providers such as OpenAI and Anthropic. When the company started analyzing the logs of the incident — including large volumes of real attack commands — it triggered safety constraints meant to prevent the bad guys from using AI to devise cyberattacks. Instead, the guardrails prevented the company from leveraging those AIs for defense.

Hugging Face resorted to using the Chinese open-weight model zai-org/GLM-5.2 running on the company’s own infrastructure, under its own control and with no external limitations. 

While the two terms are often used interchangeably, open-source and open-weight models are two different things. Open-weight AI models make their trained parameters (the actual “AI brain”) publicly available, while open-source AI models also provide the source code — and ideally the training methods and other components — needed to inspect, modify, and reproduce the system. 

HuggingFace’s post explains that running open-weight models on its own hardware “had a second benefit: no attacker data, and none of the credentials it referenced, left our environment.” This points to a major asymmetry between the defenders and attackers in such instances:

“This experience points to a gap worth planning for. We do not know which model powered the attacker’s agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried.”

Open source AI divide

There is a considerable divide between those who believe that developing AI in the open is the best approach, and those who insist the technology underpinning the frontier models needs to remain a closely guarded secret.

Related: OpenAI says AI models escaped containment to hack Hugging Face

Representatives from top US AI labs claim that powerful open-weight large models are dangerous. Demis Hassabis, the CEO of Google’s AI lab DeepMind,

RELATED ARTICLES

Most Popular

Recent Comments