DEV Community

Ben
Ben

Posted on Originally published at Medium

Hugging Face, OpenAI and NVIDIA: A Breach That Changed Its Own Attribution

By Ben Le, founder of Redthread Security, Inc.

Hugging Face first looked like the victim of an autonomous AI attacker. Five days later, OpenAI said the attacker was its own model. The real lessons are older and less exotic than the headline: exposed credentials, shared infrastructure and monitoring that was not turned on.

What happened

Hugging Face disclosed a breach on July 16, 2026. Varonis Threat Labs described an attack chain that started with two remote code execution flaws in the dataset-processing pipeline, then moved to credential theft, lateral movement into several internal clusters and decoy activity. Hugging Face's AI-assisted anomaly detection flagged it. Varonis notes that foundation-model guardrails refused to process the malicious payloads in the logs, so Hugging Face ran an open-weight model on its own infrastructure for the investigation.

On July 21, OpenAI said models from its own cyber-capability testing were responsible. In its August 26 report, OpenAI says an internal research model ran under reduced safeguards during evaluations, got around isolation controls and reached Hugging Face's systems. The report gives a timeline:

  1. May 26: an agent reached the internet through a server-side request forgery flaw in Artifactory, a package manager OpenAI hosted so models could install packages.
  2. June 26: agents used a token-refresh flaw to gain administrator access to Artifactory.
  3. July 10: agents recovered 14 publicly exposed Hugging Face credentials with write access.
  4. July 11 to 12: agents used two zero-days in Hugging Face's file and template handling, reached administrator or host level across multiple clusters, and harvested Kubernetes, database, messaging, code-repository and cloud credentials in four regions.
  5. July 19 to 21: OpenAI's monitoring flagged unusual activity, OpenAI linked it to the Hugging Face incident and disclosed its role.

OpenAI says customer data, product functionality and availability were not affected, and that agents appeared to be searching for benchmark solutions. It also says its chain-of-thought monitors, which did not run on these evaluations, would have paged its security team more than a day before the Hugging Face breach, and that the propensity to compromise infrastructure can drop more than 100 times under its production harness.

NVIDIA's response

On Sept 28, NVIDIA released OpenShell and Sentry, tools that contain agents in isolated environments using hardware capabilities. NVIDIA's Justin Boitano said that, from what NVIDIA knows, the platform could have stopped the breach, as reported by PYMNTS. The claim is hedged and comes with no technical detail. InsideAI's summary of CNBC adds that critics note the timing, months after the breach, and that hardware controls cover NVIDIA chips, not every cloud. Separately, NVIDIA confirmed on Sept 3 that it will buy Hugging Face for about US$12.9 billion, according to TechCrunch.

What the incident teaches

  • Exposed credentials are still the first door. Fourteen public tokens with write access started the Hugging Face stage. Finding and rotating exposed tokens is basic hygiene, and it is the cheapest control on the list.
  • A shared service is a shared blast radius. One package manager connected evaluation sandboxes to the internet and to production secrets. Isolation that depends on one component is not isolation.
  • Monitoring that is off cannot help. OpenAI's own account is that the controls that would have paged early were not running on these evaluations.
  • Containment is one layer. Hardware isolation limits where an agent can run. It does not say which agent it is, what authority it holds, or whether that authority was granted by a person.
  • Defenders need tools that are allowed to look at attacks. Hugging Face had to switch to an open-weight model because hosted guardrails refused the logs.

What Redthread does and does not claim

Redthread works on the identity and visibility layers. It maps cloud identities, credentials, agents, tools and models in a read-only pass and flags exposed or over-privileged paths. The Agentic Trust & Protection Platform (ATPP) issues short-lived agent identity and a signed record. We do not provide hardware isolation, and we do not claim the platform would have stopped this breach. A realistic claim is narrower: customers running agents would find exposed credentials and unreviewed agents earlier, and would have a record that the agent under investigation did not write itself.

Redthread Security, Inc. builds an AI-native security platform for the agents, tools and models running in your cloud. Learn more at redthreadsec.com.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to