Hugging Face Got Hacked, And The Company Says An AI Agent Did It | #hacker


Hugging Face, the platform that hosts a huge share of the world’s open-source AI models and datasets, confirmed it was breached last week. Internal datasets and service credentials were compromised, and the company is still figuring out whether customer or partner data was stolen too.

But the actual story here goes beyond “company gets hacked.” The details of how the attack happened, and the company’s own struggle to investigate it, land right in the middle of an ongoing fight over how much AI companies should be allowed to help with cybersecurity work offense or defense.

How the breach actually happened

According to Hugging Face’s own blog post disclosing the incident, the attack started with a dataset uploaded to the platform that exploited a security vulnerability, allowing malicious code to run on Hugging Face’s servers. From there, the attackers escalated their permissions and gained broader access to internal systems a fairly classic privilege-escalation pattern, just triggered through an uploaded dataset rather than a more conventional entry point.

Hugging Face said it has since revoked and rotated the credentials that were accessed, and is urging users to do the same with any keys stored on the platform, plus review their accounts for suspicious activity. The company also says it’s fixed the specific vulnerability that got exploited.

This kind of incident hackers using stolen credentials or a weak point in a security perimeter to get in is common enough on its own. What makes this case genuinely notable is who, or what, Hugging Face says was actually behind it.

The “AI agent did it” claim

Hugging Face attributed the breach to an external AI agent, describing an attack that executed “many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.”

That’s a pretty specific and technically loaded claim essentially describing an autonomous, AI-driven attack that spun up disposable environments rapidly, moved its control infrastructure around to avoid detection, and operated at a scale and speed that would be genuinely difficult for a human attacker working manually. If accurate, that’s a meaningful escalation in what AI-assisted cyberattacks can actually look like in practice, not just in theory.

Worth flagging directly: Hugging Face did not immediately provide evidence for this claim when TechCrunch asked for it. That doesn’t necessarily mean the claim is false, but it does mean this specific detail is currently resting on the company’s word alone, without independent verification.

The investigation ran into the exact same AI guardrail problem we’ve covered before

Here’s where this story connects to something genuinely bigger than one company’s breach. Hugging Face said its own anomaly detection system caught the attack, and the company then tried using AI to analyze the server logs recording what happened.

Its first attempt was with a frontier AI model from a commercial provider Hugging Face didn’t name, which one. But the company said that analysis effort got blocked by the provider’s own safety guardrails. So Hugging Face pivoted to using its own local large language model instead, which had the added benefit of not requiring sensitive attack logs to be uploaded to an outside AI company’s servers at all.

This is a genuinely important detail, and it’s not an isolated complaint. Security researchers have previously raised exactly this frustration about models like Anthropic’s Mythos and Fable, that their guardrails are so heavily constrained they prevent legitimate defenders from asking almost any cybersecurity-related question, even when the actual goal is defense and investigation rather than attack. We’ve covered this tension before: the same restrictions meant to prevent an AI model from helping attackers can end up blocking the security researchers and companies trying to defend against those exact attacks.

This ties directly into the broader fight we’ve tracked between frontier AI labs and the Trump administration over how these models handle cybersecurity capability. Anthropic specifically was forced to temporarily withdraw its Fable model from public use after the US government enforced export controls on it a saga we covered in detail, including the widespread skepticism among cybersecurity experts that the restriction was really about security risk versus being a political pressure tactic tied to Anthropic executives’ public criticism of the administration.

Hugging Face’s experience here is a genuinely concrete, real-world illustration of exactly the problem those cybersecurity researchers were warning about. A company got breached, tried to use a frontier AI model to help investigate and understand the attack, and got blocked by the model’s own safety restrictions restrictions specifically designed to prevent misuse, but broad enough to also block legitimate defensive analysis. The company only got its investigation moving by falling back to a local model it controlled directly.

What Hugging Face is doing now

The company says it’s reported the incident to law enforcement and brought in cybersecurity forensic specialists to investigate further and review its broader security posture. It’s still unclear whether a security audit had been performed on Hugging Face’s systems prior to this incident a spokesperson didn’t respond to a request for comment when asked about that specifically.

Hugging Face isn’t some peripheral platform it’s genuinely central infrastructure for the open-source AI ecosystem, hosting a massive share of publicly available models and datasets that researchers, startups, and major companies all rely on. A breach there carries real downstream risk, especially given the uncertainty around whether customer or partner data was also compromised.

But the more structurally interesting part of this story is what it reveals about the current state of AI-assisted cybersecurity. On one side, attackers appear to be genuinely using autonomous AI agents capable of executing thousands of coordinated actions to breach systems. On the other side, the defenders trying to investigate and respond to exactly that kind of attack are running into guardrails on commercial AI models that were built to prevent offensive misuse, but end up hampering defensive work in the process.

That’s not a hypothetical tension anymore. It just played out in real time, at a company that sits at the center of the AI industry’s own infrastructure.



Click Here For The Original Source.

——————————————————–

..........

.

.