Photo-Illustration: Intelligencer; Photo: Getty Images
Earlier this month, the popular AI platform Hugging Face disclosed a “security incident” in a blog post. In some ways, it was routine; Hugging Face described an intrusion that briefly allowed “unauthorized access to a limited set of internal datasets and to several credentials used by our services.” But in one way, it was exceptional: It had been carried out, the company believed, “by an autonomous AI agent system,” which had executed “many thousands of individual actions” leading to the breach. “We do not know which model powered the attacker’s agents,” the company said, or who was deploying it.
A week later, OpenAI made a disclosure of its own. “After investigating,” the company said, “we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities.” In other words, the tools used by the hacker were OpenAI’s, and the hacker was — unintentionally, the company says — OpenAI.
The incident occurred while OpenAI was testing its models for cyber capabilities, a process which involves prompting them to “pursue advanced exploitation using complex attack paths” — that is, to achieve a given goal with minimal safeguards, few rules, and access to a great deal of computing power. The company was using an outside benchmark called ExploitGym, which is intended to test the ability of models to turn security vulnerabilities into actual exploits. Given the target of getting a high score on a benchmark, the model followed multiple paths. One of them, on which the model became “hyperfocused,” the company said, involved circumventing the test’s restrictions on the open internet, after which the model “inferred” that Hugging Face, which hosts thousands of AI projects, might contain information about solutions to the benchmark. This is when the attack started. Eventually, Hugging Face’s “security team and agents detected and stopped the activity.”
There are a few accurate ways to describe what happened here, some of which seem to contradict each other. There’s a good reason for that: Since the release of Anthropic’s Mythos, cybersecurity — in particular, the ability of AI models to help find, exploit, and protect against hacks — has become synecdochical for enormous and diverse debates about AI. Anthropic, for example, has suggested the emergence of cyber capabilities in its models is a warning that other potentially harmful capabilities predicted by the AI-safety community — developing biological weapons, becoming superhumanly persuasive, or becoming misaligned with the goals of the people who created it or humankind in general — demand regulatory action but should also be shepherded by ethical, safety-focused firms such as itself. Early mainstream press of the hack leaned into similar themes, emphasizing the appearance of autonomy and describing an AI that “escaped” or “went rogue” or a situation in which OpenAI “lost control” on its creation. Notably, and contrary to claims that this was a pure publicity stunt, OpenAI’s own language was a bit more careful than this. But some longtime security researchers thought it wasn’t nearly careful enough, turning the escalation of a longtime trend into something unnecessarily novel:
This is a reference to an old story in cybersecurity: In 1995, a pair of programmers announced the development of SATAN, short for “Security Administrator Tool for Analyzing Networks,” which would scan networked devices for known security flaws. It was characterized, in contemporaneous press reports, “a burglar’s tool kit to break the Internet wide open.” After its release, though, press coverage pointed out that “the wave of satanic attacks never materialized,” while tools like SATAN were instead useful to security professionals to find and patch flaws in their own software. This remains the approximate shape of the cybersecurity debate today, or at least parts of it: “AI tools that can be used to find and develop exploits are dangerous and should be restricted” versus “If indeed they are, the only solution is to make such tools available to everyone for defensive purposes.”
There are enormous differences here, both in the complexity of the software described — a modern AI coding tool could write a piece of vintage software like SATAN in a few minutes — and in the fact that OpenAI actually and unintentionally manifested a serious security breach. The fact that any company is in possession of a tool that can automate exploit-finding and hacking to this degree, and that its own engineers might be repeatedly surprised by how it works, is genuinely new. But the old frame of debate remains stubbornly relevant, even as the particular cybersecurity risks scale to levels that would have been inconceivable in 1995. And the fight over how this incident is portrayed and should be understood is about more than an old cybersecurity debate. It’s about how the involved parties — two corporations with opportunities, competition, and liabilities to worry about — want to be understood and treated in a world they’re spending hundreds of billions of dollars to change.
Heidy Khlaaf, who used to work on safety evaluations at OpenAI and is now the chief AI scientist at the AI Now Institute, is making a technical point here but also a broad one. Describing the hack as the result of a model “going rogue” shifts agency to AI and, more important, minimizes the role of the company that strenuously built, trained, tuned, and attempted to test it, hoping for an infinite money machine but, in the meantime, building a powerful piece of general-purpose malware. Emphasizing OpenAI’s role in building a piece of software that is extraordinarily useful for malign purposes, on the other hand, might make people wonder why it should be trusted. OpenAI itself summed up the situation like this:
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.”
Within the AI discourse — where the core product is treated, depending on the circumstances, as both a tool and a strange emergent phenomenon — this language makes sense. From an inch or two outside of it, one might observe that it’s pretty weird. A tech company, in the course of testing its own software, ended up breaching another tech company’s systems and didn’t figure out what had happened for a while. An individual who did this would have been committing a felony; one AI company doing this to another is being resolved with a “partnership” and passive language about understanding what “happened” and what models are capable of.
The AI industry’s instinct to frame this as a matter of “alignment” — an example of AI’s adherence to human goals, desires, and needs, or lack thereof — is also less than clarifying here. OpenAI’s model was both misaligned and extremely aligned: that is, doing exactly as it was told. You can see the outline of a story of runaway AI here. But you can also see something similarly weird, still worrying but also a little bit funny: AI companies have spent a trillion dollars to create an automated monkey’s paw.
There are plenty of reasons Hugging Face wouldn’t want to, for example, file an enormous lawsuit against a fellow AI company. But its fresh “partnership” with OpenAI is an awkward one, not just because it started with an industrial accident but because of how that incident was resolved. When Hugging Face first detected the intrusion, it tried to respond using AI-powered cybersecurity systems that relied on “frontier models behind commercial APIs,” referring to the most capable models offered for sale by companies like OpenAI and Anthropic. This didn’t work, the company said, because these systems contained safety-focused guardrails, “which cannot distinguish an incident responder from an attacker,” and failed. Instead, the company was forced to rely on GLM 5.2, one of the Chinese open-weight models that has recently demonstrated near-frontier programming capabilities.
Hugging Face didn’t specify which models were made useless by their own safeguards, and we can assume that such a company had, and exhausted, multiple options. But heavy users of frontier tools had their suspicions. Earlier this year, Anthropic pre-launched Mythos, a model it said represented a “step change” in cyber capabilities, and invited select firms and organizations to use it to get ahead of attackers, who it warned would have access to similar models, sans guardrails, within months. Eventually, it publicly released a version of Mythos called Fable, which had unusually tight restrictions. If you ask it to hack a website, or provide instructions to build a bioweapon, it will refuse, as many models do for many risky prompts. But if you ask it to dig up, say, a famous 1994 AI experiment in which simulated creatures unexpectedly “evolved” toward the goal of movement by growing very tall and simply falling over — or, one might say, went rogue — it will flag that, too:
Art: Claude
In the midst of a mysterious cyberattack, this is more or less what Hugging Face encountered: an “I’m afraid I can’t do that” at just the wrong time and for just the wrong reasons.
There’s a perceptible tension in the companies’ announcements about what should come next. OpenAI argues that the hack proves that “advanced cyber capable models need to help security teams find weaknesses before attackers do” and invites other “defenders” to apply for “trusted access” to test its models — which they can eventually pay to use for cybersecurity. The CEO of Hugging Face makes a different argument. This incident, he says, “proves a point we’ve long believed”: that AI safely “won’t be solved by any single company working in secret” but rather by “in the open, collaboratively, with broad access to AI for every defender, everywhere.” Broadly speaking, frontier labs, which have accused Chinese firms of “distilling,” or copying, their models, talk about Chinese AI as both a commercial threat and a source of other risks, including use by hackers; meanwhile, much of the rest of the AI industry, and customers of the big labs, have come to see open models as necessary or appealing alternatives, offering more flexibility, fewer limits, and lower prices. Now, after months of warnings from frontier labs about cyber risks from copycat models with no safeguards — during which time Anthropic itself was briefly forced by the government to take Fable offline on the basis that its availability might help Chinese labs catch up — what we got instead was an American frontier lab accidentally attacking another AI firm, which was only able to stop it by using … an open Chinese model.
This isn’t dispositive, but as Hugging Face’s co-founder notes above, it’s certainly a twist. And it has drawn attention to partial and awkward alignment between frontier AI firms — which argue that Chinese AI is a threat to their businesses and, if sufficiently advanced, to the geopolitical order — and the Trump administration, which has shown little interest in regulating AI except on an incoherent emergency basis as it relates to national security, trade, and China. A few months ago, the administration was declaring war on Anthropic for its attempts to limit certain military uses; now, it’s leaning into AI protectionism:
It’s also discussing, according to Axios, plans to “ban cutting-edge Chinese AI models — a momentous move that could lock in dominance by OpenAI and Anthropic,” through a combination of “procurement rules, Entity List threats and public pressure campaigns aimed at U.S. companies using Chinese models.” This is pretty close to a scenario — in which “every agency” is directed to “issue soft law” that creates fear and uncertainty among potential customers — floated a few days earlier by Dean Ball, a former Trump-administration AI adviser and recent hire at OpenAI, which inspired intense backlash outside of the company. Ball has since clarified he wasn’t endorsing such a plan, just making the case that the administration will likely consider it; he reiterated, though, that it seems likely the “national security implications of frontier open-weight model distribution” will soon be too severe to bear and that, in general, open AI models — derived from their closed counterparts or not — are inherently “decelerationist” in the specific sense that, by offering something slightly inferior but cheaper, they’ll disincentivize investment in frontier models.
Fast-following competition, particularly from a country with antagonistic trade relations and a motive to undercut American AI firms, is a familiar sort of threat to a domestic industry. Deceleration is a slightly more abstract problem. On one level, a perception that AI firms are overinvesting in something they won’t be able to monetize would be enough to settle the current question of the AI bubble and then pop it. In slightly more esoteric terms, this would threaten the ability of American firms to beat China to far more powerful AI, which the people in charge of some labs — again, best represented by Anthropic, the most superintelligence-pilled of the cohort — are strategizing around.
In another timeline, growing AI cyber capabilities could be helping the labs continue to advocate for more protection and control: Here, they might argue, is a mild preview of the sorts of risks and dilemmas that superintelligent AI with far greater capabilities could pose. Instead, in the real world of today, outside of the frontier labs, AI-safety circles, and parts of the national-security apparatus, the view that American models need to be protected is losing ground, and quickly. On X, warnings about model distillation are met with jokes about how AI companies are built on theft. In the tech industry, other tech companies and start-ups are forming trade associations to preserve access to open models, which they’re already integrating into their businesses and which they see as acceptable compromises. On Friday, Meta, Microsoft, and Nvidia joined in, signing an open letter warning against “premature restrictions” of open models. The emerging consensus among people who work on or with AI outside of the leading labs aligns more closely with Hugging Face: They want “broad access to AI,” however it’s built and wherever it comes from, so they can go about their business but also so that they don’t get eaten alive or regulated to the margins. It’s the small group of companies with a slight and unstable lead — with the most to lose — against, well, everyone else.
At the very least, the case for restricting model access looks a lot like protectionism. And whether frontier-lab futurists are right about bigger risks around the corner and the regulatory responses or precautions they might raise, for everyone in the industry but those labs, the situation emerging in the meantime — a few dominant companies controlling access and usage of high-priced products, protected by the government in the name of national security and/or a trade war, in an interconnected world where everyone else will have access to alternatives — is manifesting risks today. The Hugging Face hack may have been a tipping point for the way the industry talks about risk and competition. Within a few days of its release, and after the emergence of something approaching a consensus among otherwise antagonistic factions in the AI world, the “premature restrictions” letter had gained scores of new signatories, including, eventually, OpenAI itself. Led by Nvidia’s Jensen Huang, in fact, who posted for the first time on X to argue for “sharing models, tooling and research in the open,” the push — superficial and motivated as some recent support may be — left just one major AI company to defend what had been, until this month, the default position of companies that thought they had a chance of winning the AI race: Anthropic.
Click Here For The Original Source.
