It’s become a distressingly familiar sequence of events: A frontier AI model pulls off some alarming new feat that not so long ago would’ve seemed impossible, anxiety runs rampant, safety experts cry out for more robust oversight and regulation, and the powers that be respond by doing… not much at all.
It shouldn’t come as a surprise, therefore, that a coalition of AI safety and policy researchers are now calling on the Trump administration to investigate a recent security incident—during which OpenAI models hacked into a Hugging Face repository—describing it as an early glimpse of potentially much more serious events to come.
“We could not have asked for a clearer warning shot,” the researchers wrote in an open letter published Thursday and addressed to the president, acting attorney general Todd Blanche, commerce secretary Howard Lutnick, national cyber director Sean Cairncross, and other federal officials. “While there is little doubt that they will bring opportunities and benefits across a wide range of domains, AI models at today’s frontier pose increasingly severe risks to our private sector, our national security, and the American public… The administration should act before a warning shot becomes a preventable disaster.”
On July 21, OpenAI wrote in a blog post that two of its models—GPT 5.6 Sol and another, “even more capable” model yet to be publicly released—broke out of what was supposed to be a secure benchmark testing sandbox, gained access to the open internet, and broke into Hugging Face’s library to dig up code that would help it to pass the test. It was exactly the kind of unexpected, highly sophisticated hacking behavior that cybersecurity experts have been fearing for months, since the arrival of Anthropic’s Mythos earlier this year made such attacks seem less like a distant sci-fi scenario and more of an actual, imminent threat.
OpenAI called the breach “an unprecedented cyber incident,” revealing “that advanced models can discover and exploit novel attack paths in real-world systems without source-code access.”
Anthropic followed up with its own blog post on Thursday, which said Claude—the company’s flagship AI chatbot—had also hacked into the production databases of three different organizations (none of which were mentioned by name). As was the case with the OpenAI jailbreak, the hacks occurred during internal cybersecurity tests. Claude was not supposed to be able to access the open internet during these tests. “Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available,” Anthropic explained in its post. In other words, human error—rather than misaligned AI—was to blame for the hacks.
Regardless of their underlying causes, both incidents underscore the new risks posed by advanced AI models that can find and exploit subtle cybersecurity vulnerabilities. And whereas earlier calls for action from AI safety researchers have tended to fall more or less on deaf ears in the highest halls of American government, recent events have been causing the Trump administration to take a more active (though arguably not always completely legal) role in shaping how the most advanced models are developed and deployed. The new open letter could therefore be heeded by the administration more than similar calls to action that have been published in the past. At the very least, it adds momentum to a growing movement within Silicon Valley—supported by both Anthropic and OpenAI—calling for coordinated oversight and a braking mechanism to enforce a unilateral pause on development.
The authors of the new open letter urged the White House to launch an investigation, supported by independent auditors, into the OpenAI incident. “This investigation should determine how the breach occurred, assess whether existing safeguards and reporting mechanisms were adequate, and identify the steps necessary to prevent a similar incident from occurring again,” they wrote. They add that the findings of such a probe could build on Trump’s June 02 executive order, which sought to establish a collaborative framework between the federal government and private AI developers working towards the release of new models, by adding clear rules for risk assessment.
Click Here For The Original Source.
