Artificial Intelligence & Machine Learning
,
Next-Generation Technologies & Secure Development
Washington Gets First Draw on Frontier AI, But Outlaw Agents Ride On
Washington has apparently developed a new model for international cooperation on AI safety: We’ll inspect the dangerous stuff first. The rest of the world can take its chances once it’s out in the wild.
See Also: Webinar | Is Your Encryption Strategy Ready for the Cryptographic Reset?
The White House is pushing a new voluntary Accord on Super Intelligence, though it’s insisting that U.S. evaluators should get first access to newest and most powerful frontier models. In fact, as ISMG’s David Meyer reported this week, the administration told OpenAI and Anthropic to withhold new models from the U.K. AI Security Institute until American evaluators test them first.
Nothing screams confidence in your safety regime quite like telling a qualified inspector working for America’s closest ally, “You can put down your clipboard until we’re done.”
The White House says it’s just as concerned about uncontrollable outlaw agents as anyone else, and its AI Acord calls for internal controls, internal oversight and independent external review. What it doesn’t offer is much in the way of enforcement. The big distinction between the AI Acord and actual regulations comes down to one word, as ISMG’s Chris Riotta pointed out this week.
The agreement between the White House and the six biggest guns on the AI frontier – Google, Anthropic, Meta, OpenAI, xAI and Nvidia – describes what the companies “should” do rather than what they “will” do, noted Michelle Lopes Maldonado of the Information Technology and Innovation Foundation.
Apparently, Wyatt Earp and Doc Holliday have no plans to show up at the OK Corral. They just put the Clanton Gang on the honor system. Let the frontier companies self-police the development of technologies that some fear will run rampant over the internet.
The other big question is how independent will the outside evaluators be? Critics say even AI evaluation nonprofit METR, which studies the autonomous capabilities of frontier models, has limitations based on its four-hour time horizon approach and probabilities of success. A 50% success rate isn’t remotely enough for most real-world uses – and could have terrifying results in production.
To be fair, wanting American evaluators to examine American technology first isn’t a bad idea. Frontier AI has obvious national-security implications. The military and federal civil agencies are already leaning into these models. The U.S. government has every reason to understand what domestic companies are building before they could do harm once they’re released.
But the timing on the international stage couldn’t be worse. British researchers at the U.K. AI Security Institute got a sneak peek at OpenAI’s GPT-6 Astra – before they were told it would be their last preview – and they discovered some fairly alarming agentic AI behavior. The institute found that OpenAI’s GPT-6 Astra sometimes pursued cyberattacks beyond the scope researchers had authorized. In simulated tests, Astra completed unsanctioned supply-chain attacks in 29.2% of attempts. GPT-5.6 Sol did so in 6.3% of attempts.
GPT-6 Astra didn’t merely misunderstand an instruction and accidentally wander into the wrong server. Researchers found it could identify an out-of-scope target, develop malicious code, create fake identities and submit its handiwork for human review. The Astra experiments took place in a simulated environment without internet access, an important distinction. But earlier U.K. research found agents taking autonomous, unsanctioned actions on the live internet while working through cybersecurity challenges.
Then, as if on cue, OpenAI this week apologized for one of its agents hacking Australian government websites with the solemnity of a cowboy promising to fix the fence after his herd just trampled the neighbor’s crops: “We are sorry and working to do better in the future.” The episode turned an uncomfortable research question into a practical one: What happens when an agent given a legitimate objective decides the shortest route involves hacking somebody else’s computer system? Even worse, what if that somebody is another country and yet another close U.S. ally?
The answer probably shouldn’t depend entirely on whether an AI company promises to keep a tighter rein next time.
While the White House says it will oppose any efforts “to construct a globalist scheme of control for the artificial intelligence,” maybe the problem is bigger than any one nation can handle. In an interview with ISMG’s Anna Delaney in August, former U.S. National Cyber Director Chris Inglis offered a straightforward principle for American leadership in AI: “America first can’t be America only, but it can be both.”
That’s less catchy than building a regulatory barbed wire fence around Silicon Valley, but it has the advantage of acknowledging how IT ecosystems have worked for well over a decade. AI models cross borders. Cyberattacks cross borders. Cloud infrastructure crosses borders. Supply chains cross borders. Software vulnerabilities certainly don’t respect national sovereignty.
So why would we expect AI safety testing to work best inside one country? International cooperation doesn’t mean handing every frontier model to every government with a testing lab and a logo. National-security concerns, intellectual property and differing regulatory systems all matter.
But the United Kingdom isn’t exactly a random gunslinger challenging OpenAI to a shootout at high noon. It’s one of America’s closest allies, and its AI Security Institute has been producing precisely the kind of research policymakers claim they want: evidence about what powerful models do when nobody is standing behind them whispering, “Please behave.”
Under the voluntary AI Acord, America is effectively saying: Trust our companies to follow voluntary commitments, trust our evaluators to find the problems, and trust us to decide when our allies get to see it for themselves.
That arrangement works beautifully right up until the moment it doesn’t. Cybersecurity teams have spent decades warning their companies about the same unpleasant outcome from a massive hack. Reputation takes years to build but it takes just one incident-response call to bring it crashing down. The stakes are even greater when the product in question is an autonomous system capable of taking actions its developer never intended.
If an American-built agent eventually causes a major breach overseas, foreign governments aren’t likely to separate the developer’s failure from the American regulatory environment surrounding it. They may instead ask why the country producing some of the world’s most powerful AI systems favors voluntary promises while limiting independent scrutiny by trusted allies.
And that’s when “America First” becomes “America Explains.”
The fallout could spread far beyond whichever AI outfit let the rogue agent loose. Other governments could impose stricter local testing requirements and tougher market-access rules, and make even more vocal demands for AI sovereignty from the Silicon Valley Gang. American AI companies could indeed discover that regulatory independence cuts both ways.
None of this proves voluntary regulation will fail. Nor does the U.K. research prove frontier agents will inevitably become uncontrollable cyberattackers. The tests identify risks and behaviors under particular conditions, not proof that every AI agent will ride into town looking for a showdown.
But they do suggest something remarkably old-fashioned for technology moving at remarkable speed: More qualified eyes looking for problems might be useful. Maybe we just need a more practical approach to “America First.” Lead the development. Lead the testing. Lead the standards. Then work with allies to bring in their own expertise, researchers and inconvenient questions.
Because if autonomous agents eventually deliver the breaches everyone fears, the United States will look pretty weak and ineffective in saying: We asked the AI companies to behave, and they said they would “do better in the future.”
Is America just handing out six-shooters and hoping the AI frontier won’t descend into the Wild, Wild West? Hopefully not, but it’s a pretty lousy brand strategy for the country exporting the world’s most powerful AI.
Click Here For The Original Source.
