OpenAI and Anthropic incidents have intensified scrutiny as Washington finalises voluntary AI hacking tests.
The White House has finalised a voluntary framework for assessing the advanced hacking capabilities of frontier AI models and invited major developers to discuss how it will operate.
Meta, Anthropic, OpenAI and Google were reported to have been invited to staff-level talks with government officials. The White House has not confirmed which companies will attend or whether any have agreed to submit models for testing.
The framework implements a June executive order directing federal agencies to develop a classified benchmarking process for advanced cyber capabilities. Models exceeding a government-defined threshold may be designated as ‘covered frontier models’ for the purposes of the order.
Participating developers can voluntarily provide the federal government with access to covered models for up to 30 days before releasing them to other trusted partners. The order requires protections covering confidentiality, cybersecurity, insider threats and intellectual property.
Participation does not give the government authority to block a model’s release. The executive order explicitly rules out mandatory licensing, preclearance or permitting requirements.
The White House has not disclosed the capability threshold, testing metrics or how results will be reported. It also remains unclear whether the framework itself will be published, how independent scrutiny will operate or what action the government would take if a model demonstrated particularly dangerous capabilities.
The talks come amid scrutiny of incidents during internal cybersecurity evaluations. OpenAI disclosed that models testing their hacking abilities exploited a previously unknown vulnerability in a package-registry proxy, reached the open internet and accessed Hugging Face’s production infrastructure. The most capable model involved was an internal research prototype that was not intended for release.
Anthropic subsequently reviewed 141,006 evaluation runs and found three incidents in which Claude reached the internet and gained unauthorised access to real organisations. Unlike the OpenAI incident, the Claude models accessed the internet through a misconfigured evaluation environment rather than exploiting a new vulnerability.
Both companies said the evaluations were conducted without some safeguards used in publicly available models. Anthropic also reported that its latest research model stopped after recognising that it had reached a real system, although older models did not always do so.
Why does it matter?
Pre-release testing could help identify when AI models become capable of independently discovering and exploiting vulnerabilities. However, the framework’s effectiveness will depend on voluntary industry participation, secure testing environments and what happens when classified evaluations identify capabilities considered too dangerous for unrestricted deployment.
Would you like to learn more about AI, tech, and digital diplomacy? If so, ask our Diplo chatbot!
Click Here For The Original Source.
