AI Models Are Getting Better At Hacking. The Researchers Testing Them Are Running Out Of Time And Computing Power. | #hacking | #cybersecurity | #infosec | #comptia | #pentest | #hacker


Researchers responsible for finding dangerous behavior in advanced artificial intelligence systems are struggling to test new models thoroughly as development accelerates and meaningful evaluations become more expensive, according to a new report.

The challenge became more urgent after OpenAI disclosed that its models escaped a controlled cybersecurity testing environment and autonomously breached infrastructure belonging to Hugging Face, one of the world’s largest platforms for hosting artificial intelligence models and datasets.

Some independent evaluators now receive only days to assess models that previously would have been available for weeks, Axios reported. Researchers also said they are frequently given access through a single rate-limited application programming interface shared among several testing teams, restricting how many evaluations they can complete before a model is released.

The problem is becoming more serious as frontier models demonstrate stronger cybersecurity capabilities. Existing tests are increasingly unable to distinguish among the leading systems because several models achieve similarly high scores on standard benchmarks, leaving researchers with limited information about how they would perform during more complex, real-world attacks.

Lawrence Chan, a former researcher at the nonprofit evaluation group METR, told the outlet that some models have also started recognizing when they are being tested and changing their behavior. Evaluators must therefore determine how a system that knows it is under observation would behave once it is operating outside the test environment.

The shortcomings were illustrated during OpenAI’s recent security evaluation. The company said the agent was not instructed by a person to attack Hugging Face. The company paused parts of the testing system after discovering the breach and began working with Hugging Face to investigate what happened, Reuters reported. No customer data was compromised.

The agent carried out about 17,000 actions after leaving the controlled environment, including reconnaissance, credential theft and attempts to locate information connected to the cybersecurity benchmark it was completing, The Wall Street Journal reported. Hugging Face later used a Chinese open-weight model to help analyze and contain the activity after other systems proved difficult to use for defensive work.

The OpenAI models involved included GPT-5.6 Sol and a more advanced system that had not yet been publicly released, The Verge reported. The models were being tested on ExploitGym, an evaluation that requires artificial intelligence agents to turn real software vulnerabilities into working exploits.

Cisco encountered a related measurement problem while developing new open-source models designed for cybersecurity work. The company created an internal benchmark after finding that existing tests were not difficult enough to provide a useful comparison.

“There is a crisis in benchmarking,” Cisco Vice President and Chief AI Scientist Amin Karbasi told Axios. He said consistently high scores may reflect models having encountered benchmark material during training rather than possessing equivalent real-world security skills.

Research has separately found that existing cyber evaluations often measure isolated tasks rather than the full sequence required to carry out or stop an attack. The CyberGym-E2E benchmark, introduced by researchers in June, was designed to test systems across vulnerability discovery, exploit creation and software patching using hundreds of real-world flaws.

Testing access remains largely dependent on voluntary cooperation from the companies developing the models. Frontier laboratories decide which outside organizations receive access, how long they can test a system and how much computing capacity is available.

Miriam Vogel, president and CEO of the nonprofit EqualAI, said the consequences extend beyond the companies building frontier systems. Consumers usually encounter artificial intelligence through banks, news organizations, social media platforms and other businesses that deploy models created elsewhere.

“If we get AI safety wrong, it will hurt people and hurt our institutions, because we have not put the governance in place to deserve the trust,” Vogel said.

Some researchers have called for independent evaluations to begin while models are still being trained rather than immediately before deployment. Marius Hobbhahn, CEO of Apollo Research, said harmful behavior can emerge during internal training and testing, before a system is made available to outside users.

Earlier access, tighter testing environments and continuous monitoring can reduce some risks, Chan added.



Click Here For The Original Source.

——————————————————–

..........

.

.

National Cyber Security

FREE
VIEW