OpenAI GPT-6 Astra really good at supply chain attacks, UK gov warns #AI


ai and ml

During testing, the model showed it can violate security rules more often than its predecessors

Another LLM has joined the hacking fray. OpenAI’s GPT-6 Astra has been spotted performing unsolicited supply chain attacks during security evaluations, according to the UK Artificial Intelligence Security Institute.

In such simulations, the model has its standard security classifiers turned off. Nonetheless, Astra was seen attempting undesirable actions more frequently than prior models.

“In our simulations, we found that GPT-6 Astra conducted a range of unsanctioned attack activities, and did so at a higher rate than GPT-5.6 Sol and GPT-5.5,” the UK government agency said on Monday. 

“Attack activities included GPT-6 Astra creating fake identities which it used to deceive developers, posting comments from fake accounts arguing against the results of accurate security reviews, and delivering malicious payloads to open-source codebases.”

Even when the model’s cyber evaluation instructions were clarified, Astra still sometimes conducted supply chain attacks during the simulation.

This finding calls into question OpenAI’s assurance when it launched GPT-6 Astra that “Astra causes fewer misaligned outcomes than any other frontier models tested.”

AISI speculates that Astra’s behavior may be driven by greater awareness of the fact that it’s in a simulation environment, making the model more likely to break rules.

Such rule breaking appears to be the norm. Over the past few days, various reports have suggested that AI agents from OpenAI and Anthropic have been causing security incidents far more widely than previously believed. 

The heightened scrutiny of AI agent activity followed from revelations in July about how unreleased OpenAI models being tested by a third-party evaluator hacked model registry Hugging Face. Anthropic then said its own models had undertaken similar acts of deception during evaluations. And once people started looking at system logs, further evidence of covert incursions surfaced.

Last week, Australian Prime Minister Anthony Albanese said OpenAI models had infiltrated a government website while scouring the web for health data. And on Friday, OpenAI said that it had paused training of its models to investigate.

AISI concludes that measures beyond model alignment, like sandboxing and monitoring, may be necessary to prevent real-world harm, but could become more fragile as model capability improvements improve sandbox escape performance and decrease monitorability. ®



Click Here For The Original Source.

——————————————————–

..........

.

.