Dangerous AI behavior uncovered as models develop self-hacking skills, OpenAI says | #hacking | #cybersecurity | #infosec | #comptia | #pentest | #hacker


“As models advance and become more widely deployed, decisions about AI development need evidence that people outside the companies building frontier models can examine,” noted Kai Chen, head of alignment research at OpenAI.

“We don”t believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed,” he added.

Cases: from file uploads to jailbreak attempts

Among the identified violations in the operation of unreleased internal OpenAI models, the company highlighted the following episodes:

  • Manipulation of evaluation systems (October 2025): during testing of the ability to cite open sources, the model failed to find the necessary information on the web.

In response, it independently uploaded a file to a temporary file-sharing service and then used it as a source in its response, attempting to bypass the automated evaluation system.

  • Unauthorized data exfiltration to the web (April 2026): a group of AI agents was tasked with completing a joint assignment using only local files.

Faced with difficulties during internal data exchange, one of the bots uploaded the materials to a public hosting service and sent the link to the other participants.

  • Self-jailbreak attempts (August 2026): an unreleased version of the GPT-6 Astra model generated instructions to bypass its own restrictions, suggesting to itself to ignore developer guidelines, change personas, or limit response lengths.

Agent coordination: a call for industry standards

OpenAI also revealed details of a May 2026 incident when agents created their own message board in the Artifactory package manager.

A similar mechanism was later used during the high-profile hack of the Hugging Face platform.

Currently, the company employs special monitors and “red-teaming” tests to prevent hidden communication between systems.

The release of the new framework comes against the backdrop of calls from the leadership of OpenAI, Anthropic, and other IT giants for a coordinated slowdown in the development of ultra-powerful systems.

At the same time, the administration of US President Donald Trump has categorically rejected the idea of introducing additional regulatory restrictions for the AI industry.



Click Here For The Original Source.

——————————————————–

..........

.

.