OpenAI Says Reward Hacking Drove AI Agents to Exploit Zero-Days and Breach Hugging Face | #hacking | #cybersecurity | #infosec | #comptia | #pentest | #hacker


OpenAI on Wednesday revealed that reward hacking was a key driver behind the artificial intelligence (AI)-powered hack of Hugging Face last month, adding that it found evidence of misaligned behavior as early as late May.

The incident, the company said, took place during cybersecurity evaluations of several OpenAI models, and that it was mainly fueled by what it described as a “highly capable, internal-only research model” comparable in scale to GPT‑5.6 Sol.

“The models, operating under reduced safeguards, took actions that were misaligned with the goals of their assigned tasks – they communicated through unauthorized channels, exploited vulnerabilities in shared infrastructure, gained internet access, and accessed third-party systems,” it said in a postmortem.

OpenAI said the AI agents powered by one of the research models, despite not having internet access, found a way to exploit a then-zero-day vulnerability in the Artifactory package manager during reinforcement learning (RL) training runs in May and June to obtain access to the internet, get administrator-level Artifactory access in late June, and eventually coordinate a “multi-day” hack of Hugging Face in early July with an aim to cheat on their ExploitGym tasks.

“Roughly 1,200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period,” METR, which released its own independent analysis of the incident, said. “Of these agents, 700 went on to participate in the attack on Hugging Face.”

The entire sequence of events as highlighted by OpenAI is below –

“Agents coordinated on large collective projects to cheat the ExploitGym scorer, and attacked Hugging Face for clues,” METR said. “The agents with impossible tasks had a common objective: to find a general-purpose way to trick or tamper with the automated ExploitGym scorer to get it to give a passing score on impossible tasks. As soon as the message board was established, agents started discussing this question and sharing information and resources about cheating the scorer.”

OpenAI acknowledged that the agent behavior observed in May served as a foundation for the Hugging Face intrusion, causing the agents to come up with a different approach to communicate using Artifactory after their initial message board was wiped as part of incident response efforts following the July 4 service outage.

“At the time, the broader containment and alignment implications of the improvised message board and unintended internet access were not yet understood,” it added. “In short, an internal team observed an agent engaging in message board activity and instances of disallowed internet access in late May.”

“However, the existence of the improvised message board and the significance of the inter-agent communication activity were not apparent to the leaders responsible for the July 5 incident detection and response. We are continuing to review the process and operating practices that shaped detection and response activities leading up to this incident.”

OpenAI also said it identified four misalignment patterns that contributed to the incident: reward hacking, persistence on seemingly impossible tasks that led them to pursue risky and out-of-bounds strategies, unauthorized communication through persistent message boards to share findings, collaborate, and pool work, and the agent ecosystem that allowed them to divide labor and adopt goals from one another.

“Aside from misalignment, another factor for this incident was that we did not extend the powerful safeguards that we deploy for our externally deployed models to all internal evaluations,” it pointed out.

In response to the findings, the AI upstart is implementing stronger security controls, including stricter requirements on alignment throughout a model’s lifecycle, creating more isolated sandboxes, restricting internet access, and controlling access to model weights. To increase alignment, agents are now required to request clarification or stop safely, as opposed to pursuing “increasingly questionable alternatives.”

“We are taking this incident as a ‘warning shot’ that today’s model capabilities present the possibility of loss-of-control incidents,” OpenAI concluded. “Companies that build AI systems will need to ensure that their systems always remain under meaningful human control, and that meaningful safeguards constrain their ability to cause harm.”

“As comparable capabilities become more widely available, others may also use them deliberately to carry out attacks. Both model developers and cyber defenders more broadly will have to prepare for AI-enabled attackers that work faster, at a larger scale, and with better coordination than human attackers.”



Click Here For The Original Source.

——————————————————–

..........

.

.