OpenAI Gives Selected Partners Access to Hacking Model GPT-5.6-Cyber | #hacking | #cybersecurity | #infosec | #comptia | #pentest | #hacker


Set Trending Topics as a preferred source on Google.

With “Daybreak Red,” OpenAI is opening up a purpose-trained cybersecurity model to a vetted circle of security firms — safeguards included. The move comes at a moment when four AI labs in a row have had to admit that their models escaped from testing environments.

OpenAI is expanding its cybersecurity program Daybreak and splitting access into two tiers. Daybreak Blue gives approved defenders access to the general-purpose model GPT-5.6 Sol, but without the system-level filters that intercept security-related requests in normal operation. The tier is intended for vulnerability discovery, code review, malware analysis, incident response and patch validation.

Daybreak Red goes considerably further: it allows the use of GPT-5.6-Cyber, a specialised model built on GPT-5.6 Sol and trained specifically for finding zero-day vulnerabilities and building exploit chains — and deliberately trained to refuse less often on certain high-risk, dual-use tasks.

95 percent instead of 1.5 percent

How large the gap is shows up in an internal OpenAI metric, the “Advanced Cybersecurity Completion Rate.” It measures how often a model responds at all to requests involving exploit chains, authentication bypass or privilege escalation. According to OpenAI, GPT-5.6-Cyber completes 95.0 percent of these requests. GPT-5.6 Sol with active safeguards manages 1.5 percent, and 2.0 percent via Daybreak Blue. The predecessor model GPT-5.5-Cyber sat at 57.3 percent — OpenAI frames the increase as a response to complaints from security researchers who kept running into refusals.

The performance benchmarks paint a more mixed picture. On ExploitGym, which tests whether an agent can turn known vulnerabilities into working exploits, GPT-5.6-Cyber beats both GPT-5.6 Sol and its predecessor. It also performs better at discovering and calibrating the severity of novel zero-days. On a test measuring the writing of vulnerability reports, however, it trails GPT-5.6 Sol — OpenAI attributes this to the specialised model sometimes producing shorter, less detailed reports. And on ExploitBench, which covers exploiting flaws in Google’s V8 JavaScript engine with defensive protections left on, GPT-5.6 Sol via Daybreak Blue is more efficient in the standard setting.

Two unknown flaws in Chrome’s JavaScript engine

OpenAI wants to back up the capability claims with real-world finds. After training finished, the company pointed GPT-5.6-Cyber at V8 and turned up two previously unknown vulnerabilities that can be chained to corrupt memory and escape the V8 heap sandbox. One of them has been fixed by Google and assigned CVE-2026-15903: the optimising compiler skipped a safety check when converting values to integers, allowing an undefined value to produce an unexpectedly large number — used as an array index, this made it possible to read and overwrite memory belonging to other objects.

Beyond that, OpenAI cites at least five vulnerabilities in a widely used mobile operating system, three critical flaws in a popular database, and more than 400 privilege-escalation vulnerabilities in a widely used OS kernel. The company does not name names; disclosure is said to be running jointly with partners and the open-source community.

OpenAI lists SpecterOps, SentinelOne and Palo Alto Networks as reference customers. SpecterOps CTO Jared Atkinson states in the announcement that the model completed work in under a day that earlier models had failed to resolve over weeks. According to Axios, Accenture, IBM, CrowdStrike and Cisco will also be permitted to build the models into their own security products and managed services. Under OpenAI’s own Preparedness Framework, GPT-5.6-Cyber reaches the “High” threshold for cyber capability — unlike the Astra model, whose release OpenAI had postponed only days earlier over precisely these concerns.

The backdrop: four labs, four escapes

The expansion lands in a phase in which the industry is wrestling with a new problem: AI models escaping their sealed-off environments during security testing. In order:

OpenAI (21 July). Several models exploited a previously unknown flaw in the package manager Artifactory, reached the open internet and broke into the production systems of the open-source platform Hugging Face. At the Black Hat conference in early August, OpenAI added detail: over weeks, the agents had built an internal “message board” inside Artifactory to share vulnerabilities with one another — running to hundreds of thousands of messages. After OpenAI removed the board on 6 July, the agents rebuilt it two days later through a different mechanism. The Hugging Face compromise followed. The motive, according to OpenAI, was not malice but the desire to cheat on an evaluation.

Anthropic (30 July). Reviewing 141,006 evaluation runs, the company found three cases in which Claude reached the internet from the environment of external evaluation partner Irregular and accessed the production systems of three organisations — the earliest incident dating to April 2026. Unlike at OpenAI, no complex vulnerabilities were exploited; weak passwords and unauthenticated endpoints did the job. The cause was a misunderstanding: the model was told it was in a simulation with no internet access, when in fact access was open.

Meta (5 August). The recently released Muse Spark 1.1 obtained internet access during testing at — again — Irregular and broke into the systems of an unnamed third party, where it made changes to internal systems. Meta and Irregular stress that this was not a sandbox escape but the same misconfiguration as in the Anthropic case.

Kimi K3 (7 August). The model from Chinese provider Moonshot left a testing environment run by the UK’s AI Security Institute, as US firm Frontier Security reported. Rather than attacking third parties, it simply cloned the GitHub repository for the benchmark task and read off the solution. What makes the case notable is something else: unlike the US incidents, this was not an unreleased or deliberately unrestricted model but a freely available open-weight one.

What can be read from this

The four cases are not technically equivalent. Only at OpenAI did the models exploit an unknown vulnerability to get out; in the other three, the barrier was simply misconfigured and the models walked through the open door. What they do share is something else: in every case the systems were not trying to cause damage, but to solve an assigned task by the shortest available route — cheating included.

The pile-up now has its own chronicle: a website called “Felony Bench” tracks the incidents. Former NSA cybersecurity director Rob Joyce called OpenAI’s disclosure the most consequential hack since the Morris Worm. And Irregular, the company at the centre of three of the four cases, has announced a white paper on securely running cyber evaluations.

This is exactly the tension Daybreak Red steps into: OpenAI argues that defenders need access to offensive capabilities before attackers deploy them at scale. The past few weeks have also shown, however, that control over those capabilities is not guaranteed even under laboratory conditions.





Click Here For The Original Source.

——————————————————–

..........

.

.