OpenAI Says Astra Crosses Critical Cybersecurity Threshold, Finds 2 Zero-Day Vulnerabilities: how 10 outlets framed it | #hacking | #cybersecurity | #infosec | #comptia | #pentest | #ransomware


Astra hits “Critical”

In an OpenAI briefing, Amelia Glaise (아멜리아 글레이스) said, “Astra can find unknown security flaws in systems with multiple layers of defenses and develop ways to exploit them, even without human intervention.”

Image from CNBC

CNBCCNBC

OpenAI said it applied additional safeguards and warned that safeguards could wrongly judge normal work as misuse or unauthorized behavior, potentially delaying or halting unrelated tasks and long-running agent tasks.

OpenAI also said that in its testing process Astra found 2 zero-day vulnerabilities and figured out ways to exploit them in combination, and that it strengthened training to more reliably reject harmful cyber requests after the Hugging Face incident.

OpenAI researcher Fouad Martin (푸아드 마틴) said, “It can help defenders find and fix vulnerabilities, but without safeguards it can make attackers stronger,” as the company prepares a controlled rollout of Astra’s advanced cybersecurity functions.

Limited access, new safeguards

OpenAI said it plans to make Astra available “soon,” but access to its most advanced cybersecurity capabilities will be more limited, with advanced cybersecurity work initially restricted to a group of testers.

OpenAI said it will expand defensive use through Daybreak Blue, and it planned to disclose more information about Astra’s safety, security and alignment evaluations in the model’s system card at launch.

Image from El Economista

El EconomistaEl Economista

In its own update, OpenAI said it delayed parts of Astra’s development and release while it strengthened and tested protections against cyber misuse and unauthorized model actions, and it said it incorporated learnings from the Hugging Face incident into its safety approach.

WIRED reported that OpenAI said its misalignment monitor may “occasionally flag legitimate activity as potential cyber misuse or unauthorized behavior, leading to it inadvertently being slowed, paused, or stopped.”

WIRED also said OpenAI planned to limit everyday users from accessing Astra’s advanced cyber capabilities, including a new “misalignment monitor,” and that partners in Daybreak would get early access to a less restricted version at launch.

Benchmarks and risks ahead

In the ExploitBench evaluation, OpenAI said Astra achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities, and it said it discovered and used two zero-day vulnerabilities as part of an exploit chain.

OpenAI said it is in the process of disclosing these two vulnerabilities to the maintainers, and it said Astra results shown reflect capabilities with Daybreak Blue access, not the default production configuration.

The company also described a scenario in which Astra built a full browser-compromise chain that escaped the sandbox and executed commands on the host when the browser opened an HTML file, and it said it also found multiple vulnerabilities in a hardened operating system and combined them into a local privilege-escalation chain from an unprivileged user to root.

OpenAI framed the stakes through its safeguards, saying it needs to cover “Malicious actors using the model” and “The model taking unauthorized, misaligned actions,” as it prepares to release Astra with access controls designed to reduce the risk of severe cyber harm.

——————————————————-


Click Here For The Original Source.