Better than a human at running a computer, OpenAI’s new ChatGPT 6 scored 100% on a hacking benchmark | #hacking | #cybersecurity | #infosec | #comptia | #pentest | #hacker


OpenAI has released GPT-6 Astra, an AI model the company says has achieved general artificial intelligence, and it posted a perfect 100% score on ExploitBench, a benchmark measuring how well an AI agent can turn known vulnerabilities into working exploits. That result has earned the model a “Critical” threat rating, prompting OpenAI to restrict its offensive capabilities to state actors and vetted partners through Daybreak, its cybersecurity initiative offering $1 billion in credits to eligible defenders.

OpenAI vient de lâcher GPT-6 Astra dans la nature, et le moins qu’on puisse dire, c’est que la bête ne fait pas semblant. Ce modèle décroche un score parfait de 100% sur ExploitBench, le benchmark qui mesure les capacités de hacking, tout en revendiquant une vraie percée vers l’intelligence artificielle générale. Entre la modélisation 3D sous Blender, le routage de circuits imprimés sur KiCad et une vitesse d’exécution 1,9 fois supérieure à celle de son prédécesseur sur Mind2Web, l’outil enchaîne les tâches techniques avec une aisance qui dépasse déjà celle de nombreux experts humains. Cette puissance a un prix : OpenAI classe elle-même Astra en niveau de menace “critique”, forçant la firme à verrouiller sérieusement l’accès à ses capacités offensives.

OpenAI says the AGI era has arrived

OpenAI unveiled GPT-6 Astra on September 3, 2026, just 2 days after Anthropic released Claude Fable 5.1. The company frames this sixth-generation model as the moment artificial general intelligence became real: an agent that directly controls computer systems rather than simply answering text prompts, and one that outperforms humans on complex engineering and research work.

President Greg Brockman made the claim explicit in an interview with The Washington Post. “We’ve reached a point where these models not only solve century-old math problems, but also boost the economy and improve your personal life. I think we’re there,” he said.

Faster than a human, by hours

Astra’s core advance is agentic speed. Running on the new Codex harness, it executes tasks 1.9x faster on Mind2Web. In one test, it completed a complex apartment search in 2 minutes and 54 seconds, a job that regularly took a human operator more than 6 hours.

The model also models a house in Blender and turns it into a fully playable 3D scene in Unreal Engine 5, and routes printed circuit boards in KiCad within seconds. In programming, it scores 57.9% on Terminal-Bench 4.0 and keeps structured notes instead of blindly summarizing when its context window fills, digging through old debugging sessions to retrieve exact information.

Record benchmarks, critical risk

The numbers are stark: 99.9% on ARC-AGI-3, 97.6% on FrontierMath Tier 4, and a perfect 100% on ExploitBench, OpenAI’s capability-ladder benchmark for gauging how effectively an autonomous agent can weaponize known security flaws. Tested without guardrails, Astra autonomously exploited flaws in hardened web browsers and discovered 2 previously unknown Zero-Day vulnerabilities on its own, then designed remote code execution attacks around them.

OpenAI has classified the model as a “Critical” cybersecurity threat, a first under its internal evaluation protocol. A background monitoring system now interrupts processing automatically if the AI’s reasoning drifts, a safeguard informed by the loss of control over test models on Hugging Face in July 2026.

Who gets access, and at what price

Access is limited to paying subscribers on the Plus, Pro, and Business tiers as it rolls into ChatGPT plans, plus developers through the API as `gpt-6-astra`, also distributed via Microsoft Azure and Amazon Bedrock. Pricing matches Anthropic’s rates: $10 per million input tokens and $50 per million output tokens. The most sensitive offensive capabilities remain locked away entirely, reserved for state actors and certified partners through OpenAI’s secured Daybreak program, its cybersecurity initiative with security companies and other organizations that includes $1 billion in Daybreak credits for eligible defenders.



Click Here For The Original Source.

——————————————————–

..........

.

.