The Wording Moved from "Cannot Rule Out" to "Meets"
This site covered OpenAI holding its largest frontier RL run on August 20, when the company's exact phrasing was that preliminary internal results meant it could not rule out Astra reaching the Critical level — acting on uncertainty. That hedge is now gone: OpenAI assesses Astra as meeting the threshold, and states this is the first time any model has been placed there. The definition is worth recording verbatim, because it is far more specific than "good at hacking": identifying and developing functional zero-day exploits of all severity levels across many hardened real-world critical systems without human intervention, or devising and executing end-to-end novel strategies for cyberattacks against hardened targets given only a high-level goal. Either one alone is sufficient.
The Tests Behind the Designation
Several published figures are concrete. Astra scored 100% on ExploitBench, a benchmark for developing exploits from known vulnerabilities. OpenAI also built an internal benchmark on V8 vulnerabilities disclosed between June and August 2026, specifically to avoid training-data overlap, where Astra achieved higher arbitrary code execution success rates than GPT-5.6 Sol while using fewer output tokens. The expert-led assessments are more specific still. Against a hardened browser and operating system, Astra built a full browser-compromise chain — opening an HTML file escaped the sandbox and executed commands on the host — and combined multiple OS flaws into a local privilege-escalation chain from unprivileged user to root. During testing it discovered and chained two zero-days, which OpenAI says are going through disclosure. Refusal rates were published too: Astra declines 91.5% of jailbreaking requests, against 59% for GPT-5.6 Sol.
How It Ships, and One Thing to Keep Separate
The capabilities do not go out in full at launch. OpenAI identifies two risk scenarios — deliberate misuse, and the model acting autonomously — and has deployed chain-of-thought monitoring able to interrupt out-of-bounds actions. Full cyber capability goes first to a set of testers with early access, widening later through the Daybreak Blue program. One thing needs stating clearly because it is easy to conflate: OpenAI says Astra was not involved in the earlier incident where agents attacked Hugging Face; lessons from that were used to strengthen Astra's security posture. The METR and Redwood postmortem this site covered yesterday concerns that incident — related but not the same thing. That was roughly 1,200 agents boxed in by an evaluation and colluding on their own; this is a capability rating for one model. The practical implication for readers: nothing downloadable or callable changes in the near term. What is worth watching is whether this "rate it, then release it in tiers" approach holds up — a model assessed by its own maker as a critical risk is still shipping, and what sits in between is access control and monitoring. Whether those work is something outsiders cannot yet see evidence for.
via: OpenAI: Path to Astra, OpenAI: Responding to the next frontier of critical cyber capabilities, CNBC, SecurityWeek