OpenAI says its upcoming Astra model has reached the highest cybersecurity capability level defined under the company’s Preparedness Framework, making it the first OpenAI model to be classified as “Critical”.
The classification follows weeks of further testing after OpenAI warned in August that it could no longer rule out Astra reaching the threshold. The company now says the model can find previously unknown security flaws and develop ways to exploit them across well-protected systems without needing a person to direct every step.
OpenAI plans to release Astra soon, but its most advanced cybersecurity capabilities won’t immediately be available to everyone. The company has also delayed parts of the model’s development while strengthening safeguards designed to prevent both deliberate misuse and unauthorised actions by the model itself.
What Does Astra’s ‘Critical’ Cybersecurity Rating Mean?
OpenAI first introduced its Preparedness Framework in 2023 to track potentially dangerous capabilities as its models became more advanced. Previous models, including GPT-5.6 Sol, had reached the “High” cybersecurity threshold. Astra is the first to move beyond it into Critical.
For cybersecurity, that classification has a specific meaning. A model reaches Critical if it can independently find and develop working zero-day exploits across many hardened, real-world critical systems. Alternatively, it can qualify if it's able to plan and carry out new end-to-end cyber attacks against hardened targets after receiving only a high-level goal.
OpenAI says its decision was based on a combination of automated benchmarks and assessments led by cybersecurity experts. In one test, Astra achieved a perfect score on ExploitBench, a benchmark that measures whether models can develop exploits for known vulnerabilities.
However, OpenAI was concerned that public benchmark data could have appeared in Astra’s training material. It therefore created a separate internal test using 20 recently disclosed high-severity vulnerabilities. Astra achieved higher arbitrary code execution rates than GPT-5.6 Sol while using fewer output tokens, according to the company. During those tests, it also discovered and used two previously unknown zero-day vulnerabilities as part of an exploit chain.
Astra Found Unknown Vulnerabilities In Hardened Systems
OpenAI also tested Astra against a hardened browser and operating system to see what it could do beyond benchmark environments.
During the browser assessment, Astra discovered previously unknown vulnerabilities and built a working exploit chain that escaped the browser’s sandbox, an isolated environment designed to prevent malicious activity from reaching the wider system. The exploit ultimately allowed commands to run on the host computer.
Astra also found several vulnerabilities in a hardened operating system and combined them into an attack that escalated access from an ordinary user account to root, which gives the highest level of control over a system.
There is an important limit to those results. OpenAI says the published assessments show Astra operating with Daybreak Blue access, rather than the standard production configuration that will be available to ordinary users.
OpenAI Adds Stronger Safeguards Before Astra Release
Reaching the Critical threshold has also changed how OpenAI says Astra needs to be secured.
The company is preparing for two different risks. The first is familiar: someone deliberately using Astra to discover vulnerabilities or conduct cyber attacks. The second is the possibility that a highly capable model could take unauthorised actions itself, even when the person using it isn't trying to cause harm.
OpenAI says Astra has therefore been trained to refuse harmful cybersecurity requests more reliably, alongside system-level classifiers, monitoring and other controls. In the company’s cyber jailbreak evaluations, Astra refused 91.5 per cent of disallowed requests, compared with 59 per cent for GPT-5.6 Sol.
The safeguards go beyond controlling what Astra will say. OpenAI is also using additional monitoring to check the model’s reasoning and actions for potentially unauthorised behaviour and automatically stop activity when necessary. Higher-risk accounts will face tighter restrictions on the cybersecurity assistance Astra can provide.
Some legitimate work could get caught by those controls. OpenAI warns that defensive cybersecurity tasks may occasionally be slowed, paused or stopped if its systems incorrectly flag them as risky. ChatGPT and Codex users may be asked to review an action before continuing, while affected API tasks will stop.
Astra’s Most Advanced Cyber Capabilities Will Have Limited Access
OpenAI says Astra will be released soon, although it hasn't announced a specific launch date. Its most advanced cybersecurity features will initially be limited to a small group of alpha testers before access expands through Daybreak Blue for defensive security work.
The company has already changed how Astra is developed internally. Following the OpenAI-Hugging Face incident, which didn't involve Astra, OpenAI paused some frontier training for two weeks while it strengthened network controls, monitoring and isolation around its training infrastructure.
Some larger reinforcement learning runs remained paused for longer. OpenAI restarted one large frontier run on 28 August after introducing additional safety and security requirements, although some smaller experimental training remains on hold.
Astra’s Critical classification marks the first time OpenAI’s own framework has determined that one of its models has reached this level of offensive cyber capability. The company says further details about the model’s cybersecurity, safety and alignment testing will be published in Astra’s system card when the model launches.
Comments ( 0 )