Editorial illustration for OpenAI’s Astra Model Grants Daybreak Partners Early, Less-Restricted Access
OpenAI’s Astra Model Grants Daybreak Partners Early,...
OpenAI told reporters Tuesday that its next model, code-named Astra, has crossed a line the company itself drew: it can find and exploit previously unknown software vulnerabilities without human help. That puts Astra at what OpenAI calls the "critical" cyber capability threshold in its preparedness framework, the internal rulebook meant to flag when a model's abilities carry real-world risk. It's the first time OpenAI has said one of its models hit that mark.
The company says it plans to release a public version of Astra "soon," but the model's sharpest cyber abilities won't ship to everyone at once. Those will go first to a smaller set of partners inside OpenAI's Daybreak Blue early-access program, a narrower rollout than the company's usual launch pattern.
OpenAI also disclosed that it paused parts of Astra's training for several weeks once the critical threshold became clear, along with training on a second, unnamed future model. Safety and security leaders say that pause is over. New controls are now in place, and the company says it's confident enough in them to move forward with a broader release.
In a briefing with reporters, OpenAI safety and security leaders said the company has concluded that Astra reaches the critical cybersecurity capabilities outlined in its preparedness framework, which sets thresholds and protocols for when its AI models pose new levels of risk. The company says an AI model has reached its critical cyber threshold when it can independently find and exploit previously unknown vulnerabilities in real-world software.
Why this matters OpenAI drawing a hard line between "critical" cyber capability and everything before it is a bigger deal than the Astra release itself. This is the first time the company has admitted a model crosses into territory it considers genuinely dangerous, and its answer is to hand the sharper version to Cisco, Cloudflare, and Palo Alto Networks before anyone else. That's a defensible call: infrastructure providers need lead time to patch against tools that can also be turned into attack vectors.
But it also means access to frontier capability is becoming a function of who you already are, not what you're building. For independent researchers and smaller security teams, the public release of Astra will arrive deliberately hobbled, while a handful of well-connected partners get the real thing. Worth watching: how OpenAI defines "soon," whether Daybreak Blue expands beyond its current roster, and whether this tiered-access model becomes the template for every future capability threshold the company decides is too sharp for general release.
Further Reading
- OpenAI to pause some work on AI model Astra due to security concerns - The Guardian
- OpenAI flags possible critical cybersecurity risk in upcoming model - Reuters
- OpenAI Astra model raises cyberattack concerns - CNBC
- OpenAI slows release of Astra model citing cyber capabilities - Axios
- OpenAI Pauses Development on Powerful Astra Model Over Autonomous Cyberattack Risks - Security Boulevard