Skip to main content
OpenAI's Astra model flagged as highest cybersecurity risk, with a red warning symbol over a circuit board.

Editorial illustration for OpenAI Flags Its New Astra Model at Highest Cybersecurity Risk Level

OpenAI Pauses Astra Over Critical Cybersecurity Risks

OpenAI Flags Its New Astra Model at Highest Cybersecurity Risk Level

4 min read

OpenAI has hit pause on parts of development for Astra, its next AI model, after internal tests turned up cybersecurity capabilities the company says it cannot fully account for. The decision, made "last night" according to OpenAI, marks the first time the company has flagged one of its own systems as potentially reaching "Critical," the highest tier in its Preparedness Framework. Every prior model, including GPT-5.6-Sol, topped out at "High."

The framework exists to catch AI systems capable of acting on their own in ways that could cause serious harm, and cybersecurity is one of the categories OpenAI tracks most closely. A "Critical" rating there would mean a model can independently plan and carry out cyberattacks without a human directing each step. OpenAI introduced Astra only last week, and reports had pointed to a possible release as early as next month.

Now the company says it's rolling out tighter security controls, isolated testing environments, and automated systems designed to shut down risky behavior the moment it's detected. The trigger for that response traces back to something OpenAI's own internal testers ran into first.

The decision was made "last night," according to OpenAI. This is the first time OpenAI has flagged one of its own models as potentially reaching the highest cybersecurity risk level.

Why this matters

This is the first time OpenAI has pinned "Critical" on one of its own models, and that label isn't marketing language, it's the company's own admission that Astra's cybersecurity capabilities may cross into territory where the AI can plan and run attacks without a human at the keyboard. For developers and founders building on OpenAI's stack, that's worth sitting with: the same offensive capability that trips this alarm is presumably tied to defensive and code-analysis strengths you'd want in a coding assistant. Researchers should watch what OpenAI actually changes in the "rolling out" process it mentions, not just the pause itself.

GPT-5.6-Sol topped out at "High," so whatever pushed Astra higher marks a real jump in capability, not a paperwork exercise. If Astra ships as early as rumored, the gap between "we flagged this internally" and "we shipped it anyway" will tell us more about OpenAI's actual risk tolerance than any safety framework document does. Watch the release notes closely.

Common Questions Answered

Why did OpenAI pause development on the Astra model?

OpenAI paused parts of Astra's development after internal tests revealed cybersecurity capabilities that the company cannot fully account for or control. The model demonstrated potential offensive capabilities that raised serious safety concerns during testing.

What is the significance of Astra reaching the 'Critical' risk level in OpenAI's Preparedness Framework?

This is the first time OpenAI has flagged one of its own models as potentially reaching 'Critical,' the highest tier in its Preparedness Framework, whereas all prior models including GPT-5.6-Sol only reached 'High.' The 'Critical' designation indicates that Astra may possess cybersecurity capabilities to plan and execute attacks autonomously without human intervention.

What specific cybersecurity risks does the 'Critical' classification indicate about Astra?

The 'Critical' classification suggests that Astra's cybersecurity capabilities may cross into territory where the AI can plan and run cyberattacks without requiring a human operator at the keyboard. This autonomous offensive capability represents an unprecedented risk level compared to OpenAI's previous models.

How does Astra's offensive capability relate to its defensive and code-analysis strengths?

The same technical capabilities that trigger the 'Critical' cybersecurity alarm for offensive potential are presumably tied to Astra's defensive and code-analysis strengths. This creates a dual-use dilemma where the capabilities that make the model powerful for security analysis also enable potential malicious applications.

LIVE22:57NVIDIA AI Releases NOOA: Python Framework That Turns AI Agent Into Single Class