Skip to main content
GPT-6 Astra AI, a red digital lock, and binary code representing its "critical" cybersecurity risk tier.

Editorial illustration for OpenAI's GPT-6 Astra Reaches 'Critical' Risk Tier for Cybersecurity

GPT-6 Astra Hits 'Critical' Cybersecurity Risk Tier

4 min read

OpenAI shipped GPT-6 Astra on Tuesday, five days after Anthropic put out Claude Fable 5.1. The company calls it the most intelligent and aligned model it has built, but the label that matters here is a different one: OpenAI's own safety framework now places Astra in the "critical" risk tier for cybersecurity, the first time an OpenAI model has landed there. That classification isn't cosmetic. It triggers additional internal review and disclosure requirements before the model can be used for certain tasks.

What pushed Astra into that tier is tied to the same capabilities OpenAI is marketing as a leap forward. This isn't a model that just writes longer or more polished answers. It operates a computer directly, filling out forms, editing CRM records, running QA checks on live websites, and fixing software by watching what happens on screen rather than being told step by step.

On OSWorld 2.0, a benchmark for desktop computer use, OpenAI reports Astra scoring 72.6%, ahead of Claude Opus 5 and GPT-5.6 Sol. The gap that stands out, though, isn't accuracy. It's speed, and what that speed implies about a model that can act largely on its own.

The simplest way to put it is this: Astra is built to do more, not just answer more. It can use a computer to complete tasks instead of telling you how to do them.

Why this matters

A 100% score on ExploitBench isn't a marketing stat, it's OpenAI telling us its own model can find and weaponize unknown vulnerabilities without human help. That's the literal definition of the "Critical" tier in its Preparedness Framework, and OpenAI hit publish anyway, less than a week after Anthropic's Claude Fable 5.1. For developers building on Astra's new task-execution and computer-use features, the calculus has changed: a model that finishes documents and runs long coding sessions unsupervised is also, by OpenAI's own admission, a model that can probe and exploit software on its own.

Founders shipping products on top of this need to ask what guardrails actually sit between "capable of independent exploit development" and their own infrastructure, because the answer isn't in a benchmark score. Researchers should be pushing OpenAI for specifics on containment, not just capability claims. The release cadence between labs is accelerating; the safety disclosures accompanying it need to accelerate just as fast, and right now the gap between the two is the story worth watching.

Common Questions Answered

What is the significance of GPT-6 Astra being classified in the 'critical' risk tier for cybersecurity?

This is the first time an OpenAI model has been placed in the critical risk tier, which triggers additional internal review and disclosure requirements before deployment. The critical classification indicates that Astra poses significant cybersecurity risks and requires special handling protocols before it can be used.

How does GPT-6 Astra's task-execution capability differ from previous OpenAI models?

Unlike previous models that only provide instructions, Astra is built to actually complete tasks by using a computer to execute them directly. This represents a fundamental shift from answering questions about how to do something to actively performing those tasks autonomously.

What does GPT-6 Astra's 100% score on ExploitBench reveal about its capabilities?

The perfect ExploitBench score indicates that Astra can find and weaponize unknown vulnerabilities without human assistance, which is the literal definition of OpenAI's Critical tier in its Preparedness Framework. This demonstrates the model's ability to identify and exploit security flaws autonomously, making it a significant cybersecurity concern.

When was GPT-6 Astra released relative to Anthropic's Claude Fable 5.1?

OpenAI shipped GPT-6 Astra on Tuesday, just five days after Anthropic released Claude Fable 5.1. This rapid release cycle demonstrates the accelerating pace of competition in advanced AI model development.

LIVE03:15OpenAI's GPT-6 Astra Reaches 'Critical' Risk Tier for Cybersecurity