Skip to main content
OpenAI GPT-6 Astra AI model achieves 72.6% OSWorld benchmark, demonstrating efficiency and speed.

Editorial illustration for OpenAI’s GPT-6 Astra Hits 72.6% OSWorld Benchmark, Cuts Task Time

OpenAI’s GPT-6 Astra Hits 72.6% OSWorld Benchmark, Cuts...

3 min read

OpenAI put a number on its newest model Tuesday: GPT-6 Astra clears 72.6% on the OSWorld benchmark, a test that measures how well an AI can actually operate a computer rather than just talk about it. The company is calling Astra its most intelligent and aligned model to date, but the framing matters more than the label. This isn't pitched as a chatbot. It's pitched as a system that opens a browser, edits a spreadsheet, runs terminal commands and finishes a multi-step job the way a person sitting at the keyboard would, instead of typing back instructions for a human to follow.

Access is narrow. Astra has no released weights, no self-hosting path, and no public rollout yet. It's live only inside OpenAI's Trusted Access and Daybreak programs, which puts the model in front of a small set of vetted organizations rather than the general developer pool. That gating decision, and the reasoning behind it, is where OpenAI's own language gets more specific about what kind of model this actually is and why the company isn't handing it out freely.

Codex previously used compaction, summarizing earlier turns once context filled up. That process discards the detail an agent later needs: why a fix failed, which tests ran, which requirement was added early. Astra instead keeps notes across context windows and searches back into earlier messages and tool output.

Why this matters

The benchmark race is now openly contested and openly incomparable. OpenAI's 72.6% and Anthropic's 77.9% can't be lined up side by side because they ran different OSWorld releases, which means the number that will show up in headlines and pitch decks this week is closer to marketing than measurement. For developers and founders building on computer-use agents, the more useful figure is the task-time drop, 75 minutes to 40, since that's what actually shows up in a workflow or a billing cycle.

But Astra being closed, hosted, and gated behind a "critical" cyber threshold matters just as much as any score. It means no self-hosting, no weight inspection, and dependence on OpenAI's own risk classification to decide what you're allowed to build. Researchers should treat both companies' numbers as directional, not comparable, until someone runs Astra and Fable 5.1 on the same OSWorld build.

Watch for that head-to-head test. Until it happens, "most intelligent computer-use model" is a claim each lab gets to grade on its own curve.

LIVE02:41OpenAI Launches GPT-6 Astra, Its New State-of-the-Art AI Model