Skip to main content
OpenAI's Astra AI model, a large language model, scores 72.6% on the OSWorld computer operation test.

Editorial illustration for OpenAI’s Astra AI Model Scores 72.6% on OSWorld Computer Operation Test

OpenAI’s Astra AI Model Scores 72.6% on OSWorld Computer...

3 min read

OpenAI put a number on its newest model Tuesday: 72.6% on OSWorld, the benchmark that tests whether an AI system can operate a computer the way a person would, clicking through menus, filling out forms, navigating software it wasn't specifically trained on. The model is GPT-6 Astra, and OpenAI is calling it the most capable system the company has built. President Greg Brockman went further, suggesting Astra might already meet the bar for "AGI," OpenAI's own shorthand for AI that outperforms humans at most economically valuable work.

The rollout starts narrow. Select organizations get access first through OpenAI's Daybreak program, with ChatGPT Plus, Pro, Business, and Enterprise subscribers following over the next several days, alongside API access and availability on AWS Bedrock and Microsoft Azure. Behind the number sits a training effort OpenAI says dwarfs anything it's run before: more than 100,000 GPUs at the Stargate facility in Texas, what researcher Aidan Clark described as the company's largest training run to date.

The OSWorld score matters because it's a proxy for something specific, computer use, one of the tasks OpenAI has flagged as central to what comes next.

OpenAI has shipped GPT-6 Astra, its most capable model to date. President Greg Brockman says it might already qualify as "AGI" or is at least within reach, meaning an AI system that outperforms humans at most economically valuable work by OpenAI's own definition.

Why this matters

The OSWorld jump from 65.7 to 72.6 percent matters less than the time cut, roughly 75 minutes down to 40 for the same class of tasks. That's the number that actually touches product roadmaps: if Astra holds up outside benchmark conditions, agentic workflows that felt too slow or too flaky to ship six months ago start looking viable. But we'd temper the "AGI era" framing OpenAI is reaching for here.

A 7-point gain on one benchmark, even a meaningful one for computer-use tasks, is not the same as a categorical leap, and OpenAI has an obvious incentive to declare victory given Fable 5.1 pricing is now matching them dollar for dollar. The 2.5x token cost increase over Sol is the real story for builders: teams running agents at scale need to model whether faster, more reliable task completion offsets the price jump, or whether Sol remains the better economic bet for high-volume, lower-stakes automation. Watch third-party evals once access rolls out this week, and watch whether Anthropic answers on price or capability.

LIVE22:12Meta AI's Muse Spark 1.3 Cuts Tool Calls and Tokens by ~20%