Skip to main content
GPT-6 Astra AI scores 46/100 on spatial reasoning benchmark, outperforming rivals by 34 points.

Editorial illustration for GPT-6 Astra scores 46/100 on spatial reasoning benchmark, bests rival by 34 points

GPT-6 Astra scores 46/100 on spatial reasoning...

3 min read

A new robotics benchmark just handed OpenAI's unreleased GPT-6 Astra a big win over Ai2's MolmoAct2, and the margin is hard to ignore. StationeryBench, built to test spatial reasoning through five desk-object tasks, uncapping a marker, pouring out paper clips, and passing a ruler between two robot arms among them, ran both models on identical dual-arm YAM robots across 200 trials. Astra finished 7 of 100 attempted tasks outright.

MolmoAct2 finished none. On progress scoring, Astra posted a median of 46 out of 100 against MolmoAct2's 12. The full results, along with videos and code, are posted on GitHub for anyone who wants to check the numbers themselves.

The stakes go beyond bragging rights on a niche test. OpenAI has said it wants to build consumer robots eventually, and spatial reasoning, understanding where objects sit in physical space and how to manipulate them, is the piece that's historically tripped up language models pushed into robotics. StationeryBench isn't the only benchmark flagging a jump in Astra's abilities either. Researchers watching this space are starting to ask what's driving the gap, and one Cornell AI researcher has a theory about what's under the hood.

Yoav Artzi, an AI researcher at Cornell and Google DeepMind, calls Astra a "step change in spatial reasoning." On the still-unpublished REMAP benchmark, GPT-Astra reaches accuracy close to human level, though Artzi notes that "even ASTRA doesn't get to what humans do in other scenarios."

Why this matters

A model that finishes 7 out of 100 tasks is still failing 93 out of 100 tasks. That's the number worth sitting with before anyone calls this a step change. Yes, 46 versus 12 is a real gap, and beating a zero-completion rival on uncapping a marker or passing a ruler is a legitimate signal that Astra's grasp of physical space is ahead of MolmoAct2's.

But StationeryBench is five tasks on one dual-arm setup, run 200 times, published on GitHub for anyone to check. That's a narrow slice to build a robotics roadmap on, and OpenAI's stated ambition to ship its own consumer robots means these numbers will get read as more than they are. For founders building on robotics APIs and researchers benchmarking spatial models, the useful move is to look at the actual videos and trial-by-trial data rather than the headline score, and to watch whether independent labs like Ai2 or academics such as Yoav Artzi replicate the gap on different hardware.

A 46 is progress. It's not proof of anything close to reliable manipulation yet.

LIVE18:42Anthropic CEO calls for industry-wide AI safety standards