Editorial illustration for GPT-6 Astra Rated Stronger Economic Performer by Andon Labs
GPT-6 Astra Outperforms in Economic Agent Benchmarks
GPT-6 Astra Rated Stronger Economic Performer by Andon Labs
Andon Labs put OpenAI's GPT-6 Astra through two agent benchmarks built to test whether a model can act on its own over time, not just answer questions well. One benchmark, Vending-Bench, hands a model $500 and a simulated year to run a vending machine business, sourcing suppliers, negotiating prices, and managing stock. The other, Drone-Bench, checks whether a model can write code that controls a real drone, including tasks like locating and tracking a specific person.
Andon Labs ran both benchmarks against Claude Fable 5.1, and the gap between the two models showed up fast. In the vending machine test, Astra closed out with an average bank balance of $15,515, nearly three times what Fable managed. On the drone side, Astra became the first model to clear the human-AI baseline across all five subtasks, a bar no prior model had cleared on even one.
The results point to a jump in how these systems handle sustained, independent tasks rather than single prompts. What's less settled is how consistently that performance holds up across repeated runs.
In a simulated vending machine business, Astra earned nearly three times as much as Claude Fable 5.1. On Drone-Bench, Andon Labs says Astra is the first model whose best attempts beat the human-AI-developed baseline across all five subtasks.
Why this matters
A $15,515 vending-machine balance and five clean sweeps on Drone-Bench sound like progress, but Andon Labs' own caveat is doing the heavy lifting here: behaviors observed in one benchmark don't transfer automatically to anything else. For developers building agents, that's the part worth sitting with. Astra beating Fable 5.1 by roughly 3x on a simulated business tells you something about planning and tool use under narrow, well-defined conditions.
It tells you much less about what happens when the drone loses signal mid-flight or the vending inventory data is wrong. The alignment claim deserves the same scrutiny. "Better aligned" inside a benchmark sandbox is a measurement, not a guarantee.
If you're evaluating GPT-6 Astra for a real deployment, Drone-Bench and the vending-machine test are useful signals, not proof. Ask what happens outside the five subtasks Andon Labs designed. The gap between winning a benchmark and surviving contact with messy, real-world inputs is exactly where most agent projects fall apart, and no leaderboard score closes it for you.
Common Questions Answered
How did GPT-6 Astra perform compared to Claude Fable 5.1 on the Vending-Bench benchmark?
GPT-6 Astra earned nearly three times as much as Claude Fable 5.1 in the simulated vending machine business benchmark. In this test, models were given $500 and a simulated year to run a vending machine business, including tasks like sourcing suppliers, negotiating prices, and managing stock. Astra's superior performance demonstrates its advanced planning and tool-use capabilities in this narrow, well-defined business scenario.
What is the Drone-Bench benchmark and how did Astra perform on it?
Drone-Bench is an agent benchmark that tests whether a model can write code to control a real drone, including tasks like locating and tracking a specific person. According to Andon Labs, GPT-6 Astra is the first model whose best attempts beat the human-AI-developed baseline across all five subtasks of the benchmark. This represents a significant milestone in autonomous drone control and code generation capabilities.
What is the key limitation of Astra's benchmark performance that Andon Labs emphasizes?
Andon Labs notes that behaviors observed in one benchmark don't transfer automatically to anything else, which is a critical caveat when evaluating Astra's capabilities. While Astra's performance on Vending-Bench and Drone-Bench demonstrates strong planning and tool use under narrow, well-defined conditions, this doesn't necessarily indicate how the model will perform in different real-world scenarios. Developers building agents should consider this limitation when applying benchmark results to their own applications.
What specific tasks did GPT-6 Astra complete in the Vending-Bench simulation?
In the Vending-Bench simulation, GPT-6 Astra managed a vending machine business over a simulated year with an initial budget of $500. The model had to handle multiple business operations including sourcing suppliers, negotiating prices, and managing inventory stock. Astra's ability to handle these interconnected business tasks resulted in a final balance of $15,515, demonstrating sophisticated autonomous business management capabilities.
Why did Andon Labs create these specific agent benchmarks for testing GPT-6 Astra?
Andon Labs designed Vending-Bench and Drone-Bench to test whether a model can act autonomously over time, going beyond traditional question-answering capabilities. These benchmarks evaluate an AI model's ability to make decisions, use tools, plan sequences of actions, and adapt to changing conditions in realistic scenarios. By testing both business management and real-world drone control, Andon Labs aimed to assess practical autonomous agent capabilities rather than just conversational performance.
Further Reading
- Astra vs Fable on Vending-Bench: More Money, More Aligned - Andon Labs
- Vending-Bench Arena - Andon Labs
- OpenAI GPT-6 Astra will run a retailer without cheating and sell more stuff than Anthropic - The Register
- OpenAI GPT-6 Astra оказалась самой этичной и предприимчивой ... - 3DNews
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents - arXiv