Skip to main content
AI agent teams waste tokens for minimal quality gains; a digital brain with gears and data streams.

Editorial illustration for AI agent teams waste tokens for barely measurable quality gains, research finds

AI Agent Teams Waste Tokens for Minimal Gains

• 4 min read

Vals AI ran a straightforward test this month: take two frontier models, GPT-6 Sol and Claude Opus 5.5, and pit solo performance against multi-agent teams on something called the Vibe Code Bench. The teams burned through 1.8 to 5.1 times more tokens than a single agent working alone. For that premium, buyers got almost nothing back.

Four comparisons, one statistically significant win, and that single case, GPT-6 Sol at medium reasoning, only moved the needle by 7.3 points. At maximum reasoning effort, the extra agents did nothing measurable for either model.

The pattern held up outside Vals AI's own benchmark. Anthropic's internal testing on Opus 5.5 showed the same shrinking returns as agent count rose, with jumps from ten to 100 agents barely registering after a full day of runtime. Fable 5.1 fared better on a theorem-proving task with larger teams, yet it still trailed Opus 5.5 overall, and even dipped on a separate knowledge base task when scaled up. Across the board, bigger agent teams bought speed, not better answers, a distinction that's prompted OpenAI's Noam Brown to weigh in on what multi-agent setups are actually good for.

Out of four comparisons between teams and solo agents, only one showed a statistically significant improvement: GPT-6 Sol at medium reasoning, where the team scored 7.3 points higher. At maximum reasoning, the team setup gave neither Sol nor Opus 5.5 any real advantage. The results suggest that the extra cost of agent teams isn't worth it in most cases, especially when models are already running at full compute.

Why this matters

Vals AI's numbers are a useful gut check for anyone sold on multi-agent setups as the next step up from a single capable model. Paying 1.8x to 5.1x more tokens for one statistically significant win out of four comparisons is a bad trade by any budget standard, and the Fable 5.1 result makes it worse: scaling from 30 to 100 agents on the knowledge base task actually made output slightly worse, not better. If you're a founder pricing out agent orchestration for production work, this is a reason to ask your vendor for the same kind of controlled comparison Vals AI ran, rather than taking "more agents" as a proxy for "more capability." For researchers, the Lean theorem proving result is worth sitting with: gains appeared above ten agents, yet the swarm still lost to a single Opus 5.5 instance.

That's a specific, checkable claim, not a vague trend. Until someone shows a task where team coordination beats raw model quality at comparable cost, we'd treat agent teams as a speed tool at best, not a quality upgrade worth the token bill.

Common Questions Answered

What did Vals AI's test reveal about multi-agent teams compared to solo agents on the Vibe Code Bench?

Vals AI's test found that multi-agent teams consumed 1.8 to 5.1 times more tokens than solo agents while delivering minimal quality improvements. Out of four comparisons between teams and solo agents, only one showed a statistically significant improvement: GPT-6 Sol at medium reasoning with a 7.3 point increase, making the token cost inefficient for most use cases.

Why is the token efficiency of AI agent teams important for production deployments?

Token efficiency directly impacts operational costs and resource allocation for production systems. When multi-agent setups require nearly 5 times more tokens for negligible quality gains, the financial burden becomes unsustainable for founders and organizations pricing out agent orchestration, especially when single capable models can already run at full compute.

What happened when Vals AI scaled from 30 to 100 agents on the knowledge base task?

Scaling from 30 to 100 agents on the knowledge base task actually made output slightly worse rather than better, demonstrating that more agents do not necessarily lead to improved performance. This counterintuitive result further undermines the value proposition of multi-agent team approaches for this particular benchmark.

How did GPT-6 Sol and Claude Opus 5.5 perform differently in the multi-agent team test?

GPT-6 Sol showed the only statistically significant improvement at medium reasoning with a 7.3 point advantage in team configuration, while at maximum reasoning neither GPT-6 Sol nor Claude Opus 5.5 demonstrated any real advantage from the multi-agent team setup. This inconsistency suggests that model performance gains from agent teams are unreliable and highly dependent on specific reasoning configurations.

LIVE17:56AI agent teams waste tokens for barely measurable quality gains, research finds