Editorial illustration for Sakana Claims Fugu Ultra v1.1 Outperforms Fable 5 in Own Benchmarks
Sakana's Fugu Ultra v1.1 Outperforms Claude in Benchmarks
Sakana AI put out Fugu Ultra v1.1 this week, the latest version of its router that spreads queries across a pool of top-tier public AI models rather than relying on one system. The company says the update gains up to 7.9 points over v1.0, with the sharpest jumps showing up on ProgramBench and TerminalBench 2.1. Sakana also claims v1.1 now beats Fable 5, a model that isn't even in Fugu's selection pool.
That claim is worth pausing on, and it's the kind of thing that invites scrutiny. Every number here comes from Sakana's own testing. Nobody outside the company has verified any of it yet. Pricing hasn't moved, still $5 per million input tokens and $30 per million output tokens, and Sakana says new top-tier models take about two weeks of training and evaluation before they're folded into the router's pool.
The release adds a Claude Code-compatible endpoint, letting developers call Fugu straight from the terminal. Fugu has been available since launch on OpenRouter and Vercel, though the original version landed with mixed reviews, criticized for heavy token usage, sluggish response times, and shaky output quality.
Sakana AI has released Fugu Ultra v1.1, an update to the AI router that distributes each query across a pool of publicly available top-tier models. The company claims performance gains of up to 7.9 points over v1.0, with the biggest jumps on ProgramBench and TerminalBench 2.1. Fugu v1.1 reportedly beats Fable 5, even though Fable 5 isn't part of the router's selection pool.
Why this matters A router beating a model that isn't in its own pool is a marketing sentence before it's a technical one. Sakana is comparing an aggregator against a single frontier model and calling it a win, but a router's whole job is to pick from whatever's available, so of course it can outscore any one competitor if you point it at the right benchmarks. The two-week lag before new models get added is the more useful detail here: it tells you Fugu is always trailing the frontier by design, not surpassing it.
For developers pricing out $5/$30 per million tokens against a direct API call to whatever model Fugu is routing to, that markup needs to buy something concrete, lower latency, better fallback handling, actual reliability data. None of that is in this release. Until someone outside Sakana runs ProgramBench or TerminalBench 2.1 independently, treat the 7.9-point gain as a company grading its own homework.
Worth watching if third-party benchmarks show up, or if Sakana publishes which specific models sit in the pool at any given time. Right now there's no way to check the claim, only to repeat it.
Common Questions Answered
What is Fugu Ultra v1.1 and how does it differ from traditional AI models?
Fugu Ultra v1.1 is an AI router developed by Sakana AI that distributes queries across a pool of top-tier public AI models rather than relying on a single system. This approach allows it to leverage the strengths of multiple models simultaneously, making routing decisions to optimize performance across different types of queries.
What performance improvements does Fugu Ultra v1.1 claim over the previous version?
According to Sakana AI, Fugu Ultra v1.1 gains up to 7.9 points over v1.0, with the most significant improvements appearing on ProgramBench and TerminalBench 2.1 benchmarks. These gains represent a notable update to the router's query distribution capabilities.
Why is Sakana's claim that Fugu v1.1 outperforms Fable 5 considered noteworthy?
Fable 5 is not even included in Fugu's selection pool of models, making this comparison significant from a marketing perspective. However, critics note that comparing a router that aggregates multiple models against a single frontier model may not be a fair technical comparison, since a router's job is to select the best available option for each query.
What limitation does Fugu Ultra v1.1 have regarding frontier models?
Fugu Ultra v1.1 operates with a two-week lag before new frontier models are added to its selection pool, meaning it is always trailing the cutting edge by design. This delay is built into the router's architecture and represents a fundamental constraint of how the system stays updated with the latest AI developments.
Further Reading
- Sakana Fugu vs Fable 5: Benchmarks & Verdict (2026) - AY Automate
- Sakana Fugu: Features, Benchmarks, and How It Works - DataCamp
- Sakana Fugu vs. Claude Fable 5: Benchmarks, Preise & ... - DataCamp
- Sakana「Fugu Ultra」徹底検証:ベンチマークはFable 5級 - AI - Qiita
- Sakana AI Fugu Review: Fugu Ultra vs Claude Fable 5 - Coursiv