Editorial illustration for Sakana AI launches Sakana Fugu; Fugu Ultra leads coding, reasoning and tests
Sakana AI launches Sakana Fugu; Fugu Ultra leads coding,...
The numbers do the talking: Fugu Ultra crushes four coding benchmarks, tops CharXiv Reasoning, and aces Humanity’s Last Exam. Its sibling, regular Fugu, sweeps SciCode, τ³ Banking, and Long Context Reasoning. GPT 5.5 snatches one win, MRCRv2, and that’s the only baseline left standing.
Sakana AI’s new orchestration model, Sakana Fugu, doesn’t just run another large language model. It routes tasks across a swappable pool of frontier LLMs, picking the right tool for the job. The results place Fugu models squarely alongside Anthropic’s Fable 5 and Mythos Preview, though those two stay locked away, not publicly accessible.
Early tests with nearly 500 beta users point to long, multi-step tasks as the sweet spot. One example: an AutoResearch agent that autonomously improved a small GPT’s training recipe. That’s not a headline, it’s a signal.
Fugu Ultra tops the four coding benchmarks, CharXiv Reasoning, and Humanity’s Last Exam. Regular Fugu leads SciCode, τ³ Banking, and Long Context Reasoning. GPT 5.5 wins MRCRv2, the only baseline win here.
Its Fugu models stand shoulder-to-shoulder with Anthropic’s Fable 5 and Mythos Preview. Those two are not in Fugu’s pool, since they are not publicly accessible.
Use Cases
Sakana AI ran a beta with close to 500 early users.The published examples favor long, multi-step tasks.
- AutoResearch: An agent improved a small GPT’s training recipe autonomously.
Fugu is not a leap forward in raw intelligence, it is a verdict on the architecture of AI itself. The benchmarks tell us what we already suspect: no single model owns every win. GPT 5.5 snatches one crown.
Fugu Ultra grabs the rest. Regular Fugu dominates others. The real story, though, is the orchestration.
Sakana has built a system that treats frontier models as swappable components, routing each task to its strongest engine. That is a different kind of intelligence, one that admits no monolithic answer exists. The beta users, those 500 early operators, already see the payoff.
AutoResearch autonomously tweaked a small GPT’s training recipe. That is not a parlor trick. It is a signal that Fugu can handle the messy, multi-step workflows where most LLMs stall.
The Anthropic models, Fable 5 and Mythos Preview, sit outside the pool, but Fugu does not need them to prove its point. It already matches their output on the hardest tests. This is not about one model beating another.
It is about a routing mechanism that makes the sum smarter than any single part. The future of AI might not belong to the biggest brain in the lab. It might belong to the best dispatcher on the network.
Fugu just drew the map.
Common Questions Answered
What is Sakana Fugu and how does it differ from traditional large language models?
Sakana Fugu is an orchestration model that routes tasks across a swappable pool of frontier LLMs rather than relying on a single large language model. It intelligently picks the right tool for each job by selecting from multiple models, representing a fundamentally different architectural approach to AI systems that prioritizes task-specific optimization over monolithic model design.
What benchmarks does Fugu Ultra dominate according to the article?
Fugu Ultra crushes four coding benchmarks, tops CharXiv Reasoning, and aces Humanity's Last Exam. These benchmark victories demonstrate Fugu Ultra's superior performance across diverse task categories including coding, reasoning, and comprehensive testing scenarios.
How does Sakana Fugu's performance compare to GPT 5.5 across different benchmarks?
While GPT 5.5 wins the MRCRv2 benchmark, Fugu Ultra and regular Fugu dominate the remaining benchmarks tested. Fugu Ultra captures most of the top wins while regular Fugu sweeps SciCode, τ³ Banking, and Long Context Reasoning, indicating that the orchestration approach outperforms GPT 5.5 across the majority of evaluated tasks.
What does the article suggest about the future of AI architecture based on Fugu's results?
The article argues that Fugu represents a verdict on the architecture of AI itself, suggesting that no single model can win every benchmark and that the real intelligence lies in orchestration. By treating frontier models as swappable components and routing tasks to their strongest engines, Sakana demonstrates a different kind of intelligence that acknowledges the limitations of monolithic model design.
Which benchmarks does regular Fugu excel at compared to Fugu Ultra?
Regular Fugu sweeps SciCode, τ³ Banking, and Long Context Reasoning benchmarks. While Fugu Ultra dominates coding benchmarks and reasoning tests, regular Fugu shows particular strength in scientific code evaluation, banking-related tasks, and long context processing scenarios.
Further Reading
- Sakana Fugu: A Multi-Agent Orchestration System as a Foundation Model — Sakana AI
- How Sakana trained a 7B model to orchestrate GPT, Claude and Gemini — VentureBeat
- Sakana AI Launches Fugu Multi-Agent System — Phemex News
- Sakana AI Launches Commercial Product Fugu Multi-Agent Orchestration System — KuCoin News
- Introducing Sakana Fugu AI Orchestration System Beta — LinkedIn