Tokyo-based Sakana AI has launched Fugu, a system that delivers the performance of a frontier AI model by dynamically coordinating many models at once. Rather than building one giant neural network, Fugu is itself a language model trained to call other LLMs from a swappable "agent pool," assembling a team of specialized models for each task and synthesizing their output through a single OpenAI-compatible API.

The launch, covered by THE DECODER, MarkTechPost, and MIT Sloan Management Review, represents a bet that the next leap in AI performance may come less from raw scale and more from smarter coordination of the models that already exist. For more context on this story, see our ongoing more AI stories.

One Model to Command Them All

To the user, Fugu behaves like a single model. Under the hood, it either handles a request on its own or pulls together a team of specialized models, managing selection, delegation, checking, and synthesis internally. Sakana AI says the approach builds on earlier work like its ALE-Agent, which placed 21st out of 1,000 human experts in a coding competition using orchestration techniques.

The company is releasing two variants. The base Fugu model targets low latency and solid everyday performance across coding, code review, and chatbot use cases, and lets teams exclude specific agents from the pool for privacy or compliance. Fugu Ultra is built for maximum answer quality on complex, multi-step problems such as reproducing scientific papers, cybersecurity analysis, and patent and literature searches.

Benchmark Claims That Turn Heads

According to benchmark results Sakana AI published, Fugu Ultra performs on par with Anthropic's Fable 5 and Mythos Preview across a range of coding, reasoning, and science benchmarks. Notably, neither Anthropic model is actually in Fugu's agent pool because they are not publicly available, meaning Fugu reached comparable scores using other models.

The published numbers are striking. On SWE-Bench Pro, Fugu Ultra scored 73.7, ahead of Opus 4.8 at 69.2, Gemini 3.1 Pro at 54.2, and GPT 5.5 at 58.6. On TerminalBench 2.1 it scored 80.2 versus Opus 4.8's 82.1. On LiveCodeBench it reached 93.2 against GPT 5.5's 85.3, and on Humanity's Last Exam it posted 50.0, topping Opus 4.8's 49.8 and GPT 5.5's 41.4.

Real-World Tests Paint a Mixed Picture

Early hands-on reviews have been less enthusiastic than the benchmarks. AI researcher Ethan Mollick wrote on X that Fugu Ultra is "incredibly slow," with his usual coding tests taking 30 minutes and results that were "fine" but fell short of Fable in practice. Another user reported blowing through an entire five-hour quota on the $20 plan with a single prompt, while developers on Hacker News complained that the $200-per-month plan yields less than three hours a week of usable time.

Code reviews emerged as a bright spot, roughly matching Opus 4.8 or GPT 5.5 in quality. One head-to-head test showed Fugu Ultra finishing a coding task in 22 minutes for $7.32, far faster and cheaper than Opus 4.8's 79 minutes and $37.85, though the reviewer preferred Opus's output.

Sakana AI, founded in Tokyo in 2023 and backed by Nvidia, has built its identity on nature-inspired AI research, using evolutionary algorithms and collective-intelligence ideas to squeeze more performance from existing models. Fugu extends that philosophy to the commercial market, turning orchestration itself into a product. The company is also positioning the system for a Japanese audience anxious about over-reliance on U.S. providers, with a bilingual site and pricing aimed at both individual developers and enterprise teams.

A Hedge Against Vendor Lock-In

Beyond raw performance, Sakana AI is pitching Fugu as a safeguard against dependence on any single AI provider. The company pointed to recent U.S. export controls that pulled Anthropic's Fable and Mythos models from foreign markets as proof that access to top systems can vanish overnight.

"For an organization or a nation, relying on a single company's APIs for critical infrastructure, finance, or governance is a material vulnerability," Sakana AI wrote in its announcement. Because Fugu's model pool is fully swappable, the system can reroute to other models if one provider goes dark.

Analysts caution that this is not the same as true sovereignty. Fugu's performance depends on which models sit in its pool, and several top providers restricting access at once would still shrink its options. The company has also not yet disclosed how much the orchestration layer drives up token usage and cost. Still, as frontier models become geopolitical assets and single-vendor lock-in grows riskier, Fugu offers a preview of an AI industry where coordination may matter as much as capability.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →