What changed
Sakana AI released Fugu Ultra v2 on September 11, 2026, alongside Fugu Max, behind its OpenAI-compatible API 1. The model id is fugu-ultra-v2.0, and the fugu-ultra alias now defaults to it 2. Sakana states that Fable 5, Fable 5.1 and GPT-6-Astra are not in the v2 model pool 1. Pricing for fugu-ultra-v2.0 is fixed per 1M tokens 3:
| Token type, per 1M | Standard | Context above 272K |
|---|---|---|
| Input | $5 | $10 |
| Output | $30 | $45 |
| Cached input | $0.50 | $1.00 |
A swappable pool you do not swap
Sakana’s pitch is that a swappable pool of open and specialized models protects users from vendor lock-in, API revocations and sudden service cutoffs 1. Its FAQ settles who does the swapping. The Ultra pool is fixed, while the base Fugu model lets you opt out of specific models from the console 4. Sakana does not expose which models answered a request, by design 4. I think that turns the lock-in argument around. You trade a frontier vendor you can name for a pool you cannot list, and when Sakana moves a model in or out, your outputs can shift with nothing in the response to tell you why.
No harness, no reproduction recipe
Sakana claims the best or joint-best score on five of eight benchmarks, in its own evaluation 1. OrcaRouter, a competing router that does not carry Fugu, notes that Sakana has released no evaluation harness, no per-task score grid and no reproduction recipe, and that there is no Artificial Analysis page for any Fugu model 5. It also relays that independent testers flagged latency variance: on easy tasks, the coordinator’s deliberation is pure overhead 5. With the pool hidden, I see no way to rerun those scores on your tasks.
Impact on your team
If you already call Fugu, pin fugu-ultra-v2.0 instead of the fugu-ultra alias, which now defaults to v2 2. Sakana calls the upgrade a single-line parameter change 1, and an alias makes it for you. Before any pilot, run evals on the tasks you serve and log latency next to cost for each call. I would budget on output, billed at $30 per 1M tokens and $45 above 272K of context 3. Treat Sakana’s benchmark claims as a hypothesis to test, and keep a fallback wired to a model you call directly, because the day quality drifts, you will not be able to see which models answered 4.