Sakana Fugu Ultra vs Claude Fable 5 & Mythos: Full 2026 Comparison

The headline writes itself: a Japanese startup releases an AI that "matches Claude Fable 5" without owning a frontier model. The reality is more nuanced, and for anyone deciding…

Sakana Fugu Ultra vs Claude Fable 5 & Mythos: Full 2026 Comparison
Table of contents

Sakana Fugu Ultra vs Claude Fable 5 & Mythos: Full 2026 Comparison

Last updated: June 2026

The headline writes itself: a Japanese startup releases an AI that "matches Claude Fable 5" without owning a frontier model. The reality is more nuanced, and for anyone deciding where to spend an AI budget in 2026, the nuance is the whole point. Sakana Fugu Ultra genuinely stands shoulder-to-shoulder with Anthropic's Fable 5 and Mythos Preview on several benchmarks, and it actually loses on others. More importantly, Fable 5 and Mythos sit behind US export controls and are not publicly accessible, which means for most buyers the real comparison is not Fugu versus Fable, but Fugu Ultra versus the models you can actually purchase today. This comparison gives you the honest per-benchmark scorecard, the cost-per-outcome math that subscription pages hide, and a use-case decision matrix so you can pick the right tool rather than the loudest one.

Quick answer: Fugu Ultra matches Fable 5 overall but does not beat it. Fable 5 leads on SWE-Bench Pro (86.0 vs 73.7) and Humanity's Last Exam (53.3 vs 50.0), while Fugu Ultra wins LiveCodeBench, Terminal Bench 2.1, and GPQA-Diamond. Since Fable 5 and Mythos are export-controlled, the practical choice for most buyers is Fugu Ultra versus Opus 4.8, GPT-5.5, or Gemini 3.1 Pro.

What is actually being compared

A quick orientation for readers who have not met the contenders. Sakana Fugu Ultra is the high-quality variant of Sakana Fugu, an orchestration model from Tokyo-based Sakana AI. Rather than being a single large model, it is a roughly 7-billion-parameter conductor trained with reinforcement learning to coordinate a pool of other expert models and return one answer.

Claude Fable 5 and Mythos Preview are Anthropic's frontier-class models. They are the reference points Sakana chose to measure against, and they are also, as of June 12, 2026, under US export controls and not publicly accessible. That single fact reshapes the entire comparison, because you cannot buy what you cannot access.

The remaining models in the table, Opus 4.8, GPT-5.5, and Gemini 3.1 Pro, are widely available and are the realistic alternatives most teams will weigh against Fugu Ultra.

The honest benchmark scorecard

Most coverage compresses this into "Fugu matches Fable 5." That is true at a high level and misleading in the details. Here are the published numbers, with the winner of each row marked.

Benchmark Fugu Ultra Fable 5 Opus 4.8 GPT-5.5 Gemini 3.1 Pro Winner
SWE-Bench Pro 73.7 86.0 69.2 58.6 54.2 Fable 5
Terminal Bench 2.1 82.1 80.4 74.6 78.2 70.3 Fugu Ultra
LiveCodeBench 93.2 89.8 87.8 85.3 88.5 Fugu Ultra
CharXiv Reasoning 86.6 86.1 84.2 84.1 83.3 Fugu Ultra
GPQA-Diamond 95.5 92.0 93.6 94.3 Fugu Ultra
Humanity's Last Exam 50.0 53.3 49.8 41.4 44.4 Fable 5

The pattern is clear. Fugu Ultra wins four of the six rows, including the live coding and graduate-level science reasoning tests, and it leads the entire available field on GPQA-Diamond at 95.5. But Fable 5 wins the two arguably most consequential rows: SWE-Bench Pro, the standard measure of real-world software engineering, by a substantial 86.0 to 73.7, and Humanity's Last Exam, a broad hard-reasoning test, 53.3 to 50.0. The CharXiv 86.1 figure, incidentally, is attributed to Mythos Preview rather than Fable 5.

So the accurate verdict is: Fugu Ultra matches the frontier, beats the rest of the accessible field on most tests, and trails Fable 5 specifically on practical software engineering. One important caveat from independent reviewers is that the Fable 5 and Mythos numbers are provider-reported, not the product of direct head-to-head testing, so the cross-vendor gaps should be read as approximate.

If you want to develop your own intuition for how these models differ in practice, you can try several of them through Truescho's free AI tools before committing to any subscription.

The real story: access asymmetry

Here is the part that most comparison articles bury. You cannot buy Fable 5 or Mythos. They are export-controlled and not publicly accessible, and they are not in Fugu's pool either. So a benchmark table pitting Fugu Ultra against Fable 5 is, for the typical buyer, a theoretical exercise. The practical decision is between Fugu Ultra and the models you can actually access: Opus 4.8, GPT-5.5, and Gemini 3.1 Pro.

Reframed that way, Fugu Ultra looks strong. Against the accessible field it leads SWE-Bench Pro (73.7 vs 69.2 for the next best, Opus 4.8), Terminal Bench, LiveCodeBench, CharXiv, and GPQA-Diamond. In other words, Fugu's pitch is not really "we beat Anthropic." It is "we deliver near-Fable performance using only models you can legally and publicly obtain." That is a more honest and, for many buyers, more relevant claim than the headline suggests.

This also answers a common question: why isn't Fable 5 in Fugu's pool? Because it is export-controlled and not available to orchestrate. Fugu reaches frontier-level results by coordinating other public models, not by secretly tapping the restricted ones.

Cost per outcome: the math the pricing page omits

Benchmark wins mean little without the bill attached. Fugu Ultra's pay-as-you-go pricing is $5 input / $30 output per million tokens at standard context, rising to $10 / $45 for contexts of 272K tokens and above. On paper that is competitive. The catch is fanout.

Because Fugu orchestrates multiple models per request, a single Ultra call can consume 4 to 6 times the tokens of a direct single-model call, and may spawn 3 to 5 parallel specialist calls. That multiplier is invisible on the pricing page but very visible on your bill. The heaviest tasks can cost up to about $10 per message, and some early users on the $20 subscription tier reported exhausting it in under five hours of active use.

Here is a simplified, illustrative way to think about it. Suppose a hard task would take a direct Opus 4.8 or GPT-5.5 call roughly 100K input and 50K output tokens. Through Fugu Ultra, the same task fans out, so you might actually pay for something closer to 400K to 600K tokens of combined model work, plus Sakana's orchestration margin on top, since Sakana pays full provider rates and adds its own markup. The result is that for a single hard problem where quality is paramount, that premium can be worth it, you may genuinely get a better answer. For a simple problem repeated thousands of times a day, you are paying a multiple for quality you did not need.

Workload Better choice on cost Better choice on quality Practical pick
Hard, multi-step research or security analysis Direct model Fugu Ultra Fugu Ultra
Real-world software engineering (PRs, large refactors) Direct model Fable 5 (if accessible), else close call Fugu Ultra or Opus 4.8
High-volume customer chat Direct model (Gemini/GPT) Marginal for Fugu Direct model
One-off complex tasks where being wrong is expensive Direct model Fugu Ultra Fugu Ultra
Code review on critical code Direct model Fugu Ultra Fugu Ultra

For teams that conclude direct access is the better fit, especially for high-volume work where orchestration overhead is wasted, the Truescho shop provides Claude Max, ChatGPT Plus, and Gemini subscriptions with predictable pricing.

Fugu Ultra against each accessible model

Because the headline matchup against Fable 5 is largely academic for buyers, it is worth looking at how Fugu Ultra stacks up against the three models you can actually obtain. Each comparison tells a slightly different story.

Fugu Ultra vs Opus 4.8. This is the closest practical contest for software-engineering buyers. On SWE-Bench Pro, Fugu Ultra leads 73.7 to 69.2, a meaningful but not enormous gap, and on Humanity's Last Exam the two are nearly tied at 50.0 versus 49.8. Fugu pulls clearly ahead on Terminal Bench (82.1 vs 74.6), LiveCodeBench (93.2 vs 87.8), and GPQA-Diamond (95.5 vs 92.0). The catch is cost: a direct Opus 4.8 call is a single, predictable charge, while Fugu Ultra's answer arrives with fanout attached. For one-off hard problems Fugu's quality edge is real; for repeated workloads, Opus 4.8's predictability often wins.

Fugu Ultra vs GPT-5.5. Here Fugu Ultra's advantage widens. GPT-5.5 trails substantially on SWE-Bench Pro (58.6) and Humanity's Last Exam (41.4), and Fugu leads on every coding and reasoning benchmark in the table. GPT-5.5 remains a strong, fast, well-supported general model with a mature ecosystem and predictable pricing, which keeps it the better pick for high-volume product work, but on raw hard-problem quality Fugu Ultra is ahead.

Fugu Ultra vs Gemini 3.1 Pro. Gemini 3.1 Pro is competitive on GPQA-Diamond (94.3, second only to Fugu's 95.5) and respectable on LiveCodeBench (88.5), but it lags on SWE-Bench Pro (54.2) and Terminal Bench (70.3). Gemini's strengths, large context handling and tight integration with Google's stack, are not captured in this benchmark set, so the table understates its practical appeal for certain document-heavy and multimodal workflows. For pure coding and scientific reasoning, Fugu Ultra leads.

The throughline: against the accessible field, Fugu Ultra is consistently at or near the top on benchmark quality. What it cannot offer is the cost predictability and operational simplicity of a single direct model. That trade defines the decision.

A realistic decision scenario

Picture an engineering organization choosing a default model for 2026. They run roughly three workloads: a high-volume internal chatbot (thousands of simple queries a day), automated pull-request review on a busy monorepo, and a small number of deep investigations a week (security incidents, hard architectural questions, reproducing a research result).

A single-model strategy forces one compromise across all three. Pick a frontier model and overpay massively on the chatbot; pick a cheaper model and underperform on the deep work. The orchestration-aware answer is to segment. Route the chatbot to a cheap, fast direct model where Fugu's fanout would be pure waste. Keep routine pull-request review on a mid-tier direct model. Reserve Fugu Ultra for the small slice of deep, high-stakes work where its verification stage and multi-model coordination justify the premium, and where being wrong is genuinely expensive. The same beta-tester observation, an orchestrator surfacing "more than twenty" issues where other tools flag about three, is exactly the kind of payoff that matters on a security-sensitive review and is irrelevant on a trivial diff.

This segmentation also delivers a resilience benefit. By keeping more than one provider in active use, the organization is not exposed to a single point of failure if a model is restricted, repriced, or taken offline, the very scenario the June 2026 export controls made concrete.

Use-case decision matrix

Stripping it down to recommendations:

  • Coding and code review (high stakes): Fugu Ultra's verification stage shines here. One beta tester noted that "where other tools flag about three issues, Fugu surfaced more than twenty." Worth the premium for critical code; overkill for trivial diffs.
  • Software engineering at scale (SWE-Bench-style): This is Fable 5's strongest territory (86.0), but Fable 5 is not buyable. Among accessible options, Fugu Ultra leads at 73.7, with Opus 4.8 close behind. Test both on your own repository before deciding.
  • Research reproduction and scientific work: Fugu Ultra is purpose-built for this, with strong GPQA-Diamond (95.5) and AutoResearch results. A strong pick when answer quality dominates cost.
  • Cybersecurity analysis: Another Ultra sweet spot, where multi-model verification reduces single-model blind spots.
  • High-volume chat and routine tasks: Use a direct model. Fugu's fanout makes it the wrong economics here, and the base Fugu (not Ultra) is the only orchestration variant worth considering, if at all.

Risks specific to choosing Fugu over a direct model

Three concerns should temper any decision to standardize on Fugu Ultra.

First, opacity. You cannot see which underlying model answered, and the pool composition is undisclosed. For regulated buyers who must document which system processed sensitive data, this is a real obstacle, partly mitigated by the ability to exclude specific providers from the pool.

Second, the quality ceiling. Fugu is only as good as the public models it can reach. If those models degrade or become unavailable, Fugu's advantage erodes with them. You are buying distributed dependence, not independence.

Third, latency. Coordinating, verifying, and synthesizing across multiple models is slower than a single call. For interactive products, that delay is a genuine product cost.

A fourth, quieter concern is benchmark provenance. The Fable 5 and Mythos figures Sakana compares against are provider-reported rather than independently re-run head-to-head, and the broader public numbers for Fugu Ultra are scattered across several outlets that published slightly different cuts. None of this means the figures are wrong, but it does mean a serious buyer should not treat any single benchmark table as the final word. The only number that matters for your decision is how Fugu Ultra performs on your own representative tasks, ideally measured against a direct model on the same prompts, with the real, fanout-inclusive cost recorded alongside the quality. A short, honest internal pilot will tell you more than any external leaderboard.

What to test before you commit

If you are evaluating Fugu Ultra against a direct model, run a focused trial rather than trusting the launch coverage. Pick ten to twenty tasks that genuinely represent your hardest, highest-value work, the kind where a better answer has measurable business value. Run each task through Fugu Ultra and through your leading accessible model, Opus 4.8 or GPT-5.5, on identical prompts. Record three things for every task: answer quality as judged by a domain expert, end-to-end latency, and the actual cost including Fugu's token fanout. The pattern that usually emerges is that Fugu Ultra wins on quality for a subset of genuinely hard problems and loses on cost and speed for everything else. That subset is your real use case for Fugu, and everything outside it belongs on a cheaper direct model.

Frequently asked questions

Is Sakana Fugu Ultra better than Claude Fable 5?

Not overall. Fugu Ultra matches Fable 5 and beats it on LiveCodeBench, Terminal Bench, and GPQA-Diamond, but Fable 5 leads on SWE-Bench Pro (86.0 vs 73.7) and Humanity's Last Exam (53.3 vs 50.0). Sakana itself only claims to stand shoulder-to-shoulder, not to surpass Fable 5.

Fugu Ultra vs Mythos Preview: which wins?

The two are close, both sitting at the frontier. The CharXiv Reasoning figure of 86.1 is attributed to Mythos Preview, just below Fugu Ultra's 86.6. As with Fable 5, Mythos is export-controlled and not publicly accessible, so for most buyers the comparison is theoretical rather than a real purchasing choice.

Is Fugu cheaper than Claude or GPT-5.5?

For simple tasks, usually not. Fugu's pay-as-you-go rates look competitive at $5 input / $30 output per million tokens, but each request can fan out to 4 to 6 times the tokens of a direct call, plus an orchestration margin. For complex tasks where quality matters most it can be worth it; for high-volume simple work, a direct model is cheaper.

What are Fugu Ultra's benchmark scores?

Published figures include SWE-Bench Pro 73.7, Terminal Bench 2.1 82.1, LiveCodeBench 93.2, CharXiv Reasoning 86.6, GPQA-Diamond 95.5, and Humanity's Last Exam 50.0. It also reported a best AutoResearch result of 0.9774 bits-per-byte and solved a Rubik's Cube test in 19 steps, the fewest among four models tested.

Why isn't Fable 5 in Fugu's agent pool?

Because Fable 5 is under US export controls (since June 12, 2026) and is not publicly accessible. Fugu can only orchestrate models it can reach, so it matches Fable 5's benchmark level by coordinating other available models such as GPT-5.5, Opus 4.8, and Gemini 3.1 Pro.

Is Fugu Ultra worth it for businesses?

It depends on the workload. For complex, high-stakes tasks, research reproduction, security analysis, critical code review, the quality and resilience can justify the premium. For high-volume simple queries, the token fanout makes it structurally expensive, and a direct model is the better economic choice.

Can I even access Fable 5 and Mythos in 2026?

Not publicly. Both have been restricted by US export controls since June 12, 2026, and are not generally available. This is precisely why the realistic comparison for buyers is Fugu Ultra against accessible models like Opus 4.8, GPT-5.5, and Gemini 3.1 Pro, rather than against Fable 5 directly.

Conclusion

The clean takeaway from this comparison is that "Fugu beats Claude" is the wrong frame. Fugu Ultra matches Fable 5 overall, wins four of six published benchmarks, and trails on real-world software engineering, all while Fable 5 and Mythos remain locked behind export controls and out of reach. For the buyer who can only purchase what is actually on the market, Fugu Ultra is one of the strongest accessible options, provided you respect its cost profile and reserve it for the hard, high-value problems where its multi-model verification earns its premium. For routine, high-volume work, a direct model still wins. To understand why this orchestration approach may matter beyond any single product, see our piece on AI orchestration models, and for the full primer on the model itself, read what Sakana Fugu is.

Sources