Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Fugu Max

Fugu Max is positioned as an inference-time orchestrator rather than a single pretrained model. According to its Blackbox listing, it "orchestrates our largest pool of models to push the cost–performance Pareto frontier" and "dynamically finds efficient combinations of expert agents, aiming to improve model performance and cost efficiency together." This routing-agent framing implies that Fugu Max selects and composes outputs from multiple underlying specialists on a per-query basis, which is a practical approach when latency budgets, input length, and task complexity vary widely across requests.

The same listing tags Fugu Max with reasoning, tool-use, vision, and structured-output capabilities, suggesting it is intended for multimodal assistant workloads that span text and images and benefit from function calling and JSON-constrained responses. A long context window is advertised alongside these capabilities, which fits use cases involving large documents, multi-turn agentic loops, or code-and-image reasoning. Because the design emphasis is on combining expert agents dynamically, practical strengths lie in flexible cost-quality trade-offs and broad task coverage rather than a single narrow specialization.

Vercel AI Gatewaysakana/fugu-maxfugu

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
sakana/fugu-max
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.00
Output token cost
$6.00

Limits

Output tokens
1,000,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare Fugu Max pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Fugu Max

Vercel AI Gateway

Official sourceAnnouncement

Sakana AI officially announced Fugu Max on September 11, 2026, alongside Fugu Ultra v2, positioning both as variants of the same core orchestration architecture optimized for different missions. Fugu Max is described as expanding the Pareto Efficiency Frontier by orchestrating the largest pool of open and specialized m The announcement frames orchestration as consistently outperforming isolated models, with a swappable pool of agents providing supply chain resilience by design. Fugu Max's distinct role within the family is cost-performance optimization through broader model routing, complementing Ultra v2's push for peak capability.

Vercel AI Gateway

Coverage

GIGAZINE (September 14, 2026) reports that Sakana AI announced Fugu Ultra v2 and Fugu Max as multi-model orchestration systems, with Fugu Ultra v2 outperforming GPT-6 Astra and Claude Fable 5.1 on several benchmarks. The article provides pricing tables showing Fugu Max at $2 per million input tokens, $0.25 cached input Fugu Max is described as offering a wider variety of compatible models than other series, enabling more cost-effective processing. On Terminal Bench 2.1 (agent performance measured by cost vs. score), Fugu Max achieved the highest score among comparison models while successfully keeping costs down. Both models are avai

Vercel AI Gateway

CoverageBenchmark

An explainer from eesel AI (September 14, 2026) describes Fugu Max as the fourth model in Sakana's Fugu family, following Fugu, Fugu Ultra, and Fugu Cyber, all accessible behind a single OpenAI-compatible API. The article explains that Fugu Max is "tuned for cost-performance, orchestrating the largest pool in the famil The piece clarifies the orchestration mechanism: when calling Fugu Max, the user is not talking to one network of weights but a coordinator that decides per request which model should handle the work. Sakana is characterized as a Tokyo research lab founded in 2023 by ex-Google figures, including CTO Llion Jones (a co-a

Vercel AI Gateway

Coverage

Tech Jack Solutions reports (September 13, 2026) that Sakana AI launched Fugu Max at $2 per million input tokens and $6 per million output tokens, with OpenRouter listings confirming a 1M-token context window. Sakana claims Fugu Max's output pricing is 40-60% lower than Sonnet 5, GPT 5.6 Terra, and Kimi K3, though this Both Fugu Max and Fugu Ultra v2 use the same multi-agent orchestration architecture, routing tasks across a pool of open and specialized models including NVIDIA's Nemotron family behind a single OpenAI-compatible endpoint. The article notes Fugu Ultra v2 scores 48.3 on Chartography and 74.3 on DeepSWE per Sakana's inte

OpenRouter

Coverage

DataNorth reported on September 11, 2026, that Sakana AI released Fugu Max and Fugu Ultra v2, with Fugu Max priced at $2 per million input tokens and $6 per million output tokens, which Sakana claims undercuts Sonnet 5 and Kimi K3 by 40 to 60 percent on output. Both models run as hosted APIs only and are not available The report provides a side-by-side specification table: Fugu Max has a 1M token context window, 128K maximum output, $0.25 cached input per million tokens, flat pricing regardless of context length, and no long-context surcharge. Fugu Max widens the pool with open-weight and specialist models including NVIDIA's Nemotro

OpenRouter

Coverage

The There's An AI For That directory describes Fugu Max as the cost-optimized member of Sakana AI's second-generation Fugu family of orchestration models, announced September 10, 2026. It positions Fugu Max as a learned orchestrator that routes each incoming task to the leanest model in a pool of open-weights and speci The listing reports pricing at $2 per million input tokens and $6 per million output tokens, 40-60 percent below comparable models, and claims the best overall score on six benchmarks (Terminal Bench 2.1, GPQA Diamond, AA-LCR, GDP.pdf, AutomationBench, and SWEFish). Fugu Max is accessible via an OpenAI-compatible API t

OpenRouter

Coverage

AI Weekly reported on September 11, 2026, that Sakana AI shipped Fugu Max, an orchestration engine priced at $2 per million input tokens and $6 per million output tokens that routes each request to what Sakana calls the leanest model capable of solving them. The pricing runs 40 to 60 percent below the output cost of So The report notes that Fugu Max claims the best overall score on six self-reported benchmarks including Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, and SWEFish, with no third-party evaluation cited. Sakana paired the launch with Fugu Ultra v2, a higher-capability variant built on the same routing archit

Vercel AI Gateway

Coverage

Yahoo/Forkast (September 11, 2026) frames the launch of Fugu Max v1.0 and Fugu Ultra v2 as transforming multi-agent orchestration into a standardized, API-compatible product. The article details pricing ($2/$6 per million tokens for Max, undercutting frontier models by 40-60%) and explains the underlying architecture b The piece reports that Fugu Max achieves the best overall score on six benchmarks (Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, SWEFish), while Fugu Ultra achieves best or joint-best on five of eight benchmarks. The structural argument is that when orchestration becomes model-agnostic, value migrates fr

OpenRouter

CoverageBenchmark

OpenRouter's model directory lists Sakana: Fugu Max as a distinct entry with 1M token context, $2 per million input tokens and $6 per million output tokens, dated September 11, 2026. The listing describes Fugu Max as the cost-performance model in Sakana AI's Fugu family, a learned multi-agent orchestration system that The directory entry confirms that Fugu Max supports configurable reasoning effort (high, xhigh, max), function calling, structured outputs, image and PDF input, and built-in web search and web fetch. Orchestration tokens consumed by the system are billed as standard input/output tokens. While OpenRouter is a serving ga

Videos about Fugu Max

More models around Fugu Max