AI Economics for Operators
Fugu Max Cuts Token Rates, Not Necessarily Task Cost
Sakana's Fugu Max charges 80% less per output token than Ultra v2. Count orchestration, tools, and accepted results before switching.
Sakana AI released Fugu Max and Fugu Ultra v2 on September 11, giving agent builders a cheaper orchestration tier alongside its capability-focused option. Max’s $6 per million output tokens is 80% below Ultra v2’s short-context output rate, but a lower token tariff does not establish a lower bill per accepted task.
The cheap model is a system of models
The release makes an architectural proposition rather than introducing another isolated foundation model. Sakana says Max dynamically routes work across an expanded pool of open-weight and specialized models, including NVIDIA Nemotron models. Ultra v2 targets harder multistep work. Both are available through its OpenAI-compatible API. The attraction for a builder is delegation without maintaining every routing decision; the cost is another layer whose behavior must be measured.
The arithmetic comes from two first-party disclosures. The launch gives Max’s $2 input and $6 output rates per million tokens. The console’s current pricing schedule lists Ultra v2 at $5 input and $30 output when context does not exceed 272K. Therefore, (30 − 6) ÷ 30 × 100 = 80% lower output pricing for Max. That calculation compares tariff units, not benchmark-adjusted productivity or identical generated answers.
Context changes the comparison. Ultra v2’s listed input and output prices rise above that threshold, while Max’s published rates stay fixed regardless of context length. Do not describe the 80% figure as a universal task discount: cached input, context length, internal work, and the final answer’s length all affect the bill. The clean comparison is narrower and still useful. Max deserves a trial where an expensive orchestration tier has become the default for routine work.
The billing documentation supplies the caution the launch headline does not. For Ultra, Sakana explicitly says orchestration tokens reported inside usage-detail fields represent additional real usage and enter the final price at the corresponding token rates. A client that sums only visible input and final output can undercount work. The documentation labels that explanation for Ultra; it should not be silently generalized into an undocumented Max accounting contract. Obtain a sample invoice and reconcile Max’s own returned usage before forecasting production spend.
Tools add another meter. Max’s web search and web fetch each cost $0.007 per call, according to the same pricing schedule. A model can have inexpensive output and still perform unnecessary retrieval. Conversely, a more expensive model may finish with fewer attempts. The relevant ledger records the whole task: tokens, tool charges, retries, and whether a reviewer accepted the result. None of those quantities can be replaced by the advertised output rate alone.
Sakana’s June general-availability explanation described delegation, verification, and synthesis behind a single model interface. September’s release extends that proposition toward cost efficiency. That continuity matters: the product being purchased is coordinated execution, not merely access to the cheapest member of a model pool. A transparent invoice and a reproducible success criterion are therefore part of evaluating the product, not administrative work to postpone until after adoption.
Outsourcing routing does not outsource judgment
The benchmark claims warrant testing, not dismissal. Sakana reports strong results across several tasks and identifies SWEFish as its own internal coding benchmark. Vendor-specific tasks can reflect useful work, but they are not an independent estimate of performance on a buyer’s repository. Neither a leaderboard position nor an orchestration diagram establishes how often the service will produce an acceptable migration, analysis, or review under that buyer’s constraints.
There is also a procurement distinction inside the product family. The model catalog says ordinary Fugu permits opting specific agents out of its pool, while Ultra and Max use fixed pools; specific enterprise provider configurations require contacting Sakana. A swappable architecture at the vendor does not necessarily give a customer a self-service provider exclusion. Teams with data-processing restrictions should settle that question before sending real documents, even if the initial API change is small.
Subscriptions require their own boundary. The pricing page describes monthly plans as suitable for individuals and everyday hands-on use, while pay-as-you-go tokens receive higher priority. It also says these subscriptions apply to the Sakana AI API Platform, not Sakana Chat. Do not buy a subscription expecting it to remove an unrelated consumer interface’s limits, and do not treat a personal allowance as a production capacity guarantee. The commercial mode belongs in the trial configuration.
A sensible evaluation keeps the incumbent available and sends a representative set of completed tasks through Max. Preserve the same acceptance rubric, permitted tools, and data restrictions. Record failures rather than discarding them from the cost denominator. For work that requires repair, include the repair path. This is the accounting discipline behind our earlier analysis of token mix and model pricing: unit prices become decisions only after the workload supplies the units.
Today’s AWS monitoring lead explains why evaluation itself needs a budget. That expense should be visible here too. Buying cheaper generation and then multiplying review or evaluator calls is not necessarily a mistake, but the trade needs to be intentional. Keep task-quality monitoring separate from checks that merely confirm the API returned successfully.
The strongest counterargument is straightforward: hidden coordination may save more work than it creates. If Max consistently resolves tasks that an isolated model needs several attempts to finish, its orchestration is an asset. Evidence that would strengthen the adoption case is lower invoiced cost per accepted result, stable latency, and a documented provider configuration that satisfies the customer’s obligations. Evidence that would reverse it is a worse repair burden, opaque billing, or an unacceptable fixed pool.
Switch routine workloads only after that comparison. Keep Ultra or the incumbent for tasks whose observed quality justifies the premium. Max’s tariff makes experimentation inexpensive relative to Ultra output, but Sakana’s larger claim remains a hypothesis for each operator to test: a better coordinator can beat a bigger model. The invoice and the accepted result—not the brand of model hidden inside—decide whether it did.
Sources
- Sakana AI — September 11 Max and Ultra v2 release, architecture, and Max rates
- Sakana AI console — token prices, orchestration accounting, tools, and subscriptions
- Sakana AI console — model pools, configuration constraints, and supported versions
- Sakana AI — June Fugu general availability and coordinated execution