skip to content
The Weighted Average

Models & Open Source

MiMo V2.6 Undercuts Grok Output Pricing by 6.9x

MiMo V2.6 Pro costs $0.87 per million output tokens against Grok 4.7's $6. Test the ordinary API before buying speed or self-hosting.

a close up of a building with many windows
a close up of a building with many windows. Photograph by Willem Chan

Xiaomi’s September 22 MiMo V2.6 release prices ordinary Pro output at $0.87 per million tokens, while opening the new model series for developers. Against Grok 4.7’s base $6 output rate, that is a 6.9x unit-price gap: a reason to test Xiaomi’s ordinary API, not evidence that it completes an equivalent coding job 6.9 times more cheaply.

Cheap output is a route, not a verdict

The arithmetic is deliberately narrow: $6 ÷ $0.87 = 6.8966, rounded to 6.9x. Both inputs are dollar prices per million output tokens retrieved from the providers’ current documentation. The comparison uses Grok’s lower-context global rate and MiMo’s ordinary Pro rate. It excludes cached input, tool charges, alternate service tiers, and any difference in how much output each model generates to finish a task.

The Xiaomi announcement lists Pro’s uncached input at $0.435 per million tokens. Its cached-input rate is $0.0036. Xiaomi says the series retains V2.5 API prices, so this is a capability update at an existing tariff, not a new across-the-board price cut. That distinction matters for current users: their decision is whether the new model improves accepted work at the existing rate, not whether a cheaper token automatically warrants migration.

Speed changes the comparison. Xiaomi lists $8.70 per million output tokens and $4.35 uncached input for Pro UltraSpeed, and claims up to 20x faster output. The output price is ten times the ordinary Pro tariff and above Grok’s base $6 rate. Do not turn the ordinary model’s price advantage into a recommendation for the premium route. The fast option buys a different service characteristic and needs its own latency measurement.

That is also why this article does not calculate dollars per second saved. Xiaomi supplies an up-to speed claim, not a matched trace from the reader’s application. Time to useful completion includes tool calls, retries, and checking the result; output generation speed alone does not establish the value of the premium. For interactive work, measure the delay users actually experience. For unattended work, require a reason to pay more merely to finish generation sooner.

Availability is broader than one interface. Xiaomi’s release names its API platform, AI Studio, MiMo Code, MiMo Desktop, and OpenRouter. Vercel’s AI Gateway also publishes a MiMo V2.6 Pro model page. That provides another integration path, but not permission to assume all routes have identical terms, model settings, or accounting. Record the exact provider route in an evaluation so a subsequent deployment tests the same thing.

The official Pro-RL model card describes a one-million-token context window and text, image, video, and audio modalities. Those capabilities make the release relevant to multimodal agent builders as well as coding teams. They do not make every modality equally cheap or every long-context request equally effective. The output-rate calculation is a bounded comparison, not a complete multimedia application budget.

Open weights are not a small-machine promise

Self-hosting is a separate purchasing decision. The same model card describes a sparse mixture-of-experts architecture with 1.02 trillion total parameters and 42 billion activated. Activated parameters describe work selected during inference; they are not a replacement for the total model when planning deployment. The card also publishes distributed serving examples. An open release creates control and experimentation options without making the infrastructure free.

For most teams deciding this quarter whether to add a low-price candidate, the ordinary hosted API is the cleaner first experiment. It separates the quality question from the serving-engine question. Teams with a specific hosting requirement can evaluate the released checkpoint independently, but should not import the hosted tariff as if it were their own achievable infrastructure cost. No retrieved source supplies a universal self-hosted price for this model.

Xiaomi’s benchmark disclosures are promising but need careful labeling. Its model card reports 71.9 on DeepSWE v1.1 and 34.9 on Terminal Bench 4.0 for MiMo V2.6 Pro. The launch article separately discusses training-progress measurements. These are vendor-published evaluations with their own settings, not our reproduced results. Do not splice the strongest score from one disclosure into the cheapest deployment configuration and call the combination independently verified.

The same restraint governed our earlier Qwen Omni analysis: inexpensive model units can coexist with an expensive workflow. Today’s Grok 4.7 lead explains how context and regional settings alter even one model’s tariff. Across vendors, the comparison adds still more variables—tokenization, reasoning, tool reliability, and the amount of human correction required before a result is accepted.

A useful pilot therefore keeps the repository tasks, tests, and approval rules fixed. Compare ordinary MiMo Pro with the incumbent on accepted changes, total billed tokens, elapsed time, and interventions. Retain failures in the denominator. A cheap successful demonstration does not describe the cost of a queue containing difficult tasks, retries, and rejected patches. These are recommended evaluation criteria; this article has not run that bake-off.

The strongest case against switching is that a larger number of attempts or more review could consume the token-price advantage. The strongest case for it is that ordinary API rates leave meaningful room for a capable alternative to earn a place in routing. Neither can be settled by price alone. Repeatable accepted outcomes at lower total cost would justify expansion; persistent repair work, unacceptable latency, or unsuitable data-handling terms would reverse it.

The verdict is to add the ordinary API to a bounded evaluation, not replace the default immediately. Buy UltraSpeed only after measuring the value of its latency, and investigate self-hosting only when control or scale warrants a separate infrastructure project. Xiaomi has made output inexpensive enough to question incumbency. It has not made quality assurance optional.

Sources