AI Economics for Operators
Spark X2.5 Doubles Output Cost, Not Every Bill
Spark X2.5 needs about 2.14 uncached input tokens per output token to beat X2 on an unchanged workload at its launch rates.
iFLYTEK’s September 7 release of Spark X2.5 lowers the price of uncached input but doubles the output-token rate relative to Spark X2. At the launch tariff, an unchanged workload needs about 2.14 input tokens for every output token before the input saving offsets the output premium; the new model is not an automatic cheaper replacement.
Reconstructed on September 7, 2026, from records available by September 7; this holiday edition’s discovery window covers September 3–7.
The upgrade changes which tokens matter
The tariff is unusually revealing. IT Home’s dated launch report lists X2.5 at ¥1.6 per million input tokens, ¥0.24 per million cache-hit tokens, and ¥6 per million output tokens. The official iFLYTEK model catalog lists the predecessor, Spark X2, at ¥3 per million input and output tokens. Its X2.5 entry marks the new rates as a limited-time half-price promotion, against undiscounted input and output rates of ¥3.2 and ¥12. Buyers should distinguish a launch discount from a durable generational price cut.
The break-even calculation joins the predecessor’s official tariff to the new rates in the September 7 report. Let input and output volumes be measured in millions of tokens, excluding cache hits. X2 costs 3 × input + 3 × output; X2.5 costs 1.6 × input + 6 × output. Equating them gives (6 − 3) ÷ (3 − 1.6) = 2.14 input tokens per output token, rounded. Above the exact ratio, X2.5 has a lower token bill; below it, X2 does. This is a price crossover for equal token volumes, not a forecast of equal performance or equal tokenization.
Spark X2.5 doubles X2's output-token price
Yuan per million output tokens · September 7, 2026; X2.5 promotional rate
That distinction is useful for teams with measured workloads. A system that reads substantial material and returns a short answer has a different cost profile from an agent that generates long explanations, code, or intermediate results. The release changes the relative price of those activities. An operator can therefore identify promising migration candidates from existing input and output counts before running a quality evaluation. A blanket model-string replacement ignores the very information that determines whether the promotion saves money.
The capability announcement is broader than the price table. The official catalog describes a mixture-of-experts model with 293 billion total parameters and 30 billion active parameters, a 256K context window, and coverage of more than 200 languages. It says the model was trained on a fully domestic Chinese platform and improves coding and agent capabilities. Those are vendor descriptions, not independently reproduced evidence that a particular production task will improve.
AIbase’s same-day account also distinguishes the flagship release from the smaller X2.5 edge models. The shared family name should not collapse different deployment choices into one benchmark. A hosted flagship, a locally deployed small model, and a previous-generation API can differ in quality, latency, operational responsibility, and data handling even when their naming suggests an orderly upgrade ladder.
Our earlier analysis of DeepSeek’s API margin and price increase examined the buyer’s exposure to changing tariffs. Spark supplies another reason to preserve a workload-level cost ledger. The relevant unit is not the cheapest input token on the rate card. It is an accepted task, with the tokens and retries that task actually consumes.
A promotion is a pilot condition, not a contract
The first adopters should be existing Spark customers whose uncached input-to-output ratios exceed the crossover and whose data can remain within their approved deployment arrangement. They have a concrete hypothesis to test: the input saving may pay for the higher output rate without degrading accepted results. Customers with output-heavy workloads should demand a quality or completion-rate improvement large enough to justify the premium, rather than assuming a newer generation must be cheaper.
Caching changes the arithmetic and deserves its own experiment. The ¥0.24 cache-hit rate is not interchangeable with the ¥1.6 uncached rate in the launch report. The break-even above deliberately excludes cache hits because the correct comparison requires observed eligibility and hit rates for both deployments. Treating every repeated token as a billable cache hit would manufacture a saving the evidence does not establish. Capture the service’s actual usage fields before extending the calculation.
The promotion creates another clear failure condition. At the catalog’s displayed undiscounted rates, X2.5 charges more than X2 for both input and output. There is then no positive uncached token mix that makes the unchanged workload cheaper on token prices alone. Better task quality could still justify the new model, but the commercial argument would have changed. A procurement owner should record the applicable rate, the promotion’s terms, and the process for receiving notice of a change; the retrieved listing does not supply a dependable end date to put into a financial forecast.
The implementation bill is less visible but no less real. Retesting tool calls, validating structured responses, comparing multilingual output, and checking operational controls consume engineering and review time. The launch sources do not quantify those costs, so there is no defensible universal payback period. Keep the initial migration reversible and retain the old model as a comparison until accepted-task cost, rather than raw completion volume, improves.
Today’s lead on OpenAI’s distinction between agent activity and research progress makes the same accounting demand at a different scale. More work performed by an agent can mean more useful output, more attempted work, or simply more consumption. Here, even an unchanged output volume becomes more expensive. A quality comparison must accompany the price comparison if the team wants to claim an economic gain.
Evidence that would change the verdict is straightforward: a representative evaluation showing fewer retries, better accepted results, and lower total cost at the customer’s actual token mix. Evidence against migration would be longer outputs, more corrections, a loss of the promotion, or operational requirements the new service cannot meet. Spark X2.5 deserves a measured trial for input-heavy workloads. It does not deserve a universal upgrade instruction disguised as a discount.