skip to content
The Weighted Average

AI Economics for Operators

Gemini 3.8 Flash's Cheap Output Expires in January

Gemini 3.8 Flash's January output rate is 25% above Grok 4.6 today. Budget the expiry and measure accepted tasks before switching.

architectural photography of glass building
architectural photography of glass building. Photograph by Christian Ladewig

Google’s September 2 announcement of Gemini 3.8 Flash pairs a more deliberate coding model with an introductory price that expires at year-end. Its scheduled January output rate will stand 25% above Grok 4.6’s current rate: Google’s $7.50 per million output tokens, divided by xAI’s $6, minus one. That makes this week’s rollout decision less about capturing a discount than proving the work remains economical after it disappears.

The discount has a scheduled reverse gear

This is not a September 8 launch. The decision now is whether to promote the September 2 release into production while its promotional price can obscure the longer-term bill. Google lists $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, followed by $1.50 and $7.50 from January 1, 2027. Those are the Gemini API’s standard paid-tier rates, not a negotiated enterprise contract or a subscription allowance.

The comparison with Grok is deliberately narrow. xAI’s Grok 4.6 announcement lists $2 input and $6 output per million tokens; its faster variant costs twice as much and is excluded here. Gemini’s introductory rates undercut both ordinary Grok rates. January’s schedule splits the result: Gemini retains cheaper input but becomes more expensive on output. Neither vendor has promised that today’s competitive ranking will survive until then. This is a budget stress test against a retrieved price, not a prediction of xAI’s next tariff.

Gemini's output discount becomes a 25% premium

Standard $/1M output tokens. Gemini intro ends Dec 31; January rate is scheduled. Grok held at today's price.

Gemini 3.8 · introGrok 4.6 · currentGemini 3.8 · Jan ’27$0$3$6$9$3.75$6$7.5
Gemini · introGrok 4.6 · currentGemini · Jan ’27$0$3$6$9$3.75$6$7.5
Google Gemini API pricing; xAI Grok 4.6 announcement · Sep 8, 2026

The original calculation gives buyers a useful boundary without inventing a representative workload. Let I and O denote identical volumes of uncached input and billed output, measured in millions, for each model. January Gemini costs 1.50I + 7.50O; today’s Grok costs 2I + 6O. Equality requires 1.50O = 0.50I, or an output-to-input ratio of 1:3 , using the Google schedule and xAI rate card in its announcement. Above that ratio, Gemini’s January token subtotal is higher; below it, cheaper input compensates for dearer output.

That is an accounting boundary, not a routing rule. It excludes cache reads, storage, tools, batch discounts and different service tiers. Google’s output price explicitly includes thinking tokens, so visible answer length is not the right output counter. More importantly, equal prompts do not guarantee equal token volumes across models. The ratio tells a finance team where to investigate; it cannot tell an engineering team which model finishes a job more cheaply.

The distinction extends our earlier analysis of Gemini 3.7 Flash’s introductory price cut, rather than announcing that older discount again. Google’s current pricing page gives 3.7 the same standard rates and expiry as 3.8. Staying on the predecessor can avoid unnecessary reasoning work, but it does not escape the posted January rate change. The new decision concerns how much additional work 3.8 buys, not whether its model name preserves a permanent markdown.

Buy fewer failed tasks, not cheaper-looking tokens

Google supplies the strongest warning against a blanket switch. Its launch post says 3.8 can take extra reasoning steps and call tools iteratively, using more tokens on complex tasks. It explicitly recommends lower effort or the still-supported 3.7 Flash when compute efficiency is the primary constraint. A stronger model can therefore raise spending even before the introductory rate ends. That is not a defect if the extra work produces an accepted result; it is a defect in any forecast that assumes the token count stays fixed.

There is a concrete integration check, too. Google’s latest-model developer guide describes low, medium and high thinking effort, with medium the default, and says the minimal setting is unsupported and returns an error. An operator should inspect the request configuration before changing model identifiers. A migration that inherits an incompatible effort setting is not a free upgrade, whatever the price column says.

The strongest case for switching is nevertheless serious: extra verification may replace failed runs and manual repairs. Google positions the model for long-horizon software engineering and multi-step enterprise workflows. If those are precisely where your current agent stalls, paying for more reasoning can reduce the cost of an accepted task. A token-price premium alone cannot refute that case. Nor can a vendor’s benchmark claim establish it for your repository, approval process or latency budget.

Make the trial answer that narrower question. Replay representative tasks under the current acceptance rules; record billed input, billed output including thinking, tool charges, rejected attempts and human correction time separately. Price each observed trajectory at both the promotional and scheduled rates. Do not convert review minutes into dollars until finance supplies an actual labor-cost input, and do not quietly lower the quality threshold to make the new model win.

The evidence that would change the verdict is a repeatable reduction in total spending per accepted result at January rates, with no unacceptable latency or review burden. Conversely, longer traces without more accepted outcomes would justify keeping efficiency-first work on 3.7 or another tested route. A revised official price schedule would change the budget calculation; it would not erase the need for that evaluation.

For this quarter, the operator instructions are straightforward:

  • Complex-agent owners: test 3.8 where failures are costly; authorize broader use only after measuring accepted results at the scheduled rate.
  • High-volume routine workloads: retain the existing route until additional reasoning earns its cost. Check effort compatibility before any migration.
  • Procurement and finance: carry both price schedules, preserve an alternative model path, and distinguish token charges from the full operating bill.

The same discipline underlies today’s analysis of local-versus-cloud inference routing: placement is a decision about the whole job. Gemini 3.8 deserves an evaluation, not a permanent default bought on a temporary price.

Sources