skip to content
The Weighted Average

AI Economics for Operators

Gemini 3.7 Flash Halves Google's Workhorse Price

Google's Gemini 3.7 Flash cuts output prices 50% three weeks after 3.6, while DeepSWE scores jump 16 points. The workhorse war is now about cost per task.

A laptop screen displaying colorful lines of code
A laptop screen displaying colorful lines of code. Photograph by Mohammad Rahmani

Google shipped Gemini 3.7 Flash today at $0.75 per million input tokens and $3.75 per million output tokens—half the list price of the 3.6 Flash it replaces, which launched exactly three weeks ago. The speed of the markdown matters more than its size: a vendor that reprices its workhorse twice in one month is telling you where the agent economy’s margin now lives, and operators routing production traffic should re-run their cost-per-task math this week, not next quarter.

The fastest markdown in the workhorse wars

The raw performance claims are substantial. Google reports 3.7 Flash at 65.3% versus 49.0% for 3.6 Flash on DeepSWE v1.1, 43.6% versus 34.4% on FrontierCode 1.1 Main, an Elo of 1,588 versus 1,538 on the WebDev Arena leaderboard, 34.0% versus 22.0% on the GDP.pdf document-comprehension eval, and 30.4% versus 17.0% on AutomationBench, its enterprise-workflow measure. These are vendor-reported numbers on launch day, and the usual discount applies. But the direction is consistent across five unrelated harnesses, which is more than most point releases offer, and the claims target exactly the workflows—debugging, issue resolution, multi-step tool use—that decide production agent bills.

The price is the story. When 3.6 Flash arrived on July 21 at $1.50 and $7.50 per million tokens, we read it as Google turning Gemini Flash into an agent factory—a routing portfolio where cost per completed job displaces cost per token. Three weeks later, Google has halved the workhorse tier itself. The introductory rate holds through the end of the year, per the announcement’s own footnote, which makes this a five-month pricing commitment rather than a weekend promotion. It also retroactively discounts every architecture decision made on 3.6 Flash economics in July.

The trajectory deserves a sentence of its own. When 3.6 Flash launched in July, its $7.50 output rate undercut the $9 that 3.5 Flash commanded; today, three weeks later, the rate is $3.75. That is a 58% decline from the price the Flash line carried this summer, and each step arrived with higher measured quality rather than as a clearance sale. Model vendors have cut prices before, but usually on aging inventory. Google is discounting its newest workhorse on the day it ships, which is the behavior of a company buying market share, not one clearing shelves.

Context explains the urgency. Reuters reporting on Alphabet’s August 5 leadership overhaul describes a flagship Gemini model delayed two months after internal testing showed it trailing OpenAI and Anthropic in coding—the first commercially decisive capability. The Economic Times’ account of the reshuffle records Demis Hassabis moving to a chair and chief scientist role, Koray Kavukcuoglu taking over DeepMind, two original Gemini technical co-leads departing to found a competitor, and Sergey Brin personally pressing staff to accelerate. A company that cannot yet win the frontier can still win the invoice. Flash is where Google has chosen to fight, and price is the weapon it can deploy fastest.

Sixty-two percent cheaper per point

Here is the arithmetic Google did not print. Divide output price by DeepSWE score and you get a rough cost per unit of measured coding ability. For 3.6 Flash: $7.50 across 49.0 points is $0.153 per point per million output tokens. For 3.7 Flash: $3.75 across 65.3 points is $0.057 per point. That is a 62% decline in three weeks, from the same vendor, on the same benchmark family. Input tokens, caching, retries, and tool-call overhead all modify the real bill, but the list-price trajectory is unambiguous—and it compounds with the token-efficiency gains Google claims, since a model that needs fewer output tokens per task multiplies the saving.

Google halved its workhorse price in three weeks

List price per 1M output tokens, Gemini Flash; 3.7 rate is introductory through end of 2026

3.7 Flash3.6 Flash$0$2$4$6$8$7.5$3.75
3.7 Flash3.6 Flash$0$2$4$6$8$7.5$3.75
Google blog · Jul–Aug 2026

The competitive frame sharpens it further. xAI’s Grok 4.6, launched yesterday at $2 and $6 per million tokens, scores 65.9% on DeepSWE—statistically a sibling of 3.7 Flash’s 65.3%—while matching GPT-5.6 Sol on the Artificial Analysis composite index. At $6 of output per point, Grok costs $0.091; Gemini’s introductory rate is 37% cheaper per DeepSWE point for equivalent measured performance. The Grok 4.5 price war already foreshadowed this: token efficiency, not leaderboard rank, is the battleground, and Google just moved the front line.

The stranger inversion involves DeepSeek. DeepSeek’s pricing page lists V4 Pro at $0.435 input and $0.87 output today—but on August 16 at 16:00 UTC it switches to peak and off-peak billing, with peak output at $3.96 per million and off-peak at $1.98. At peak hours, Google’s introductory Flash rate undercuts DeepSeek’s flagship on output price; off-peak, DeepSeek reclaims the cheap seat, and its cache-hit input rate of $0.0036 per million remains absurdly low for repetition-heavy workloads. The conclusion is not that Gemini beats DeepSeek on price. It is that the word “cheap” now requires a time of day, a cache-hit rate, and a workload shape—and that Google has inserted itself into a conversation DeepSeek used to own outright.

The enterprise-workflow numbers may matter more than the coding ones. AutomationBench’s jump from 17.0% to 30.4% is nearly a doubling, and GDP.pdf’s rise from 22.0% to 34.0% targets exactly the document-dense work—contracts, filings, research packets—that knowledge-work agents choke on. A model that fails seven of ten enterprise workflows is a demo; one that fails six starts to look like infrastructure. Google is clearly optimizing for the crossover point where a workhorse becomes trustworthy enough to run unattended, because unattended is where token volumes explode.

The input side of the ledger tells a subtler story. At $0.75 per million input tokens, 3.7 Flash halves the 3.6 rate too, but input prices across the industry have been low enough for long enough that the real cost driver in agent workloads is output: reasoning traces, code, retries, and the verbose intermediate artifacts agents produce before anyone sees an answer. A router that sends classification and extraction to a cheap tier while reserving output-heavy repair work for 3.7 Flash now optimizes against a materially different price surface than the one architects designed against in July.

Distribution completes the play. 3.7 Flash lands simultaneously in Google Antigravity, AI Studio, and the Gemini Enterprise Agent Platform, and it now powers Gemini Spark, the consumer agent, in the more than 160 countries where Spark operates. Google is amortizing one model across developer, enterprise, and consumer surfaces—which is exactly how you subsidize an introductory price while the flagship recovers.

The case against switching

Three honest caveats. First, every benchmark cited above comes from Google’s launch materials. The 3.7 Flash model card adds safety detail—updated CBRN and cyber-offense safeguards—but independent reproduction of the coding scores will take weeks. If Artificial Analysis or the benchmark owners measure a smaller delta, the 62% figure compresses proportionally. A halved price with flat quality is still a good deal; it is just not a rout, and procurement teams should not write the rout into forecasts yet.

Second, the introductory clock. Google says the rate holds “through the end of the year,” which means every routing decision made this autumn carries a January repricing risk. If 3.7 Flash reverts toward 3.6 levels, the cost-per-point advantage halves overnight. Operators should treat the rate as a five-month option, not a new equilibrium, and negotiate committed-use discounts against the post-promotion price, not the sticker. The vendors who remember introductory pricing longest are the ones who built on it.

Third, the strategic read cuts both ways. A vendor slashing its workhorse price weeks after a leadership purge, with its flagship delayed on coding gaps, may be buying time rather than earning share. If the delayed flagship arrives strong, Flash’s economics were a bridge; if it slips again, they were a moat-digging exercise funded by Search margin. There is also a deflationary spiral risk for buyers to consider: when every vendor cuts prices every few weeks, the rational procurement posture is perpetual waiting, and the teams that actually ship agents pay an opportunity cost for sitting on their hands. The evidence that would settle the strategy question: independent coding evals of 3.7 Flash, the flagship’s ship date, and whether the introductory rate survives into Q1. Until then, the correct posture is to harvest the discount without anchoring the roadmap to it.

What to do before the clock runs out

One more number belongs in the planning meeting. If 3.7 Flash’s introductory rate holds through December 31, a workload burning 100 million output tokens a month saves $375 monthly against 3.6 Flash list prices—modest for a startup, but multiply it across a fleet of agents and a full fiscal year and the markdown funds headcount. Price cuts at this layer of the stack are not consumer promotions; they are margin transfers from model vendors to the companies building on top of them.

The quarter’s decision tree is unusually concrete:

  • High-volume agent builders on 3.6 Flash: migrate this month. The API shape is unchanged, the evals point one direction, and the saving is 50% of output spend before any quality gain. Budget the win; do not spend it on longer chains yet.
  • Teams evaluating Grok 4.6 for long-running agents: benchmark both on your own trajectories. List prices say Gemini is 37% cheaper per DeepSWE point, but Grok’s agentic reinforcement training may convert to fewer retries on multi-hour tasks—measure cost per accepted result, not per token.
  • DeepSeek shops: model the August 16 peak window—01:00 –04:00 and 06:00 –10:00 UTC—before assuming the status quo. Batch-heavy workloads should shift off-peak; latency-sensitive traffic now has a genuine Google alternative at peak hours for the first time.
  • Everyone: watch how the savings land downstream. Cheaper workhorse inference is what lets an AI roll-up like Thrive Holdings’ $2 billion enterprise bet claim 36x help-desk speedups, and it is why customers like Harvey—now raising at a $15.5 billion valuation—appear in Google’s own launch materials. When the input gets cheaper, the application layer gets more valuable.

The workhorse has become the loss leader. The operator who treats this as a commodity price war and nothing more will misread it: it is a land grab for the default slot in every agent router, priced to end on December 31. The winners of a workhorse war are never the vendors fighting it—they are the builders who route traffic through the cheapest capable model this quarter and keep the freedom to reroute the next.

Sources