skip to content
The Weighted Average

Agentic Engineering

Cursor's Router Ships a Commit for $4.63

Cursor's model router reports $4.63 per commit against $12.69 on Fable 5 — a 2.7x spread that turns model choice into a routing decision, not a preference.

Code editor displaying React source code on a dark screen
Code editor displaying React source code on a dark screen. Photograph by Juanjo Jaramillo

Cursor has published the first credible price tag on a unit engineering leaders actually manage: the merged commit. In its launch of Cursor Router, the company reports a cost per commit of $4.63 in Balance mode and $6.76 in Intelligence mode, against $7.34 on Claude Opus 4.8 and $12.69 on Claude Fable 5, measured across online A/B tests on millions of live requests. The spread between the cheapest routed commit and the most expensive frontier commit is 2.7x — for work the company says lands at comparable user satisfaction.

The diagnosis behind the product is the more useful disclosure. Roughly 60 percent of Cursor developers pick a single model as a daily driver, so routine work runs at frontier prices and, in the company’s words, “AI spend growing much faster than output quality.” The router is a classifier trained on 600,000-plus live requests that reads each query’s complexity and domain before a model runs, and Cursor claims early-access enterprises saved 30 to 50 percent against routing everything to Opus 4.8.

The number the vendor did not compute

Combine Cursor’s commit prices with independent spend data and you get an annualized figure nobody published. Jellyfish’s August 2026 engineering report, drawn from 276,000 engineers and 99 million pull requests, puts top-decile AI spend at $813 per developer per month. If that spend were entirely Fable-5-priced commits at $12.69, it buys about 64 commits a month; at Balance-mode’s $4.63, the same budget buys about 176. That is roughly 112 additional commits per heavy user per month, or a 2.7x throughput ceiling on identical spend — the routing decision expressed as delivery, not as a discount.

The comparison set matters as much as the spread. Cursor reports that Auto Intelligence mode “lands near Fable on user satisfaction of output at about 60% lower cost for teams,” while lifting satisfaction about 15 percent over Opus 4.8 at nearly the same cost, and that Auto Balance clears Opus 4.8 on satisfaction at about 36 percent lower cost. Those are self-reported margins from a vendor selling routing, but they are measured on production traffic rather than a benchmark suite, which is the harder and more honest test.

Two caveats sit on that arithmetic. Cursor’s cost-per-commit figures come from its own A/B tests using its own satisfaction metric, and the $813 is a distribution tail rather than a typical developer. But the direction survives both: Jellyfish separately finds throughput rises with token-consumption decile and then flattens at the top, which is exactly the pattern you would expect if the heaviest spenders were buying frontier tokens for routine work.

Cursor’s own framing is honest about why it measures this way. The company argues offline evals are “limited by their small size, their distance from real-world usage, and the difficulty of reducing success to a rubric.” Offline evals “omit the extra cache-miss cost that comes from switching models,” so the company trained on a dataset where routing causes cache misses and reports savings net of them. Anyone building an in-house gateway should copy that discipline: a router that looks free in an eval harness can be expensive in a conversation.

What a router does not fix

Routing is a cost lever, not a capability lever, and its savings are bounded by the price of the floor model. That floor keeps dropping — GLM-5.3-Flash now sells Opus-class coding scores at a 42x blended price gap — which argues that the biggest routing gains are still ahead of the frontier-vendor routers, not inside them. Jellyfish measures open-weight model selection at under 2 percent of engineers but roughly doubling over six weeks, with Kimi at 51.8 percent and GLM at 51.2 percent of that small population. A router restricted to frontier vendors is optimizing inside the expensive tier of a market whose cheap tier is improving faster.

There is a demand-side ceiling too. Ramp’s August AI Index reports that Anthropic’s flagship Fable 5 took only 6 percent of tokens and 11.4 percent of dollars businesses spent with Anthropic in its first month — evidence that buyers were already refusing the top tier before any router made refusing it automatic. Routing formalizes a judgment the market had made by hand.

There is also a governance cost. A router that silently swaps models between turns makes reproducibility harder, complicates incident forensics when an agent misbehaves, and pushes model choice from an engineer’s decision into an admin policy. Cursor lets admins pin modes and block specific models, which is the right control surface, but somebody has to own it. The same discipline this paper argued for when Claude Code’s auto mode became a control plane applies here: log which model produced which diff, or you cannot debug the harness.

The vendor-risk dimension is newly sharp. OpenAI said Friday it plans to stop supplying models to Cursor from November 12 following the SpaceX acquisition, telling Reuters it cannot be confident SpaceX will use its technology within its terms of service. A router optimizing across a model pool is only as good as the pool: remove GPT-5.6 Sol, which Cursor’s own comparison uses as its cost-matched benchmark, and the Pareto frontier the product sells moves. Anthropic said it would increase compute for Claude in Cursor, which narrows the gap but concentrates the dependency — an outcome that follows directly from SpaceX turning Cursor into a vertical developer channel and from today’s lead on how policy shocks reprice the infrastructure underneath all of it.

The verdict: measure your own cost per merged commit before buying anyone’s router, because that ratio — not tokens, not seats — is the only number that connects AI spend to shipped work. If your median engineer is running a frontier model on lint fixes, the 2.7x is sitting on your invoice already. What would change this call is evidence that routed commits get reverted more often; Cursor reports a keep-rate metric but has not published keep rates by routing mode.

Sources