AI Economics for Operators
Copilot's New Tiers Do Not Cap Model Spending
GitHub adds three Auto routing tiers, but each bills the selected model. A 10% discount still leaves Astra output at $45 per million tokens.
GitHub’s September 14 Copilot update adds three Auto routing tiers, not three fixed-price model packages. Even with the paid-plan 10% discount, GPT-6 Astra’s default-context output rate would be $45 per million tokens when selected through Auto: the new choice changes routing preferences, not the need to measure spending.
Pick a preference, not a price ceiling
Efficiency favors cost, Balance weighs cost against quality and latency, and Intelligence favors quality. GitHub says all three use the same available model set and evaluate the individual prompt. A small task can therefore reach a smaller model even under Intelligence. Conversely, Efficiency is not a promise to use the cheapest model regardless of whether it can solve the task. The labels describe an objective, not an entitlement to a particular engine.
The rollout covers Visual Studio Code, Copilot CLI, and the GitHub Copilot app. That scope matters when prescribing a team default. The Auto documentation distinguishes task-aware routing from availability-oriented routing: some other IDE integrations select for reliability and availability without offering the same tier controls. A policy that works in one developer’s editor is not proof that every client exposes the identical decision.
The arithmetic starts with the Copilot model rate card. It lists GPT-6 Astra output at $50 per million tokens in the default context tier, with input at $10, cached input at $1, and cache writes at $12.50. Combine that published output price with the separate launch announcement’s 10% Auto discount: $50 × 0.90 = $45. This is a normalized component rate, not the cost of a completed change, a guaranteed routing destination, or a claim that a single response emits a million tokens.
A request can incur input and caching charges alongside output. The rate card also distinguishes longer-context pricing. Neither selecting Intelligence nor obtaining the discount removes those terms. The practical implication is to compare the billed token mix of accepted work, rather than announcing a ten-percent project saving from a ten-percent component discount. A different route can change both the price and the quantity consumed.
GitHub converts usage into AI credits, with one credit equal to $0.01. Included allowances and additional usage depend on the plan. That accounting layer should remain separate from the router: exhausting an allowance and choosing an expensive model are related operational events, but they are not the same setting. The documentation also preserves a separate legacy request-based model for some existing annual subscribers. Teams should establish which billing system their account actually uses before applying token-rate arithmetic to an invoice.
There is a useful design detail in the Auto guide: routing follows natural cache boundaries. GitHub says switching models mid-session has shown higher cost without sufficient quality gains. That makes this more than a dropdown that picks whichever model is cheapest at the instant of a request. It also means a useful evaluation should preserve realistic session history. Isolated prompts may miss the cache behavior that determines the price of an extended engineering task.
Make the router earn its discretion
The first adopters should be teams doing varied work whose developers currently choose one model for everything. Run a bounded comparison across maintenance, explanation, and complex implementation tasks already representative of the repository. Keep acceptance criteria unchanged. Record selected model, billed usage, latency, review effort, and whether the result was accepted. GitHub exposes the selected model beside or within responses in the supported clients, so model visibility can be part of the review rather than an inference from the prose style.
Do not score the trial by tokens alone. A cheaper response that needs more retries or human correction may cost more per accepted change. Equally, a more expensive response can be justified if it avoids substantial repair work. This is the same distinction in our Fugu analysis of token rates versus completed-task cost. The router changes the mixture; the engineering organization still has to define success.
The supported-model reference describes policy and evaluation-model boundaries. Available models can change, and Auto remains subject to plan, administrator, residency, and compliance restrictions. Individual users can disable evaluation models. GitHub specifically warns that evaluation models may perform worse on security-related or other prompt categories. A routing preference should never substitute for the organization’s permitted-model policy or its code-review obligations.
That is also the strongest argument against an indiscriminate Efficiency default. Quality failures can be sparse, expensive, and poorly represented in an average latency chart. The right test includes the work where a bad answer would be costly, not only docstrings that make the router look economical. Conversely, Intelligence should not become a ceremonial setting that nobody evaluates because its name sounds safer. Both choices need observed outcomes.
Today’s Cornelis analysis separates utilization claims from delivered infrastructure value. Copilot poses a smaller version of that procurement problem: optimization is a mechanism, not evidence that the resulting system is cheaper. Here the cost of testing includes model consumption and reviewer time; the launch provides no universal migration saving or fixed tier budget to subtract from those costs.
Switch routine work after Auto produces comparable accepted results at a better measured cost or response time. Retain explicit model selection where reproducibility, policy, or task-specific evidence makes it preferable. Reverse the rollout if model churn, retries, or review effort consume the apparent saving. Expand it if those costs remain controlled across real sessions. The new tiers are useful because they express intent; they become trustworthy only when the bill and the accepted diff agree.