skip to content
The Weighted Average

AI Economics for Operators

Copilot's Agent Bill Matches a Seat at 3,000 Credits

Microsoft's new Copilot adds metered Code and Autopilot work. At 3,000 credits, usage spending equals a $30 seat-month before base licensing.

Desktop computer and laptop beside plants and a bright office window
Desktop computer and laptop beside plants and a bright office window. Photograph by Alesia Kazantceva

Microsoft’s September 25 Copilot redesign puts Code and Autopilot on usage-based billing, so administrators should review spending policies before expanding access. At the published pay-as-you-go rate, 3,000 Copilot Credits cost $30—as much as one month of the annually billed Microsoft 365 Copilot seat, before the separate qualifying Microsoft 365 license.

A new workspace, a second bill

The interface makes the change look like a product tour. Home combines Chat and Cowork, Code builds applications, and Autopilot is a persistent agent that can keep working without waiting for another prompt. The purchasing decision is less cosmetic: which work belongs inside a predictable subscription, and which work can accumulate a separate consumption bill while the user is elsewhere?

Microsoft explicitly divides those categories. Its announcement places everyday Chat and Office productivity inside the user subscription, while Cowork, Code, Autopilot, long-running agent capabilities, and selected frontier models use usage-based billing. Home and Code are entering the Frontier rollout over the coming weeks; Autopilot is expanding to private preview at month-end. These are not universal production entitlements that every licensed tenant can use today.

The price comparison combines two independent Microsoft records. The enterprise pricing page lists Microsoft 365 Copilot at $30 per user per month, paid yearly. The Copilot Credits licensing guide sets pay-as-you-go at $0.01 per credit. Therefore $30 ÷ $0.01 = 3,000 credits. That is a budget equivalence, not an included allowance, a task quota, or a prediction that a user will consume that amount.

For an existing subscriber, this gives finance a useful reference point. Whenever an individual workflow consumes 3,000 pay-as-you-go credits, its metered spend has added another seat-month’s price. Nothing about that observation proves the workflow is uneconomic. A valuable completed task may justify considerably more. It does establish why a seat-count forecast cannot be the entire agent budget.

This is an expansion of an existing billing model, not a surprise conversion of every Copilot feature into a meter. Microsoft’s June 16 Cowork general-availability announcement already required a Copilot subscription plus usage charges. The new decision concerns more creation and autonomous work entering that model, alongside a more unified interface. Procurement should preserve that history rather than describing yesterday’s launch as the invention of paid agent execution.

The distinction also separates this story from GitHub Copilot’s model-routing preferences and discounts. Similar branding does not make the billing units, administrators, or entitlements interchangeable. A coding team’s GitHub configuration cannot be assumed to govern Microsoft 365 Code or Autopilot. Ask which product is running the work before asking whether “Copilot” has a budget. Put the product, billing account, and responsible team on the pilot approval so the eventual invoice can be reconciled with the work that was authorized.

The policy can expand before the budget does

The most consequential setting may already exist. Microsoft’s usage-based billing overview says spending policies enable “Auto-apply new services” by default. Newly supported services can inherit a policy as they become available. Administrators can turn that setting off when they want to review additions first. That makes policy scope a live purchasing choice, not merely a technical preference buried beneath the launch.

Automatic coverage can be sensible. A mature team may already have owners, tested limits, and an approval process that should apply to new workloads without creating another queue. But a policy designed for a narrow Cowork pilot may not express the intent of a broader Code or persistent-agent rollout. The recommendation is to inspect the actual policy, not to assume that every tenant is exposed or that automatic coverage is inherently wrong.

The detailed cost-management documentation describes organization and user limits, threshold notifications, and credit-request routing. Those controls should be configured around an accountable workflow owner. Alerts answer who gets told; limits and access rules answer what can continue. Before production use, test how the selected service behaves at its configured boundary rather than inferring interruption semantics from the existence of a dashboard.

There is a second budget trap in prepayment. The licensing guide offers 300,000 prepaid credits at a 5% discount, with unused credits expiring after the annual term. At the published one-cent pay-as-you-go rate, the undiscounted amount is $3,000; the discounted purchase is $2,850. To match that spend on pay-as-you-go, the organization must actually use 285,000 credits. Thus 95% utilization is the simple break-even point, before financing costs or negotiated terms.

Prepayment should follow measured use, not a desire to make the budget look controlled. A discounted pool can lower the unit price while raising the cost of useful work if much of it expires. Conversely, a stable, observed workload may justify buying ahead. The calculation tells finance what utilization must be defended; it does not decide whether every department should share one pool or how internal chargeback should work.

Application hosting broadens the question further. Microsoft’s Copilot Managed Runtime announcement introduces public-preview hosting inside the Microsoft 365 tenant boundary. It describes Entra identity, organizational policies, Git-backed source and version control, and a central app inventory. That could reduce duplicated platform work. It does not remove the need to own an application’s lifecycle after the original prompt has finished.

The same boundary appears in today’s Row Zero analysis of spreadsheet capacity and enterprise entitlements. A familiar interface makes adoption easier, but the operator still has to identify the permissions, hosting, model use, and commercial terms behind it. “Built in” is an interface description, not a complete cost model.

Cheap execution can still be expensive work

The strongest case for the redesign is integration. Microsoft says Code can build tenant-hosted internal tools and Autopilot has its own identity, memory, computer, and workspace. An organization already operating inside that environment may prefer a common governance surface to separate deployments assembled for each team. That preference is reasonable when the shared controls actually cover the intended work and its failure modes.

The strongest counterargument is that the same convenience can make poorly defined work easier to scale. A persistent agent that keeps finding more to do needs a stopping condition tied to a business outcome. A spending limit contains exposure; it does not establish whether the next action is useful. The workflow owner should define completion, escalation, and review before measuring how many tasks the agent can start.

Nor does the published credit price reveal a task price. Microsoft’s licensing guidance identifies models, runtime, context, and tools as Cowork cost components. A change in any component can alter consumption. The launch does not provide a universal fixed tariff for an Autopilot supplier review or a Code-generated application. Quoting one would require inventing a workload or treating an example as a promise.

Historical comparisons deserve restraint too. In June, Microsoft reported Cowork cost per prompt averaging 30–40% less than Claude Cowork with its Microsoft 365 connector in internal testing. The published methodology covered 125 test runs across 12 prompts using Opus 4.8. Those boundaries make the claim interpretable. They do not prove the same saving for September’s Code, Autopilot, or a different model mix.

A local evaluation should therefore preserve the work, not merely the prompt. Use a representative task with an existing acceptance standard. Record the final artifact, credit consumption, elapsed time, reviewer effort, and any repeated attempts. Count a rejected result as consumed work rather than quietly dropping it from the denominator. This is a proposed measurement method, not a claim that Microsoft exposes every desired field in one export.

Our analysis of Strands’ benchmark cost per accepted pass makes the same distinction between a low component price and a successful result. For Microsoft 365 workflows, the acceptance condition may be a reconciled workbook or a reviewed supplier pack rather than passing code tests. The denominator changes; the need to define it does not.

There are also reasons to delay. A team with a stable process may find that preview access, permissions work, and output review consume more effort than the automation saves. Another may need an audit or data-access behavior that is not demonstrated in its tenant. Those are legitimate deployment constraints, not resistance to progress. Evidence that would overturn the cautious verdict is repeatable accepted output at a lower total operating cost under the required controls.

Give every agent an owner and an exit

The first adopters should be organizations with licensed users, well-bounded recurring work, and administrators who can observe consumption. Their next step is a controlled expansion of an existing workflow, not an organization-wide invitation to automate anything. Code is especially worth testing where a small internal application has a clear owner and a known maintenance burden; Autopilot needs an equally clear definition of when background work should stop.

Separate migration approval from model enthusiasm. Microsoft’s launch roadmap places several capabilities in Frontier or private preview. A preview can answer whether the approach fits. It should not be described internally as a completed replacement for an established production process. Keep the existing route available until the new one meets the same acceptance and recovery requirements.

Finance should start with pay-as-you-go when usage is poorly understood, then revisit prepayment using measured annual demand. The minimum prepaid tier’s 95% utilization threshold is a useful discipline precisely because it is not glamorous. It asks whether the organization expects to use what it buys. That question should be answered before a discount becomes a reason to invent more agent work.

Infrastructure buyers face a parallel problem in Nscale’s distinction between closing cash and later financing commitments. Promised resources and useful resources belong in different columns. With Copilot, licensed access, allocated credits, consumed credits, and accepted outcomes are likewise separate facts. Keeping them separate is how a pilot becomes an operating system rather than a collection of impressive demos.

The rollout checklist is short:

  • Microsoft 365 administrators: inspect existing spending policies, especially automatic coverage of new services. Confirm eligible users, the named owner, and the behavior at configured limits before enabling broader work.
  • Finance and licensing teams: budget the subscription, qualifying base license, and metered usage separately. Use 3,000 credits per $30 as a comparison point, not an allowance; buy annual pools only against defensible utilization.
  • Workflow owners: define the accepted artifact, review responsibility, and stopping rule. Include failed attempts and human correction when deciding whether automation pays.
  • Developers and platform teams: test the managed runtime’s deployment, permission, inventory, and recovery behavior on a bounded internal app before treating hosting as solved.

The verdict is to adopt selectively, not to avoid the platform. Microsoft has made the commercial boundary unusually explicit: routine assistance and delegated execution are different kinds of spending. Buyers should be equally explicit about the evidence that earns expansion. An agent budget needs a stopping rule before it needs a bigger model.

Sources