skip to content
The Weighted Average

AI Economics for Operators

OpenAI Starts Charging Only When the Agent Wins

OpenAI now bills some large accounts per completed task. At Salesforce's own 70% resolve rate, a $2 resolution carries just $0.43 of metered work.

A woman hands a payment card to a man across a counter
A woman hands a payment card to a man across a counter. Photograph by Helcim Payments

OpenAI has begun letting some of its largest customers pay only when its AI finishes the job, according to reporting by Kevin McLaughlin and Amir Efrati for The Information, summarized by TNW as a quiet shift to outcome-based pricing for select accounts. The arrangement is unannounced, the customers unnamed, and the prices unknown. What is knowable is the arithmetic underneath it, and that arithmetic says the vendor is taking a bet it has already run the numbers on: at Salesforce’s own published action prices and its own published resolve rate, each $2 resolution carries roughly $0.43 of metered work.

That gap is the whole story. Outcome pricing looks like generosity — the buyer stops paying for failure — but it is a margin structure, and the vendor sets the per-success price high enough that the failures it absorbs are already funded. Procurement teams reading “you only pay when it works” this quarter should read it as “we have priced our error rate, and you are paying for it in the successes.”

The number the vendors published without meaning to

Salesforce sells the two halves of the calculation on the same website. Its Agentforce pricing page lists Flex Credits at $500 per 100,000 credits, with each agent action costing 20 credits — ten cents an action — and its own worked case-management example uses three actions per case, which the page prices at $0.30. The same page lists Help Agent Resolutions at $2. Now bring in the success rate. Announcing the pay-per-resolution model, Salesforce said that on its own support site “Agentforce has handled 4.3 million inquiries and resolved 70 percent of them,” a line captured in Salesforce Ben’s account of the pricing shift.

Divide the metered cost of delivery by the success rate and you get the vendor’s real cost basis per billable event: $0.30 ÷ 0.70 = $0.43. Against a $2 charge, that is a 4.7x spread on the metered layer alone — before infrastructure, support, and the tokens Salesforce buys rather than meters. Nobody published that ratio. It is the number that tells a buyer whether an outcome price is a discount or a repackaged premium.

Run the same operation on Intercom, whose terms are unusually explicit. Intercom’s Fin outcome documentation sets one charge of $0.99 per resolved conversation, $9.99 for a sales qualification, and one outcome maximum per conversation regardless of how many actions Fin takes. Apply the same 70% resolve rate and the effective bill per attempted conversation is $0.99 × 0.70 = $0.69 — a third below the $2-per-conversation meter Salesforce originally charged for every 24-hour session whether or not anything got fixed. Outcome pricing is genuinely cheaper per attempt at these rates. It is also, per success, several times the marginal cost of the work.

The definitions are where the money hides. Intercom bills an “assumed resolution” when a customer simply leaves after Fin’s last answer, and does not bill when Fin detects frustration and escalates. Those two rules point in opposite directions: one converts silence into revenue, the other converts detected failure into a free escalation. A buyer who does not audit which bucket its traffic lands in is not buying outcomes; it is buying a classifier’s opinion of outcomes.

The seat-based comparison has quietly collapsed underneath all of this. Zendesk’s plan pricing still starts at $19 per agent per month, which is the number outcome contracts are implicitly benchmarked against, and a support desk resolving even 30 conversations a month per equivalent agent blows past it at a dollar a resolution. The seat was never a measure of work; it was a measure of headcount, and it survived because nobody could count the work. Agents can count it, which is precisely why the pricing model changed within eighteen months of the technology becoming reliable enough to bill on.

Why the largest model vendor moved now

OpenAI’s core business still sells capacity by the token, and its public API price list shows what that costs: $2.00 per million input tokens and $12.00 per million output on gpt-5.6-terra, $0.20 and $1.20 on gpt-5.6-luna, plus $10.00 per thousand web-search calls and container time billed by the minute. Every one of those lines charges for attempts. An agent that tries eight times and succeeds once bills eight times, and the customer eats the seven. The extreme case is already on the record: one developer running a hundred agents in parallel accumulated $1.3 million of OpenAI tokens in thirty days, a bill generated by attempts rather than results.

Anthropic publishes the buyer-side version of the same problem. Its Claude Code cost guidance reports that across enterprise deployments the average is around $13 per developer per active day and $150–250 per developer per month, with 90% of users under $30 a day. Every dollar in that range is an attempt-priced dollar: it accrues whether the change lands or gets reverted. Set it against the $2 resolution and the contrast is stark — a support vendor now charges only for finished work, while the coding tool with the clearest completion signal in software, a merged commit, still charges by the token.

That asymmetry became a procurement problem the moment agents started running unattended. This paper’s arithmetic on what a coding agent actually costs per merged commit found the same structure from the buyer’s side: the sticker price per token is not the price of the work, because the denominator is successes, not calls. Outcome pricing simply moves that denominator onto the vendor’s invoice.

The market had already moved without OpenAI. Futurum Group’s May analysis of outcome and hybrid pricing found 27% of enterprise software buyers prefer outcome-based structures and 43% prefer consumption, with fewer than one in five still preferring per-seat. Research director Keith Kirkpatrick’s conclusion — “Outcome-based pricing is becoming a market standard” — matters less than his structural point: vendors offering seats alone are now disqualified before evaluation begins. Zendesk went furthest, restricting billing to resolutions verified by an LLM evaluation within 72 hours, with committed rates TNW puts at roughly $1.20 to $1.50. The pattern is not confined to support software either: the same week’s disclosure that OpenAI holds $5.5 billion of SB Energy warrants for signing a 20-year lease shows the industry rewriting who bears risk at every layer of the stack, from a data-center lease priced at $688 million a gigawatt down to a single resolved ticket.

For OpenAI the shift is defensive as much as commercial. Enterprises that cannot forecast a bill run pilots forever, and the paper has tracked that stall directly: only 6% of firms report a measurable earnings impact from AI despite widespread deployment. A price that arrives only on success removes the CFO’s strongest objection at exactly the point where deals die. It also transfers real risk. Under token billing the buyer funds every failed attempt; under outcome billing the vendor does, which is why per-resolution prices cluster near a dollar rather than a cent. It is also why the vendor’s own financial position becomes the buyer’s concern: absorbing failures requires a balance sheet, and the majority of two hyperscalers’ reported pretax profit is now paper gains rather than operating cash.

Where this blows up

Three failure modes, in order of how quickly a buyer will meet them.

The first is definitional drift. A resolution is one of the few AI outputs anyone can define, and even there Intercom needs a page of rules to distinguish resolutions from abandonments, spam, and procedure handoffs. The agentic work OpenAI is actually selling — multi-step research, code changes, workflow execution — has no equivalent field in a database. When “completed” is a judgment call, the party writing the judgment owns the invoice. The paper’s earlier look at Meta’s per-token business-agent billing hit the same wall from the other direction: comparing $0.16 of messages against $0.99 of outcome is not a comparison at all unless both sides agree what got finished.

The second is adverse selection on difficulty. A vendor that eats failures has an incentive to route hard cases to humans early, and a buyer measuring only cost per resolution will not notice that its containment rate is quietly falling while its bill stays flat. The honest metric is cost per case entering the queue, not cost per case the vendor was willing to bill.

The third is abuse in both directions. Salesforce Ben’s editor-in-chief Peter Chittum names the vendor’s exposure — token and compute costs must stay low enough and success rates high enough for the model to survive — while a buyer can, in principle, drive interactions that never resolve. Neither side has published a fraud rate, which means neither side has an audited one.

A fourth risk deserves its own sentence because it is the one buyers cannot audit at all: concentration. When the vendor absorbs failure cost, the vendor’s own compute economics decide whether the price survives contact with a bad quarter, and those economics are not disclosed. The paper’s estimate of what Anthropic pays per megawatt of rented compute exists precisely because labs do not publish the input that determines whether their commercial promises are durable. An outcome contract is a bet on a cost structure you cannot see.

There is also a quieter risk in the accounting. A vendor booking revenue only on success has every reason to define success generously in year one and tighten later, exactly as Anthropic did when it framed a headline increase to Claude Code limits that landed as a reduction against the status quo — the structure examined in today’s brief on the 17% weekly-limit cut. Price changes in this market arrive as improvements.

What to do before the next renewal

The evidence that would change this verdict is specific: a published, audited definition of task completion for non-support agent work, or a vendor disclosing its success rate alongside its per-success price. Absent either, treat outcome pricing as a financing structure rather than a discount, and negotiate accordingly.

  • Compute the vendor’s implied cost basis before signing. Take the metered price of the actions the task requires, divide by the vendor’s own published success rate, and compare to the per-outcome price. Salesforce’s numbers give $0.43 against $2; if a vendor will not supply both inputs, that refusal is the answer.
  • Price the attempt, not the success. At a 70% resolve rate, Intercom’s $0.99 becomes $0.69 per attempted conversation. That is the number to compare against your current per-seat or per-token spend, because attempts are what your traffic generates.
  • Audit the billable-outcome definition line by line. Ask specifically how silence, escalation, partial completion, and reopened cases are classified, and require the log that shows which conversations were billed and why. Intercom publishes this; assume a vendor that does not is holding the ambiguity on purpose.
  • Cap the blended rate, not the unit rate. A per-success price with no volume ceiling converts a successful deployment into an uncapped bill. Negotiate a monthly maximum tied to case volume entering the queue.
  • Keep a token-billed fallback path live. For any workload where completion is a judgment call rather than a database field — code changes, research, multi-system workflows, or an agent runtime that just shipped 16,000 pull requests in one release — outcome pricing is currently unpriceable, and the open-weight route still prices attempts honestly if the outcome contract goes bad.

The largest model vendor selling results instead of capacity is the clearest sign yet that the token era has an expiry date. It is not, however, a gift. Someone still pays for every failed attempt; the only question this quarter is whether it shows up on your invoice as a line item or as a markup.

Sources