skip to content
The Weighted Average

AI Economics for Operators

ChatGPT Ads' 15.3% Edge Needs an Attribution Audit

OpenAI adds visual ads and reports a 15.3% acquisition-cost edge, but conflicting measurement docs make a matched reporting test essential.

Blank white street advertising panel beside a city tram stop
Blank white street advertising panel beside a city tram stop. Photograph by bram naus

Audit the conversion denominator before raising the ChatGPT budget. OpenAI’s October 5 visual-ad announcement reports a 15.3% lower attributed acquisition cost for WeightWatchers than its blended paid-search benchmark, while the company’s current reporting pages disagree about whether the headline conversion total includes people who saw an ad without clicking it. The case study may be sound. The documentation conflict still makes copying its result into a forecast premature.

The new placement comes with an old accounting problem

The creative expansion is real but not immediately universal. OpenAI plans to test visual ads during image generation later this month in the United States with an initial advertiser group. It says the ads will be labeled, separate from generated images, and independent of answers. For a marketer, that creates a placement worth evaluating; it does not establish that existing campaign results transfer to the new format.

The operating decision is whether to fund a controlled acquisition test or expand a proven channel. Those are different budget requests. A test buys information about a particular offer, audience, creative, and measurement setup. Expansion assumes that relationship survives more spending. The announcement supports the former much more comfortably than the latter, especially when the campaign being promoted and the placement being introduced are not the same evidence.

OpenAI’s measurement-partnership briefing distinguishes attribution from experiments estimating additional business. It describes WorkMagic’s geographic study for Dose as estimating 2.3 times the incremental orders captured by last-click attribution. That is a different result from a lower attributed CPA. Neither percentage should substitute for the other in a forecast: one compares methods for recognizing a channel’s effect, while the other compares acquisition costs credited to channels.

This is progress on the weakness identified in the archive’s earlier examination of ChatGPT advertising economics. Measurement infrastructure gives a buyer a more useful question than audience size: what additional profitable business does the placement produce? The newly announced integrations help assemble evidence. They do not erase the distinction between a credited purchase and a purchase caused by an advertisement.

A practical buying rule follows. Keep the pilot’s authorization separate from its success threshold. Specify the completed business outcome, the eligible customer population, the reporting settings, and the period after which the result will be evaluated. Decide the acceptable acquisition cost before seeing the dashboard. Otherwise, an attractive number can become the definition of success after the fact.

The same mistake appears outside advertising. Today’s Telekom analysis separates gross savings targets from token costs, and the Ghost hardware brief separates a purchase price from delivered agent value. In each case, the vendor supplies a useful input. The buyer still has to construct the denominator and account for the missing costs. Advertising now offers more measurement machinery, which makes that responsibility more important, not less.

A dashboard can refresh before its answer exists

The most actionable derived number is a clock ratio. OpenAI’s advertiser basics say impressions, clicks, and click-through rate refresh approximately every 15 minutes. Its results guide allows 24–48 hours for attributed conversions to appear. Divide 24 and 48 hours by a quarter-hour: 96–192 traffic-refresh intervals fit inside the stated conversion-processing window. This is arithmetic from two published guides, not a measurement of every account’s delay.

The implication is straightforward. A repeatedly changing traffic counter does not mean the acquisition result has matured. The basics guide also warns that spend can lag by 7–8 hours. Recent clicks, incomplete spend, and delayed conversions can therefore describe different stages of processing even when displayed together. A media buyer who treats the newest row as a complete financial observation can reward noise or shut off a campaign before its outcome arrives.

Reporting settings add another dimension. The results guide permits click windows of 7, 14, or 30 days and a view-through window of either zero or 1 day. Its current definition includes eligible view-through outcomes in the conversion total when that option is enabled. The Ads reporting API documentation likewise describes combined click-and-view totals, with omitted settings defaulting to 30-day click attribution, one-day view attribution, and the date of the credited ad interaction.

The 24–48-hour processing allowance does not close the attribution window: a later purchase can still be credited to an earlier ad interaction.

Those defaults matter to an engineer maintaining a reporting pipeline. An omitted parameter is still a measurement choice. A historical comparison should record the windows and reporting clock alongside the result. A migration review should compare explicit click-only and combined exports against the same source events, rather than accept an unchanged field name as proof that the field’s meaning stayed constant.

There is a useful sensitivity test for the headline case study. A 15.3% reduction leaves cost per credited acquisition at 0.847 of its comparator. At fixed spending, the equivalent increase in credited acquisitions is 1 ÷ 0.847 − 1, or approximately 18.1%. In other words, changing only the credited outcome count by that amount could produce the same apparent CPA reduction. This is not an allegation about WeightWatchers, nor a claim that view-through attribution inflated its result. It is the size of the measurement effect an auditor must be able to exclude before interpreting the published advantage as transferable efficiency.

Do not turn that sensitivity into another benchmark chart. The observed comparison and the hypothetical denominator change are not independent campaign results. Their purpose is to define a falsifiable review question: does the advantage survive when both channels use the agreed outcome definition, comparable attribution settings, and equally mature reporting periods?

The documentation disagreement is a test requirement

Two official descriptions cannot be reconciled by enthusiasm. The Measurement Pixel guide says view-through conversions remain outside the Conversions total, with CPA and optimization still based on clicks. The Conversions API guide repeats that description. Those statements conflict with the reporting documentation and results guide retrieved for this article. Today’s measurement blog also describes advanced optimization learning from eligible click-through and view-through events. Public documentation alone does not establish which behavior every account currently receives.

The correct response is to test the account and ask for clarification, not declare the entire product broken. Save a dated export with explicit settings, retain the matching API response, and compare click-through, view-through, and total counts. Record whether the dashboard and API agree. Escalate unexplained differences with the exact configuration attached. Until that reconciliation is complete, describe a reported CPA as a configured attribution result rather than a settled economic fact.

Event collection needs its own audit. OpenAI’s conversion-measurement guidance recommends using the same event ID when a purchase is sent through both browser and server paths, allowing deduplication. That recommendation changes implementation work: the order system should provide a stable identity for the completed purchase, and integration tests should establish that retries or dual delivery do not create an additional counted business outcome.

The supported-event specification distinguishes a completed order from a checkout start, lead, and trial. It also specifies monetary amounts in the currency’s minor units. A technically successful event submission is therefore not enough. Engineering and finance should agree which event represents the business objective and verify that its amount reconciles to the order ledger. Otherwise, an accurate attribution system can optimize an inaccurately defined outcome.

The strongest counterargument is that last-click reporting can miss genuine influence. The Dose experiment gives a reason to investigate that possibility, and dismissing every view-through conversion would be as lazy as treating every attributed conversion as causal. A person can be persuaded by an impression and purchase later through another route. A useful evaluation must allow that hypothesis while testing it against an appropriate comparison group.

Evidence that would change the spending verdict is clear: a matched campaign analysis whose advantage persists under explicit attribution settings, reconciled order data, and a credible incremental-lift design. Evidence that would weaken it includes disappearing gains under matched windows, unstable counts after processing, or a result driven by customers who were already likely to buy. These are proposed decision tests, not findings about the cited advertisers. The archive’s AI billing analysis made the same distinction between association and causation; attaching a percentage to a channel does not settle the counterfactual.

Buy a measurement result before buying more reach

The engineering bill belongs in the pilot budget. Browser instrumentation, server events, order reconciliation, consent behavior, reporting exports, and experiment analysis require ownership. No fetched source supplies a defensible universal implementation price or workload estimate, so inventing one would make the investment memo less useful. Ask the team for its own scoped estimate and compare that expense with the information the pilot is expected to produce.

Start with the conversion-tracking setup that connects data sources, event settings, and campaign goals. Treat those as distinct review items. An event can exist without being the goal management thinks it is measuring. Have someone outside the integration implementation read the exported configuration and explain what a counted conversion actually means. That review is inexpensive compared with scaling spend on the wrong outcome.

Keep optimization and payment separate as well. The bidding documentation says conversion campaigns optimize for the selected action but bill for valid clicks. A conversion objective is not a promise to charge only for completed sales. Set the spending limit and the acquisition acceptance threshold independently, then check both. Improving the optimizer’s information may help delivery, but it cannot remove the need to reconcile the invoice with the campaign’s business result.

Placement review remains another independent gate. OpenAI’s advertising policies describe restrictions on sensitive conversations and unsuitable contexts. A buyer with additional brand requirements should verify the available controls against those requirements before launch. The announced independent suitability pilots are an avenue for stronger evidence; their existence should not be presented as completed certification of a particular campaign.

For this quarter, the operating checklist is deliberately narrow:

  • Marketing owners should authorize a bounded pilot. Specify the offer, business outcome, acceptable acquisition cost, and decision date before increasing spend. Preserve a comparison that can reveal whether sales are additional.
  • Data engineers should make attribution settings explicit. Export click-only and combined results, retain the reporting clock, and reconcile totals after the documented processing interval. Resolve the disagreement between documentation and account behavior before treating historical rows as comparable.
  • Finance should price the whole experiment. Include campaign charges and the actual implementation and analysis work. A conversion objective changes what the system pursues; it does not make payment contingent on the outcome.
  • The joint team should define the reversal condition. Scale only if the advantage survives matched measurement and acceptable customer economics. Stop or redesign the test when it cannot answer that question.

OpenAI has made its advertising proposition more testable. That is a meaningful improvement for buyers who can instrument a disciplined experiment. The next useful number is not another impressive percentage in a launch post. It is a reconciled result from the buyer’s own campaign, with a definition that survives being read by engineering, marketing, and finance together.

Sources