skip to content
The Weighted Average

AI Economics for Operators

Pareto's Cheap Tokens Need a Retention Review

Union Alpha now points to paid Pareto. A normalized token basket costs 83.3% less than Astra, but no training does not mean no retention.

A monitor, keyboard, mouse, and game controllers on a white desk
A monitor, keyboard, mouse, and game controllers on a white desk. Photograph by Sebastian Bednarek

Developers testing Union Alpha should review their purchasing and data terms before extending the experiment: OpenRouter’s former stealth listing now directs users to Unbiased Pareto. Its current prices imply an 83.3% lower bill than GPT-6 Astra for an equal fresh-token basket, but the listing explicitly allows provider retention of prompts and completions while prohibiting their use for training.

The free experiment now has a counterparty

The news is more consequential than a model changing its badge. September 18 coverage of Union Alpha’s abbreviated free preview described demand overwhelming available capacity and an early move to paid access. That account made the launch interesting. The current marketplace pages determine what a developer can reasonably purchase now. Do not preserve the original preview’s assumptions in an application simply because the first test worked.

OpenRouter lists Pareto at $2.50 per million input tokens and $7.50 per million output tokens, with $0.25 per million cached-input tokens. For a normalized basket containing one million fresh input tokens and one million output tokens, the total is $10. This is a comparison unit, not an assumed production workload. It deliberately excludes caching, tools, marketplace payment charges, and differences in how much work either model needs to finish a task.

The comparison source is OpenRouter’s GPT-6 Astra listing, which quotes standard rates of $10 input and $50 output per million tokens. The same basket therefore costs $60. Subtracting $10 from $60 and dividing by $60 gives 83.3%. The calculation combines the two live model listings; it does not import a benchmark claim from launch coverage or assert that their tokens buy interchangeable intelligence.

That distinction makes the figure useful rather than merely flattering. A software team can take its own fresh-input, cached-input, and output counts and substitute the quoted rates. It must then compare accepted results. A lower tariff is an invitation to run that test, not evidence that the cheap model can absorb every demanding workflow currently assigned to a more expensive one.

The sharpest discrepancy concerns data, not dollars. Early reporting described prompts and completions as neither stored nor used for training. The current Union Alpha description says the provider may retain them, but does not use them for training. Those are materially different promises. A prohibition on one use of data does not determine whether a copy exists, who can access it, or when it disappears.

The Stealth Program agreement also distinguishes OpenRouter’s routing role from the model provider’s responsibilities. Read it as the historical preview’s governing context, not as a substitute for the paid offering’s applicable terms. The route has acquired a disclosed supplier and a new purchasing destination. Procurement should record both instead of assuming that an anonymous trial automatically became an approved production service.

Buy the result, then check where the inputs went

The first candidates for a Pareto pilot are teams with repeatable coding or research tasks, approved test material, and an existing way to judge the answer. Keep customer secrets and restricted repositories outside the experiment until the actual paid route’s handling terms are accepted. That is not a claim that Pareto is uniquely risky. It is the ordinary consequence of a trial whose public description has changed in a consequential way.

Our earlier analysis of frontier-model data custody as a purchasing feature supplies the durable framework: price, capability, and custody belong in the same decision. Today’s Qwen media-agent budget analysis makes the related engineering point. An appealing component tariff does not settle the total cost or the permissions of the workflow surrounding it.

There is another integration surface, but not necessarily another independent supplier. Cloudflare documents Union Alpha through its AI API, including an OpenAI-compatible request shape and usage responses. That establishes a documented access path. It does not establish that the model identifier will persist, that the same price applies there, or that changing gateways changes the underlying model’s data policy. Confirm those details on the route being purchased.

The current Pareto marketplace page says it has one provider. That matters more for a production dependency than the number of websites through which a developer can reach it. Two front doors are not proof of two independent inference backends. A fallback should therefore be tested as a genuinely separate working path rather than assumed from the presence of a router in the architecture.

Migration costs remain outside the normalized basket. Engineers must check tool-call behavior, truncation, refusal handling, and recovery after an interrupted request. Reviewers must decide whether an output actually satisfies the task. These are proposed acceptance checks, not observed Pareto failures. The tariff comparison offers no evidence about their outcomes, and this publication has not run a head-to-head production benchmark.

The strongest counterargument to caution is straightforward: a narrow, low-risk workflow may deliver acceptable results at a substantially lower price without a large migration. That is a good reason to test promptly. It is not a reason to pretend a cached rate applies to fresh inputs, or that no-training language means zero retention. Those shortcuts would corrupt both the financial estimate and the approval record.

Change the verdict when representative tasks show lower total cost per accepted result, the purchased endpoint behaves reliably, and the supplier’s written handling terms fit the data involved. Reject migration when extra attempts consume the price advantage, output quality increases review work, or retention remains incompatible with the workload. Pareto has earned a price-sensitive evaluation, not an automatic transfer of production trust.

Sources