Policy & Geopolitics
Cheap Claude Tokens Cost $4.62 Once You Verify
Chinese proxies resell Claude at about 10% of list, but an audit found identity checks failed 45.83% of the time — repricing the discount sharply.
The cheapest Claude tokens in the world are not cheap once you price the uncertainty. Chinese developers routinely buy access to Anthropic’s models through overseas API proxies at roughly 10% of official pricing, according to an analysis by Oxford China Policy Lab researcher Zilan Qian published as a study of how cheap Claude tokens are sold in China. Against Anthropic’s published rate of $25 per million output tokens for Opus 5 on the Claude pricing page, that implies about $2.50 per million output tokens — an eye-watering discount on the same nominal model.
Then apply the audit. A systematic comparison of third-party “shadow API” endpoints against official ones found identity verification failures in 45.83% of fingerprint tests and performance divergence reaching 47.21%, per the peer-reviewed study “Real Money, Fake Models: Deceptive Model Claims in Shadow APIs”. If fewer than 55% of requests reach the model the label promises, the expected cost of a million output tokens from the advertised model is not $2.50 but $2.50 ÷ 0.5417 = $4.62 — and the remaining spend buys an unidentified substitute. That is this story’s derived figure, and it converts a 90% discount into something closer to 82%, before any consideration of what happens to the prompts.
The discount is a data trade, not a price
Qian’s analysis describes the economics plainly: operators call it “one fish, three meals.” The first margin is arbitrage on access, assembled from bulk-registered accounts farming free credit, resold unused quota, education and corporate discount arbitrage, and single high-tier subscriptions carved among many users through hourly token quotas. The second is degradation — proxies substituting cheaper model tiers for premium ones, a practice the Chinese developer community calls “diluting,” plus cache discontinuity from rotating accounts that forces users to pay full price for context that would otherwise be nearly free.
Anthropic runs what is probably the strictest access regime of any major provider — phone-number, credit-card, and billing-address checks, bans on companies majority-owned from unsupported regions, and live-selfie identity verification for selected users — and the market persists anyway, as The Decoder’s summary of the grey-market analysis sets out. Enforcement that raises the cost of circumvention without eliminating it produces a professionalized supply chain rather than compliance.
The third meal is the one that matters for anyone weighing a discounted gateway anywhere in the world: the logs are the product. Every request through a proxy exposes the full prompt, the full response, tool calls, and in coding workflows a great deal of surrounding repository context. Qian reports that Chinese developers describe the token business as customer acquisition and the log harvest as the actual margin, and notes that datasets of Claude Opus 4.6 reasoning outputs with no clear provenance already circulate publicly on Hugging Face. She is careful that there is no proof operators systematically sell such data — the claim is that rock-bottom prices become viable when logs carry value, which is a structural argument rather than an accusation.
The shadow-API audit gives that structure independent teeth. Its authors identified 17 such endpoints used across 187 academic papers, one of them accumulating 5,966 citations, meaning the substitution problem has already contaminated published research that assumed it was measuring a frontier model. If a reproducibility crisis can grow inside a peer-reviewed literature, a procurement team evaluating a discounted reseller has no structural advantage.
The governance reading for buyers anywhere
None of this is a China-only phenomenon, and that is the point. Any intermediary sitting between an application and a model provider can substitute the model, retain the traffic, or break cache continuity, and the customer’s only defense is verification they actually perform. Legitimate aggregators operate under transparent enterprise agreements — the layer whose value this paper priced when Stripe’s $7.5 billion OpenRouter acquisition turned routing into a payments asset. The grey market’s existence is a reminder of what that layer is worth when it is trustworthy, and what it costs when it is not.
Three controls follow for any team buying inference through a reseller. Fingerprint the model you are paying for, periodically and automatically, with a fixed probe set whose expected outputs you record — the audit’s method is reproducible in an afternoon. Price the discount against verified delivery rather than list, using the same arithmetic above; a 90% discount that delivers the named model half the time is a different product. And treat prompt confidentiality as a contractual term with named retention limits, because a gateway that never states its retention policy has one anyway.
There is also a supply-side reading for the labs. Qian argues that access restrictions have historically created profitable circumvention markets, and that requests arriving through proxies show the provider the proxy’s account rather than the end user, which weakens the abuse-detection systems that depend on per-account behavioral patterns. That is an argument about the limits of enforcement, and it sits alongside the enforcement gap this paper examined when export controls proved to miss remote compute access. It is also the mirror image of the hardware story in today’s lead on Nvidia’s memory-driven server increase: when the legitimate cost of compute rises, the incentive to find an illegitimate discount rises with it.
What could break this read: the 45.83% figure comes from an audit of three representative shadow APIs among 17 identified, not a census, and the proxies studied may not be the ones any given buyer would encounter. The 10% price point is likewise a reported market rate rather than a published tariff. The verdict holds anyway, because it does not depend on the exact numbers — it depends on the structure, and the structure is that an unverifiable intermediary can charge less precisely because it can deliver something else. Evidence that would change it: a large-scale audit finding shadow endpoints deliver the labeled model reliably. Nothing in the current literature suggests that result is coming.