Models & Open Source
A Free Anonymous Model Gives Away $12.32 a Call
Ox Alpha offers a 1M-token window free on OpenRouter. Priced at GPT-5.6 Sol rates, one max-length call would cost $12.32 — and nobody will say who pays.
An anonymous model called Ox Alpha is serving a 1,048,576-token context window with a 131,072-token maximum output, free, on OpenRouter, which describes it on the Ox Alpha model page as “a reasoning model designed for coding, sustained agentic work, and production workloads” and lists no provider. The listing says the model is “developed and operated by a third-party provider who has chosen to remain anonymous during this preview,” per TechCrunch’s report on the stealth release.
Price what is being given away. Fill that window once and generate the maximum output, and at OpenAI’s published long-context rates for GPT-5.6 Sol — $8 per million input tokens and $30 per million output on the OpenAI API pricing page — a single call would bill $8.39 in input plus $3.93 in output, or $12.32. At Anthropic’s Opus 5 rates of $5 and $25 on the Claude pricing page, the same call runs $8.52. Somebody is absorbing frontier-scale inference costs to put an unnamed model into production codebases, and that is the interesting fact, not the benchmark.
Free preview, unnamed counterparty
The provenance question has produced confident answers in opposite directions, which is itself informative. Speculation initially settled on Chinese labs, with a Wccftech analysis pointing to Zhipu on the grounds that it live-tested GLM-5 under a similar alpha alias, then updating to suggest Microsoft’s MAI lineage after observers noted the model appears to use the cl100k_base tokenizer. TechCrunch records Stripe CEO Patrick Collison calling the model “very impressive,” AI analyst Andrew Curran reporting that early consensus had already dissolved, and Reddit threads asserting Chinese origin and non-Chinese origin with equal confidence.
Wccftech also reports the operator claims capacity for 100 trillion tokens per day during the free window. That claim is worth converting rather than repeating: at GPT-5.6 Sol’s long-context output rate, 100 trillion tokens a day would carry roughly $3 billion of daily list-price value; even at GPT-5.6 Luna’s discounted $1.80 long-context output rate, it is about $180 million a day. Either figure is implausible as sustained giveaway economics, which means the number is marketing, the token mix is overwhelmingly cheap input, or the throughput ceiling is theoretical. None of those readings supports treating the endpoint as production capacity.
Test lane, not dependency
The operator rule here is narrow and firm: an anonymous endpoint is a benchmark target, never a production dependency. You cannot audit a data-handling policy for a provider with no name, cannot escalate an incident to a party that has not identified itself, and cannot price a migration away from a service with no published rate card. Route evaluation traffic to it if the long-context behaviour is interesting; route nothing that carries customer data, proprietary source, or a service-level promise.
The million-token window is precisely what makes that discipline hard to hold. Long-context coding agents are expensive at every other provider, and a free window eight to twenty times the size of a standard 200,000-token context invites exactly the workload teams most want to offload: whole-repository reasoning, multi-file refactors, long transcripts of prior agent runs. Those payloads are also the ones carrying the most proprietary context per request. The free tier is cheapest where the data is most valuable, which is an uncomfortable coincidence even if it is only a coincidence.
There is a second cost that does not show on any invoice. Benchmarking against an endpoint that may be withdrawn without notice produces results you cannot reproduce, cite internally, or defend at a review six weeks later. If the evaluation is worth running, run it against the anonymous endpoint and at least one named model on the same tasks, so the comparison survives the preview window closing.
This paper priced the adjacent version of that risk when discounted Claude tokens turned out to cost $4.62 once you verified what actually served them — the same structural problem, since a request to an unidentified backend can be served by any model at all, and the customer’s only defence is verification they perform themselves. The cheap-token trade and the free-token trade differ in price and in nothing else that matters.
There is a legitimate counter-case. Stealth previews are a normal way to gather feedback without leaderboard pressure, OpenRouter is a well-known distribution surface now being acquired by Stripe, and a model this generous on context is genuinely useful for the long-horizon coding work the listing advertises. If the provider names itself with a published price and a data policy, everything above becomes a footnote to a launch.
Until then, three things would change the verdict: a named provider, a post-preview rate card, and a third-party benchmark run on the public endpoint. Absent all three, the sensible position is the one today’s lead recommends for frontier tiers generally — measure cost per completed task rather than cost per token — with the added condition that a model you cannot name has no place in that comparison’s production column. The archive’s earlier reading of harness-versus-model cost swings on identical work applies here too: a free model inside an inefficient loop is not free, and a free model you cannot identify is not measurable.