skip to content
The Weighted Average

Models & Open Source

GLM-5.3-Flash Sells Opus-Class Coding at 1/42

Z.ai's open-weight GLM-5.3-Flash nearly matches Claude Opus 4.8 on its own coding bench at a blended $0.24 per million tokens — a 42x price gap.

Laptop screen displaying colorful lines of source code
Laptop screen displaying colorful lines of source code. Photograph by Markus Spiske

Z.ai has published open weights for GLM-5.3-Flash, a 320-billion-parameter model with just 18 billion active parameters, and priced it at $0.15 per million input tokens and $0.50 per million output. On the company’s own coding evaluation it scores 29.0 against Claude Opus 4.8’s 29.5. Blend those rates at a typical three-to-one input-output ratio and the Flash model costs $0.2375 per million tokens against Opus 4.8’s $10.00 — a 42x gap for a half-point difference on the vendor’s chosen benchmark.

That derived multiple is the number to carry into a routing decision. The inputs are public: Z.ai’s pricing page lists $0.15 and $0.50 for GLM-5.3-Flash, currently halved to $0.075 and $0.25 under a promotion running to September 9; Anthropic’s platform pricing lists Opus at $5 input and $25 output. Weight each 75% input and 25% output and divide. Even at Z.ai’s list price rather than the promotion, one Opus token buys forty-two Flash tokens.

What the weights actually contain

The architecture explains the price rather than excusing it. Z.ai’s model card describes a hybrid of sparse and linear attention — a first for the GLM series — plus Manifold-Constrained Hyper-Connections and a 30-trillion-token multimodal pre-training corpus. The developer documentation is more specific about what that buys: against GLM-5.3, the Flash variant cuts attention computation by 3.01x and KV cache size by 4.44x. Long-context serving is where inference bills concentrate, and this model was designed against that line item rather than against a leaderboard.

The benchmark claims come with the usual asterisk that they are the vendor’s. Z.ai reports 63.4 versus GLM-5.2’s 46.2 on DeepSWE v1.1 and 48.8 versus 26.2 on AutomationBench, and says the model scores 57 on the Artificial Analysis Intelligence Index at $0.045 per task at discounted rates. The Code Bench comparison against Opus 4.8 was run in Claude Code 2.1.207 at max effort — a harness choice that favors nobody in particular but was still selected by the party publishing the result.

Provenance is now settled, which was not true a week ago. This paper analyzed the anonymous endpoint when it was serving a million-token window free with no named operator, and argued it was a benchmark target rather than a production dependency because you cannot escalate an incident to a company that will not identify itself. Z.ai has since confirmed it tested the model anonymously as ox-alpha on OpenCode and OpenRouter, and notes that all of that traffic was served on Chinese AI chips. The provenance objection is resolved; the weights are MIT-licensed and downloadable.

Where the 42x holds, and where it evaporates

The gap is real but it is not free money. Three things break it.

First, the benchmark parity is one lab’s evaluation of its own model. A half-point on Z.ai Code Bench is not a half-point on your repository. Independent throughput is also thin: OpenRouter’s listing for the larger GLM-5.3 shows 33 tokens per second and 4.23 seconds of latency from a single provider, which is not competitive with frontier serving. A model that costs a fortieth as much per token and runs at a third the speed changes the calculus for an interactive agent, though not for a batch pipeline.

Second, the promotional half-price expires September 9. Build a cost model on $0.075 and you will rebuild it in two weeks. Use list.

Third, single-provider risk replaces provenance risk. The weights are MIT-licensed and servable on SGLang, vLLM, and Transformers, so the escape hatch exists — but exercising it means running a 320B-parameter model yourself, and the 18B active count reduces compute per token without reducing the memory you must hold. That is a real infrastructure project, not a config change.

The routing conclusion is narrow and useful. Any workload where a coding agent burns tokens on retrieval, log reading, or multi-file context — the bulk of agentic token spend — should be tested against Flash immediately, because a 42x spread survives a lot of quality loss. Anything customer-facing, latency-sensitive, or contractually bound to a data-handling regime stays where it is until independent evaluations exist. The same discipline this paper applied to DeepSeek’s vision API at five cents per 600 screenshots applies here: cheap tokens are a procurement fact, not an architecture decision.

What would change the verdict: an independent Terminal-Bench or SWE-bench run confirming near-Opus performance on third-party harnesses, or a second serving provider at competitive throughput. What would kill it: post-promotion pricing that drifts toward GLM-5.3’s $1.40 and $4.40, which would cut the spread from 42x to roughly 4.7x and turn a routing revolution into a discount. Today’s lead explains the pressure behind those frontier rates — Anthropic is paying $16.3 million per megawatt-year for the capacity that serves them, and that cost has to land somewhere.

Sources