skip to content
The Weighted Average

Models & Open Source

Tencent's Hy4 Charges 5x GLM for a 2% Edge

Hy4 preview beat GLM-5.3 by 0.07 points on Tencent's own expert evaluation while costing 5.6x more per input token. Buy the model, not the leaderboard.

Electronic circuit board with a prominent blue microchip
Electronic circuit board with a prominent blue microchip. Photograph by Brecht Corbeel

Tencent open-sourced Hy4 preview this week — 770 billion total parameters, 49 billion active, a context window above 1 million tokens — and published the comparison that matters more than the parameter count. In a blind internal evaluation of 203 engineering tasks rated by 163 Tencent experts, Hy4 preview scored 2.99 out of 4.00, against 2.92 for GLM-5.3 and 2.94 for Kimi K3. Tencent set API pricing at $0.834 per million input tokens and $2.501 per million output.

The size jump is real. Tencent’s previous flagship, Hy3, shipped in July at 295 billion parameters with a 256,000-token context, as Simon Willison noted in his side-by-side of the two releases; Hy4 is 2.6 times larger with four times the context and a 1.56TB download. Generation-over-generation, Tencent calls this its largest measured capability gain.

Now price the gap. GLM-5.3 lists at $1.40 per million input and $4.40 output on OpenRouter’s model catalog, but its Flash sibling — the one most teams actually route to — runs $0.075 input and $0.25 output. Against full GLM-5.3, Hy4 is 40% cheaper per input token. Against the median of the two GLM tiers, roughly $0.149, Hy4 costs about 5.6 times more per input token for a 0.07-point edge, a 2.4% improvement in Tencent’s own rating. Nobody published that ratio, and it is the only number in this launch that changes a routing decision.

The self-improvement claim is the real story

Buried under the benchmarks is a disclosure with longer legs. Tencent says Hy4 preview “contributed to its own development process,” proposing approaches, running experiments, and iterating on training methods, data strategies, evaluation frameworks, and low-level operators. It claims the model autonomously analyzed bottlenecks in its own inference system and lifted end-to-end throughput by 31.8% over baseline through operator fusion and communication optimization.

That is a vendor’s unaudited claim about its own model, and it should be read as one. But the mechanism is checkable in a way most capability claims are not: throughput improvements show up in serving cost, and the model card on Hugging Face publishes enough architecture detail — 78 layers, 256 routed experts, gated sparse attention with cross-layer index reuse, a native speculative-decoding layer — for third parties to attempt replication. If a 31.8% inference gain from model-directed optimization holds up outside Tencent, the compounding argument gets a data point that surveys cannot supply.

The honest caveat comes from Tencent itself, which ships the preview with known issues: the model spends longer than necessary reasoning through complex tasks and over-verifies its own work. Both failure modes inflate output tokens, which is where the price is. A model that thinks 40% longer at 5.6x the input price is not a bargain because a rating table says 2.99.

The chat template makes the control surface unusually blunt. Hy4 exposes only two reasoning settings — high, the default, and no_think — with no intermediate effort tier, so a team fighting over-verification has one lever and it is binary. Compare that with the graduated effort controls frontier vendors now ship, and the operational cost of the missing middle setting is easy to underestimate: workloads that need moderate reasoning either pay for maximum reasoning or get none.

What this changes about routing

The launch lands in a market where only about 6% of AI-using US businesses touch a model-serving platform at all, and where the ones that do skew heavily Chinese. Hy4 does not change that denominator. It changes the top of the open-weight menu, and it does so days after GLM-5.3-Flash showed it could sell Opus-class coding performance at a fraction of the price.

Architecture explains part of the price. Hy4’s attention uses gated sparse attention with cross-layer index reuse, and the model card lists 256 routed experts with 8 activated per token plus a shared expert, alongside a native speculative-decoding layer. That design keeps active parameters at 49 billion out of 770 billion, which is what makes sub-dollar input pricing possible on a model this size. It also means serving economics depend heavily on how well a given host implements the sparse path — a reason to benchmark on the provider you will actually use rather than on the model in the abstract.

Three operator moves follow. First, ignore the aggregate score and run Hy4 against your own task distribution — Tencent’s evaluation is 203 tasks chosen by Tencent, co-designed with CodeBuddy and WorkBuddy, and a 0.07-point margin is inside the noise band of any expert-rated rubric. Second, measure tokens consumed per completed task, not price per token; the over-verification behavior Tencent admits to will show up there or nowhere. Third, take the free window seriously as a testbed: Hy4 is free on WorkBuddy and CodeBuddy for two weeks, and Hy3 access on both platforms runs free until September 30, which is a zero-cost way to build the comparison before committing traffic.

What would change the verdict: independent benchmarks on long-context coding that reproduce the GLM and Kimi margins outside Tencent’s rubric, or serving-cost data confirming the 31.8% throughput gain. Both are measurable within a quarter. Until then, this is a strong model with a thin published lead and an unaudited story about optimizing itself — worth testing, not worth migrating for: whoever publishes the benchmark also chose it. The same skepticism belongs on the security side of a self-optimizing model, where today’s lead finds that the interval between a hint about a bug and an automated probe has collapsed to minutes.

Sources