skip to content
The Weighted Average

Compute & Market Power

OpenAI's Jalapeño Turns Latency Into a Power Bill

OpenAI's first Jalapeño benchmarks claim 1.5-1.9x more tokens per kilowatt than Nvidia's best. At interactive speed the gap becomes 104x — and a cost line.

A silicon wafer with a grid of iridescent microchip dies under shallow focus
A silicon wafer with a grid of iridescent microchip dies under shallow focus. Photograph by Laura Ockel

OpenAI published the first measured results for Jalapeño, its custom inference chip, and the headline is not speed — it is watts. Across three open-weight models, OpenAI’s first-results post claims 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the best commercially available systems. Convert those ratios into the only number a finance team recognises, and the picture sharpens: at the interactive decoding speed Nvidia’s GB300 can barely sustain, the electricity to serve a million DeepSeek R1 tokens works out to about $0.21 on the Nvidia system and $0.002 on Jalapeño. Power was never the dominant line in an inference bill. This is the first credible argument that it could become one.

The benchmark that reframes the question

Jalapeño was tested on SemiAnalysis’s InferenceX benchmark, a public harness that measures the full path of serving a request rather than raw matrix throughput. OpenAI ran it against GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T — none of them OpenAI models, which is the point. Results were normalised by each accelerator’s published package power: 700 watts for Jalapeño, 1,200 for GB200, 1,400 for GB300. OpenAI adds that measured sustained draw stayed at or below 550 watts on the tested workloads, so the normalisation is, if anything, unkind to its own chip.

The peak numbers are the modest half of the story. On GPT-OSS 120B, Jalapeño delivered 85,448 mixed tokens per second per kilowatt against 44,960 for a GB200 system. On DeepSeek R1 it posted 19,641 against 11,781 for GB300, and on Kimi K2.5, 18,195 against 11,862. Call it a half to a doubling — meaningful, not revolutionary, and roughly what a purpose-built ASIC should manage against a general-purpose GPU.

Jalapeño's peak efficiency lead is real but modest: 1.5x to 1.9x

Peak mixed tokens per second per kilowatt, package TDP normalised, 8k/1k prompts

Jalapeño Nvidia GB200 / GB300

DeepSeek R1 670BGPT-OSS 120BKimi K2.5 1TJalapeñoNvidiaJalapeñoNvidiaJalapeñoNvidia020K40K60K80K100K19.6K11.8K85.4K45K18.2K11.9K
DeepSeekGPT-OSSKimiJalapeñoNvidiaJalapeñoNvidiaJalapeñoNvidia050K100K19.6K11.8K85.4K45K18.2K11.9K
OpenAI first-results appendix; SemiAnalysis InferenceX · Aug 2026

The interesting divergence appears when you hold the user experience fixed. Nvidia’s GB300 bottoms out at 5.90 milliseconds between tokens on DeepSeek R1 — about 169 tokens per second per user. Jalapeño reaches 1.43 milliseconds, or 700 tokens per second per user. Force the Nvidia system to the interactive speed it can just about hit and its efficiency collapses: 118 mixed tokens per second per kilowatt, against 12,258 for Jalapeño at the same responsiveness. That is the 104x figure, and it is a statement about architecture, not marketing. General-purpose GPUs pay for low latency by abandoning batching; a chip designed around keeping KV cache local does not.

Nine months separated Jalapeño’s first design from tapeout, according to OpenAI, inside a roughly sixteen-month program with Broadcom that began in mid-2024 and reached fabrication in November 2025, as The Decoder reconstructed from the Hot Chips session. OpenAI and Broadcom announced the chip’s existence last October; the gap between announcement and audited numbers has been short by silicon standards.

The comparison Nvidia would prefer is already on its own website. Nvidia’s published Vera Rubin NVL72 specification claims up to 10x more tokens per megawatt than GB200 NVL72 and one-tenth the cost per million tokens on interactive agentic reasoning, with 3,600 PFLOPS of NVFP4 inference per rack and HBM4 memory. Set that against Jalapeño’s 1.9x over GB200 and the ranking flips on paper — which is exactly why SemiAnalysis insists Rubin, not Blackwell, is the honest benchmark, and why its finding that the two land at roughly parity on total cost per token is the most important sentence in the coverage.

The kilowatt is now a line item you can price

Here is the arithmetic behind this article’s headline figure, stated so you can redo it. A system delivering 19,641 mixed tokens per second per kilowatt produces 19,641 × 3,600 = 70.7 million tokens per kilowatt-hour. The EIA’s May 2026 electric power table puts the average US industrial price at 8.71 cents per kilowatt-hour. Divide: $0.00123 of electricity per million tokens on Jalapeño at peak. The GB300 comparison system, at 11,781 tokens per second per kilowatt, yields 42.4 million tokens per kilowatt-hour and $0.00205 per million — a difference of eight ten-thousandths of a dollar. At peak throughput, power is noise.

Now redo it at 169 tokens per second per user, the interactive operating point. Jalapeño’s 12,258 tokens per second per kilowatt becomes 44.1 million tokens per kilowatt-hour and $0.00197 per million tokens. The GB300’s 118 becomes 425,000 tokens per kilowatt-hour and $0.205 per million tokens — a hundredfold jump, from a fifth of a cent to twenty cents. For a service billing frontier tokens at dollars per million, twenty cents of electricity is still not the binding constraint. For a service selling cheap tokens at DeepSeek’s off-peak $0.66 per million output tokens, a fifth of the revenue evaporating into the meter is a different business.

That is the mechanism worth internalising: agentic workloads move buyers toward the operating point where efficiency curves diverge most. An agent that chains forty tool calls cannot amortise latency the way a chat completion can, because every step waits on the one before it. OpenAI makes exactly this argument, and the benchmark’s structure — evaluating at matched user experience rather than matched batch size — encodes it. The paper’s own framing of a “Pareto frontier” is a polite way of saying the tradeoff between throughput and interactivity is the product decision, and Jalapeño moves it.

The second-order effect lands on software. OpenAI says it brought three open-weight models to high performance in two months using Codex, and that AI-generated implementations of selected GPT-OSS attention and mixture-of-experts blocks ran 1.5 to 1.8 times faster than human-expert versions. SemiAnalysis drew the obvious conclusion, writing that “the CUDA moat is potentially dead given how fast OpenAI can bring up new models on their silicon.” If porting cost falls, the switching cost that has protected Nvidia’s margins falls with it — a thread this paper has been pulling on since Nvidia’s 15% server price rise concealed a 3.6x memory bill.

The ways this benchmark flatters itself

Start with provenance. OpenAI produced the numbers; SemiAnalysis verified some runs on-site. That is better than a slide deck and worse than an independent lab. Richard Ho, OpenAI’s head of hardware, told reporters that “the bottom line is that the results show a very, very significant performance advance over state of the art,” in TechCrunch’s account of the press call — a vendor characterising a vendor benchmark.

The comparison hardware is the sharper objection. Blackwell is shipping today; Nvidia’s Vera Rubin platform, which shares Jalapeño’s HBM4 memory generation, is the like-for-like rival, and SemiAnalysis notes that on total cost of ownership per token the two come out roughly even. Jalapeño also declined optimisations its rivals used, including multi-token prediction and speculative decoding — which cuts both ways, since it means both headroom for OpenAI and an apples-to-oranges asterisk on today’s ratios.

Then there is the small matter of existence at scale. Ho estimated deployment “in very small volumes” at the end of 2026 with meaningful volume in 2027, and SemiAnalysis reports the chip has not moved past engineering samples. Rubin racks are being delivered to customers now: an Indian operator has just placed a binding order for 9,000 of them. A 1.5x efficiency edge available in eighteen months competes against a 1.0x edge available this quarter, discounted by whatever Nvidia ships in between.

Independent reconstructions of the Hot Chips session have landed in the same place. NextBigFuture’s summary of the announcement repeats the 1.5-to-1.9x and 2.1-to-4.1x figures without additional verification, and The Verge’s account of the benchmark release frames the results as OpenAI’s own. That uniformity is itself a caution: every outlet is quoting one dataset, produced by the party with the most to gain, partially spot-checked by a firm whose benchmark is the instrument.

Finally, watts are not the bill. Power is a minority of inference cost against amortised silicon, memory, networking, and the operator’s margin — which is why the peak-throughput gap of eight ten-thousandths of a dollar per million tokens changes nobody’s spreadsheet. The interactive gap matters because it is enormous in ratio terms, not because twenty cents is large. What would change the verdict: an independent InferenceX run against a Vera Rubin NVL72 at matched precision, with both systems permitted their best decoding optimisations, plus a disclosed unit price for Jalapeño capacity. Absent those, this is a strong engineering result with an unpriced product attached.

What to do before the 2027 ramp

The near-term consequence is not a purchase; it is a negotiating position. OpenAI has demonstrated that a first-generation ASIC, designed in nine months with help from its own models, can reach the efficiency frontier — and every hyperscaler with a silicon program now has a public benchmark to point at. That pressure arrives on top of a capital cycle already running hot, with data center capex needing 32% compound growth to reach $3 trillion and Chinese labs posting API gross margins above 80% on far cheaper inference.

The operator checklist:

  • If you serve interactive agents, start measuring tokens per second per user, not tokens per second. The entire Jalapeño argument lives at the operating point where your users wait on sequential steps; a fleet tuned for batch throughput will not see the difference and will not capture it either.
  • If you are signing multi-year inference capacity this quarter, price the option to switch. Jalapeño will not be buyable in volume until 2027, but its existence is a reason to resist term lengths beyond that horizon, and to insist on the right to reprice against published efficiency benchmarks.
  • If you run open-weight models, budget the porting work down, not up. Two months to bring three unplanned models to high performance on new silicon — with AI-written kernels beating human ones on selected blocks — is a direct challenge to the assumption that CUDA lock-in is permanent.
  • If you buy on power, get the operating point in writing. A quoted efficiency figure at peak throughput and the same figure at your latency target can differ by two orders of magnitude, as this benchmark’s own tables demonstrate.

Watch for three things: an independent Vera Rubin comparison, the first disclosed price for Jalapeño-served tokens, and whether OpenAI’s inference gross margin — the reason it built this at all, as the company’s own note about improving operating leverage concedes — visibly moves in 2027. Until then, treat the 104x as physics, the 1.5x as procurement, and the deployment date as the only number that binds. This paper has argued before that OpenAI’s compute overhead is becoming a disclosed budget line; Jalapeño is the company trying to shrink the other side of that ledger.

Sources