skip to content
The Weighted Average

Agentic Engineering

GPU Kernel Benchmarks Can Be 55.8 Points Too Lenient

A new verifier finds a 55.8-point gap between loose GPU-kernel checks and contract tests; agent builders need adversarial correctness gates.

A black flat-screen computer monitor on a desk
A black flat-screen computer monitor on a desk. Photograph by Mohammad Rahmani

A GPU kernel can be fast enough to win a benchmark and wrong enough to corrupt a training run. A new preprint audited 2,638 machine-generated kernels and found that 62.1% carried at least one contract violation, while 39.5% failed a tolerance-free gate; the central lesson is a 55.8-point gap between a loose benchmark’s green light and a stronger correctness test. The paper’s result is narrow, not a verdict on every AI-written kernel, but it changes what an engineering team should mean by “passed.”

The immediate decision is simple: keep cheap benchmark checks for screening, but do not promote generated GPU code into a training or inference path until it survives adversarial contracts. Teams building kernel-generating agents, custom CUDA paths, or performance autotuners should add the stricter layer this quarter. The cost is CI time and test maintenance. The cost of skipping it is a speedup that silently changes the answer.

The green light was too cheap

The paper, “A Contract-Grade Verifier for LLM-Generated GPU Kernels”, starts from a familiar bargain. A generated kernel replaces a PyTorch operation, runs on a few randomized inputs at a fixed shape, and passes if its output is close enough to a reference. The same loop can then time the candidate and advertise a speedup. It is fast, reproducible, and useful for sorting a large number of attempts. It is also a narrow window into what the program promises.

The authors audited a public Dr. Kernel/KernelGYM corpus. They ran 3,134 kernels from selected operator classes—matmul, attention, softmax, scan, normalization, convolution, and reduction—then excluded toolchain and compile artifacts. The 2,638 remaining kernels had already been accepted by the source system’s own correctness-and-speed harness. That denominator matters: the experiment is not “we found defects in random code.” It is “we found defects inside code that had already received a benchmark pass.”

The benchmark family is not foolish. The KernelBench repository explicitly separates correctness from performance and defines fast_0 as the share of tasks that are correct, while higher fast_p measures require both correctness and a speedup threshold. Its evaluation guide also tells users to treat unusually good results with suspicion and says the project does not endorse third-party results. The problem is not that a screening harness exists. The problem is asking a screening harness to certify a production contract it was never designed to express.

That contract is larger than “the output was close on these inputs.” A kernel should preserve exceptional values when the reference produces NaN or infinity, behave consistently across repeated runs, work when shapes change, respect declared accumulation precision, avoid illegal aliasing, and fail loudly rather than return a plausible number after a device-side error. The KernelBench evaluation code is a useful record of the conventional loop. The PyTorch allclose definition makes the hidden bargain explicit: approximate equality depends on absolute and relative tolerances. Those tolerances are appropriate for some floating-point comparisons; they cannot, by themselves, prove that every legal input and execution path behaves correctly.

The engineering context is getting less forgiving. Triton’s programming model lets developers write specialized GPU programs without hand-authoring every low-level instruction, while the CUDA programming guide exposes a much larger surface of memory, synchronization, precision, and device behavior. Agents can now search that surface quickly. The bottleneck moves from producing a kernel to proving that the kernel is the same program under the conditions a real workload will encounter.

That is the systems version of the archive’s earlier AI-written C++ runtime-tax finding: generated code can look productive while shifting the bill into a later quality layer. In GPU work, the late bill is not merely slower execution. It can be an incorrect gradient, an invalid reduction, or a model that trains on corrupted state.

A correctness gap with a denominator

The paper’s arithmetic makes the argument harder to dismiss. The standard harness accepted 2,472 of 2,638 kernels, or 93.7%. The contract verifier found that 62.1% had at least one violation, leaving 37.9% with no reported violation under that battery. Subtract the two acceptance shares—93.7 − 37.9—and the result is a 55.8 percentage-point gap. This is not a claim that 55.8% of all kernels fail in production. It is the measured distance between two definitions of “correct” on the same accepted set.

A loose benchmark leaves a 55.8-point correctness gap

Share of 2,638 kernels under two correctness definitions

Loose benchmark acceptedNo contract violation found0%20%40%60%80%100%93.7%37.9%
Loose benchmark acceptedNo contract violation found0%50%100%93.7%37.9%
Shah & Shrestha · arXiv 2608.12700 · Aug 2026

The direction of disagreement is more revealing than the headline rate. The paper reports 1,487 kernels that passed the benchmark check but failed the contract verifier; 958 of those failed a tolerance-free gate. The paper’s 1,043 tolerance-free failures are the full corpus total, while 958 is the subset inside that one-way disagreement cell. Only 14 went the other way. A stricter test that simply disliked every approximation would produce more disagreement in both directions. This nearly one-way split suggests a blind spot in what the cheap check probes: it routinely misses behavior outside its fixed-shape, approximate-equality slice.

The authors report several defenses against the obvious objection that their verifier is merely overzealous. It agreed with KernelBench’s own correctness code on 98.5% of 1,030 paired cases. Seven of seven positive-control kernels passed the applicable contract gates. A stratified hand audit separated genuinely broken cases from tolerance-dependent and out-of-scope cases. Under a hardened per-dtype benchmark variant, 1,263 kernels still passed the benchmark check and failed the verifier. None of that makes the preprint an independent certification authority, but it does make “the checker was just stricter” an incomplete rebuttal.

The paper’s differential also gives platform teams a useful design principle. Do not throw away the fast screen. A cheap fixed-shape test can reject obvious failures before a more expensive battery runs. Instead, make the tests sequential: benchmark checks answer “is this candidate worth investigating?” Contract checks answer “may this candidate become a dependency?” A generated kernel that wins by 3× but fails a shape or non-finite-value contract is not a speedup. It is a rejected artifact with a misleading benchmark score.

The distinction matters even more when the generated code is part of an agent loop. An agent may explore hundreds of variants, keep the fastest one, and hand a human a polished diff. If the acceptance signal rewards a narrow input distribution, the agent learns to optimize the signal rather than the program’s real obligations. The same dynamic appears in agent-harness evaluation: a score is only as meaningful as the environment and failure modes it exposes.

What the verifier catches—and what it does not

The contract battery has 12 adversarial gates. They cover value behavior, exceptional values, ordering and determinism, precision, aliasing and crashes, gradients, and resource metadata. In the forward-only headline audit, not every gate applied to every row; the paper identifies seven load-bearing gates. That qualification is important. The result is a measurement of a defined testing envelope, not a magical proof that a kernel is correct everywhere.

The modal defect is especially consequential: a candidate can turn a reference NaN or infinity into an ordinary finite number. A visual test may look fine because the ordinary number is plausible. A training loop may continue for hours before the resulting model drifts. Other failures arise when a kernel behaves at its native shape but breaks at a longer sequence, accumulates in a lower precision than the reference, or produces a different answer on repeated runs. The paper’s full methods and failure analysis explain why each class requires a different test rather than one tighter tolerance.

Blackwell makes the point concrete without making the story about one chip. The paper’s inward-facing contribution is a native Blackwell backward for a gated-linear-recurrence family, and it discusses a 512-column Tensor Memory constraint. The related Mamba issue records how a resource-management problem can force a training path onto a much slower fallback. A kernel can therefore be numerically close and still be operationally wrong: it may compile into an illegal configuration, deadlock, or quietly select a path whose economics invalidate the original decision.

The best counterargument is scope. This is a v1 arXiv preprint, not a peer-reviewed production incident report. The sample is a selected subset of a public corpus, the headline audit is forward-oriented, and the environment is tied to a B200 with specified PyTorch and Triton versions. The original KernelBench paper describes a benchmark for generated CUDA and DSL kernels, not a universal production test. A percentage from this audit should not be pasted into a reliability forecast for every GPU, model, compiler, or workload.

The counterargument does not rescue the old workflow; it tells buyers what evidence to request next. The thesis would weaken if an independent team reran the verifier on a broader corpus, across multiple GPU generations and compiler stacks, and found that the violations rarely correlate with production failures. It would strengthen if the same pattern appeared in gradient kernels, multi-GPU collectives, and long-running training jobs. Those are the next tests—not another leaderboard screenshot.

Put contracts in the fast path

The practical response is not to stop using agents or hand-write every kernel. It is to move correctness from the end of the process into the agent’s definition of done. Start with a cheap screen for compilation, speed, and ordinary outputs. Then apply adversarial contracts to every candidate that might be merged, published, or used to generate a training artifact.

The first boundary belongs in the generator. Give the agent explicit requirements for NaN and infinity propagation, shape generality, deterministic repeated runs, accumulation dtype, aliasing, and failure behavior. The PyTorch numerical-accuracy guidance is a useful reminder that floating-point behavior is contextual; a contract should state which differences are acceptable instead of hiding them inside a single universal tolerance.

The second boundary belongs in CI. Keep the benchmark score, but label it “screened.” Add a contract score that names each gate, stores failing inputs, and blocks promotion when a tolerance-free property breaks. Run the full battery on a sample of exploratory candidates if compute is scarce; run it on every release candidate. The exact schedule is a local economics decision, but the semantic boundary should not be ambiguous.

The third boundary belongs in procurement and review. If a vendor or internal agent claims a kernel speedup, ask for the reference operator, input-shape envelope, dtype rules, exceptional-value tests, repeatability checks, GPU and compiler versions, and the failure corpus. A number without those details is a demonstration, not a deployable performance commitment. The NVIDIA photonics analysis in today’s edition makes the same point at the network layer: infrastructure claims become procurement evidence only when the measurement conditions and fallback path are explicit.

  • Teams generating CUDA or Triton kernels should add contract gates before the next benchmark sweep. Keep the fast screen, but block merges on non-finite propagation, shape variation, determinism, precision, aliasing, and crash tests.
  • Platform owners should record the rejected artifact, not just the winning speedup. The failure distribution is part of the model of the generator; without it, an agent can optimize toward a permissive harness.
  • Model and agent evaluators should separate “passes the harness” from “safe to depend on.” The Faraday replication study and GitHub’s agent-app rollout show adjacent versions of the same rule: a workflow is only as trustworthy as the evidence and approval boundary around its automation.

The new preprint does not prove that AI-written GPU kernels are unusable. It proves something more actionable: a benchmark pass can be dramatically cheaper than a correctness contract. Builders who keep those meanings separate can still harvest agent-generated performance. Builders who collapse them will eventually ship a number that was never true.

Sources