skip to content
The Weighted Average

Wire

Kex-bench finds a PoC-dependent exploit gap

Kex-bench found the strongest evaluated coding-agent setup solved 5.0% of 20 Windows tasks without a reference proof of concept, versus 68.9% of 45 tasks when a reference PoC was supplied. The arXiv benchmark paper measures controlled exploit-primitive generation, not full compromise; keep this governance-level finding separate from operational exploit detail. Security teams should test whether agent evaluations depend on privileged reference artifacts, alongside the archive’s warning about frontier cyber-model capability claims.