skip to content
The Weighted Average

Wire

CWE-Bench v1 leaves patching at 68%

Collinear AI’s CWE-Bench v1 gives coding agents 120 held-out audit-and-patch tasks across 73 weakness classes, and the leading models solve only 68% at pass@1. The benchmark’s public leaderboard adds a three-judge panel to a deterministic exploit-and-regression gate, but keeps the evaluation set private, so the score is a deployment signal rather than a reproducible public test. Teams granting agents security write access should measure both patch success and preserved behavior against their own repositories; the archive’s agent-evaluation governance analysis explains why completion alone is not a safety case.