skip to content
The Weighted Average

Wire

SWE-bench audit flags 13.6% misalignment

PAIChecker’s authors found 13.6% of SWE-bench Verified instances contain some form of pull-request/issue misalignment across five patterns and 11 scenarios, according to the accepted ASE 2026 paper. Their multi-agent checker reached 92.12% binary accuracy on SWE-Gym and 91.67% on SWE-bench Multilingual. That puts a meaningful caveat under model rankings such as Claude Opus 4.7’s SWE-bench lead: teams should audit whether a benchmark’s prompt and patch oracle describe the same job before treating a score delta as agent progress.