Wire
SWE-Bench ProMax caps agents at 41.2%
SWE-Bench ProMax gives the best tested coding agent a 41.2% resolve rate on 170 multilingual refactoring tasks averaging 11.4 modified files and 261.6 changed lines. The arXiv paper spans Python, Java, TypeScript, Go, C, C++, and Rust, and says its expert-curated tasks are designed to avoid flawed tests and training-set patch reproduction. Engineering teams should file the score as a reminder that multi-file behavior-preserving work remains unsaturated, then benchmark their own repositories rather than extrapolating from easier coding leaderboards, alongside the archive’s harness-focused evaluation analysis.