skip to content
The Weighted Average

Wire

AACR-Bench seals 57 code-review tasks

AACR-Bench Harbor released 57 sealed pull-request review tasks with 462 reference findings for evaluating code-review agents. The canonical benchmark package grades submitted findings by path, side, line, and semantic match; its five-task canary also recorded 40% to 60% invalid trials for two DeepSeek/Pi configurations, versus zero for Luna and Terra, while explicitly treating that evidence as preliminary. For teams following the shift from raw benchmark scores to accepted agent work, the useful artifact is the reproducible task-and-verifier boundary—not the tiny canary ranking.