Wire
AACR-Bench seals 57 code-review tasks
AACR-Bench Harbor released 57 sealed pull-request review tasks with 462 reference findings for evaluating code-review agents. The canonical benchmark package grades submitted findings by path, side, line, and semantic match; its five-task canary also recorded 40% to 60% invalid trials for two DeepSeek/Pi configurations, versus zero for Luna and Terra, while explicitly treating that evidence as preliminary. For teams following the shift from raw benchmark scores to accepted agent work, the useful artifact is the reproducible task-and-verifier boundary—not the tiny canary ranking.