Wire
Agent benchmarks mis-score 15.3% of failures
A manual audit found 15.3% of 150 failure verdicts across five computer-use agent benchmarks were wrong: 10.7% were evaluator false negatives and 4.7% came from broken tasks. The research preprint says brittle scripted oracles can reject valid alternative paths or miss decisive visual evidence, making one success-rate number conceal faults in task construction, observation, scoring, and reporting. Teams evaluating computer-use agents whose economics depend on reliable task completion should manually review a sample of both failures and tasks before choosing a vendor—the benchmark may be grading its own defects.