Wire
WhatWorkedBench tests agents across 36 tasks
Researchers introduced WhatWorkedBench with 36 tasks, 30 data sources, eight workflow types, and 1,248 configuration records to test whether research agents can predict how component changes affect results, not merely optimize a final score, according to the preprint. With eight new measurements, Gaussian-process modeling lifted reported effect recovery from 0.632 to 0.698 in one Flash cohort and from 0.621 to 0.720 in another; results remain preprint evidence, not a production guarantee. Teams building experiment-running agents should log response surfaces and measurement budgets, not only best-run scores, extending the archive’s benchmark-cost analysis.