Wire
HarnessOpt-Bench scores 111 agent-harness runs
HarnessOpt-Bench ran 111 scored trials of five frontier LLM optimizers improving agent harnesses across four downstream tasks under fixed evaluation budgets. The preprint reports that optimizer-model choice separated outcomes more than the coding harness used, while native harnesses were not consistently better; builders should therefore hold out test cases and meter optimization budgets rather than tune prompts against the scoreboard, extending the whole-system benchmarking lesson from Octobench to harnesses that rewrite themselves.