Wire
Harness-R1 lifts agent success 9.3 points
Harness-R1 raised Qwen3.5-9B’s average task success from 44.3% to 53.6% by training a separate 9B model to convert failure trajectories into validated patches for the executable agent runtime, the authors report in the new Harness-R1 preprint. The gain held after model fine-tuning, when harness edits moved success from 59.2% to 64.2%, reinforcing the warning from Octobench’s five-task harness swing that orchestration is part of the measured system. Agent teams should keep failed trajectories: they may train the wrapper, not just the model.