skip to content
The Weighted Average

Wire

RecoveryBench finds 85% agreement hiding 1.4% success

RecoveryBench recorded 85.4% agreement for one GPT-5.6 Luna recovery setup while held-out success reached only 1.4%, according to a new arXiv evaluation of agent failure recovery. The authors ran 864 episodes across 12 failed Terminal-Bench checkpoints and found agreement was driven by all-zero outcomes; selecting by mean outcome reached 6.5% held-out success versus 1.3% when selecting by agreement. Teams measuring multi-agent coordination overhead should report absolute success beside stability scores, or a dashboard can reward agents for failing together.