skip to content
The Weighted Average

Wire

Reusable skills lift coding-agent recall up to 180%

A global skill-evolution framework improved coding-agent recall by 31.8% to 180% on bug-revealing test generation, according to a new evaluation of OpenHands and mini-SWE-agent. The study covered 108 real bugs across nine open-source projects plus 500 industrial bug reports, and its internal deployment reported a 61.4% F1 gain; those bounded tasks do not establish general coding performance. For teams using the six-agent trial budget as a stop-loss, the practical signal is to replay proposed skill updates against historical failures before adding them to a shared bank.