skip to content
The Weighted Average

Wire

SuperScout cuts coding-agent cost per solve fivefold

A 7-billion-parameter repository scout helped a four-model coding-agent system solve 159 of 266 Python SWE-bench Pro tasks—one more than the best single fixer—at about one-fifth the total cost per solve. The SuperScout preprint says scouting cost less than half a cent of GPU time per task, but its no-router ablation tied the routed system, suggesting that the sandbox-verified handoff—not clever model selection—carried the result. Teams confronting the review bottleneck behind stacked agent pull requests should first test a cheap scout that reproduces the issue and strips false claims before paying for a sophisticated router.