Wire
Argus reaches 78% on SWE-Bench Pro
Argus reached about 78% on SWE-Bench Pro, versus 59% for Direct Copilot, while using 1.41 times as many aggregate tokens in the authors’ GPT-5.5 experiments. The August 5 preprint describes a fixed-weight runtime whose manager, planner, engineer, and reviewer retain only reviewed state; later benchmark waves used 21% fewer solve-input tokens and 15% less active workflow time than startup waves. Builders should file this beside evidence that a harness alone can swing five coding tasks: durable state and verification may buy substantial accuracy, but the paper’s 41% token premium and unreproduced benchmark claims still demand a costed trial.