skip to content
The Weighted Average

Wire

Web search cuts benchmark accuracy by 8 points

Web search cut benchmark accuracy by as much as 8 percentage points in a 4,812-response audit of 401 prompts, while repeated answers disagreed on up to 21% of prompts. The comparison also found ChatGPT’s interface less accurate than the API on both tested benchmarks when search was disabled, so teams should vary interface, search state, and repeated runs rather than trust one API-only score—a practical extension of the case for production-shaped agent regression tests.