Wire
SWE-Serve finds a production-correctness gap
NVIDIA and UC Berkeley researchers introduced SWE-Serve, a 53-task benchmark where the best agent configuration reached 75% mean pass@1. The SWE-Serve paper tests 11 models across 31 model-effort configurations; on 19 tasks with end-to-end coverage, production-serving checks cut the apparent pass rate from 69.4% to 45.9%. Builder takeaway: pair local tests with live serving checks—green CI can still conceal broken inference behavior, a useful complement to the Strands harness cost-per-pass analysis.