Wire
Human edits cut coding-agent success 7.7 points
Plausible user edits made during an agent run cut the average resolve rate of nine coding models by 7.7 percentage points on SWE-bench Verified, according to the SWE-Touch preprint and benchmark. The agents often kept conflicting code or overwrote it without re-inspecting the repository and running targeted tests, a practical failure mode for the collaborative workflows described in the coding-agent fork between goals and outcomes. Teams sharing a workspace with agents should treat file-change detection, conflict reconciliation, and affected-test reruns as required harness behavior rather than model niceties.