skip to content
The Weighted Average

Wire

Orchard-SWE reaches 73% with 3B active parameters

Microsoft Research’s Orchard-SWE reached 69.7% on SWE-bench Verified, or 73.0% with value-model reranking, using about 3 billion active parameters. The open Orchard framework and training release distilled 107,000 agent interactions and trains models inside deployment harnesses such as Codex instead of a simplified proxy. After Octobench found a five-task swing around identical model weights, Orchard gives builders another reason to evaluate the environment and harness as trainable parts of the agent rather than neutral packaging.