skip to content
The Weighted Average

Wire

Echoverse lifts a 9B computer agent 30.6 points

Echoverse raised a 9B computer-use model’s average across 14 evaluation splits from 36.5% to 67.1% after training it on just 12 stateful synthetic environments. The Microsoft Research preprint reports that shallow replicas cut live-site accuracy from 80% to 75%, while deep environments raised it to 85%; repairing one environment also lifted its trained model from 16.2% to 38.5%. For builders evaluating computer-use systems against production economics, the implication is sharp: spend simulation budget on behavioral depth, grounded graders, and iterative repairs before multiplying near-identical sandboxes.