skip to content
The Weighted Average

Wire

Surge's DAYJOB benchmark stays below 25%

Surge AI’s DAYJOB benchmark reports that the strongest tested agents completed only 24.7% of Healthcare assignments and 23.9% of Finance assignments across 130 expert-built tasks. The DAYJOB release deliberately gives agents short, underspecified requests, an average of 19.8–25.7 input files, and work estimated at 19.6–21.6 human hours, rather than a clean prompt with a prescribed output. For builders, treat “agent can execute” and “agent can scope work” as different capabilities: keep human review on ambiguous, long-horizon workflows until an evaluation set measures both. Opus 5.5’s output-budget test shows why agent economics still depend on completed work, not token price alone.