Wire
Agent evaluations find 54% variance in repeat runs
A new benchmark found that repeating the same AI-agent configuration accounted for about 54% of outcome variance across four scientific coding tasks, after releasing more than 18,000 trajectories. In the paper on configurable agent systems, researchers also found task information mattered more than model size or time budget, while a dedicated verification tool changed behavior more than a prompt asking the agent to self-check. Teams comparing agent vendors should report run variance and configuration—not just one success rate—and the archive’s sampling-cost analysis for agent evaluation shows how to budget that measurement.