skip to content
The Weighted Average

Wire

Budget-aware agents top out at 7.3%

EcoAgent-Bench found tested tool-API agents reached at most 7.3% economic consistency across 304 tasks that price actions and impose explicit budgets. The paper’s primary results say agents often stopped before necessary escalation or overspent on cheap tasks, while moving GPT-5.4 across a budget threshold changed its escalation rate from 0% to just 3%. The practical corollary to a bounded six-agent trial budget is stark: builders should encode spend ceilings and escalation rules in the harness rather than expect a capable model to infer economical behavior from a dollar limit.