Wire
Argo-Bench leaves data agents at 34.8%
Argo-Bench’s paper found that the strongest of 14 frontier and open-weight models scored at least 95 on only 34.8% of its 210 enterprise data tasks, averaging 59.5 points. The benchmark simulates an 81-million-order food-delivery business across 235 ERP tables and 7.5 billion rows, then scores actions by their consequences rather than SQL string match. Builders should test agents in stateful, consequence-bearing environments before trusting clean text-to-SQL scores; the archive’s evaluation-cost analysis covers the measurement burden.