skip to content
The Weighted Average

Wire

DataSpace caps data agents at 66.34%

DataSpace tested six frontier multimodal models and five agent harnesses across 410 analytics tasks, but the best system reached only 66.34% accuracy. Its 7,439-artifact workspaces mix CSV, JSON, SQLite, Markdown, PDF, and video, while changing only the harness produced a 15.36-point spread with the same backbone. For teams building on the case for governing agent evaluations as production systems, the result says model selection alone cannot certify a data agent: test the complete harness against joins and multimodal evidence before it touches consequential analytics.