skip to content
The Weighted Average

Wire

AgentWorld tops out at 52% on long-horizon teams

AgentWorld’s new multi-agent benchmark covers 200 tasks—100 human-annotated plus 100 variants—with teams of 3–20 agents coordinating for more than 50 interaction rounds; the best tested model reached only 52.0% task success. The arXiv paper adds Causal Collaboration Effectiveness to distinguish useful team actions from mere completion, and attributes failures to communication breakdowns, role confusion, and lost shared plans. Builders should treat long-horizon orchestration as an unproven reliability layer rather than assume adding agents raises throughput; the archive’s Business Arena reliability analysis makes the single-agent versus multi-agent tradeoff concrete.