Wire
Multilingual agents lose up to 18.4 points
OmnilingualGAIA2 finds an 8.8–18.4 pass@3-point gap across ten languages when it evaluates seven frontier and open-weight agents on a multilingual version of GAIA2. The arXiv paper says the gap concentrates on tool orchestration rather than quantitative reasoning, does not close with model scale, and is predominantly model-driven (55%) after accounting for a 6.4% translation-contamination floor. Teams deploying agents outside English should add localized tool-use tests before claiming parity, extending the archive’s agent-evaluation governance beyond English leaderboards.