skip to content
The Weighted Average

Wire

MANTA lifts multi-agent scores 5.8 points

MANTA scored 74.0 across five agent benchmarks, 5.8 percentage points above the strongest baseline, by changing agent roles, links, execution order, and information visibility while a task was running. The authors’ preprint reports three repeated runs over 30 questions per benchmark with Gemma 4 31B, and says the adaptive system used fewer tokens than the other multi-agent systems evaluated. Builders following the shift from fixed workflows to orchestration layers should test dynamic topology against a static graph, but treat this preprint’s 150-question benchmark sample as a pilot signal rather than a production guarantee.