skip to content
The Weighted Average

Wire

MLPerf puts edge agents in 16.5GB of VRAM

MLCommons closed submissions for its first edge-agentic MLPerf benchmark with a 27B model compressed into about 16.5GB of VRAM, versus roughly 54GB at BF16, according to the MLPerf Inference v6.1 specification. The reference pairs a 32K context with 1,007 tool-calling turns; reasoning cut accuracy from 86.23% to 78.19% and made the run about 60% longer. For builders weighing local inference against cloud-scale capacity, the emerging edge-agent bar is not tokens per second alone but memory, tool accuracy, and end-to-end turn latency measured together.