skip to content
The Weighted Average

Wire

MLPerf sets a 69% agent-training target

MLCommons is adding its first LLM post-training workload to MLPerf Training v6.1, with Qwen 3.5 397B trained against software-engineering tasks and a 69% pass@4 target, according to MLPerf’s benchmark specification. The benchmark runs agents through 251 validation problems and measures time-to-quality rather than raw token throughput. Builders evaluating post-training stacks should file this as a shift toward measuring whether agent training solves code, not merely how quickly GPUs emit it; AWS’s Strands harness cost-per-pass analysis offers nearby context on what those evaluation loops cost.