skip to content
The Weighted Average

Wire

Document-review agents close half the expert gap

A structured multi-agent reviewer raised rule-intensive document-review performance from a best standalone-model score of 0.3280 to 0.5094 CMCS, against experts at 0.6640, on GB/T-Bench’s 7,306 traceable errors across 488 standards documents. The framework splits global inspection, targeted diagnosis, deterministic rule scans, and verification into specialized skills; operators automating compliance review should borrow that decomposition but retain human sign-off, consistent with the production-control lesson from privileged agent evaluations.