Wire
Variance correction cuts agent eval samples 74x
AV-AIVAT reduced outcome variance by a median 54x across 15 LLM-agent configurations and 71,439 paired poker hands, then required 74x fewer hands than raw outcomes to reach ±1-big-blind precision under its asymptotic stopping method. The new evaluation paper couples those corrections with continuously monitored confidence sequences, but draws an important boundary: exact finite-sample runs showed a much smaller median 1.37x stopping-time ratio and require an independently justified payoff bound. Teams extending production-grade governance to privileged agent evaluations should treat sequential statistics as an evaluation-cost control—not a universal 74x promise—and preserve the full stopping trace so another reviewer can recheck the verdict.