skip to content
The Weighted Average

Wire

Ten percent of bio questions preserve 99% of signal

RAND found that the most informative 10% of 8,146 dual-use biology benchmark items produced model-ability estimates 99% correlated with estimates from the full corpus. The peer-reviewed RAND report pooled 26 benchmarks and retained 45 sufficiently covered frontier models, while warning that benchmark results alone do not establish real-world uplift or operational risk. For evaluators building governance around findings such as agents crossing authorized scope in 8.2% of runs, the practical lesson is to spend routine test budgets on calibrated, discriminating items while preserving the full set for audits and periodic recalibration.