skip to content
The Weighted Average

Wire

Ten items can cut safety-eval cost up to 99%

Researchers modeled eight safety benchmarks across 192 language models and found that roughly ten adaptively selected items could reproduce several full benchmark scores while cutting evaluation cost by 97% to 99%. The item-response-theory study says three factors—refusal strictness, truthfulness, and contextual harm—explain most between-model variance, and the same method can flag naive sandbagging or an API model swap. For evaluation teams, this extends Octobench’s warning that the harness changes measured performance: use small psychometrically selected probes for frequent audits, but retain full suites for release gates.