Models & Open Source
Ai2 Says 90% of Your Eval Questions Are Waste
BenchMIRT audits benchmarks question by question and finds 10% of items reproduce the full ranking — and that WMDP measures reasoning.
The Allen Institute for AI released BenchMIRT this week, a method that audits language-model benchmarks one question at a time, and its headline result is a cost claim disguised as a measurement claim: across 16 benchmarks, keeping only the most informative 10% of questions generally preserved the same picture of which models were stronger or weaker, per Ai2’s BenchMIRT write-up. The analysis covered more than 34,000 questions and 100 open-weight models.
For any team paying to run evals on every release candidate, that is a directly actionable number. A curated tenth of an eval suite that reproduces the ranking of the full suite cuts the recurring inference bill on evaluation by roughly an order of magnitude — and evaluation is one of the few AI line items that scales with your release cadence rather than your traffic.
What the audit actually found
BenchMIRT borrows multidimensional item response theory from psychometrics, estimating both a model’s strength on latent capabilities and each question’s difficulty and discrimination. Trained on results from 100 models across six general-reasoning benchmarks — MMLU-Pro, GPQA, MATH, BBH among them — and ten safety benchmarks from Ai2’s Olmo 3 suite, it recovered two dominant dimensions, safety and general reasoning, without being told which benchmark measured what. Ai2 says re-running the analysis from scratch surfaced the same two dimensions, and publishes the code on GitHub alongside a technical report.
The artifacts are unusually complete for a methods release. Ai2 published the per-question estimates as a Hugging Face data collection alongside the code, which means a team can inspect the discrimination scores for benchmarks it already runs rather than re-deriving them. That matters because the finding is not transferable as a rule of thumb — which 10% of items carry the signal is specific to each benchmark, and the only way to know yours is to run the estimator over your own results matrix.
Two findings should change how a team reads its own scorecard. BBQ, a bias benchmark conventionally grouped with safety, aligned more strongly with general reasoning — meaning a low BBQ score may reflect comprehension difficulty rather than unsafe behavior. WMDP, which tests dangerous dual-use knowledge in biology, chemistry, and cybersecurity, also tracked reasoning rather than safety, and inversely: stronger general reasoning was associated with lower WMDP scores, because the benchmark rewards refusal. A team using WMDP as a safety gate is partly measuring how good the model is at reasoning, in the wrong direction.
HarmBench splits internally. Its standard and contextual harmful-request items align with safety, while its copyright items — requests to reproduce song lyrics — align with general reasoning. Averaging those into a single number is how a benchmark quietly stops measuring the thing its name promises. BenchMIRT’s other quantitative claim is predictive: it correctly forecast whether a model would answer an unseen question 79% of the time, against 70% for the naive approach of assuming per-question performance matches the model’s overall benchmark score — a 9-point gain that is the mechanism behind the 10% subsetting result.
The trade-offs Ai2 names itself
The limitations are disclosed and they are real. Every model in the training set was released by March 2025, so the analysis says nothing about how these dimensions behave on current frontier systems — a serious gap in a week when Anthropic shipped a model claiming 55.8% on Terminal-Bench 4.0 and disclosed that its own production safeguards zeroed out some benchmark tasks. The dimensions BenchMIRT discovers also depend on the benchmark set it is given; a different mix could surface different latent capabilities, which means the safety/reasoning split is a property of these 16 evals, not of language models.
Ai2 is also candid about the dual-use hazard in its own tool. The same question-level estimates that identify a benchmark’s most informative safety items could be used to remove them, producing a weakened evaluation an unsafe model passes. The team argues existing tooling already permits similar trimming and that transparency is worth the risk. That is a defensible call and an uncomfortable one, arriving in the same week the frontier labs are gating capability tiers behind vetted access precisely because evaluation results are becoming commercial artifacts.
There is a subtler methodological caveat in how the safety suite was assembled. Ten of the sixteen benchmarks come from Ai2’s own Olmo model program, so the safety dimension is defined largely by one lab’s view of what safety evaluation should cover. That is not a flaw in the mathematics; it is a reminder that latent dimensions inherit the priors of the eval set, and a team whose risk surface is prompt injection or tool misuse rather than harmful-content refusal should expect a different factor structure entirely.
One more caveat matters for the cost pitch. If the goal is ranking models by predicted performance on randomly held-out items, Ai2 says the plain benchmark average performs slightly better than BenchMIRT. The method’s advantage is resolution, not raw predictive power — which makes the 10% subsetting result a claim about preserving capability estimates, not about preserving every use of a score.
The operator verdict: run this against your internal eval suite before your next model migration, not instead of it. If a tenth of your questions reproduces your ranking, the savings are immediate; if your suite is smaller and domain-specific, the more valuable output is discovering which of your questions measure nothing. The archive’s work on outcome-based agent regression loops made the same argument from the opposite direction: an eval is only as good as the discrimination of its individual items. Evidence that would change the call is a replication on post-2025 frontier models — until then, treat the two-dimension result as a hypothesis about older open weights.