skip to content
The Weighted Average

Wire

Stanford questions 56 AI benchmarks

Stanford researchers found that tests claiming to measure the same AI capability often disagreed across 56 widely used benchmarks, the university’s summary of the research says. Builders should treat leaderboard movement as a hypothesis, not a deployment signal, until benchmark construct validity and cross-test agreement survive independent checks—an important companion to Anthropic’s evaluation-budget warning.