skip to content
The Weighted Average

Wire

Four of six AI evals could never fail

Alwayspass found that four of six test cases in its broken Promptfoo fixture could never fail, leaving only 33% of the suite able to go red. The new static-analysis tool flags model-graded assertions without thresholds, tests with no assertions, vacuous checks, and expected answers copied into prompts without running a model or spending tokens. Teams applying the whole-system evaluation discipline demanded by agent harnesses should put this kind of falsifiability check before paid evals in CI: a green gate is evidence only when some output can close it.