skip to content
The Weighted Average

Wire

CRED benchmark reaches 99% recall on policy errors

Project APE’s CRED benchmark reports 99% recall for Codex 5.6 Sol on its single-error pool, up from roughly 30% for the best models at the end of 2024, across 200 injected errors in 100 reproducible papers. The updated benchmark says the cheapest 50%-recall verification cost fell about 90× in a year, but recall degrades when papers contain multiple errors and confidence is not calibrated enough to remove humans. Teams building research or compliance agents should price a verifier plus human review, not treat a benchmark score as permission to automate sign-off; the harness-evaluation case for separating model score from workflow reliability makes the same distinction.