skip to content
The Weighted Average

Wire

SciReview makes 100 errors pass a precision gate

SciReview pairs 100 expert-authored conceptual errors with adversarial passages that are plausible to criticize but correct, then tests three frontier systems as scientific reviewers. The August 10 RelSciFM workshop record reports that the raw metric rewards the model that flags most aggressively, but a strict zero-false-positive gate reverses the ranking and leaves that model with the smallest share of verified catches. For operators building on the case for harness-level grounded evaluation, the implication is to score precision and abstention beside recall before an automated reviewer can reject work or trigger costly remediation.