skip to content
The Weighted Average

Wire

BIABench caps bioimage agents at 0.19

BIABench tests general-purpose and biology-specific agents on 16 end-to-end analyses reconstructed from published studies; on tasks adding 3D or a time axis, no agent scored above 0.19. The arXiv benchmark paper reports repeated-run variability and says process scores could not distinguish correct from wrong runs without ground truth. Scientific-AI builders should keep raw-data checks and task-specific ground truth in the loop; the archive’s benchmark-validity analysis shows why a score is not deployment evidence.