Wire
C5R's SciUniverse tests AI lab work
C5R’s SciUniverse benchmark gives frontier models 92 scientific tasks across 17 task families, with its best Pass@1 result at 45.3% and average model/API inference cost of $40.61 per attempt. The official benchmark runs models through Facility-0 or its digital twin across chemistry, biology, and materials science, testing instrument control, protocol adaptation, facility management, and interpretation of real measurements—not clean-data question answering. For builders, file the gap between knowing science and doing it: evaluate physical constraints and recovery from failed experiments before granting an agent authority over lab equipment or human operators. Enveda’s clinical-evidence gate for AI drug discovery makes same diligence point on downstream evidence.