Wire
MentalHealthBench scores AI in 1,215 conversations
OpenAI’s MentalHealthBench evaluates models on 1,215 synthetic mental-health conversations using rubrics written by more than 80 licensed psychologists and psychiatrists across more than 20 countries and 19 languages. The benchmark paper spans non-acute, high-acuity, and emergency cases; its top task-clipped score was 57.3% for GPT-6 Astra, making it a diagnostic measure rather than clinical clearance. Builders shipping health features should pair it with the archive’s clinical-evidence standard for AI assistants and treat context-seeking and urgency calibration as release gates.