ChatGPT News 09/23/2026 AI Rating: Medium

OpenAI Releases MentalHealthBench, Built With 80+ Licensed Mental Health Experts

#OpenAI#ChatGPT#MentalHealth#AISafety#Benchmark

OpenAI released MentalHealthBench, an open-source evaluation framework for how AI systems handle mental health conversations, built to reflect realistic usage rather than only emergency scenarios.

Details

  • Development: created with more than 80 licensed mental health experts from 22 countries, representing 19 languages and nearly 20 subspecialties; experts reviewed synthetic conversations and built scoring rubrics weighted from -10 to +10, rewarding beneficial behavior and penalizing harmful behavior, with each conversation assessed by at least three experts and criteria kept only when two or more agreed
  • What it measures: ten dimensions of model behavior β€” including safety, context-seeking, preserving user agency, and giving actionable guidance β€” across four user personas (adults, teens 13-17, caregivers, and clinicians) and three acuity levels (everyday non-acute conversations, high-acuity situations, and emergencies requiring urgent real-world support)
  • Results: OpenAI published comparisons across frontier models as of September 2026, showing steady improvement over time, with newer models scoring better on context-seeking and overall performance
  • Availability: released openly for independent research, examination, and further development by the research community

What happened next

The benchmark gives OpenAI and outside researchers a shared, expert-validated way to measure a specific safety dimension β€” mental health conversations β€” that’s hard to evaluate with generic capability benchmarks, arriving as part of the same week’s broader push on AI safety transparency alongside the misalignment reporting framework and third-party assessment principles.