OpenAI introduced MentalHealthBench on Wednesday, an open benchmark of 1,215 synthetic mental health conversations designed to measure how AI systems respond in realistic scenarios — from everyday well-being questions to urgent crises — with every scenario reviewed and scored by licensed clinicians.
The release matters well beyond OpenAI's own models. More than one billion people use ChatGPT each week, according to the company, and AI chatbots have quietly become a first stop for people in emotional distress. Until now, the industry has lacked a rigorous, shared yardstick for that behavior. For more context on this story, see our ongoing AI industry coverage.
Built with more than 80 clinicians in 22 countries
OpenAI co-created the benchmark with a cohort of more than 80 licensed psychologists and psychiatrists across 22 countries, who collectively speak 19 languages and represent nearly 20 mental health subspecialties, according to coverage of the announcement by Unite.AI. The full release contains 5,262 expert-authored rubric criteria.
Each conversation was reviewed by at least three experts through a three-stage process: two clinicians independently authored weighted criteria, a third adjudicated and refined them, and only criteria agreed upon by at least two experts — and not contradicted by a third — were retained. Each criterion targets a single aspect of a model's response and carries a weight from -10 to +10, with larger absolute values indicating greater clinical importance.
The American Psychological Association's CEO, Dr. Arthur Evans, was quoted in the announcement saying that mental health exists on a continuum, and that AI systems engaging people across that range need grounding in both clinical science and lived experience.
What is in the benchmark
The 1,215 conversations are deliberately varied in severity and in who is asking for help:
- Non-acute conversations account for 53.5 percent of the dataset, high-acuity conversations for 18.2 percent, and emergent conversations for 28.3 percent.
- Four user profiles are represented: adults at 68.1 percent, teens at 21.2 percent, clinicians at 5.8 percent, and caregivers at 4.9 percent.
- The conversations were generated synthetically using privacy-preserving techniques the research paper describes as similar to its Clio methodology, aiming to reflect real-world ChatGPT mental health usage patterns without exposing real users.
- Seventy tasks — 5.8 percent of the benchmark — carry prior user context, such as a recent loss in the family.
- Beyond English, the benchmark includes 105 Spanish, 54 Hindi, 34 Arabic, and 29 Portuguese conversations, with additional conversations in German, Italian, Persian, Indonesian, Turkish, and Chinese.
How the models scored
Evaluations were scored by an automated grader — GPT-5.6 Sol running at high reasoning effort — with four independently sampled judgments per task, and results can be decomposed across ten expert-defined behavioral axes including context seeking, empathy, urgency calibration, and reality testing.
In OpenAI's reported results, GPT-6 Astra scored highest at 57.3 percent, followed by GPT-6 Sol at 53.9 percent, Anthropic's Claude Opus 5.5 at 52.4 percent, and GPT-6 Luna at 50.2 percent. Older generations trailed well behind: GPT-4o from March 2025 scored 32.1 percent and Gemini 2.5 Pro scored 29.5 percent, with all figures carrying 95 percent confidence intervals.
Two reference points frame those numbers. Completions written with full access to the grading rubrics scored 99.0 percent — a sanity check on the evaluation's noise ceiling — while completions written by clinicians themselves scored just 38.5 percent, largely because clinicians write short, in-person-style responses rather than comprehensive chatbot replies.
OpenAI said the results show steady improvement across model generations, while highlighting room to improve in seeking appropriate context and calibrating urgency. Notably, the paper reports a tradeoff: models that are cautious on emergent, high-risk conversations can be slower to engage on everyday ones, and vice versa.
Users and experts disagree on what "good" looks like
Alongside the benchmark, OpenAI ran a separate analysis with 44 adults who had used AI for mental health or emotional support, representing 16 countries and 14 languages. Participants rated model responses and wrote their own criteria, though their review was limited to non-acute conversations to avoid exposing them to distressing material.
User and expert rubrics aligned on just 25.7 percent of total rubric weight, with 1.0 percent directly contradictory. Users emphasized practical next steps and tone; experts placed greater weight on gathering relevant context and carefully interpreting ambiguous situations. OpenAI's conclusion: user feedback is a coherent and complementary signal, but not interchangeable with clinical and safety guidance.
An auditable diagnostic, not a leaderboard
OpenAI is explicit that MentalHealthBench is an auditable diagnostic tool rather than a definitive leaderboard, and that its language comparisons are descriptive — they cannot isolate the effect of language from differences in acuity, topic, culture, or user profile. The dataset is available for download, and each example carries a canary string so researchers can detect if the data leaks into future training corpora.
The company also pointed to related efforts, including research grants for AI mental health work, expert convenings with the Partnership on AI, and assistance for Transluce's independent mental health evaluation effort.
The caveats are real: the conversations are synthetic, grading is automated, and OpenAI's own models top a benchmark OpenAI designed. But by releasing the expert rubrics, the grading methodology, and the data openly, OpenAI has given regulators, clinicians, and competitors a common instrument for a domain where — as the company's own weekly usage numbers suggest — the stakes keep rising.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →