OpenAI introduced MentalHealthBench on Wednesday, an open benchmark of 1,215 synthetic mental health conversations designed to measure how AI systems respond in realistic scenarios — from everyday well-being questions to urgent crises — with every scenario reviewed and scored by licensed clinicians.
The release matters well beyond OpenAI's own models. More than one billion people use ChatGPT each week, according to the company, and AI chatbots have quietly become a first stop for people in emotional distress. Until now, the industry has lacked a rigorous, shared yardstick for that behavior. For more context on this story, see our ongoing AI industry coverage.
Built with more than 80 clinicians in 22 countries
OpenAI co-created the benchmark with a cohort of more than 80 licensed psychologists and psychiatrists across 22 countries, who collectively speak 19 languages and represent nearly 20 mental health subspecialties, according to coverage of the announcement by Unite.AI. The full release contains 5,262 expert-authored rubric criteria.
Each conversation was reviewed by at least three experts through a three-stage process: two clinicians independently authored weighted criteria, a third adjudicated and refined them, and only criteria agreed upon by at least two experts — and not contradicted by a third — were retained. Each criterion targets a single aspect of a model's response and carries a weight from -10 to +10, with larger absolute values indicating greater clinical importance.
The American Psychological Association's CEO, Dr. Arthur Evans, was quoted in the announcement saying that mental health exists on a continuum, and that AI systems engaging people across that range need grounding in both clinical science and lived experience.
What is in the benchmark
The 1,215 conversations are deliberately varied in severity and in who is asking for help:
- Non-acute conversations account for 53.5 percent of the dataset, high-acuity conversations for 18.2 percent, and emergent conversations for 28.3 percent.
- Four user profiles are represented: adults at 68.1 percent, teens at 21.2 percent, clinicians at 5.8 percent, and caregivers at 4.9 percent.
- The conversations were generated synthetically using privacy-preserving techniques the research paper describes as similar to its Clio methodology, aiming to reflect real-world ChatGPT mental health usage patterns without exposing real users.
- Seventy tasks — 5.8 percent of the benchmark — carry prior user context, such as a recent loss in the family.
- Beyond English, the benchmark includes 105 Spanish, 54 Hindi, 34 Arabic, and 29 Portuguese conversations, with additional conversations in German, Italian, Persian, Indonesian, Turkish, and Chinese.
How the models scored
Evaluations were scored by an automated grader — GPT-5.6 Sol running at high reasoning effort — with four independently sampled judgments per task, and results can be decomposed across ten expert-defined behavioral axes including context seeking, empathy, urgency calibration, and reality testing.
In OpenAI's reported results, GPT-6 Astra scored highest at 57.3 percent, followed by GPT-6 Sol at 53.9 percent, Anthropic's Claude Opus 5.5 at 52.4 percent, and GPT-6 Luna at 50.2 percent. Older generations trailed well behind: GPT-4o from March 2025 scored 32.1 percent and Gemini 2.5 Pro scored 29.5 percent, with all figures carrying 95 percent confidence intervals.
Two reference points frame those numbers. Completions written with full access to the grading rubrics scored 99.0 percent — a sanity check on the evaluation's noise ceiling — while completions written by clinicians themselves scored just 38.5 percent, largely because clinicians write short, in-person-style responses rather than comprehensive chatbot replies.
OpenAI said the results show steady improvement across model generations, while highlighting room to improve in seeking appropriate context and calibrating urgency. Notably, the paper reports a tradeoff: models that are cautious on emergent, high-risk conversations can be slower to engage on everyday ones, and vice versa.
Users and experts disagree on what "good" looks like
Vedle srovnávacího testu provedl OpenAI samostatnou analýzu se 44 dospělými, kteří používali AI pro duševní zdraví nebo emocionální podporu, zastupujících 16 zemí a 14 jazyků. Účastníci hodnotili modelové odpovědi a psali svá vlastní kritéria, i když jejich kontrola byla omezena na neakutní konverzace, aby nebyli vystaveni stresujícímu materiálu.
Rubriky uživatelů a expertů odpovídaly pouze 25,7 procentům celkové váhy rubrik, přičemž 1,0 procenta si přímo protiřečily. Uživatelé kladli důraz na praktické další kroky a tón; odborníci kladli větší důraz na shromažďování relevantních souvislostí a pečlivé interpretování nejednoznačných situací. Závěr OpenAI: zpětná vazba od uživatelů je koherentní a doplňkový signál, který však nelze zaměnit s klinickými a bezpečnostními pokyny.
Kontrolovatelná diagnostika, ne výsledková tabulka
OpenAI jasně uvádí, že MentalHealthBench je auditovatelný diagnostický nástroj spíše než definitivní žebříček a že jeho jazyková srovnání jsou popisná – nemohou izolovat účinek jazyka od rozdílů v ostrosti, tématu, kultuře nebo profilu uživatele. Soubor dat je k dispozici ke stažení a každý příklad nese kanárkový řetězec, takže výzkumníci mohou zjistit, zda data unikají do budoucích tréninkových korpusů.
Společnost také poukázala na související úsilí, včetně výzkumných grantů pro práci v oblasti duševního zdraví AI, odborných setkání s Partnerstvím pro AI a pomoci při nezávislém hodnocení duševního zdraví společnosti Transluce.
Upozornění jsou skutečná: konverzace jsou syntetické, hodnocení je automatizované a vlastní modely OpenAI jsou na vrcholu benchmarku navrženého OpenAI. Otevřeným zveřejněním odborných rubrik, metodologie hodnocení a dat však OpenAI poskytla regulačním orgánům, lékařům a konkurentům společný nástroj pro doménu, kde – jak naznačují vlastní týdenní čísla používání společnosti – sázky stále rostou.
---
Stay Ahead of AIZískejte nejnovější zprávy, analýzy a průlomové informace o umělé inteligenci – vše na jednom místě.
Přečtěte si další zprávy o AI →