OpenAI introduced MentalHealthBench on Wednesday, an open benchmark of 1,215 synthetic mental health conversations designed to measure how AI systems respond in realistic scenarios — from everyday well-being questions to urgent crises — with every scenario reviewed and scored by licensed clinicians.
Реліз важливий далеко за межі власних моделей OpenAI. More than one billion people use ChatGPT each week, according to the company, and AI chatbots have quietly become a first stop for people in emotional distress. До цього часу в галузі не вистачало чіткого спільного критерію такої поведінки. Щоб дізнатися більше про цю історію, перегляньте наш поточний висвітлення індустрії ШІ.
Створено з більш ніж 80 клініцистами в 22 країнах
OpenAI co-created the benchmark with a cohort of more than 80 licensed psychologists and psychiatrists across 22 countries, who collectively speak 19 languages and represent nearly 20 mental health subspecialties, according to coverage of the announcement by Unite.AI. Повний випуск містить 5262 критерії рубрики, розроблені експертами.
Each conversation was reviewed by at least three experts through a three-stage process: two clinicians independently authored weighted criteria, a third adjudicated and refined them, and only criteria agreed upon by at least two experts — and not contradicted by a third — were retained. Each criterion targets a single aspect of a model's response and carries a weight from -10 to +10, with larger absolute values indicating greater clinical importance.
The American Psychological Association's CEO, Dr. Arthur Evans, was quoted in the announcement saying that mental health exists on a continuum, and that AI systems engaging people across that range need grounding in both clinical science and lived experience.
Те, що є в тесті
1215 розмов навмисно відрізняються за ступенем серйозності та тим, хто просить допомоги:
- Non-acute conversations account for 53.5 percent of the dataset, high-acuity conversations for 18.2 percent, and emergent conversations for 28.3 percent.
- Представлено чотири профілі користувачів: дорослі – 68,1 відсотка, підлітки – 21,2 відсотка, клініцисти – 5,8 відсотка та опікуни – 4,9 відсотка.
- The conversations were generated synthetically using privacy-preserving techniques the research paper describes as similar to its Clio methodology, aiming to reflect real-world ChatGPT mental health usage patterns without exposing real users.
- Сімдесят завдань — 5,8 відсотка контрольного показника — містять попередній контекст користувача, наприклад нещодавню втрату в сім’ї.
- Beyond English, the benchmark includes 105 Spanish, 54 Hindi, 34 Arabic, and 29 Portuguese conversations, with additional conversations in German, Italian, Persian, Indonesian, Turkish, and Chinese.
Як оцінили моделі
Evaluations were scored by an automated grader — GPT-5.6 Sol running at high reasoning effort — with four independently sampled judgments per task, and results can be decomposed across ten expert-defined behavioral axes including context seeking, empathy, urgency calibration, and reality testing.
Згідно з результатами OpenAI, GPT-6 Astra отримав найвищий результат — 57,3 відсотка, за ним йдуть GPT-6 Sol — 53,9 відсотка, Claude Opus 5.5 від Anthropic — 52,4 відсотка та GPT-6 Luna — 50,2 відсотка. Older generations trailed well behind: GPT-4o from March 2025 scored 32.1 percent and Gemini 2.5 Pro scored 29.5 percent, with all figures carrying 95 percent confidence intervals.
Дві точки відліку обрамляють ці числа. Завершення, написані з повним доступом до оціночних рубрик, набрали 99,0 відсотка — перевірка розумності на стелі шуму оцінки — тоді як завершення, написані самими клініцистами, набрали лише 38,5 відсотка, головним чином тому, що клініцисти пишуть короткі відповіді в особистому стилі, а не вичерпні відповіді чат-бота.
У OpenAI заявили, що результати демонструють постійне покращення в поколіннях моделей, водночас висвітлюючи можливості для вдосконалення пошуку відповідного контексту та калібрування терміновості. Примітно, що в документі повідомляється про компроміс: моделі, які обережно ставляться до аварійних розмов із високим ризиком, можуть повільніше вступати в повсякденні, і навпаки.
Користувачі та експерти не погоджуються щодо того, як виглядає «добре».
Alongside the benchmark, OpenAI ran a separate analysis with 44 adults who had used AI for mental health or emotional support, representing 16 countries and 14 languages. Participants rated model responses and wrote their own criteria, though their review was limited to non-acute conversations to avoid exposing them to distressing material.
User and expert rubrics aligned on just 25.7 percent of total rubric weight, with 1.0 percent directly contradictory. Users emphasized practical next steps and tone; experts placed greater weight on gathering relevant context and carefully interpreting ambiguous situations. OpenAI's conclusion: user feedback is a coherent and complementary signal, but not interchangeable with clinical and safety guidance.
An auditable diagnostic, not a leaderboard
OpenAI is explicit that MentalHealthBench is an auditable diagnostic tool rather than a definitive leaderboard, and that its language comparisons are descriptive — they cannot isolate the effect of language from differences in acuity, topic, culture, or user profile. The dataset is available for download, and each example carries a canary string so researchers can detect if the data leaks into future training corpora.
The company also pointed to related efforts, including research grants for AI mental health work, expert convenings with the Partnership on AI, and assistance for Transluce's independent mental health evaluation effort.
The caveats are real: the conversations are synthetic, grading is automated, and OpenAI's own models top a benchmark OpenAI designed. But by releasing the expert rubrics, the grading methodology, and the data openly, OpenAI has given regulators, clinicians, and competitors a common instrument for a domain where — as the company's own weekly usage numbers suggest — the stakes keep rising.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →