OpenAI introduced MentalHealthBench on Wednesday, an open benchmark of 1,215 synthetic mental health conversations designed to measure how AI systems respond in realistic scenarios — from everyday well-being questions to urgent crises — with every scenario reviewed and scored by licensed clinicians.
The release matters well beyond OpenAI's own models. More than one billion people use ChatGPT each week, according to the company, and AI chatbots have quietly become a first stop for people in emotional distress. Until now, the industry has lacked a rigorous, shared yardstick for that behavior. For more context on this story, see our ongoing AI industry coverage.
Built with more than 80 clinicians in 22 countries
OpenAI co-created the benchmark with a cohort of more than 80 licensed psychologists and psychiatrists across 22 countries, who collectively speak 19 languages and represent nearly 20 mental health subspecialties, according to coverage of the announcement by Unite.AI. The full release contains 5,262 expert-authored rubric criteria.
Each conversation was reviewed by at least three experts through a three-stage process: two clinicians independently authored weighted criteria, a third adjudicated and refined them, and only criteria agreed upon by at least two experts — and not contradicted by a third — were retained. Each criterion targets a single aspect of a model's response and carries a weight from -10 to +10, with larger absolute values indicating greater clinical importance.
The American Psychological Association's CEO, Dr. Arthur Evans, was quoted in the announcement saying that mental health exists on a continuum, and that AI systems engaging people across that range need grounding in both clinical science and lived experience.
What is in the benchmark
The 1,215 conversations are deliberately varied in severity and in who is asking for help:
- Non-acute conversations account for 53.5 percent of the dataset, high-acuity conversations for 18.2 percent, and emergent conversations for 28.3 percent.
- Four user profiles are represented: adults at 68.1 percent, teens at 21.2 percent, clinicians at 5.8 percent, and caregivers at 4.9 percent.
- The conversations were generated synthetically using privacy-preserving techniques the research paper describes as similar to its Clio methodology, aiming to reflect real-world ChatGPT mental health usage patterns without exposing real users.
- Seventy tasks — 5.8 percent of the benchmark — carry prior user context, such as a recent loss in the family.
- Beyond English, the benchmark includes 105 Spanish, 54 Hindi, 34 Arabic, and 29 Portuguese conversations, with additional conversations in German, Italian, Persian, Indonesian, Turkish, and Chinese.
How the models scored
Evaluations were scored by an automated grader — GPT-5.6 Sol running at high reasoning effort — with four independently sampled judgments per task, and results can be decomposed across ten expert-defined behavioral axes including context seeking, empathy, urgency calibration, and reality testing.
In OpenAI's reported results, GPT-6 Astra scored highest at 57.3 percent, followed by GPT-6 Sol at 53.9 percent, Anthropic's Claude Opus 5.5 at 52.4 percent, and GPT-6 Luna at 50.2 percent. Older generations trailed well behind: GPT-4o from March 2025 scored 32.1 percent and Gemini 2.5 Pro scored 29.5 percent, with all figures carrying 95 percent confidence intervals.
Two reference points frame those numbers. Completions written with full access to the grading rubrics scored 99.0 percent — a sanity check on the evaluation's noise ceiling — while completions written by clinicians themselves scored just 38.5 percent, largely because clinicians write short, in-person-style responses rather than comprehensive chatbot replies.
OpenAI said the results show steady improvement across model generations, while highlighting room to improve in seeking appropriate context and calibrating urgency. Notably, the paper reports a tradeoff: models that are cautious on emergent, high-risk conversations can be slower to engage on everyday ones, and vice versa.
Users and experts disagree on what "good" looks like
Di samping penanda aras, OpenAI menjalankan analisis berasingan dengan 44 orang dewasa yang telah menggunakan AI untuk kesihatan mental atau sokongan emosi, mewakili 16 negara dan 14 bahasa. Peserta menilai respons model dan menulis kriteria mereka sendiri, walaupun ulasan mereka terhad kepada perbualan bukan akut untuk mengelak daripada mendedahkan mereka kepada bahan yang menyusahkan.
Rubrik pengguna dan pakar diselaraskan pada hanya 25.7 peratus daripada jumlah berat rubrik, dengan 1.0 peratus secara langsung bercanggah. Pengguna menekankan langkah dan nada praktikal seterusnya; pakar meletakkan lebih berat pada pengumpulan konteks yang relevan dan dengan teliti mentafsir situasi samar-samar. Kesimpulan OpenAI: maklum balas pengguna ialah isyarat yang koheren dan saling melengkapi, tetapi tidak boleh ditukar ganti dengan panduan klinikal dan keselamatan.
Diagnostik boleh diaudit, bukan papan pendahulu
OpenAI secara eksplisit bahawa MentalHealthBench ialah alat diagnostik yang boleh diaudit dan bukannya papan pendahulu yang pasti, dan perbandingan bahasanya adalah deskriptif — mereka tidak boleh mengasingkan kesan bahasa daripada perbezaan dalam ketajaman, topik, budaya atau profil pengguna. Set data tersedia untuk dimuat turun, dan setiap contoh membawa rentetan kenari supaya penyelidik dapat mengesan jika data bocor ke dalam korpora latihan masa hadapan.
Syarikat itu juga menunjukkan usaha berkaitan, termasuk geran penyelidikan untuk kerja kesihatan mental AI, pertemuan pakar dengan Perkongsian mengenai AI, dan bantuan untuk usaha penilaian kesihatan mental bebas Transluce.
Kaveat adalah nyata: perbualan adalah sintetik, penggredan adalah automatik, dan model OpenAI sendiri mengatasi penanda aras yang direka OpenAI. Tetapi dengan mengeluarkan rubrik pakar, metodologi penggredan dan data secara terbuka, OpenAI telah memberikan pengawal selia, doktor dan pesaing instrumen biasa untuk domain di mana — seperti yang dicadangkan oleh nombor penggunaan mingguan syarikat sendiri — kepentingannya terus meningkat.
---
Kekal Mendahului AIDapatkan berita, analisis dan penemuan terkini AI — semuanya di satu tempat.
Baca lebih banyak berita AI →