OpenAI introduced MentalHealthBench on Wednesday, an open benchmark of 1,215 synthetic mental health conversations designed to measure how AI systems respond in realistic scenarios — from everyday well-being questions to urgent crises — with every scenario reviewed and scored by licensed clinicians.

The release matters well beyond OpenAI's own models. More than one billion people use ChatGPT each week, according to the company, and AI chatbots have quietly become a first stop for people in emotional distress. Until now, the industry has lacked a rigorous, shared yardstick for that behavior. For more context on this story, see our ongoing AI industry coverage.

Built with more than 80 clinicians in 22 countries

OpenAI co-created the benchmark with a cohort of more than 80 licensed psychologists and psychiatrists across 22 countries, who collectively speak 19 languages and represent nearly 20 mental health subspecialties, according to coverage of the announcement by Unite.AI. The full release contains 5,262 expert-authored rubric criteria.

Each conversation was reviewed by at least three experts through a three-stage process: two clinicians independently authored weighted criteria, a third adjudicated and refined them, and only criteria agreed upon by at least two experts — and not contradicted by a third — were retained. Each criterion targets a single aspect of a model's response and carries a weight from -10 to +10, with larger absolute values indicating greater clinical importance.

The American Psychological Association's CEO, Dr. Arthur Evans, was quoted in the announcement saying that mental health exists on a continuum, and that AI systems engaging people across that range need grounding in both clinical science and lived experience.

What is in the benchmark

The 1,215 conversations are deliberately varied in severity and in who is asking for help:

  • Non-acute conversations account for 53.5 percent of the dataset, high-acuity conversations for 18.2 percent, and emergent conversations for 28.3 percent.
  • Four user profiles are represented: adults at 68.1 percent, teens at 21.2 percent, clinicians at 5.8 percent, and caregivers at 4.9 percent.
  • The conversations were generated synthetically using privacy-preserving techniques the research paper describes as similar to its Clio methodology, aiming to reflect real-world ChatGPT mental health usage patterns without exposing real users.
  • Seventy tasks — 5.8 percent of the benchmark — carry prior user context, such as a recent loss in the family.
  • Beyond English, the benchmark includes 105 Spanish, 54 Hindi, 34 Arabic, and 29 Portuguese conversations, with additional conversations in German, Italian, Persian, Indonesian, Turkish, and Chinese.

How the models scored

Evaluations were scored by an automated grader — GPT-5.6 Sol running at high reasoning effort — with four independently sampled judgments per task, and results can be decomposed across ten expert-defined behavioral axes including context seeking, empathy, urgency calibration, and reality testing.

In OpenAI's reported results, GPT-6 Astra scored highest at 57.3 percent, followed by GPT-6 Sol at 53.9 percent, Anthropic's Claude Opus 5.5 at 52.4 percent, and GPT-6 Luna at 50.2 percent. Older generations trailed well behind: GPT-4o from March 2025 scored 32.1 percent and Gemini 2.5 Pro scored 29.5 percent, with all figures carrying 95 percent confidence intervals.

Two reference points frame those numbers. Completions written with full access to the grading rubrics scored 99.0 percent — a sanity check on the evaluation's noise ceiling — while completions written by clinicians themselves scored just 38.5 percent, largely because clinicians write short, in-person-style responses rather than comprehensive chatbot replies.

OpenAI said the results show steady improvement across model generations, while highlighting room to improve in seeking appropriate context and calibrating urgency. Notably, the paper reports a tradeoff: models that are cautious on emergent, high-risk conversations can be slower to engage on everyday ones, and vice versa.

Users and experts disagree on what "good" looks like

ਬੈਂਚਮਾਰਕ ਦੇ ਨਾਲ, ਓਪਨਏਆਈ ਨੇ 16 ਦੇਸ਼ਾਂ ਅਤੇ 14 ਭਾਸ਼ਾਵਾਂ ਦੀ ਨੁਮਾਇੰਦਗੀ ਕਰਦੇ ਹੋਏ, ਮਾਨਸਿਕ ਸਿਹਤ ਜਾਂ ਭਾਵਨਾਤਮਕ ਸਹਾਇਤਾ ਲਈ AI ਦੀ ਵਰਤੋਂ ਕਰਨ ਵਾਲੇ 44 ਬਾਲਗਾਂ ਦੇ ਨਾਲ ਇੱਕ ਵੱਖਰਾ ਵਿਸ਼ਲੇਸ਼ਣ ਚਲਾਇਆ। ਭਾਗੀਦਾਰਾਂ ਨੇ ਮਾਡਲ ਜਵਾਬਾਂ ਨੂੰ ਦਰਜਾ ਦਿੱਤਾ ਅਤੇ ਉਹਨਾਂ ਦੇ ਆਪਣੇ ਮਾਪਦੰਡ ਲਿਖੇ, ਹਾਲਾਂਕਿ ਉਹਨਾਂ ਦੀ ਸਮੀਖਿਆ ਉਹਨਾਂ ਨੂੰ ਦੁਖਦਾਈ ਸਮੱਗਰੀ ਦੇ ਸਾਹਮਣੇ ਆਉਣ ਤੋਂ ਬਚਾਉਣ ਲਈ ਗੈਰ-ਤੀਬਰ ਗੱਲਬਾਤ ਤੱਕ ਸੀਮਿਤ ਸੀ।

ਉਪਭੋਗਤਾ ਅਤੇ ਮਾਹਰ ਰੁਬਰਿਕਸ ਕੁੱਲ ਰੂਬਰਿਕ ਭਾਰ ਦੇ ਸਿਰਫ਼ 25.7 ਪ੍ਰਤੀਸ਼ਤ 'ਤੇ ਇਕਸਾਰ ਹੋਏ, 1.0 ਪ੍ਰਤੀਸ਼ਤ ਸਿੱਧੇ ਤੌਰ 'ਤੇ ਵਿਰੋਧੀ ਹਨ। ਉਪਭੋਗਤਾਵਾਂ ਨੇ ਵਿਹਾਰਕ ਅਗਲੇ ਕਦਮਾਂ ਅਤੇ ਟੋਨ 'ਤੇ ਜ਼ੋਰ ਦਿੱਤਾ; ਮਾਹਿਰਾਂ ਨੇ ਸੰਬੰਧਿਤ ਸੰਦਰਭ ਨੂੰ ਇਕੱਠਾ ਕਰਨ ਅਤੇ ਅਸਪਸ਼ਟ ਸਥਿਤੀਆਂ ਦੀ ਧਿਆਨ ਨਾਲ ਵਿਆਖਿਆ ਕਰਨ 'ਤੇ ਜ਼ਿਆਦਾ ਭਾਰ ਪਾਇਆ। ਓਪਨਏਆਈ ਦਾ ਸਿੱਟਾ: ਉਪਭੋਗਤਾ ਫੀਡਬੈਕ ਇੱਕ ਸੁਮੇਲ ਅਤੇ ਪੂਰਕ ਸੰਕੇਤ ਹੈ, ਪਰ ਕਲੀਨਿਕਲ ਅਤੇ ਸੁਰੱਖਿਆ ਮਾਰਗਦਰਸ਼ਨ ਨਾਲ ਪਰਿਵਰਤਨਯੋਗ ਨਹੀਂ ਹੈ।

ਇੱਕ ਆਡਿਟ ਯੋਗ ਡਾਇਗਨੌਸਟਿਕ, ਲੀਡਰਬੋਰਡ ਨਹੀਂ

ਓਪਨਏਆਈ ਸਪੱਸ਼ਟ ਹੈ ਕਿ ਮੈਂਟਲਹੈਲਥਬੈਂਚ ਇੱਕ ਨਿਸ਼ਚਤ ਲੀਡਰਬੋਰਡ ਦੀ ਬਜਾਏ ਇੱਕ ਆਡਿਟ ਕਰਨ ਯੋਗ ਡਾਇਗਨੌਸਟਿਕ ਟੂਲ ਹੈ, ਅਤੇ ਇਹ ਕਿ ਇਸਦੀ ਭਾਸ਼ਾ ਦੀ ਤੁਲਨਾ ਵਰਣਨਯੋਗ ਹੈ - ਉਹ ਤੀਬਰਤਾ, ਵਿਸ਼ਾ, ਸੱਭਿਆਚਾਰ, ਜਾਂ ਉਪਭੋਗਤਾ ਪ੍ਰੋਫਾਈਲ ਵਿੱਚ ਅੰਤਰ ਤੋਂ ਭਾਸ਼ਾ ਦੇ ਪ੍ਰਭਾਵ ਨੂੰ ਅਲੱਗ ਨਹੀਂ ਕਰ ਸਕਦੇ ਹਨ। ਡੇਟਾਸੈਟ ਡਾਊਨਲੋਡ ਕਰਨ ਲਈ ਉਪਲਬਧ ਹੈ, ਅਤੇ ਹਰੇਕ ਉਦਾਹਰਨ ਵਿੱਚ ਇੱਕ ਕੈਨਰੀ ਸਤਰ ਹੁੰਦੀ ਹੈ ਤਾਂ ਜੋ ਖੋਜਕਰਤਾ ਇਹ ਪਤਾ ਲਗਾ ਸਕਣ ਕਿ ਕੀ ਡੇਟਾ ਭਵਿੱਖ ਦੀ ਸਿਖਲਾਈ ਕਾਰਪੋਰਾ ਵਿੱਚ ਲੀਕ ਹੁੰਦਾ ਹੈ।

ਕੰਪਨੀ ਨੇ ਸੰਬੰਧਿਤ ਯਤਨਾਂ ਵੱਲ ਵੀ ਇਸ਼ਾਰਾ ਕੀਤਾ, ਜਿਸ ਵਿੱਚ AI ਮਾਨਸਿਕ ਸਿਹਤ ਕਾਰਜ ਲਈ ਖੋਜ ਗ੍ਰਾਂਟਾਂ, AI 'ਤੇ ਭਾਈਵਾਲੀ ਨਾਲ ਮਾਹਰ ਸੰਮੇਲਨ, ਅਤੇ Transluce ਦੇ ਸੁਤੰਤਰ ਮਾਨਸਿਕ ਸਿਹਤ ਮੁਲਾਂਕਣ ਯਤਨਾਂ ਲਈ ਸਹਾਇਤਾ ਸ਼ਾਮਲ ਹੈ।

ਚੇਤਾਵਨੀਆਂ ਅਸਲ ਹਨ: ਗੱਲਬਾਤ ਸਿੰਥੈਟਿਕ ਹਨ, ਗਰੇਡਿੰਗ ਸਵੈਚਲਿਤ ਹੈ, ਅਤੇ ਓਪਨਏਆਈ ਦੇ ਆਪਣੇ ਮਾਡਲ ਓਪਨਏਆਈ ਡਿਜ਼ਾਈਨ ਕੀਤੇ ਬੈਂਚਮਾਰਕ ਵਿੱਚ ਸਿਖਰ 'ਤੇ ਹਨ। ਪਰ ਮਾਹਰ ਰੁਬਰਿਕਸ, ਗਰੇਡਿੰਗ ਵਿਧੀ ਅਤੇ ਡੇਟਾ ਨੂੰ ਖੁੱਲ੍ਹੇ ਤੌਰ 'ਤੇ ਜਾਰੀ ਕਰਕੇ, ਓਪਨਏਆਈ ਨੇ ਰੈਗੂਲੇਟਰਾਂ, ਡਾਕਟਰਾਂ, ਅਤੇ ਪ੍ਰਤੀਯੋਗੀਆਂ ਨੂੰ ਇੱਕ ਡੋਮੇਨ ਲਈ ਇੱਕ ਸਾਂਝਾ ਸਾਧਨ ਦਿੱਤਾ ਹੈ ਜਿੱਥੇ - ਜਿਵੇਂ ਕਿ ਕੰਪਨੀ ਦੇ ਆਪਣੇ ਹਫਤਾਵਾਰੀ ਵਰਤੋਂ ਸੰਖਿਆਵਾਂ ਦਾ ਸੁਝਾਅ ਹੈ - ਦਾਅ ਵਧਦਾ ਰਹਿੰਦਾ ਹੈ।

---

AI ਤੋਂ ਅੱਗੇ ਰਹੋ

ਨਵੀਨਤਮ AI ਖਬਰਾਂ, ਵਿਸ਼ਲੇਸ਼ਣ ਅਤੇ ਸਫਲਤਾਵਾਂ ਪ੍ਰਾਪਤ ਕਰੋ - ਸਭ ਇੱਕ ਥਾਂ 'ਤੇ।

ਹੋਰ AI ਖਬਰਾਂ ਪੜ੍ਹੋ →