OpenAI 于周三推出了 MentalHealthBench,这是一个包含 1,215 次综合心理健康对话的开放基准,旨在衡量人工智能系统在现实场景中的反应 - 从日常福祉问题到紧急危机 - 每个场景都由有执照的临床医生进行审查和评分。
此次发布的重要性远远超出了 OpenAI 自己的模型。据该公司称,每周有超过 10 亿人使用 ChatGPT,人工智能聊天机器人已悄然成为情绪困扰者的第一站。到目前为止,该行业还缺乏针对这种行为的严格、共享的标准。有关此故事的更多背景信息,请参阅我们正在进行的 人工智能行业报道。
由 22 个国家的 80 多名临床医生共同打造
根据 Unite.AI 的公告报道,OpenAI 与来自 22 个国家的 80 多名持有执照的心理学家和精神病学家共同创建了该基准,他们总共讲 19 种语言,代表近 20 个心理健康亚专业。完整版本包含 5,262 条专家撰写的评分标准。
每次对话均由至少三名专家通过三阶段流程进行审查:两名临床医生独立制定加权标准,第三名临床医生对其进行裁决和完善,只有至少两名专家同意且没有与第三名专家相矛盾的标准才会被保留。每个标准针对模型响应的一个方面,权重从 -10 到 +10,绝对值越大表明临床重要性越高。
公告中援引美国心理学会首席执行官阿瑟·埃文斯 (Arthur Evans) 博士的话说,心理健康存在一个连续体,而让整个范围内的人们参与的人工智能系统需要以临床科学和生活经验为基础。
基准测试中有什么
1,215 次对话的严重程度和寻求帮助的对象故意有所不同:
- 非急性对话占数据集的 53.5%,高敏锐度对话占 18.2%,紧急对话占 28.3%。
- 代表了四种用户类型:成人占 68.1%,青少年占 21.2%,临床医生占 5.8%,护理人员占 4.9%。
- 研究论文描述的对话是使用类似于 Clio 方法的隐私保护技术综合生成的,旨在反映现实世界的 ChatGPT 心理健康使用模式,而不暴露真实用户。
- 70 项任务(占基准的 5.8%)带有先前的用户背景,例如最近失去家庭。
- 除了英语之外,基准测试还包括 105 个西班牙语、54 个印地语、34 个阿拉伯语和 29 个葡萄牙语对话,以及德语、意大利语、波斯语、印度尼西亚语、土耳其语和中文的其他对话。
模型如何评分
评估由自动评分器(以高推理能力运行的 GPT-5.6 Sol)进行评分,每个任务有四个独立采样的判断,结果可以在十个专家定义的行为轴上分解,包括上下文搜索、同理心、紧急度校准和现实测试。
在 OpenAI 报告的结果中,GPT-6 Astra 得分最高,为 57.3%,其次是 GPT-6 Sol(53.9%)、Anthropic 的 Claude Opus 5.5(52.4%)和 GPT-6 Luna(50.2%)。老一代产品远远落后:2025 年 3 月的 GPT-4o 得分为 32.1%,Gemini 2.5 Pro 得分为 29.5%,所有数据的置信区间均为 95%。
这些数字由两个参考点构成。完全访问评分标准的完成分数为 99.0%(这是对评估噪音上限的健全性检查),而临床医生自己编写的完成分数仅为 38.5%,这主要是因为临床医生写的是简短的、面对面的回复,而不是全面的聊天机器人回复。
OpenAI 表示,结果显示各代模型的稳步改进,同时强调在寻求适当的背景和校准紧迫性方面还有改进的空间。值得注意的是,该论文报告了一个权衡:对紧急、高风险对话持谨慎态度的模型可能会更慢地参与日常对话,反之亦然。
用户和专家对于“好”是什么样子存在分歧
Alongside the benchmark, OpenAI ran a separate analysis with 44 adults who had used AI for mental health or emotional support, representing 16 countries and 14 languages. Participants rated model responses and wrote their own criteria, though their review was limited to non-acute conversations to avoid exposing them to distressing material.
User and expert rubrics aligned on just 25.7 percent of total rubric weight, with 1.0 percent directly contradictory. Users emphasized practical next steps and tone; experts placed greater weight on gathering relevant context and carefully interpreting ambiguous situations. OpenAI's conclusion: user feedback is a coherent and complementary signal, but not interchangeable with clinical and safety guidance.
An auditable diagnostic, not a leaderboard
OpenAI is explicit that MentalHealthBench is an auditable diagnostic tool rather than a definitive leaderboard, and that its language comparisons are descriptive — they cannot isolate the effect of language from differences in acuity, topic, culture, or user profile. The dataset is available for download, and each example carries a canary string so researchers can detect if the data leaks into future training corpora.
The company also pointed to related efforts, including research grants for AI mental health work, expert convenings with the Partnership on AI, and assistance for Transluce's independent mental health evaluation effort.
The caveats are real: the conversations are synthetic, grading is automated, and OpenAI's own models top a benchmark OpenAI designed. But by releasing the expert rubrics, the grading methodology, and the data openly, OpenAI has given regulators, clinicians, and competitors a common instrument for a domain where — as the company's own weekly usage numbers suggest — the stakes keep rising.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →