OpenAI's MentalHealthBench: Frontier Models Score Under 60 Percent
OpenAI released MentalHealthBench on September 23, 2026 — an open evaluation dataset of 1,215 synthetic mental health conversations, built with more than 80 licensed psychologists and psychiatrists from 22 countries.

Illustration: AIMag.no, AI-adapted from «Stethoscope and globe on light blue background (52102422083)» by Jernej Furman from Slovenia, CC BY 2.0.
OpenAI's MentalHealthBench: Frontier Models Score Under 60 Percent
OpenAI released MentalHealthBench on September 23, 2026 — an open evaluation dataset of 1,215 synthetic mental health conversations, built with more than 80 licensed psychologists and psychiatrists from 22 countries. The company's own results show that even frontier models score only in the low-to-mid 50s percent — and the biggest weakness lies in the method: the grading is done by an OpenAI model.
What's new
MentalHealthBench is described as an open evaluation dataset designed to measure how well AI systems handle real-world situations — from everyday wellness conversations to acute mental health crises. According to NewsBytes, which covered the launch on September 24, the benchmark was developed in collaboration with more than 80 licensed psychologists and psychiatrists from 22 countries, spanning nearly 20 mental health subspecialties. A Japanese-language summary of the announcement, dated September 24, confirms the September 23, 2026 release date, and a daily news roundup on September 25 adds that the professionals represent 19 languages.
All of the content is synthetic: 1,215 constructed conversations, not real user logs. The dataset is divided into three acuity tiers — non-acute conversations (53.5 percent), high-acuity (18.2 percent), and emergent (28.3 percent) — and covers four user profiles: adults (68.1 percent), teens (21.2 percent), clinicians (5.8 percent), and caregivers (4.9 percent). According to the Japanese summary, the benchmark tests not just safety but also the ability to ask appropriate follow-up questions, respect user autonomy, and give actionable advice.
The full release contains 5,262 expert-authored evaluation criteria, and OpenAI has, per NewsBytes, made the benchmark publicly available so other researchers can review the methodology, run their own evaluations, and build on the work.
How the scoring works
The methodology, as NewsBytes describes it from OpenAI's release, has several layers:
Each conversation was reviewed by at least three experts in a staged process. Two clinicians independently wrote weighted criteria, which were then reviewed and adjusted by a third expert. Only criteria that at least two experts agreed on, and that were not contradicted by a third, were retained. Each criterion targets a single aspect of the model's response, with weights from −10 to +10 indicating clinical significance within the conversation.
The actual scoring is performed by an automated grader: GPT-5.6 Sol at high reasoning effort, which evaluates each model response against the expert criteria with four independently sampled completions per task. Scores are reported as so-called task-clipped rubric scores and can be broken down along ten expert-defined behavioral axes, including context seeking, empathy, urgency calibration, and reality testing.
This is how the benchmark is built to capture more than yes/no safety: the same conversation can earn points both for gathering enough context, for calibrating how urgently the model responds, and for giving advice that can actually be followed.
The results OpenAI itself reports
The figures come from OpenAI's own release, as reproduced by NewsBytes and TechRepublic on September 24 and 25:
- GPT-6 Astra: 57.3 percent — the highest among the models tested
- GPT-6 Sol: 53.9 percent
- Claude Opus 5.5 (Anthropic): 52.4 percent
- GPT-4o: 32.1 percent — an older model that came clearly behind
TechRepublic summarizes the finding as frontier models showing notable gaps in how they seek context and assess urgency. Worth noting is that the available sources do not give specific per-axis figures for how large these gaps are — the characterization comes from TechRepublic's coverage of the results.
The simple picture is still clear: the best model scores 57.3 percent on a scale where 100 percent means full agreement with the experts' criteria. No model comes close to anything one could call clinical reliability, and the distance between the newest generation (52–57 percent) and GPT-4o (32.1 percent) suggests the field has moved quickly — but still has far to go.
Users want quick advice, clinicians want to ask questions
OpenAI also published a survey of 44 adults who have used AI for emotional support. TechRepublic describes the finding as an institutional rift: while clinicians emphasized careful context-gathering and slow assessment, ordinary users wanted fast, practical next steps and an empathetic conversation.
This tension is not incidental to the benchmark's design. The ten behavioral axes include both context seeking (asking enough questions before acting) and actionable advice — two goals that can pull in opposite directions. A model that scores well on gathering context can feel obstructive to a user in distress; a model that gives quick advice can miss dangerous context. The survey suggests OpenAI is aware that the clinical standard and user expectations do not necessarily point the same way.
The limitations: who holds the measuring stick?
The most principled question about MentalHealthBench is independence. All reported scores were produced by GPT-5.6 Sol — an OpenAI model — evaluating responses against criteria in a release published by OpenAI. No third-party replication is documented in the available source material. The benchmark is at least open, so other researchers can in principle run their own evaluations; whether anyone does, and whether the results then agree, is an open question.
Another limitation concerns real-world validity. TechRepublic points out that the scores do not show how often a model actually helps someone in a genuine mental health conversation. The dataset consists of synthetic conversations written for testing purposes, and the Japanese summary notes an explicit limitation from the announcement: a single score cannot represent the quality of all consultations, which encompass cultural and individual differences.
OpenAI's own framing is also worth noting. Dr. Declan Grabb, head of the company's mental health safety research, told The Deep View: "ChatGPT is not a therapist, and is not here to replace a clinician."
Why it matters
MentalHealthBench is first and foremost a measurement tool, not a safety certification. What it documents is that even with more than 80 clinical experts and more than 5,000 weighted criteria, frontier models' ability to conduct a good mental health conversation measures in the low-to-mid 50s percent — according to the company that built the test.
For users of AI chatbots, the practical conclusion is simple: even the best models do not handle these conversations reliably enough to replace clinical assessment, as Grabb's statement confirms. For the industry, the interesting part is the methodology — an open, expert-built standard for an arena where both the potential for harm and for benefit is high. And for those tracking AI companies' self-evaluation, this case is an example working in both directions: a genuine attempt at transparency with expert involvement, and at the same time a reminder that when the measuring stick is homemade, the results are best read as the company's own reporting — not as independent verification.
Sources
- OpenAI Mental Health AI Test Finds Context and Urgency Gaps — www.techrepublic.com
- [AI News 2026/09/25] OpenAI's AI accesses Australian government site without authorization|柴葉 剛|小説を書く人 — note.com
- OpenAI's new benchmark assesses AI's mental health conversation skills — www.newsbytesapp.com
- OpenAI Releases 'MentalHealthBench' to Measure AI Mental Health Responses, and More | AI News Summary for September 24, 2026|シン| AIに興味のある休職者 — note.com