In a post on X, @OpenAI describes MentalHealthBench as a new, openly released benchmark for evaluating AI conversations about mental health.

What MentalHealthBench is designed to cover

OpenAI says the benchmark is designed to span the “full spectrum” of mental-health conversations people bring to AI, from everyday support to more acute crisis scenarios. The company contrasts this scope with other mental-health benchmarks, which it says often focus on emergencies.

The announcement does not specify the benchmark’s tasks or how it assesses different kinds of conversations. That leaves important details about what a score represents and how to interpret performance unanswered.

Clinician input and open release

OpenAI says more than 80 mental health clinicians contributed input during the benchmark’s development. The company says it is releasing the benchmark openly so researchers can examine its methods, run their own evaluations and build on the work. The announcement does not describe the clinicians’ specific roles or provide terms for accessing or reusing the benchmark.

OpenAI presents the release as a way to demonstrate continued improvement by frontier models in realistic mental-health conversations. Benchmark results are evaluation outcomes; by themselves, they do not establish that a model is clinically effective or safe in real-world care.

What the accompanying graphics show

One graphic, titled “Overall model performance,” displays percentage scores for named model entries, with error bars. The visible scores range from 29.5% to 57.3%. The graphic does not define the score or error bars, or explain the evaluation procedure, so the figures alone do not show what a difference means. Although the post describes improvement over time, the chart itself is not a time series.

A second graphic, titled “Visualizing the space of mental health benchmarks,” presents four panels: real mental-health conversation summaries, existing mental-health benchmarks, OpenAI internal safety evaluations and MentalHealthBench. The panels show plotted points and legends naming examples or categories. The graphic does not explain its axes or how to interpret distances between points, so it illustrates how OpenAI groups these evaluation sources without establishing how comprehensive or clinically reliable any of them is.

Sources